DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over GUIZILINI et al. (US Pat. Pub. No. 20220292837 “Guizilini”) in view of YAMAMOTO et al. (US Pat. Pub. No. 20210112228 “Yamamoto”).
Regarding claim 13 Guizilini teaches, An electronic apparatus (Fig. 6) for three-dimensional (3D) reconstruction of an object by using view synthesis (“[0041] FIG. 2C illustrates an example of a 3D reconstruction 240 of the scene 202 according to aspects of the present disclosure. [0056]….. The view synthesis module 508 may generate the reconstructed image 510 based on the estimated depth and the estimated pose. The view synthesis module 508 may also be referred to as a scene reconstruction network”), the electronic apparatus comprising:
memory configured to store one or more instructions; and at least one processor configured to execute the one or more instructions, wherein the one or more instructions, when executed by the at least one processor (“[0006] Another aspect of the present disclosure is directed to an apparatus. The apparatus having a memory, one or more processors coupled to the memory, and instructions stored in the memory”), cause the electronic apparatus to:
obtain source images of a scene including an object (Fig. 5 shows source image 504 is input “[0056]….. the source image 502 may be an image at time step t−1. The view synthesis module 508 may generate the reconstructed image 510 based on the estimated depth and the estimated pose.”),
However Guizilini is silent about generate a target viewpoint based on a spatial distribution of source viewpoints corresponding to the source images, generate a target image corresponding to the target viewpoint by performing the view synthesis;
Yamamoto teaches generate a target viewpoint based on a spatial distribution of source viewpoints corresponding to the source images, generate a target image corresponding to the target viewpoint by performing view synthesis (“[0057] First, the communication unit 7 receives the image data transmitted by the imaging apparatus 2 (step S0). [0099] Camera position information indicates a position or a direction of each of the imaging apparatuses in the real space. For example, the camera position information may be information indicating a spatial position of each of the imaging apparatuses with reference to a predetermined position in the real space. The camera position information may be information indicating an imaging direction of each of the imaging apparatuses with respect to a predetermined direction. The camera position information may be information indicating an imaging angle of each of the imaging apparatuses.
[0142]…… More specifically, the viewpoint image generation unit 18 generates the target viewpoint image data from image information of the 360 video included in the intermediate 360 video set, the information required to obtain the target viewpoint image through synthesis”);
Guizilini and Yamamoto are analogous art as both of them are related to image processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Guizilini by generating a target viewpoint based on a spatial distribution of source viewpoints corresponding to the source images, generating a target image corresponding to the target viewpoint by performing view synthesis as taught by Yamamoto.
The motivation for the above is to have a 3d model seen from a specific viewpoint.
Guizilini modified by Yamamoto teaches generate a 3D model of the object by 3D reconstruction based on the source images and the target image (Guizilini “[0041] FIG. 2C illustrates an example of a 3D reconstruction 240 of the scene 202 according to aspects of the present disclosure. The 3D reconstruction may be generated from the depth map 220 as well as a pose of the target image 200 and a source image”).
Claim 1 is directed to a method claim and its steps are similar in scope and functions performed by the elements of apparatus claim 13 and therefore claim 1 is also rejected with the same rationale as specified in the rejection of claim 13.
Claim(s) 2-4 and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Guizilini modified by Yamamoto as applied to claims 1 and 13 above, and further in view of Shen et al. (US Pat. Pub. No. 20180260668 “Shen”).
Regarding claims 2 and 14 Guizilini modified by Yamamoto teaches generating a temporary target image corresponding to the target viewpoint by performing the view synthesis (Yamamoto “[0131] As a specific example of (4) above, for example, the intermediate 360 video set generation unit 10 selects and refers to the input image data (pixel at the position x.sub.cn,m on the image I.sub.vk) representing the image captured by the camera with the highest resolution relative to the imaging target (corresponding to the pixels at the position x.sub.vk,m of the image I.sub.vk) (or the camera closest to the imaging target) and generates intermediate viewpoint image data indicating the image of the imaging target. Alternatively, for example, the intermediate 360 video set generation unit 10 refers in preference to the input image data (pixel at the position x.sub.cn,m on the image representing the image captured by the camera”) but is silent about evaluating a quality of the temporary target image; and generating the target image by reperforming the view synthesis based on a result of the evaluating the quality of the temporary target image.
Shen teaches generating a temporary target image corresponding to target viewpoint by performing view synthesis; evaluating a quality of the temporary target image; and generating the target image by reperforming the view synthesis based on a result of the evaluating the quality of the temporary target image (“[0056] Harmonized images generated in accordance with training the image harmonizing neural network may be referred to herein as training harmonizing images. The generated harmonized images can be compared to a reference image to facilitate training of the image harmonizing neural network. In this regard, the image harmonizing neural network can be modified or adjusted based on the comparison such that the quality of subsequently generated harmonized images increases. Such training helps to maintain realism of the colors used when adjusting the foreground region(s) to match the background region, or vice versa, during the harmonization process. As used herein, a reference image refers to the unedited image prior to any alteration to synthesize a training composite image. Such a reference image is used as a standard, or ground-truth, for evaluating the quality of a harmonized image generated from a training composite image by the image harmonizing neural network. These ground-truth images can be referred to as reference images and/or reference harmonized images”);
Shen and Guizilini modified by Yamamoto are analogous art as both of them are related to image processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Guizilini modified by Yamamoto by generating a temporary target image corresponding to target viewpoint by performing view synthesis; evaluating a quality of the temporary target image; and generating the target image by reperforming the view synthesis based on a result of the evaluating the quality of the temporary target image as taught by Shen.
The motivation for the above is to generate a better and accurate target image.
Regarding claims 3 and 15 Guizilini modified by Yamamoto and Shen teaches wherein the generating the target image by reperforming the view synthesis based on the result of the evaluating the quality of the temporary target image comprises: adjusting a processing cost of the view synthesis based on the result of the evaluating the quality of the temporary target image; and generating the target image by performing the view synthesis with the adjusted processing cost (Shen “[0056] Harmonized images generated in accordance with training the image harmonizing neural network may be referred to herein as training harmonizing images. The generated harmonized images can be compared to a reference image to facilitate training of the image harmonizing neural network. In this regard, the image harmonizing neural network can be modified or adjusted based on the comparison such that the quality of subsequently generated harmonized images increases. Such training helps to maintain realism of the colors used when adjusting the foreground region(s) to match the background region, or vice versa, during the harmonization process. As used herein, a reference image refers to the unedited image prior to any alteration to synthesize a training composite image. Such a reference image is used as a standard, or ground-truth, for evaluating the quality of a harmonized image generated from a training composite image by the image harmonizing neural network. These ground-truth images can be referred to as reference images and/or reference harmonized images”).
Regarding claim 4 Guizilini modified by Yamamoto and Shen teaches wherein the view synthesis is based on a deep learning network, and wherein the processing cost of the view synthesis is adjusted by changing the processing cost of the deep learning network ((Shen “[0095] These comparisons can be used at block 708 where the neural network system can be adjusted using the determined loss functions. Errors determined using loss functions can be used to minimize loss in the neural network system by backwards propagation of such errors through the system. As indicated in FIG. 7, the foregoing blocks may be repeated any number of times to train the neural network system (e.g., using different training images and corresponding reference information for each iteration).”).
Claim(s) 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Guizilini modified by Yamamoto as applied to claims 1 and 13 above, and further in view of TAKESHITA et al. (US Pat. Pub. No. 20130202220 “Takeshita”).
Regarding claim 5 and 16 Guizilini modified by Yamamoto is silent about wherein the generating the target image corresponding to the target viewpoint by performing the view synthesis comprises: generating source depth images from the source images; generating masked source depth images by performing object masking on the source depth images; and generating the target image by performing the view synthesis by using the source images and the masked source depth images.
Takeshita teaches generating source depth images from the source images; generating masked source depth images by performing object masking on the source depth images; and generating the target image by performing the view synthesis by using the source images and the masked source depth images (Fig. 4 ABSTRACT “A mask correcting unit corrects an externally set mask pattern. A depth map processing unit processes a depth map of an input image for each of a plurality of regions designated by a plurality of mask patterns corrected by the mask correcting unit. An image generation unit generates an image of a different viewpoint on the basis of the input image and depth maps processed by the depth map processing unit”);
Takeshita and Guizilini modified by Yamamoto are analogous art as both of them are related to image processing.
Therefore it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Guizilini modified by Yamamoto by generating source depth images from the source images; generating masked source depth images by performing object masking on the source depth images; and generating the target image by performing the view synthesis by using the source images and the masked source depth images as taught by Takeshita.
The motivation for the above is to prevent artifacts thereby improve image quality.
Allowable Subject Matter
Claims 6-12 and 17-20 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claims 6 and 17 the best available prior art fails to expressly teach “generating a grid map on a coordinate system; matching the source viewpoints with cells of the grid map based on coordinate values of the source viewpoints on the coordinate system; and generating the target viewpoint to be matched to any one cell among the cells of the grid map that are not matched with the source viewpoints”.
Claims 7-12 and 18-20 are also objected by virtue of dependency.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Ahn et al. (US Pat. No.11756197) teaches Image generation and evaluating quality of the image using neural network and adjust cost of the neural network.
TANG et al. (US Pat. Pub. No. 20240395768) reconstructing 3d model based on source image and target image.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAPTARSHI MAZUMDER whose telephone number is (571)270-3454. The examiner can normally be reached 8 am-4 pm PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said Broome can be reached at (571)272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAPTARSHI MAZUMDER/ Primary Examiner, Art Unit 2612