Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Applicant’s election filed on April 07, 2026 is acknowledged. Currently claims 1-14 are pending.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on January 29, 2024 is in compliance with the provisions of 27 CFR 1.97. Accordingly, the information disclosure statement is being considered and attached by the examiner.
Election/Restrictions
Applicant's election with traverse of Species I (claims 1-14) in the reply filed on April 07, 2026 is acknowledged. The traversal is on the ground(s) that the Restriction Requirement fails to identify any different electronic resources to be used, any different search strategies that would be required, or any different search queries to be used. This is not found persuasive because the species or groupings of patentably indistinct species require a different field of search (e.g., searching different classes/subclasses or electronic resources, or employing different search strategies or search queries), such as Species I includes searching G06N 3/0895 (weakly supervised learning, e.g. semi-supervised or self-supervised learning) while Species II includes searching G06V 10/72 (data preparation, e.g. statistical preprocessing of image or video features).
The requirement is still deemed proper and is therefore made FINAL. Thus, the examiner will consider the elected claims 1-14.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-5, 7-12, and 14 are rejected under 35 U.S.C. 103 as being unpatentable of Sugano et al., US 20210368206 A1, (hereinafter “Sugano”) in view of Petersen et al., US 20240378698 A1, (hereinafter “Petersen”) in further view of Li et al., US 20250117995 A1, (hereinafter “Li”).
Regarding claim 1, Sugano teaches an electronic device comprising:
at least one memory configured to store a motion model for determining three-dimensional (3D) coordinates of pixels within a reference image based on estimated depths of the pixels within the reference image ([0165] “In a computer, a central processing unit (CPU) 101, a read only memory (ROM) 102, and a random access memory (RAM) 103 are mutually connected via a bus 104.”) ([0048] “The three-dimensional data acquisition unit 21 includes an image acquisition unit 44 and a three-dimensional model generating unit 43. The image acquisition unit 44 acquires a plurality of camera images in which a subject is imaged from a plurality of viewpoints, and also acquires a plurality of pieces of active depth information indicating a distance from another plurality of viewpoints to the subject.” wherein a motion model is a three-dimensional model and estimated depths of the pixels are pieces of active depth information); and
at least one processing device configured to ([0162] “Moreover, a program may be processed by a single CPU or processed in a distributed manner by a plurality of CPUs or graphics processing units (GPUs).”):
generate a ([0048] “Then, the three-dimensional model generating unit 43 generates three-dimensional model information representing a three-dimensional model of the subject on the basis of the plurality of camera images and the plurality of pieces of active depth information,” wherein a motion model is the three-dimensional model);
perform 3D to two-dimensional (2D) reprojection of the pixels within the reference image based on the motion model to generate a reprojected image view corresponding to a shifted camera perspective ([0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the motion model is a three-dimensional model);
generate a([0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images); and
([0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images).
Sugano does not specifically disclose generate a ground truth optical flow map for displacement of the pixels.
However, Petersen teaches generate a ground truth optical flow map for displacement of the pixels ([0055] “Using supervised learning as an illustrative example, a training dataset can include one or more videos and ground truth for the training can include optical flow maps for frames of the video. A loss (e.g., mean-squared error (MSE) or other loss) can be determined based on a comparison of estimated optical flow maps output by the flow estimation engine 514 with the ground truth optical flow maps.”) ([0056] “For instance, a dense optical flow can be computed between frames to generate optical flow vectors for each pixel in a frame, which can be included in a dense optical flow map. In some cases, the optical flow map can include vectors for less than all pixels in a frame. In some examples, the optical flow vector for a pixel can be a displacement vector (e.g., indicating horizontal and vertical displacements, such as horizontal (x-) and vertical (y-) displacements) showing the movement of a pixel from the previous low-resolution frame 505 to the input low-resolution frame 504.”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include a ground truth optical flow map of Petersen in the reprojection model of Sugano to approximate the two-dimensional motion field and depict pixel movement between frames, thereby allowing for more accurate 3D to 2d reprojection.
Sugano in view of Petersen does not specifically disclose generate an occlusion mask, the occlusion mask corresponding to occluded pixels; and
perform occlusion region inpainting of the occluded pixels to generate an inpainted image view.
However, Li teaches generate an occlusion mask, the occlusion mask corresponding to occluded pixels ([01965] “At operation 1505, the system obtains additional training data including an occlusion mask. In some cases, the operations of this step refer to, or may be performed by, a training component as described with reference to FIG. 5.”); and
perform occlusion region inpainting of the occluded pixels to generate an inpainted image view ([0033] “Furthermore, according to some aspects, the image generation system identifies an occlusion area for a modified view of the image based on the image and the depth map and uses the image generation machine learning model to generate a modified image by inpainting the occlusion area.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use techniques such as occlusion masks and inpainting of Li in the 3D to 2D reprojection model of Sugano in view of Petersen to identify regions of images that cannot be reliably reconstructed from the source data due to obstructions in the image. Using these techniques of Li aids in decreasing noise, improving training accuracy, and overall increasing the accuracy of reprojection.
Regarding claim 2, Sugano in view of Petersen and Li teaches the electronic device of claim 1, wherein:
the occlusion mask for the reference image comprises a reference occlusion mask (Li - [0195] “At operation 1505, the system obtains additional training data including an occlusion mask. In some cases, the operations of this step refer to, or may be performed by, a training component as described with reference to FIG. 5.” wherein a reference occlusion mask is a training occlusion mask);
the inpainted reprojected image view comprises a first of at least one inpainted reprojected image view (Li - [0033] “Furthermore, according to some aspects, the image generation system identifies an occlusion area for a modified view of the image based on the image and the depth map and uses the image generation machine learning model to generate a modified image by inpainting the occlusion area.”) (Sugano - [0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images); and
the at least one processing device is further configured to (Sugano - [0162] “Moreover, a program may be processed by a single CPU or processed in a distributed manner by a plurality of CPUs or graphics processing units (GPUs).”):
warp the at least one inpainted reprojected image view based on the ground truth optical flow map to generate one or more registered frames (Li - [0033] “Furthermore, according to some aspects, the image generation system identifies an occlusion area for a modified view of the image based on the image and the depth map and uses the image generation machine learning model to generate a modified image by inpainting the occlusion area.”) (Sugano - [0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images) (Petersen - [0055] “Using supervised learning as an illustrative example, a training dataset can include one or more videos and ground truth for the training can include optical flow maps for frames of the video. A loss (e.g., mean-squared error (MSE) or other loss) can be determined based on a comparison of estimated optical flow maps output by the flow estimation engine 514 with the ground truth optical flow maps.”) (Petersen - [0057] “A warping engine 516 of the video model 500 can warp a previously-upsampled frame 503 (e.g., upsampled by the video model 500) to generate a warped previously-upsampled frame 517 using the determined optical flow between the input low-resolution frame 504 and the previous low-resolution frame 505.”); and
generate a ground truth view occlusion mask using the reference image and the one or more registered frames (Li - [0196] “According to some aspects, the training component generates the occlusion mask by applying a mask to randomly selected pixels of a ground-truth image” wherein a ground truth view occlusion mask is the occlusion mask applied to a ground-truth image) (Petersen - [0057] “A warping engine 516 of the video model 500 can warp a previously-upsampled frame 503 (e.g., upsampled by the video model 500) to generate a warped previously-upsampled frame 517 using the determined optical flow between the input low-resolution frame 504 and the previous low-resolution frame 505.”).
The motivation for combining Sugano, Petersen, and Li is the same motivation as used for claim 1.
Regarding claim 3, Sugano in view of Petersen and Li teaches the electronic device of claim 1, wherein:
the at least one processing device is further configured to train an image registration model to generate a predicted flow map based on the reference image and the inpainted reprojected image view (Sugano - [0162] “Moreover, a program may be processed by a single CPU or processed in a distributed manner by a plurality of CPUs or graphics processing units (GPUs).”) (Petersen - [0055] “Using supervised learning as an illustrative example, a training dataset can include one or more videos and ground truth for the training can include optical flow maps for frames of the video. A loss (e.g., mean-squared error (MSE) or other loss) can be determined based on a comparison of estimated optical flow maps output by the flow estimation engine 514 with the ground truth optical flow maps.” wherein a predicted flow map is an estimated optical flow map) (Li - [0033] “Furthermore, according to some aspects, the image generation system identifies an occlusion area for a modified view of the image based on the image and the depth map and uses the image generation machine learning model to generate a modified image by inpainting the occlusion area.”) (Sugano - [0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images); and
the at least one processing device is configured to apply a supervised loss function between the predicted flow map and the ground truth optical flow map (Sugano - [0162] “Moreover, a program may be processed by a single CPU or processed in a distributed manner by a plurality of CPUs or graphics processing units (GPUs).”) (Li - [0179] “A loss function refers to a function that impacts how a machine learning model is trained in a supervised learning model. For example, during each training iteration, the output of the machine learning model is compared to the known annotation information in the training data. The loss function provides a value (the “loss”) for how close the predicted annotation data is to the actual annotation data. After computing the loss, the parameters of the model are updated accordingly and a new set of predictions are made during the next iteration.”) (Petersen - [0055] “Using supervised learning as an illustrative example, a training dataset can include one or more videos and ground truth for the training can include optical flow maps for frames of the video. A loss (e.g., mean-squared error (MSE) or other loss) can be determined based on a comparison of estimated optical flow maps output by the flow estimation engine 514 with the ground truth optical flow maps.” wherein a predicted flow map is an estimated optical flow map).
The motivation for combining Sugano, Petersen, and Li is the same motivation as used for claim 1.
Regarding claim 4, Sugano in view of Petersen and Li teaches the electronic device of claim 3, wherein the at least one processing device is further configured to:
warp the inpainted reprojected image view to the reference image based on the predicted flow map to generate a predicted image registration (Li - [0033] “Furthermore, according to some aspects, the image generation system identifies an occlusion area for a modified view of the image based on the image and the depth map and uses the image generation machine learning model to generate a modified image by inpainting the occlusion area.”) (Sugano - [0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images) (Petersen - [0055] “Using supervised learning as an illustrative example, a training dataset can include one or more videos and ground truth for the training can include optical flow maps for frames of the video. A loss (e.g., mean-squared error (MSE) or other loss) can be determined based on a comparison of estimated optical flow maps output by the flow estimation engine 514 with the ground truth optical flow maps.” wherein a predicted flow map is an estimated optical flow map) (Petersen - [0057] “A warping engine 516 of the video model 500 can warp a previously-upsampled frame 503 (e.g., upsampled by the video model 500) to generate a warped previously-upsampled frame 517 using the determined optical flow between the input low-resolution frame 504 and the previous low-resolution frame 505.” wherein a predicted image registration is the warped frame); and
remove or inpaint one or more occluded regions within the predicted image registration based on a ground truth view occlusion mask to generate a registered frame for the predicted image registration (Li - [0033] “Furthermore, according to some aspects, the image generation system identifies an occlusion area for a modified view of the image based on the image and the depth map and uses the image generation machine learning model to generate a modified image by inpainting the occlusion area.”) (Li - [0195] “At operation 1505, the system obtains additional training data including an occlusion mask. In some cases, the operations of this step refer to, or may be performed by, a training component as described with reference to FIG. 5.”) (Petersen - [0057] “A warping engine 516 of the video model 500 can warp a previously-upsampled frame 503 (e.g., upsampled by the video model 500) to generate a warped previously-upsampled frame 517 using the determined optical flow between the input low-resolution frame 504 and the previous low-resolution frame 505.” wherein a predicted image registration is the warped frame).
The motivation for combining Sugano, Petersen, and Li is the same motivation as used for claim 1.
Regarding claim 5, Sugano in view of Petersen and Li teaches the electronic device of claim 4, wherein the at least one processing device is configured to apply a self-supervised loss function between the reference image and the registered frame for the predicted image registration during the training of the image registration model (Sugano - [0162] “Moreover, a program may be processed by a single CPU or processed in a distributed manner by a plurality of CPUs or graphics processing units (GPUs).”) (Li - [0179] “A loss function refers to a function that impacts how a machine learning model is trained in a supervised learning model. For example, during each training iteration, the output of the machine learning model is compared to the known annotation information in the training data. The loss function provides a value (the “loss”) for how close the predicted annotation data is to the actual annotation data. After computing the loss, the parameters of the model are updated accordingly and a new set of predictions are made during the next iteration.”) (Petersen - [0057] “A warping engine 516 of the video model 500 can warp a previously-upsampled frame 503 (e.g., upsampled by the video model 500) to generate a warped previously-upsampled frame 517 using the determined optical flow between the input low-resolution frame 504 and the previous low-resolution frame 505.” wherein a predicted image registration is the warped frame) (Petersen - [0055] “Using supervised learning as an illustrative example, a training dataset can include one or more videos and ground truth for the training can include optical flow maps for frames of the video. A loss (e.g., mean-squared error (MSE) or other loss) can be determined based on a comparison of estimated optical flow maps output by the flow estimation engine 514 with the ground truth optical flow maps.” wherein a predicted flow map is an estimated optical flow map).
The motivation for combining Sugano, Petersen, and Li is the same motivation as used for claim 1.
Regarding claim 7, Sugano in view of Petersen and Li teaches the electronic device of claim 1, wherein:
the occlusion mask is based on the reference image and the reprojected image view (Li - [0196] “According to some aspects, the training component generates the occlusion mask by applying a mask to randomly selected pixels of a ground-truth image”) (Sugano - [0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images); and
the occluded pixels in the reprojected image view correspond to one or more portions of a first object in the reference image that were occluded by a second object closer than the first object based on a depth map corresponding to the estimated depths of the pixels within the reference image (Sugano - [0049] “The two-dimensional image conversion processing unit 22 for example performs two-dimensional image conversion processing in which the three-dimensional model represented by the three-dimensional model information supplied from the three-dimensional data acquisition unit 21 is subjected to perspective projection from a plurality of directions and converted into a plurality of two-dimensional images.” wherein the reprojected image view is the converted plurality of two-dimensional images) (Li - [0035] “In the example, the image generation system computes a camera view of the image and of the depth map, and shifts the camera view to obtain a modified view for an occluded image and an occluded depth map. Each of the occluded image and the occluded depth map include “missing” pixels lacking information caused by occlusion from the modified view. The image generation system uses the image generation model to generate a modified image and a modified depth map including an alternative view of the cute puppy depicted in the image by inpainting the missing pixels of the occluded image and the occluded depth map.”).
The motivation for combining Sugano, Petersen, and Li is the same motivation as used for claim 1.
Regarding claim 8, the claim recites similar limitations to claim 1 but in the form of a method. Therefore, claim 8 recites similar limitations to claim 1 and is rejected for similar rationale and reasoning (see the analysis for claim 1 above).
Regarding claim 9, the claim recites similar limitations to claim 2 but in the form of a method. Therefore, claim 9 recites similar limitations to claim 2 and is rejected for similar rationale and reasoning (see the analysis for claim 2 above).
Regarding claim 10, the claim recites similar limitations to claim 3 but in the form of a method. Therefore, claim 10 recites similar limitations to claim 3 and is rejected for similar rationale and reasoning (see the analysis for claim 3 above).
Regarding claim 11, the claim recites similar limitations to claim 4 but in the form of a method. Therefore, claim 11 recites similar limitations to claim 4 and is rejected for similar rationale and reasoning (see the analysis for claim 4 above).
Regarding claim 12, the claim recites similar limitations to claim 5 but in the form of a method. Therefore, claim 12 recites similar limitations to claim 5 and is rejected for similar rationale and reasoning (see the analysis for claim 5 above).
Regarding claim 14, the claim recites similar limitations to claim 7 but in the form of a method. Therefore, claim 14 recites similar limitations to claim 7 and is rejected for similar rationale and reasoning (see the analysis for claim 7 above).
Allowable Subject Matter
Claims 6 and 13 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMANDA PEARSON whose telephone number is (703)-756-5786. The examiner can normally be reached Monday - Friday 9:00 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached on (571)- 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AMANDA H PEARSON/Examiner, Art Unit 2666
/MING Y HON/Primary Examiner, Art Unit 2666