DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-7 and 10-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yi US20210150678 (hereinafter “Yi”) in view of Dang et al “A perceptual image completion approach based on hierarchical optimization scheme” (hereinafter “Dang”).
Regarding claim 1, Yi teaches a computer-implemented method for image augmentation, the method comprising (see paragraph 0033, a machine learning technique implemented on a computer device for image inpainting [augmenting]):
obtaining, by a computing system comprising one or more processors (see paragraph 0033 and 0064, a computer device for image inpainting with the use of a processor that implements the computation module and the inpainting subsystem), input data and one or more input masks, wherein the input data is descriptive of data with one or more occlusions, and wherein the one or more input masks are associated with the one or more occlusions (see paragraph 0074 and Figure 3, the obtaining of an inpainting mask and a high resolution images [input data] as an input. The high-resolution image has areas to be inpainted [occlusions]. The inpainting mask indicates portions to be inpainted [interpreted as occluded areas]. The inpaintings are objects to be removed or missing portions in the frame, 0034);
processing, by the computing system (see paragraph 0033, a machine learning technique implemented on a computer device for image inpainting), the input data and the one or more input masks with an image augmentation model to generate intermediate data (see paragraph 0079 and Figure 3, the inpainting generator[augmentation model] processes the low resolution image [down sampled from the original high resolution input, 0075] and the inpainting mask to generate a low resolution inpainted image and attention score [intermediate data]), wherein the intermediate data is descriptive of the input data with mask data associated with the one or more occlusions replaced by predicted data (see paragraph 0079 and 0090-0091, the low resolution inpainted image and attention scores are based on the coarse inpainting output [predicted data] that is added to the low resolution image to replace the mask area);
processing, by the computing system (see paragraph 0033, a machine learning technique implemented on a computer device for image inpainting), the intermediate data and the one or more input masks with a texture transfer block to generate refined output data (see paragraph 0083 and Figure 3, the low resolution inpainted image and the attention score outputted from the inpainting generator are subsequently processed in steps 308-316 to generate the high resolution inpainted image with the use of the attention transfer module 308 [texture transfer module]), wherein the texture transfer block augments the masked data based at least in part on the intermediate data (see paragraph 0104-0105, the attention-transfer module 308 used the attention score [part of the intermediate data] to aggregate the contextual residual information to the masked region and produce the aggregated residual image),
providing, by the computing system (see paragraph 0033, a machine learning technique implemented on a computer device for image inpainting), the refined output data as an output (see paragraph 0084 and Figure 3, the output high-resolution inpainted image).
Yi does not teach the texture transfer block generates a plurality of downsampled versions of the input data and determines shift data for each of the plurality of downsampled versions of the input data, wherein the shift data is descriptive of source data in the input data utilized for replacing occlusion data, wherein respective shift data for at least one or more of the plurality of downsampled versions is based at least in part on the intermediate data.
Dang teaches a computer-implemented method for image augmentation, the method comprising (see section 1, an image inpainting method):
obtaining input data, wherein the input data is descriptive of data with one or more occlusions (see section 2.1 and Table 1, input of image I [original image] with an inpainting region into the algorithm, wherein the image is to be inpainted),
the texture transfer block generates a plurality of downsampled versions of the input data (see section 2.1, generate a Gaussian pyramid [a set of images with various levels of detail] from the input image by downsampling the images) and determines shift data for each of the plurality of downsampled versions of the input data (see section 2.3 and Table 1, generate a shift map [offset map] for the levels of the pyramids [the levels of the pyramid are downsampled images]), wherein the shift data is descriptive of source data in the input data utilized for replacing occlusion data (see section 2.3 and Equation 7, the offset map [shift map] defines the relationship between the pixels to be inpainted [occlusion data] and the pixels in the known regions [input data]), wherein respective shift data for at least one or more of the plurality of downsampled versions is based at least in part on the intermediate data (see section 2.3, an offset map [shift map] is generated from the inpainting of the lowest resolution image [intermediate data]); and
providing, by the computing system, the refined output data as an output (see section 5, a final output image with improved visual quality [refined output data]).
Yi and Dang are analogous art because they are from the same field of endeavor of computer vision and image processing with the focus of image inpainting using intermediate data of low-resolution inpainted image before outputting an improved quality image.
Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Yi to generate a plurality of downsampled versions of the input data as taught by Dang. The motivation for doing so would have been to allow for a top-down completion method that allows the lowest resolution level to restore the damaged region to allows for accounting human perception at the higher resolution (Dang, Section 1).
Regarding claim 2, Yi and Dang teach the method of claim 1.
Yi teaches processing the intermediate data and the one or more input masks with the texture transfer block to generate the refined output data (see paragraph 0083 and Figure 3, The low resolution inpainted image and the attention score outputted from the inpainting generator are subsequently processed in steps 308-316 to generate the high resolution inpainted image with the use of the attention transfer module 308 [texture transfer module]
Yi does not teach generating an image pyramid comprising the plurality of downsampled versions of the input data, wherein the plurality of downsampled versions of the input data are descriptive of lower resolution versions of the input data with the one or more occlusions masked.
Dang teaches generating an image pyramid comprising the plurality of downsampled versions of the input data, wherein the plurality of downsampled versions of the input data are descriptive of lower resolution versions of the input data with the one or more occlusions masked (see section 2.1, generate a Gaussian pyramid [a set of images with various levels of detail] from the input image [see Figure 12, the image to be inpainted can include masked occlusions] by down sampling the images).
Regarding claim 3, Yi and Dang teach the method of claim 2.
Dang teaches the shift data comprises a set of integer vectors (see equation 7 and section 2.4, the offset map is expressed as
(
∆
x
,
∆
y
)
for each pixel which is a 2-dimensional vector, the limits of each coordinate of the shift are [-a, a], yielding
(
2
a
+
1
)
2
possible labels [the inclusion of the limits of the each coordinate and the number of possible limits interpreted as the shift vectors are integers]).
Regarding claim 4, Yi and Dang teach the method of claim 1.
Dang teaches the texture transfer block comprises a plurality of upsample shifts (see section 2.4, offset map values are upscaled to match the image at a higher resolution).
Regarding claim 5, Yi and Dang teach the method of claim 4.
Dang teaches the plurality of upsampled shifts comprise scaling the shift data based at least in part on a resolution difference between downsampled versions (see section 2.4, the offset-map values are upscaled to match the image at a higher resolution. As the offset map values are upscaled to match the next higher-resolution, the factor to upscale the values need to be based on the difference between the current levels resolution and the resolution of the higher-resolution level).
Regarding claim 6, Yi and Dang teach the method of claim 1.
Yi teaches the one or more input masks identify masked pixels that correspond to one or more objects (see paragraph 0074, the inpainting mask indicates the pixels of the original high-resolution image to be impinged [paragraph 0033-0034, the image inpainting methods may be applied to the removal or repositioning of an object in the high-resolution image, thus the method is inpainting objects of the image and the inpainting mask is of the object as the objects are areas to be inpainted]).
Regarding claim 7, Yi and Dang teach the method of claim 6.
Yi teaches the one or more occlusions comprise the one or more objects (see paragraph 0074, the inpainting mask indicates the pixels of the original high-resolution image to be inpainted [occluded areas].paragraph 0033-0034, the image inpainting methods may be applied to the removal or repositioning of an object in the high-resolution image, thus the method is inpainting objects of the image and the inpainting mask is of the object as the objects are areas to be inpainted).
Regarding claim 10, Yi and Dang teach the method of claim 1.
Dang teaches the refined output data comprises a refined augmented image (see paragraph 0084 and Figure 3, the output high-resolution inpainted [augmented] image).
Regarding claim 11, Yi teaches a computing system for image augmentation, the system comprising (see paragraph 0033, a machine learning technique implemented on a computer device for image inpainting):
one or more processors (see paragraph 0064, the use of a processor that implements the computation module and the inpainting subsystem);
and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising (see paragraph 0062, the execution device [the neural network processor may be provided in the execution device, 0064] may implement code [instructions] from the storage system to perform the processing):
obtaining input data and one or more input masks, wherein the input data is descriptive of data with one or more occlusions, and wherein the one or more input masks are associated with the one or more occlusions (see paragraph 0074 and Figure 3, the obtaining of an inpainting mask and a high-resolution image [input data] as an input. The high-resolution image has areas to be inpainted [occlusions]. The inpainting mask indicates portions to be inpainted [interpreted as occluded areas]);
processing the input data and the one or more input masks with an image augmentation model to generate intermediate data (see paragraph 0079 and Figure 3, the inpainting generator[augmentation model] processes the low resolution image [downsampled from the original high resolution input, 0075] and the inpainting mask to generate a low resolution inpainted image and attention score [intermediate data]), wherein the intermediate data is descriptive of the input data with mask data associated with the one or more occlusions replaced by predicted data(see paragraph 0079 and 0090-0091, the low resolution inpainted image and attention scores are based on the coarse inpainting output [predicted data] that is added to the low resolution image to replace the mask area);
processing the intermediate data and the one or more input masks with a texture transfer block to generate refined output data (see paragraph 0083 and Figure 3, The low resolution inpainted image and the attention score outputted from the inpainting generator are subsequently processed in steps 308-316 to generate the high resolution inpainted image with the use of the attention transfer module 308 [texture transfer module]), wherein the texture transfer block augments the masked data based at least in part on the intermediate data (see paragraph 0104-0105, the attention-transfer module 308 used the attention score [part of the intermediate data] to aggregate the contextual residual information to the masked region and produce the aggregated residual image),
providing the refined output data as an output (see paragraph 0084 and Figure 3, the output high-resolution inpainted image).
Yi does not teach the texture transfer block generates a plurality of downsampled versions of the input data and determines shift data for each of the plurality of downsampled versions of the input data, wherein the shift data is descriptive of source data in the input data utilized for replacing occlusion data, wherein respective shift data for at least one or more of the plurality of downsampled versions is based at least in part on the intermediate data.
Dang teaches obtaining input data, wherein the input data is descriptive of data with one or more occlusions (see section 2.1 and Table 1, input of image I [original image] into the algorithm, wherein the image is to be inpainted),
the texture transfer block generates a plurality of downsampled versions of the input data (see section 2.1, generate a Gaussian pyramid [a set of images with various levels of detail] from the input image by downsampling the images) and determines shift data for each of the plurality of downsampled versions of the input data (see section 2.3 and Table 1, generate a shift map [offset map] for the levels of the pyramids [the levels of the pyramid are downsampled images]), wherein the shift data is descriptive of source data in the input data utilized for replacing occlusion data (see section 2.3 and Equation 7, the offset map [shift map] defines the relationship between the pixels to be inpainted [occlusion data] and the pixels in the known regions [input data]), wherein respective shift data for at least one or more of the plurality of downsampled versions is based at least in part on the intermediate data (see section 2.3, an offset map [shift map] is generated from the inpainting of the lowest resolution image [intermediate data]); and
providing the refined output data as an output (see section 5, a final output image with improved visual quality [refined output data]).
Yi and Dang are analogous art because they are from the same field of endeavor of computer vision and image processing with the focus if image inpainting using intermediate data of low-resolution inpainted image before outputting an improved quality image.
Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Yi to generate a plurality of downsampled versions of the input data as taught by Dang. The motivation for doing so would have been to allow for a top-down completion method that allows the lowest resolution level to restore the damaged region to allows for accounting human perception at the higher resolution (Dang, Section 1).
Regarding claim 12, Yi and Dang teach the system of claim 11.
Dang teaches the texture transfer block constructs an image pyramid of the input data at different resolutions (see section 2.1, generate a Gaussian pyramid [a set of images with various levels of detail] from the input image).
Regrading claim 13, Yi and Dang teach the system of claim 12.
Dang teaches generating the image pyramid comprises (see section 2.1, the construction of the Gaussian pyramid of images):
processing the input data and the one or more input masks to generate a first downsampled image, wherein the first downsampled image is generated by weighting pixels of the input image to reduce the resolution by half (see section 2.1 and 2.4, the input image with the inpainting region is down sampled. To upscale between offset map levels of the pyramid the values doubled which is interpreted as the scaling factor between resolution of the pyramid level is 2 to be upscaled and thus 0.5 to down sampled);
and processing the first downsampled image to generate the second downsampled image (see section 2.1, the creation of the pyramid is based on the down sampling of the preceding level of the pyramid).
Regarding claim 14, Yi and Dang teach the system of claim 13.
Dang teaches the operations further comprise:
processing the first downsampled image in the image pyramid to generate a first output, wherein the first output comprises first shift data, wherein the first shift data is descriptive a set of first vectors associated with the copying and placing of pixels for replacing occlusion pixels (see section 2.1 and Table 1, the down sampling the original image to construct a gaussian pyramid of images and using the pyramid of image shift/offset maps [the offset map defining the relationship between the pixels to be inpainted and the pixels in the known region. The offset map is used to keep track of the coping process, see section 2.2.2-2.3]. See equation 7 and section 2.4, the offset map is expressed as
(
∆
x
,
∆
y
)
for each pixel which is a 2-dimensional vector);
generating upsampled first shift data based at least in part on the first shift data and a resolution difference between the first downsampled image and a second downsampled image (see section 2.4, the offset-map values are upscaled to match the image at a higher resolution. As the offset map values are upscaled to match the next higher-resolution, the factor to upscale the values need to be based on the difference between the current levels resolution and the resolution of the higher-resolution level); and
wherein the refined output data is generated based at least in part on the upsampled first shift data (see section 5, the smooth final output image is achieved from the offset map which is produced and interpolated [interpolated using nearest neighbor algorithm and the upscaled, see section 2.4] for higher resolution images).
Regarding claim 15, Yi and Dang teach the system of claim 11.
Dang teaches the texture transfer block constructs an image pyramid of the input data at different sizes (see section 2.1, the construction of a gaussian pyramid of the image to be inpainted where the number of pyramid levels is linked to the original size of the image and the minimum size is adjusted based on the width and height of the original image. Thus, it is interpreted that the pyramid can be constructed of input image at different sizes as the pyramid parameters are adjusted based on the size of the input image).
Regarding claim 16, Yi and Dang teach the system of claim 11.
Yi does not teach the texture transfer block comprises a non-machine-learning block.
Dang teaches the texture transfer block comprises a non-machine-learning block (see abstract and section 2, the framework is based on a greedy strategy and a global optimization strategy using exemplar-based method. Dang makes no reference to a machine learning block and additionally contrasts the framework and method from the relevant art of dictionary learning approaches in section 1).
Claim 17 is analogous to claims 1 and 11, and thus similar analyzed and rejected as claims 1 and 11.
Regarding claim 18, Yi and Dang teach the one or more non-transitory computer-readable media of claim 17.
Yi teaches the texture transfer block transfers high resolution pixels to one or more occlusion areas (see paragraph 0083 and Figure 3, the aggregated residual image of the attention transfer module is added to the low frequency inpainted inside mask area. The low resolution inpainted image and the attention score output from the inpainting generator is subsequently processed in steps 308-316 to generate the high resolution inpainted image with the use of the attention transfer module 308 [texture transfer module]).
Regarding claim 19, Yi and Dang teach the one or more non-transitory computer-readable media of claim 17.
Yi the refined output data is provided for display via a user interface (see paragraph 0060 and 0143 the execution device is provided with an I/O interface [the I/O interface allows interaction between the user and the execution device], the execution device outputs the high-resolution inpainted image to the user).
Regarding claim 20, Yi and Dang teach the one or more non-transitory computer-readable media of claim 17.
Yi teaches the operations further comprise:
storing the refined output data to a memory of a user computing device (see paragraph 0084, the high frequency inpainted image may be stored in the data storage of the execution device. The system may be implemented on a personal computing device [see 0034], thus the data storage of the execution device is on the personal computing device).
Claim 8-9 is rejected under 35 U.S.C. 103 as being unpatentable over Yi in view of Dang in view of Li US20210241432 (hereinafter “Li”).
Regarding claim 8, Yi and Dang teach the method of claim 1.
Yi nor Dang teach the one or more input masks comprise a bystander mask and a main subject mask.
Li teaches the one or more input masks comprise a bystander mask and a main subject mask (see paragraph 0090 and Figure 7 and 8, the identification of two main object masks A1 and B1, which may have different priorities leading to a target mask [main subject mask]. It is interpreted that the non-target mask could be considered the bystander before being re-masked and combined with the background mask see paragraph 0071).
Li, Yi, and Dang are analogous art because they are from the same field of endeavor of computer vision and image processing with the focus of segmented regions to be enhanced or modified to obtain a final target image.
Before the effective filling date of the invention, it would have been obvious to one of ordinary skill in the art to modify Yi and Dang to include bystander and a main subject mask as taught by Li. The motivation for doing so would have been to give priority to different objects based on the categories the object belongs to (Li, paragraph 0090).
Regarding claim 9, Yi, Li, and Dang teach the method of claim 8.
Li teaches the main subject mask is subtracted from the bystander mask to generate a distractor mask for processing (see paragraph 0090 and Fig 7-8, the determination of a plurality of object masks and a background mask [distractor mask]. The background mask is the remaining area without the object masks as seen in Figure 7 and 8. See paragraph 0100 the areas are then processes to obtain a target image).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see the attached 892-notice of references cited.
Contact information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to EMILY R. HAUK whose telephone number is (571)272-5966. The examiner can normally be reached M-F 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/EMILY R HAUK/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669