DETAILED ACTION
This action is in response to the application filed on October 21st, 2024. Claims 1-15 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on October 21st, 2024 is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 7, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over “Fourier Space Losses for Efficient Perceptual Image Super-Resolution” (herein after referred to by its primary author, Fuoli) in view of US20240179428 (herein after referred to by its primary author, Ater) and “High-Resolution Metalens Imaging with Sequential Artificial Intelligence Models” (herein after referred to by its primary author, Hsu).
In regards to claim 1, Fuoli teaches an image restoration device comprising: a memory that stores an original image captured by a camera and an artificial intelligence model; and a processor that trains the artificial intelligence model (Fuoli Section 4.2.1 “Our losses significantly surpass all three RankSRGAN models in both restoration metrics PSNR/SSIM and even achieve the highest FID score. Only the NIQE and PI optimized models have slightly higher LPIPS scores, which however comes with a 2.4x higher runtime on GPU.” Examiner note: GPU’s are known to contain a processor and memory), wherein the artificial intelligence model includes: an image restoration model that generates a restored image based on input data (Fuoli Figure 2 “Spatial Domain” G(x)); a discrimination model that Fourier-transforms the restored image generated by the image restoration model and distinguishes the Fourier-transformed restored image (Fuoli Figure 2 “Fourier Domain” and Figure 3), and the processor is configured to: perform adversarial learning of the image restoration model and the discrimination model (Fuoli Section 3.3.1 “In order to further boost the perceptual quality we employ a GAN training scheme with two types of GAN-architectures, applied in spatial and Fourier domain.” Examiner note: GAN stand for general adversarial network) and generate a restored image for a new original image captured by the camera from the image restoration model that has completed training (Fuoli Figure 2 Examiner note: Once the image restoration model of this disclosure has employed training, it can be applied to any input image).
Fuoli does not teach an image restoration model that crops the original image into preset patches and generates a restored image based on input data in which coordinate information of the patches is embedded into a cropped patch image. Furthermore, while not required by the claim, Fuoli does not teach an image restoration device which takes an image taken by a camera including a metalens. This limitations is required by the remaining independent claims and the invention described in the specification, and therefore will be shown to be taught for this claim.
Ater teaches an image restoration model that crops the original image into preset patches (Ater Paragraph [0060] “Steps S504 and S505 may then be repeated for other receptive patches of the input sample image. For example, if the sample image is divided into sixteen patches, and step S506 determines that only the first patch has been operated on so far, the method of FIG. 10 resumes to step S404 to operate on the second patch.”) and generates a restored image based on input data in which coordinate information of the patches is embedded into a cropped patch image (Ater Figure 7; Paragraph [0053] “Features of the patch and coordinates of the patch may be input to CoordConv layers 410. The CoordConv layers 410 may take an input patch from a RAW image and concatenate the x-y coordinates layers.”).
Ater is considered to be analogous to the claimed invention because they are in the same field of image enhancement. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the system of Fuoli to include the teachings of Ater, to provide the advantage of a flexible system which can learn to restore the input image regardless of position (Ater Paragraph [0044] “In an embodiment, the CNN is of type CoordConv, which aims to solve spatially variant problems by modifying equivariant convolutional layers. It works by giving the convolution access to its own input coordinates through the use of extra coordinate channels. CoordConv allows networks to learn either complete translation invariance or varying degrees of translation dependence, as required by the end task.”).
Furthermore, Hsu teaches an image restoration model that generates a restored image based on input data from a camera containing a metalens (Hsu Figure 4); and the processor is configured to generate a restored image for a new original image captured by the camera from the image restoration model that has completed training. (Hsu Figure 5).
Hsu is considered to be analogous to the claimed invention because they are both in the same field of metalens image enhancement. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the system of Fuoli in view of Ater to include the teachings of Hsu, to provide the advantage of efficient improvements in overall image quality (Hsu Page 11615 Right Column “The Autoencoder and CodeFormer sequential models efficiently repaired the images, resulting in significant enhancements in the overall image quality. Furthermore, the steps taken in designing the metalens used in this study were relatively straightforward and, as such, so were both the fabrication and subsequent implementation, making it a promising approach for a range of practical applications within imaging technology.”)
In regards to claim 7, Fuoli in view of Ater and Hsu teaches a display that outputs a restored image (Ater Paragraph [0101] “Thereafter, the application processor 1200 may read the encoded image signal from the internal memory 1230 or the external memory 1400, decode the encoded image signal, and display image data generated based on a decoded image signal.”) and renders obvious the remaining limitations as in the consideration of claim 1.
In regards to claim 12, Fuoli in view of Ater and Hsu teaches an image restoration method comprising: comparing the Fourier-transformed restored image with the correct image through the discrimination model, and based on the comparison result, causing the image restoration model to perform adversarial learning so that the image restoration model generates a high-frequency restored image from a low-frequency original image (Fuoli Figure 2 Examiner note: This figure shows generating an adversarial loss in the Fourier Domain, that loss is then used to train the network to train the image transformation model) and renders obvious the remaining claim limitations as in the consideration of claim 1.
Claims 2-3, 8-9, and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Fuoli in view of Ater and Hsu as applied to the claims above, and further in view of “meshgrid: 2-D and 3-D grids” (herein after referred to as MATLAB) and “CrackFormer: Transformer Network for Fine-Grained Crack Detection”.
In regards to claim 2, Fuoli in view of Ater and Hsu teaches the image restoration device according to claim 1, wherein the processor is configured to: generate the coordinate information based on a middle pixel of the patch image (Ater Paragraph [0053] “In an embodiment, the prediction 460 is an error value of a center pixel of the patch, but is not limited thereto. For example, the error value could correspond to any pixel of the patch. Features of the patch and coordinates of the patch may be input to CoordConv layers 410. The CoordConv layers 410 may take an input patch from a RAW image and concatenate the x-y coordinates layers. “); and generate the input data by concatenating the coordinate information converted into 2D coordinate data to the patch image (Ater Figure 7 “Concatenate Channels”).
Fuoli in view of Ater and Hsu does not teach convert[ing] the coordinate information into 2D coordinate data through a meshgrid method; and generat[ing] the input data by concatenating the coordinate information converted into 2D coordinate data to the patch image through a 1x1 convolution layer.
However, MATLAB teaches convert[ing] the coordinate information into 2D coordinate data through a meshgrid method (MATLAB “meshgrid” method Examiner note: The meshgrid method could be used to create the “i coordinate” and “j coordinate” of Ater Figure 7).
Fuoli in view of Ater and Hsu teaches an image restoration device which fuses coordinate information with a cropped patch image. The coordinate information is represented as two matrices each containing one half of the coordinate information, one containing the j (or x) coordinates, and the other containing the i (or y) coordinates. The claimed device differs from this by using a meshgrid method to create the coordinate information. MATLAB teaches a meshgrid(x,y) method which is a known matlab function which outputs 2 matrices, one containing x coordinate information and the other containing y coordinate information (MATLAB “2-D Grid” example). One of ordinary skill in the art could have substituted the method of creating the coordinate information of Fuoli in view of Ater and Hsu with the method of creating the coordinate information of MATLAB, and the results of this substitution would have been predictable because the coordinate information would be in the same form, and therefore any further processing could be performed on either coordinate information. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have used MATLAB’s meshgrid method to create the coordinate information used in Fuoli in view of Ater and Hsu.
Furthermore, Liu teaches generat[ing] the input data by concatenating the coordinate information converted into 2D coordinate data to the patch image through a 1x1 convolution layer (Liu Figure 3 “Position embedding” Examiner note: The input patch image passes through a 1x1 convolutional layer and is concatenated with the position embedding).
Fuoli in view of Ater, Hsu, and MATLAB teaches an image restoration device which fuses coordinate information with a cropped patch image. The claimed device differs from this by using a 1x1 convolutional layer to concatenate the patch image and the coordinate information. Liu teaches a method of concatenating a positional embedding to a patch image passed through a 1x1 convolutional layer. One of ordinary skill in the art could have substituted the method of concatenating the coordinate information of Fuoli in view of Ater, Hsu, and MATLAB with the method of concatenating the coordinate information of Liu, and the results of this substitution would have been predictable because both concatenation method result in a patched image with coordinate information embedded. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have used Liu’s concatenation method to concatenate the coordinate information used in Fuoli in view of Ater, Hsu, and MATLAB.
In regards to claim 3, Fuoli in view of Ater, Hsu, MATLAB, and Liu teaches the image restoration device according to claim 2, wherein the processor inputs the input data output from the 1x1 convolution layer into a first neural network whose output data has the same size as the input data, and outputs the restored image from the first neural network (Hsu Figure 4 Examiner note: The position embedded patch image of Fuoli in view of Ater, Hsu, MATLAB, and Liu would be used as input to the image transformation of figure 4).
In regards to claim 8, Fuoli in view of Ater, Hsu, MATLAB, and Liu renders obvious the claim limitations as in the consideration of claim 2.
In regards to claim 9, Fuoli in view of Ater, Hsu, MATLAB, and Liu renders obvious the claim limitations as in the consideration of claim 3.
In regards to claim 13, Fuoli in view of Ater, Hsu, MATLAB, and Liu renders obvious the claim limitations as in the consideration of claim 2.
In regards to claim 14, Fuoli in view of Ater, Hsu, MATLAB, and Liu renders obvious the claim limitations as in the consideration of claim 3.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Fuoli in view of Ater, Hsu, MATLAB, and Liu as applied to the claims above, and further in view of “Learning Enriched Features for Real Image Restoration and Enhancement” (herein after referred to by its primary author, Zamir).
In regards to claim 4, Fuoli in view of Ater, Hsu, MATLAB, and Liu teaches the image restoration device according to claim 3, wherein the first neural network includes a CNN (Convolution Neural Network) model (Hsu Page 11617 Left Column “For the encoder component, input images of size 256 × 256 × 3 (pixel × pixel × color channel) were employed. The encoder consists of four convolutional operations, each using a 3 × 3 kernel, a stride of 2, and a rectified linear unit activation function.”).
Fuoli in view of Ater, Hsu, MATLAB, and Liu does not teach wherein the first neural network includes a CNN (Convolution Neural Network) model including at least one of MIRNet, MPRNet, and NAFNet, or a Transformer model including at least one of Restormer and Uformer.
However, Zamir teaches wherein the first neural network includes a CNN (Convolution Neural Network) model including at least one of MIRNet, MPRNet, and NAFNet, or a Transformer model including at least one of Restormer and Uformer (Zamir Figure 1; Figure 1 Description “The proposed network MIRNet is based on a recursive residual design.”).
Fuoli in view of Ater, Hsu, MATLAB, and Liu teaches an image restoration device which uses a convolutional neural network (CNN) to perform image restoration. The claimed device differs from this by using a more specific form of CNN, MIRNet, MPRNet, and NAFNet specifically. Zamir teaches that performing image restoration using MIRNet is known. One of ordinary skill in the art could have substituted the CNN of Fuoli in view of Ater, Hsu, MATLAB, and Liu with the MIRNet of Zamir, and the results of this substitution would have been predictable because the CNNs would simply perform the same function they would alone, image restoration. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have used Zamir’s MIRNet to perform the image restoration of Fuoli in view of Ater, Hsu, MATLAB, and Liu.
Claims 5-6, 10-11 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Fuoli in view of Ater, Hsu, MATLAB, and Liu as applied to the claims above, and further in view of “Occluded offline handwritten Chinese character recognition using deep convolutional generative adversarial network and improved GoogLeNet” (herein after referred to by its primary author, Li).
In regards to claim 5, Fuoli in view of Ater, Hsu, MATLAB, and Liu teaches the image restoration device according to claim 3, wherein the discrimination model discriminates whether the Fourier-transformed restored image is true or false, and the processor compares the Fourier-transformed restored image with the correct image through the discrimination model, and based on the comparison result, causes the image restoration model to perform adversarial learning so that the image restoration model generates a high-frequency restored image from a low-frequency original image (Fuoli Figure 2 Examiner note: This figure shows generating an adversarial loss by comparing the correct image and the enhanced image in the Fourier Domain, that loss is then used to train the network to train the image transformation model).
Fuoli in view of Ater, Hsu, MATLAB, and Liu does not teach wherein the discrimination model includes a CNN model including at least one of GoogleNet, AlexNet, and VGG Network.
However, Li teaches wherein the discrimination model includes a CNN model including at least one of GoogleNet, AlexNet, and VGG Network that discriminates whether the Fourier-transformed restored image is true or false (Li Figures 1 & 2).
Fuoli in view of Ater, Hsu, MATLAB, and Liu teaches a discrimination model which discriminates a Fourier-transformed image from the correct image. The claimed device differs from this by using a CNN including at least one of GoogleNet, AlexNet, or VGG Network. Li teaches performing adversarial learning using a discrimination model including GoogleNet. One of ordinary skill in the art could have substituted the discrimination model of Fuoli in view of Ater, Hsu, MATLAB, and Liu with the discrimination model including a CNN of Li, and the results of this substitution would have been predictable because both elements are used to perform adversarial learning during the training stage. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have used Li’s discriminator model including GoogleNet to perform the adversarial learning of Fuoli in view of Ater, Hsu, MATLAB, and Liu.
In regards to claim 6, Fuoli in view of Ater, Hsu, MATLAB, Liu, and Li teaches the image restoration device according to claim 5, wherein the processor trains the discrimination model so that it outputs a preset first reference value or more for the correct image and outputs a preset second reference value or less for the restored image (Fuoli Figure 3; Section 3.3.1 “The transformed real and fake scores ρ/Φ are then evaluated with the sigmoid cross-entropy GAN-objective in (12).” Examiner note: The first reference value would be the real score, and the second reference value would be the fake score).
In regards to claim 10, Fuoli in view of Ater, Hsu, MATLAB, Liu, and Li renders obvious the claim limitations as in the consideration of claim 5.
In regards to claim 11, Fuoli in view of Ater, Hsu, MATLAB, Liu, and Li renders obvious the claim limitations as in the consideration of claim 6.
In regards to claim 15, Fuoli in view of Ater, Hsu, MATLAB, Liu, and Li renders obvious the claim limitations as in the consideration of claim 6.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CALEB LOGAN ESQUINO whose telephone number is (703)756-1462. The examiner can normally be reached M-Fr 8:00AM-4:00PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at (571) 270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CALEB L ESQUINO/ Examiner, Art Unit 2677
/ANDREW W BEE/ Supervisory Patent Examiner, Art Unit 2677