DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The Amendment filed on May 21th, 2026 has been entered. Claims 1, 5–7 and 9–10 are currently pending. Claims 1 and 5 have been amended. Claims 2-4 and 8 have been canceled.
Response to Arguments
Applicant's arguments filed 05/21/2026 have been fully considered. The arguments are moot with respect to the previous 35 U.S.C. § 103 rejection of Claim 1 over Kupyn in view of Wang alone, because Applicant has amended Claim 1 to include the limitations of original claims 2-4. However, Applicant's arguments asserting that the newly amended Claim 1 is not rendered obvious by the prior art are not persuasive, as explained below.
Applicant's arguments, see pages 6–8 of the Remarks, state that amended claim 1 recites a unitary, sequential two-stage attention architecture that cannot be separated into or equated with a conventional Feature Pyramid Network and an ECA module. Applicant further argues that Kupyn does not disclose parallel multi-branch pyramid convolution, per-scale channel attention, Softmax normalization across scales, weighted multiplication, or weighted concatenation of the resulting feature maps. These arguments are not persuasive because the rejection does not rely on Kupyn alone to teach the newly incorporated limitations. Kupyn is relied upon for the FPN-based motion-deblurring generator, including obtaining a blurred image, producing a clear image, extracting and aggregating multi-scale convolutional features, and training using adversarial and content losses. Zhang is relied upon for the multi-branch pyramid convolution and attention operations now incorporated into amended claim 1. In particular, Zhang teaches generating sub-feature maps using parallel convolutional branches having different kernel sizes, obtaining a channel-attention weight for each sub-feature map, applying Softmax normalization across the attention weights, multiplying each sub-feature map by its corresponding normalized attention weight, and concatenating the weighted sub-feature maps to obtain a fused feature map. Zhang further teaches incorporating this PSA structure into a ResNet-style residual block.
Applicant further argues, see page 6 of the Remarks, that the claimed "global average pooling + 1D convolution" module is not equivalent to Wang's ECA module because Applicant's module allegedly operates only after multi-scale Softmax-weighted fusion, forming a "two-stage cascaded attention" not disclosed by Wang. This argument is not persuasive because the rejection does not require Wang to disclose the upstream Softmax fusion step; Wang was cited specifically and solely for the claimed GAP-then-1D-convolution channel interaction mechanism, i.e.,
y
=
g
X
followed by
ω
=
σ
C
1
D
k
y
[Wang, Sec. 3.2.2, Eq. 1, 8–9], which corresponds directly to the claimed
y
=
g
X
'
and
ω
=
σ
C
o
n
v
1
D
k
y
steps. That Wang's ECA module is applied to a single-scale feature map in Wang's own disclosure does not patentably distinguish the claim, because the rejection relies on Kupyn (as modified by Zhang) to supply the multi-scale fused feature map
X
'
as the input to the GAP/1D-convolution operation taught by Wang; combining a known channel-attention technique with a known multi-scale fused feature map, each performing its established function, is precisely the type of combination contemplated by KSR Int'l Co. v. Teleflex Inc., 550 U.S. 398 (2007).
Applicant's arguments, see pages 7–8 of the Remarks, further state that there would have been no motivation to combine Kupyn, Wang, and Zhang because the references allegedly concern different technical fields, including motion deblurring, image classification, object detection, and image segmentation. These arguments are not persuasive. Each reference concerns convolutional-neural-network feature extraction and feature recalibration for computer-vision processing. Kupyn expressly applies an FPN originally developed for object detection to a motion-deblurring generator and teaches that its FPN-based architecture may use different plug-in backbone networks. Zhang teaches that its PSA/EPSA block is a flexible component for improving multi-scale feature representation, while Wang teaches that its ECA module is a lightweight component for improving channel-feature discrimination through local cross-channel interaction. One of ordinary skill in the art would therefore have been motivated to incorporate Zhang's multi-scale pyramid attention into Kupyn's deblurring feature-extraction pathway and to apply Wang's lightweight local channel-interaction technique to Zhang's resulting fused feature map in order to improve the representation of blur-related features at different spatial scales, emphasize informative fused channels, suppress less useful channels, and maintain low computational complexity. The modification would have amounted to the predictable use of known convolutional-network components according to their established functions, with a reasonable expectation of success.
Applicant additionally argues that the claimed combination achieves unexpected technical effects, including reduced artifacts, improved texture recovery, increased object-detection accuracy after deblurring, and lightweight computation. However, Applicant has not provided objective evidence, such as comparative testing, an expert declaration, or ablation data commensurate in scope with the claims, establishing that the alleged results are unexpected or attributable to the claimed combination. Moreover, the alleged benefits are consistent with the expected purposes of the cited teachings: Kupyn teaches improving deblurring quality and efficiency through multi-scale feature aggregation; Zhang teaches improving multi-scale feature representation through pyramid attention and recalibration; and Wang teaches improving channel discrimination with negligible computational overhead. Accordingly, the unsupported assertions of unexpected results do not overcome the prima facie case of obviousness, and do not substitute for objective evidence of unexpected results. See MPEP § 716.02(b). Each reference continues to perform its known function (multi-scale fusion, per-scale attention weighting, and channel-wise interaction) in a manner that yields predictable results.
Because amended claim 1 now incorporates the limitations of former claims 2–4, the previous rejection of claim 1 over Kupyn in view of Wang has been superseded in its prior form. An updated rejection of amended claim 1 under 35 U.S.C. § 103 is made over Kupyn in view of Wang and Zhang, using the same teachings previously applied to claims 1–4. The corresponding rejections have likewise been updated as follows:
Claims 1 and 9–10 are rejected under 35 U.S.C. § 103 as being unpatentable over Kupyn in view of Wang and Zhang.
Claim 5 is rejected under 35 U.S.C. § 103 as being unpatentable over Kupyn in view of Wang and Zhang, further in view of Wei.
Claims 6–7 are rejected under 35 U.S.C. § 103 as being unpatentable over Kupyn in view of Wang and Zhang, further in view of Kupyn '18 and Luo.
The previous rejections of canceled claims 2–4 and 8 are withdrawn as moot.
Applicant's amendments to claims 1 and 5 have been fully considered. The limitations incorporated into amended claim 1 from former claims 2–4 are addressed in the updated rejections by Zhang, which was previously cited and applied to those limitations in the Non-Final Office Action. The necessity to modify and reorganize the rejections of amended claim 1, its corresponding claims 9–10, and dependent claims 5–7 was directly necessitated by Applicant's substantive amendment incorporating former claims 2–4 into independent claim 1.
Accordingly, this action is properly made final in accordance with MPEP § 706.07(a).
Based on these facts, this action is made FINAL.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1 and 9–10 are rejected under 35 U.S.C. §103 as being unpatentable over Kupyn (Kupyn et al, DeblurGAN-v2: Deblurring (Orders-of-Magnitude) Faster and Better, 2019) in view of Wang (Wang et al, ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks, 2020), further in view of Zhang (Zhang et al, EPSANet: An Efficient Pyramid Squeeze Attention Block on Convolutional Neural Network, 2021).
Regarding claim 1, Kupyn teaches a method for image motion deblurring, comprising:
obtaining a motion-blurred image to be deblurred;
( [Sec. 3, "DeblurGAN-v2 Architecture"]: Kupyn discloses DeblurGAN-v2 to restore a sharp image
I
S
from a single blurred image
I
B
, via the trained generator. )
inputting the obtained blurred image into a pre-constructed and pre-trained image motion deblur model based on a multi-scale feature fusion module
( ["Abstract"], [Sec. 1, "Introduction"] & [Sec. 3.1 , "Feature Pyramid Deblurring"]: Kupyn teaches inputting the blurred image into a trained generator that restores a sharp image, and further teaches introducing a Feature Pyramid Network (FPN) as a core building block in the generator. )
wherein, the image motion deblur model is obtained through extracting characteristic information of different spatial scales and frequencies through the multi-scale feature fusion module for feature fusion,
( [Sec. 3.1 , "Feature Pyramid Deblurring"]: Kupyn teaches extracting and compressing / fusing characteristic information of different spatial scales through the multi-scale feature fusion module because Kupyn explains that FPN is introduced to incorporate multi-scale features, and that the architecture takes five final feature maps of different scales, upsamples them to the same input size, and concatenates them into one tensor containing semantic information on different levels; and in [Sec. 3.3 , "Double-Scale RaGAN-LS Discriminator"], Kupyn further teaches training the model using a hybrid loss including adversarial loss and content loss. )
wherein the constructed image motion deblur model comprises: a convolutional layer for preliminary feature extraction, a plurality of residual blocks with the same structure, and a convolutional layer for image reconstruction;
( [Sec. 3.1, “Feature Pyramid Deblurring”]: Kupyn teaches the FPN module comprises “a bottom-up pathway” that is “the usual convolutional network for feature extraction” for the entire backbone network and further teaches that existing CNNs for image deblurring “typically refer to ResNet-like structures”; Kupyn also teaches that its architecture takes five final feature maps of different scales as output and “additionally add[s] two upsampling and convolutional layers at the end of the network to restore the original image size and reduce artifacts”. A convolutional layer for preliminary feature extraction is the entry point of the "Bottom-Up" pathway; In Kupyn's FPN, blocks provide the different "levels" (C2 through C5) that the pyramid uses is a plurality of residual blocks with the same structure. )
the residual block comprises the multi-scale feature fusion module
( [Sec. 3.1 , "Feature Pyramid Deblurring"]: Kupyn teaches that DeblurGAN-v2's FPN backbone outputs five final feature maps of different scales that are upsampled and concatenated into one tensor containing semantic information on different levels. )
Kupyn teaches the claimed FPN-based multi-scale feature extraction and fusion for motion deblurring, but does not expressly disclose performing local cross-channel interaction on the fused feature map using one-dimensional convolution, where Wang teaches:
“and a local channel information interaction module” and “and the local channel information interaction module”;
( [Sec. 3, "Proposed Method"] & [Sub-Sec. 3.2.2, "Local Cross-Channel Interaction"]: Wang teaches an efficient channel attention (ECA) module that learns channel attention from aggregated convolution features without dimensionality reduction and captures local cross-channel interaction. )
and exchanging a fused feature map with local channel information in an one-dimensional convolution manner through the local channel information interaction module
( [Sec. 3, "Proposed Method"], [Sub-Sec. 3.2.2, "Local Cross-Channel Interaction"] & [Sub-Sec. 3.3, "ECA Module for Deep CNNs"]: Wang teaches that the weight of a channel is calculated by considering interaction between that channel and its
k
adjacent channels [Eq. 6], and further teaches that such strategy can be readily implemented by a fast 1D convolution with kernel size of
k
[Eq. 7], specifically
ω
=
σ
C
1
D
k
y
, where
C
1
D
indicates 1D convolution [Eq. 8]; Wang also teaches that, after aggregating convolution features, the ECA module performs 1D convolution followed by a sigmoid function to learn channel attention. )
the multi-scale feature fusion module comprises
( [Fig. 2], [Sub-Sec. 3.2.2, “Local Cross-Channel Interaction”] & [Sub-Sec. 4.2.1, “MobileNetV2”]: Wang teaches a channel attention mechanism layer because ECA is an efficient channel-attention module; Wang further teaches that, after global average pooling, local cross-channel interaction is implemented by fast 1D convolution, specifically
ω
=
σ
C
1
D
k
y
, where
C
1
D
indicates 1D convolution. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Kupyn’s multi-scale feature-fusion deblurring model by applying Wang’s ECA module to the fused convolutional feature map. Kupyn seeks improved deblurring quality through aggregation of multi-scale features, while Wang teaches that global average pooling followed by lightweight 1D convolution efficiently captures local cross-channel interaction and adaptively recalibrates convolutional features. The modification would predictably emphasize informative blur-related channels and suppress less useful channels, thereby improving feature discrimination and restoration quality with negligible additional computational cost and a reasonable expectation of success.
Kupyn [as modified by Wang] teaches multi-scale feature fusion followed by local cross-channel recalibration, but does not expressly disclose the recited pyramid-convolution branches, per-scale Softmax attention weighting, and weighted concatenation, where Zhang teaches:
the multi-scale feature fusion module comprises a pyramid convolutional layer
( [Sec. 1, “Introduction”] & [Sec. 3.2, “PSA Module”]: Zhang teaches a “multi-scale pyramid convolution structure” that processes the input at multiple scales, and further teaches that the SPC module performs multi-branch feature extraction using multi-scale convolutional kernels in a pyramid structure, and in [Fig. 4 & 5], shows parallel convolution branches with different kernel sizes and group sizes whose outputs are concatenated. )
wherein the step of extracting characteristic information of different spatial scales and frequencies through the multi-scale feature fusion module for feature fusion comprises:
obtaining an initial feature map X;
( [Page. 4, Sec. 3.2, “PSA Module”], & [Fig. 4]: Zhang teaches that the SPC module extracts the spatial information of the input feature map in a multi-branch way, where the input feature map is provided to parallel branches for subsequent multi-scale processing. )
conducting feature extraction on the obtained initial feature map X under different spatial scales and frequencies by using multiple types of convolutional kernels in the pyramid convolutional layer to obtain a plurality of sub-feature maps expressed as:
F
i
=
C
o
n
v
k
i
×
k
i
G
i
X
in the formula,
F
i
∈
R
C
'
×
H
×
W
represents the i-th sub-feature map obtained from the initial feature map X after passing through the i-th type of convolution kernel, i = 0,1,2, …, S - 1; S represents the type of convolution kernel; R represents the feature domain, C', H, and W respectively represent the number of channels, height, and width of the sub-feature maps, Conv represents the convolution operation; kᵢ×kᵢ represents the size of the i-th kernel; Gᵢ represents the calculation parameter for the number of channels in the i-th type of convolutional kernel, expressed as follows:
G
i
=
2
k
i
-
1
2
,
k
i
>
3
.
1
,
k
i
=
3
( [Sec. 3.2, “PSA Module”, Eq. (3) & (4)] & [Fig. 4]: Zhang expressly teaches that the SPC module uses multi-scale convolutional kernels in a pyramid structure, with feature maps generated functions are defined. )
using the channel attention mechanism layer to obtain channel attention weights of each sub-feature map, and using the Softmax normalization function to calibrate the channel attention weights of each sub-feature map, the expression is:
Z
i
=
S
E
F
i
a
t
t
i
=
S
o
f
t
m
a
x
Z
i
=
e
x
p
Z
i
∑
i
=
0
S
-
1
e
x
p
Z
i
in the formula,
Z
i
∈
R
C
'
×
1
×
1
is the channel attention weight of the i-th sub-feature map, SE represents the channel attention mechanism;
a
t
t
i
represents the normalized channel attention weight of the i-th sub feature map;
( [Sec. 3.2, “PSA Module”, Eq. (6-9)] & [Fig. 3 & 4]: Zhang teaches that the channel-wise attention vectors with different scales are obtained as
Z
i
=
S
E
W
e
i
g
h
t
F
i
, where
Z
i
∈
R
C
'
×
1
×
1
, and further teaches that a soft assignment weight is given by
a
t
t
i
=
S
o
f
t
m
a
x
Z
i
, thereby teaching the recited channel attention mechanism and Softmax normalization function to calibrate the channel attention weights of each sub-feature map. )
multiplying each sub-feature map with its corresponding normalized channel attention weight, and concatenating the multiplied feature maps using concatenation operation to obtain the fused feature map, expressed as:
Y
i
=
F
i
⊙
a
t
t
i
X
'
=
C
a
t
Y
0
Y
1
Y
2
…
Y
S
-
1
in the formula, Yᵢ represents the i-th sub-feature map with channel attention weights,
⊙
represents multiplication of channels; X' represents the fused feature map, and Cat represents concatenation operation;
( [Sec. 3.2, “PSA Module”, Eq. (10-11)] & [Fig. 3]: Zhang teaches multiplying the re-calibrated weight
a
t
t
i
of the multi-scale channel attention with the feature map
F
i
of the corresponding scale to obtain
Y
i
, and further teaches that the refined output is obtained by concatenation as
O
u
t
=
C
a
t
Y
0
Y
1
⋯
Y
S
-
1
. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to further modify Kupyn [as modified by Wang] by incorporating Zhang’s PSA multi-scale pyramid-convolution structure into the deblurring feature-fusion pathway. Zhang teaches extracting parallel multi-scale feature maps using different convolution kernels, recalibrating the respective feature maps with Softmax-normalized channel-attention weights, and concatenating the weighted feature maps. Applying this known structure to Kupyn’s multi-scale deblurring generator would predictably provide richer blur-feature representation across different spatial scales before Wang’s local cross-channel interaction, thereby improving feature discrimination and restoration quality with a reasonable expectation of success.
Kupyn [as modified by Wang and Zhang] further teaches the subsequent local channel information interaction performed on the fused feature map, as follows:
wherein the step of exchanging the fused feature map with local channel information in the one-dimensional convolution manner comprises:
obtaining the fused feature map output by the multi-scale feature fusion module;
( [Sec. 3.2, “PSA Module”, Eq. (11)] & [Fig. 3]: Zhang teaches that, after the multi-scale feature maps are weighted, “the process to obtain the refined output can be written as
O
u
t
=
C
a
t
Y
0
Y
1
⋯
Y
S
-
1
”, corresponding to obtaining the fused feature map output by the multi-scale feature fusion module. )
using the global average pooling layer to perform a global average pooling operation on the fused feature map, the expression is:
y
=
g
X
'
=
1
W
H
∑
m
=
1
,
n
=
1
W
,
H
X
m
n
'
in the formula,
g
X
'
represents the global average pooling of the fused feature map,
W
,
H
respectively represents the width and the height of the fused feature map
X
'
,
X
m
n
'
represents the pixel values in the m-th row and the n-th column of the fused feature map
X
'
, and y represents the output;
( Wang, in [Sec. 3.2, “Efficient Channel Attention (ECA) Module”, Eq. (1) and surrounding text; Fig. 2 / introductory discussion]: Wang teaches that, for the output
X
of a convolution block,
y
=
g
X
, where
g
X
=
1
W
H
∑
i
=
1
,
j
=
1
W
,
H
X
i
j
, and further teaches that
g
X
is channel-wise global average pooling (GAP). )
the output y after global average pooled is interacted with the local channel information through the one-dimensional convolutional layer, expressed as:
ω
=
σ
C
o
n
v
1
D
k
y
in the formula,
ω
represents the channel attention weight after interaction,
C
o
n
v
1
D
k
represents the one-dimensional convolution kernel, k is the size of the convolution kernel, and
σ
represents the Sigmod activation function;
( [Sub-Sec. 3.2.2, “Local Cross-Channel Interaction”, Eq. (9)] & [Fig. 2]: Wang teaches that, after channel-wise global average pooling, ECA captures local cross-channel interaction and can be efficiently implemented by fast 1D convolution, specifically
ω
=
σ
C
1
D
k
y
, where C1D indicates 1D convolution. )
multiplying the channel attention weight
ω
with the fused feature map
X
'
to assign channel attention to the fused feature map
X
'
to obtain an information interaction feature map
X
'
'
.
( [Fig. 2] & [Sec. 3.2.2, “Local Cross-Channel Interaction”]: Wang teaches that, given the aggregated features obtained by global average pooling, ECA generates channel weights by fast 1D convolution and applies those weights to the feature map by element-wise product. )
Regarding claims 9–10. The rationale provided for claim 1 is incorporated herein. In addition, the method for image motion deblurring claim 1 corresponds to the electronic device of claim 9, as well as the computer-readable storage medium of claim 10, and performs the steps disclosed herein. Therefore, the claims are all rejected.
Claim 5 is rejected under 35 U.S.C. §103 as being unpatentable over Kupyn [as modified by Wang and Zhang] in view of Wei (Wei et al, Dynamic scene deblurring and image de-raining based on generative adversarial networks and transfer learning for Internet of vehicle, 2021).
Regarding claim 5, Kupyn [as modified by Wang and Zhang] teaches the method for image motion deblurring according to claim 4, wherein the output of the residual block after processing the obtained initial feature map
X
is the result of adding the initial feature map
X
and the information interaction feature map
X
'
'
;
( Zhang, in [Sec. 1 “Introduction”], [Sec. 3.2 “PSA Module”] & [Fig. 5]: Zhang teaches that the proposed EPSA block is obtained by replacing the 3×3 convolution with the PSA module in the bottleneck blocks of the ResNet, and Fig. 5 shows the residual block with a “+” skip/add operation. The EPSA block can be easily added as a plug-and-play component into a well-established backbone network, and significant improvements on model performance can be achieved. )
Kupyn [as modified by Wang and Zhang] discloses residual-block add/skip idea in the EPSA/ ResNet-style block for feature map X; however, it does not cleanly teach the requirement that the image-reconstruction portion comprises three convolutional layers where Wei teaches:
the convolutional layer for image reconstruction comprises three convolutional layers, and the output of the residual block after processed by the three convolutional layers is added to the obtained blurred image to obtain the final output clear image.
( Wei, [Abstract] & [Sec. 3, “Method”]: Wei teaches a residual block containing three 256-channel convolutional layers and further teaches a GAN-based motion-deblurring generator with a global skip connection block. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to further modify Kupyn [as modified by Wang and Zhang] by incorporating Wei’s known three-convolution residual reconstruction arrangement and global input-to-output skip connection. Kupyn [as modified by Wang and Zhang] already employs residual feature processing and image restoration, while Wei teaches using three convolutional layers and adding the reconstructed residual output to the blurred input in a GAN-based motion-deblurring generator. The modification would predictably improve recovery of blur-related image details and stabilize residual learning, with a reasonable expectation of success.
Claims 6–7 are rejected under 35 U.S.C. §103 as being unpatentable over Kupyn [as modified by Wang and Zhang] in view of Kupyn'18 (Kupyn et al, DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial Networks, 2018), further in view of Luo (Luo et al, Comparison and Benchmarking of AI Models and Frameworks on Mobile Devices, 2020).
Regarding claim 6, Kupyn [as modified by Wang and Zhang] teaches the method for image motion deblurring according to claim 1, wherein the training method of the image motion deblur model comprises:
Kupyn [as modified by Wang and Zhang]'s DeblurGAN-v2 is based on a relativistic conditional GAN with a double-scale discriminator and introduces FPN into the generator; however, DeblurGAN-v2 fails to disclose the specific training recipe and the specific loss formulation where Kupyn'18's DeblurGAN teaches:
( Kupyn’18, [Sec. 5, “Training Details”]: Kupyn’18 teaches that DeblurGAN was trained on random crops of size 256×256 from GoPro training dataset images. [Kupyn’18, Sec. 6.1, “GoPro Dataset”]: Kupyn ’18 further teaches that the GoPro dataset consists of 2103 pairs of blurred and sharp images and reports results on the GoPro test dataset of 1111 images. )
inputting the training set into the constructed image motion deblur model to reconstruct the deblurred image through multi-scale feature fusion and local channel information interaction;
( Kupyn [as modified by Wang] teaches the constructed image motion deblur model of claim 1, including multi-scale feature fusion and local channel information interaction. Wherein Kupyn’18, [Fig. 4], teaches that the generator network takes the blurred image as input and produces the estimate of the sharp image during training. )
conducting a supervised training on the model according to the testing set using a loss function based on adversarial loss and content loss;
( [Abstract]; [Sec. 1, “Introduction”]; [Sec. 3.2, “Network architecture”]; [Fig. 4]: Kupyn ’18 teaches supervised training of DeblurGAN using a multi-component loss function including Wasserstein GAN with gradient penalty (WGAN-GP) and perceptual / content loss, and teaches that the total loss consists of the WGAN loss from the critic and the perceptual loss. [Sec. 6.1, “GoPro Dataset”]: Kupyn’18 further teaches use of a separate GoPro test dataset for evaluation. Therfore, Kupyn’18 teaches supervised training using a loss function based on adversarial loss and content loss together with separate training and testing sets. )
repeating the training process to update the network parameters of the optimized model until the loss function converges or reaches the preset number of training iterations, and stop training to obtain the final optimized image motion deblur model.
( [Sec. 5, “Training Details”]: Kupyn’18 teaches iterative optimization of the generator and critic using Adam, with an initial learning rate of 10⁻⁴ for both generator and critic, and further teaches that after the first 150 epochs the learning rate is linearly decayed to zero over the next 150 epochs, thereby teaching repeating the training process to update network parameters through preset training iterations until training is completed to obtain the final optimized model. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to train Kupyn [as modified by Wang and Zhang] using Kupyn ’18’s known DeblurGAN training regimen, including random 256×256 crops, separate training and testing datasets, adversarial and perceptual losses, and iterative parameter updates over a preset number of epochs. Kupyn ’18 concerns the closely related predecessor motion-deblurring GAN and teaches established training techniques for producing a trained generator from blurred and sharp image pairs. Applying those known techniques to the modified DeblurGAN-v2 architecture would have been a routine and predictable implementation choice, with a reasonable expectation of successfully obtaining the optimized deblurring model.
Kupyn [as modified by Wang, Zhang and Kupyn'18] still fails to disclose where Luo teaches:
compressing the collected images in the dataset into images with a resolution of 360×360 and randomly cropping the image to
( Luo, [Sec. VII "Technical Description, Sub-section B “Data preprocessing”]: Luo teaches a known image-preprocessing technique in which an image is cropped according to the shortest side to form a square intermediate image to 360x360, before resizing to the model’s required input size. Luo proves that 360×360 is just an intermediate square crop before resizing to the final network input. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to further modify Kupyn [as modified by Wang, Zhang, and Kupyn ’18] by using Luo’s known 360×360 square-image preprocessing before Kupyn ’18’s random 256×256 training crop. Luo teaches standardizing source images to a square intermediate size before final model-input preparation. Applying that conventional preprocessing step would predictably provide dimensionally uniform images for subsequent random cropping and network training, constituting a routine implementation choice with a reasonable expectation of success.
Regarding claim 7, Kupyn [as modified by Wang, Zhang, Kupyn '18 and Lou] teaches the method for image motion deblurring according to claim 6, wherein the loss function based on the adversarial loss and the content loss is as follows:
L
t
o
t
a
l
=
L
a
d
v
+
λ
L
c
o
n
t
e
n
t
in the formula,
L
t
o
t
a
l
represents the loss function,
L
a
d
v
represents the adversarial loss,
L
c
o
n
t
e
n
t
represents the content loss,
λ
is a content loss coefficient; the adversarial loss
L
a
d
v
is expressed as WGAN-GP as below:
L
a
d
v
=
E
x
~
∼
P
g
D
x
~
-
E
x
∼
P
r
D
x
+
λ
E
x
^
∼
P
x
^
∇
x
^
D
x
^
2
-
1
2
in the formula,
D
represents discriminator,
x
represents clear image,
x
~
represents network output image,
x
^
represents random image,
x
^
=
ϵ
x
~
+
1
-
ϵ
x
,
ϵ
~
U
0,1
;
P
x
^
represents that a distribution of image samples uniformly sampled along a straight line between a pair of points
U
0,1
sampled from clear image distribution
P
r
and network output image distribution
P
g
;
( [Sec. 1, “Introduction”], [Sec. 2.2, “Generative adversarial networks”], [Sec. 3.1, “Loss function”], [Sec. 3.2, “Network architecture”], [Eq. (3-6)]: Kupyn’18 teaches that DeblurGAN is based on a conditional GAN and a multi-component loss function, and expressly teaches use of Wasserstein GAN with gradient penalty, further teaching that during training the critic network is WGAN-GP and that the total loss includes the WGAN loss from the critic. )
the content loss adopts perceptual loss, expressed as:
L
c
o
n
t
e
n
t
=
∑
ϕ
x
-
ϕ
(
x
~
)
2
2
in the formula,
ϕ
represents the pre-trained VGG19 network.
( [Sec. 3.1 “Loss function”]: Kupyn’18 teaches that, instead of raw-pixel L1 (MAE) / L2 (MSE) loss, DeblurGAN adopts perceptual loss, which is an L2 loss based on the difference of the generated and target image CNN feature maps, and further teaches that the perceptual loss is the difference between the VGG-19 conv3.3 feature maps of the sharp and restored images, where the VGG19 network is pretrained on ImageNet, thereby teaching the claimed content loss based on a pre-trained VGG19 network. )
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEN KUDO
Examiner
Art Unit 2671
/KEN KUDO/Examiner, Art Unit 2671
/VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671