Detailed Action
This Office Action is in response to the remarks entered on 06/22/2026. Claims 2, 10, 16 and 19 have been canceled. New claims 22-24 have been entered herein. Claims 1, 3-9, 11-15, 17-18 and 20-24 are presently pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 9, 15 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Fukui et al. (“Attention Branch Network: Learning of Attention Mechanism for Visual Explanation”, 2019, hereinafter ‘Fukui’) in view of Zheng et al. (“Learning Multi-Attention Convolutional Neural Network for Fine-Grained Image Recognition”, 2017, hereinafter ‘Zheng’) and further in view of KARANAM & WU (US 20210090289 A1, hereinafter ‘Karanam’).
Regarding claim 1, Fukui teaches:
A computing system comprising: ([Fukui, page 1, right col, 2.2. Attention mechanism, line 1-2] and [page 4, left col, 4. Experiments, line 1 – right col, line 11] implies that the method is implemented using a generic computer)
generating, via a feature extraction network, based on an input image, one or more of a first feature map or a second feature map, wherein the feature extraction network comprises a first neural network including a plurality of convolution layers, wherein the first feature map is obtained from a last layer of the plurality of convolution layers; ([Fukui, page 2, Fig 2. (a); left col, 3. Attention Branch Network, line 1 – right col, line 7] The feature extractor receives the input image xi and generates the feature map g(xi). The feature extractor contains multiple convolution layers and extract feature maps from an input image)
generating, via a feature importance network, a feature importance vector wherein the feature importance network comprises a second neural network; ([Fukui, page 2, Fig. 1] and [page 3, right col, 3.1. Attention branch, line 1 – page 4, left col, line 9] The Attention Branch which receives the generated feature map from the feature extractor is the second neural network. [Fukui, page 3, left col, 3rd para, lines 1-6] shows that the weights are generated based on training)
generating, via an activation function, an attention map based on a weighted sum of the feature importance vector and the first feature map; ([Fukui, page 3, left col, line 2-9 and line 23-41] discloses that the attention map is generated by multiplying the weighted sum of the K x h x w feature map by the weight (i.e., feature importance vector) at the last fully-connected layer. K feature maps are generated and combined to generate the attention map. Each feature maps in the K feature map can be interpreted as the 1st feature map and the 2nd feature map. [page 3, left col, 3rd para, lines 16-19] shows the sigmoid function which is the activation function. [page 2, left col, 2nd para, lines 6-10] discloses that effective weights are extracted by using ABN for visual explanation, which indicates that the w is a feature importance vector)
determining a classification output based on combining the attention map and one or more of the first feature map or the second feature map; and ([Fukui, page 2, Fig. 2; right col, line 3-7] The perception branch outputs the probability of each class (classification output) by receiving the feature map from the feature extractor and attention map. [page 3, left col, 3.2. Perception branch, line 1 – right col, line 6] discloses combining the attention map and feature maps to generate classification probabilities)
However, Fukui does not specifically disclose:
a processor; and a memory coupled to the processor, the memory storing instructions which, when executed by the processor, cause the computing system to perform operations comprising:
generating, via a feature importance network, a feature importance vector based on combining the input image and the first feature map
generating a feature visualization image by overlaying the attention map onto the input image.
Zheng teaches:
a processor; and a memory coupled to the processor, the memory storing instructions which, when executed by the processor, cause the computing system to perform operations comprising: ([Zheng, page 5219, right col, lines 1-5] shows that the method is computer vision technique which requires a processor and memory)
generating, via a feature importance network, a feature importance vector based on combining the input image and the first feature map ([Zheng, page 5222, left col, lines 8-47] The weights are used to cluster discriminative parts (important part). The features W and the input image X is combined to generate the weight vector (i.e., feature importance vector) di(X) as shown in equation (3). The attention map is generated based on the equation (4) based on the d_j weights. W is feature because the paragraph discloses that inputs of
f
i
(
)
are input convolutional features)
It would have been obvious before the effective filing date of the claimed invention to a person
having ordinary skill in the art to apply the method of generating an attention map based on a weighted sum of the feature importance vector and the first feature map of Zheng to improve the feature importance learning system of the present invention. The suggestion and/or motivation for doing so is to improve both efficiency and classification performance of the system by permitting the creation of refined and extremely compact sets that retain the most meaningful features.
However, Fukui in view of Zheng does not specifically disclose:
generating a feature visualization image by overlaying the attention map onto the input image.
Karanam teaches:
generating a feature visualization image by overlaying the attention map onto the input image. ([Karanam, 0030] and [0042] collectively disclose the output module overlaying an attention map onto the input image)
It would have been obvious before the effective filing date of the claimed invention to a person
having ordinary skill in the art to apply the method of overlaying the attention map onto the input image of Karanam to improve the feature importance learning system of the present invention. The suggestion and/or motivation for doing so is to assist debugging process and improve reliability of the machine learning model by visualizing the inference process of the machine learning model.
Claim 9 is a method claim which recites the same feature as the claim 1, and is rejected for at least the same reasons.
Regarding claim 15, Fukui in view of Zheng teaches:
At least one non-transitory computer readable medium comprising instructions which, when executed by a computing system, cause the computing system to perform operations comprising ([Zheng, page 5219, right col, lines 1-5] shows that the method is computer vision technique which requires a processor and memory)
Claim 15 is a non-transitory computer readable medium claim which recites the same feature as the claim 1, and is rejected for at least the same reasons.
Regarding claim 21, Fukui teaches:
wherein the activation function comprises a rectified linear unit function. ([Fukui, page 2, Figure 2(a)] The Attention Branch network includes ReLU, which denotes the rectified linear unit function)
Claims 3, 5-6, 11, 13, 17, 22-23 are rejected under 35 U.S.C. 103 as being unpatentable over Fukui in view of Zheng in view of Karanam and further in view of Zhou et al. (US 20230026811 A1, hereinafter ‘Zhou’).
Regarding claim 3, Fukui in view of Zheng and further in view of Karanam teaches:
The computing system of claim 1.
However, Fukui in view of Zheng and further in view of Karanam does not specifically disclose:
wherein the second feature map is obtained from an intermediate layer, other than the last layer, of the plurality of convolution layers.
Zhou teaches:
wherein the second feature map is obtained from an intermediate layer, other than the last layer, of the plurality of convolution layers. ([Zhou, 0091 and 0093] discloses inputting the input image into a shallow feature extractor and generating a plurality of intermediate feature maps. The concatenation module 306 generates a combined feature map by combining the plurality of intermediate feature maps and the last feature map 303M)
It would have been obvious before the effective filing date of the claimed invention to a person
having ordinary skill in the art to apply the method of obtaining and combining intermediate feature maps of Zhou to improve the feature importance learning system of the present invention. The suggestion and/or motivation for doing so is to improve the accuracy of the machine learning model output by averaging out the noise contained in the feature map.
Regarding claim 5, Fukui teaches:
The computing system of claim 3, wherein combining the attention map and one or more of the first feature map or the second feature map comprises: ([Fukui, page 3, left col, 3.2. Perception branch, line 6 - right col, Equation (2), line 9] The feature maps
g
c
(
x
i
)
are combined with attention maps
M
(
x
i
)
using a dot-product operation (element-wise multiplication function) )
generating an output map by combining, via an attention mechanism, the attention map and one or more of the first feature map or the second feature map; and ([Fukui, page 3, left col, 3.2. Perception branch, line 6 - right col, Equation (2), line 9] The feature maps
g
c
(
x
i
)
(includes the fist and the second feature map) are combined with attention maps
M
(
x
i
)
using a dot-product operation (element-wise multiplication function) )
applying an activation function to the output map. ([Fukui, page 3, left col, 3.2. Perception branch, line 1-6; page 2, Fig. 2(c)] and [page 4, left col, line 27-34] collectively discloses that the perception branch receives the converted feature map converted by applying attention map
M
t
(
x
)
at specific task t. The perception branch is a neural network classifier which consists of a plurality of activation functions)
Regarding claim 6, Fukui teaches:
The computing system of claim 5, wherein generating an attention map based on a weighted sum of the feature importance vector and the first feature map comprises: ([Fubuki, page 3, left col, line 2-9 and line 23-41] discloses that the attention map is generated by multiplying the weighted sum of the K x h x w feature map (the first feature map) by the weight (feature importance vector) at the last fully-connected layer. K feature maps are generated and combined to generate the attention map. Each feature maps in the K feature map are interpreted as the 1st feature map and the 2nd feature map. Claim 6 tells that the weights are the feature importance vector)
computing a specific weighted sum
Σ
k
=
1
N
w
k
F
M
k
, wherein weights
w
k
are derived from respective coefficients of the feature importance vector, and
F
M
k
is a k-th channel of the first feature map; and ([Fubuki, page 3, left col, line 2-9 and line 23-41] discloses that the attention map is generated by multiplying the weighted sum of the K x h x w feature map (the first feature map) by the weight (feature importance vector) at the last fully-connected layer. K feature maps are generated and combined to generate the attention map. Each feature maps in the K feature map are interpreted as the 1st feature map and the 2nd feature map. Claim 6 tells that the weights are the feature importance vector)
applying an activation function to a result of the specific weighted sum; and ([Fubuki, page 3, left col, line 2-9 and line 23-41] The aggregated K feature maps are normalized by the sigmoid function)
wherein the attention mechanism comprises an equation
F
O
=
F
L
⊗
(
1
+
A
M
)
, wherein
F
O
is the output map,
F
L
is the one or more of the first feature map or the second feature map,
A
M
is the attention map, and
⊗
denotes an element-wise multiplication function. ([Fukui, page 3, left col, 3.2. Perception branch, line 6 - right col, Equation (2), line 9] The feature maps
g
c
(
x
i
)
are combined with attention maps
M
(
x
i
)
using a dot-product operation (element-wise multiplication function) )
Claim 11 is a method claim which recites the same feature as the claim 3, and is rejected for at least the same reasons.
Claim 13 is a method claim which recites the same feature as the claim 5, and is rejected for at least the same reasons.
Claim 17 is a non-transitory computer readable medium claim which recites the same feature as the claim 3, and is rejected for at least the same reasons.
Claim 22 is an apparatus claim which recites the same feature as the apparatus claim 5, and is rejected for at least the same reasons.
Claim 23 is an apparatus claim which recites the same features as the apparatus claim 6, and is rejected for at least the same reasons.
Claims 4, 12 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Fukui in view of Zheng in view of Karanam and further in view of Yu et al. (Yu et al. “Region Normalization for Image Inpainting”, 2020, hereinafter ‘Yu’).
Regarding claim 4, Fukui in view of Zheng and further in view of Karanam teaches:
The computing system of claim 1.
However, Fukui in view of Zheng and further in view of Karanam does not specifically disclose:
wherein combining the input image and the first feature map comprises:
generating an intermediate image by applying one or more of a downsize function or a greyscale function to the input image;
generating an intermediate feature map by applying a normalize function to the first feature map; and
generating a masked image by multiplying, via element-wise multiplication, the intermediate image and the intermediate feature map.
Yu teaches:
wherein combining the input image and the first feature map comprises: ([Yu, page 12736, right col, Figure 3 (b)] The Figure 3 discloses combining the (MaxPool, AvgPool) image and the normalized input feature)
generating an intermediate image by applying one or more of a downsize function or a greyscale function to the input image; ([Yu, page 12736, Figure 3; left col, 3.3 Learnable Region Normalization, line 14 – right col, line 23] discloses generating a feature (image) by applying max-pooling and average-pooling which are downsampling (downsizing) functions)
generating an intermediate feature map by applying a normalize function to the first feature map; and ([Yu, Figure 3; page 12736, right col, line 8-18] discloses generating a normalized feature (intermediate feature map) by normalizing the input feature F)
generating a masked image by multiplying, via element-wise multiplication, the intermediate image and the intermediate feature map. ([Yu, Figure 3; page 12736, right col, line 8-18] The Region normalization result and
γ
are combined using matrix multiplication. The symbol ⊗ denotes matrix multiplication which is an element-wise multiplication)
It would have been obvious before the effective filing date of the claimed invention to a person
having ordinary skill in the art to apply the method of masking the input image by multiplying the image and the intermediate feature map of Yu to improve the feature importance learning system of the present invention. The suggestion and/or motivation for doing so is to improve the accuracy of the machine learning model output by emphasizing relevant features thereby help visualizing the feature importance more effectively.
Claim 12 is a method claim which recites the same feature as the claim 4, and is rejected for at least the same reasons.
Claim 18 is a non-transitory computer readable medium claim which recites the same feature as the claim 4, and is rejected for at least the same reasons.
Claims 7-8, 14, 20 and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Fukui in view of Zheng in view of Karanam and further in view of Hwang et al. (Hwang et al. “Aircraft Detection using Deep Convolutional Neural Network for Small Unmanned Aircraft Systems”, 2018, hereinafter ‘Hwang’).
Regarding claim 7, Fukui in view of Zheng and further in view of Karanam teaches:
The computing system of claim 1.
However, Fukui in view of Zheng and further in view of Karanam does not specifically disclose:
wherein the input image comprises an image of at least a portion of an aircraft or an aircraft component, and wherein the classification output comprises a determination of at least one of an identification of or a state of the aircraft or the aircraft component.
Hwang teaches:
wherein the input image comprises an image of at least a portion of an aircraft or an aircraft component, and wherein the classification output comprises a determination of at least one of an identification of or a state of the aircraft or the aircraft component. ([Hwang, page 4, D. Training the deep convolutional neural network model, line 1-10] and [Hwang, page 5, IV. Aircraft detection test results, line 1 – page 6, line 8] collectively disclose performing aircraft detection task using the neural network model)
It would have been obvious before the effective filing date of the claimed invention to a person
having ordinary skill in the art to apply the method of inputting image with at least a portion of an aircraft or an aircraft component and generating a classification output of Hwang to implement the feature importance learning system of the present invention. The suggestion and/or motivation for doing so is to improve the aircraft industry by automating the aircraft identification process.
Regarding claim 8, Fukui in view of Karanam and further in view of Hwang teaches:
The computing system of claim 1, wherein at least one of the first neural network or the second neural network is implemented by an artificial intelligence (AI) accelerator. ([Hwang, page 4, D. Training the deep convolutional neural network model, line 1-10] discloses performing the aircraft detection task using a neural network model implemented using a GPU accelerator that accelerates the training process)
Claim 14 is a method claim which recites the same feature as the claim 7, and is rejected for at least the same reasons.
Claim 20 is an apparatus claim which recites the same feature as the claim 7, and is rejected for at least the same reasons.
Claim 24 is an apparatus claim which recites the same feature as the claim 8, and is rejected for at least the same reasons.
Response to Arguments
Arguments regarding 35 U.S.C. 103 Rejections
Arguments: Applicant asserts that Fukui discloses that CAM is a part of the attention branch, and VGGNet and ResNet are a part of the perception branch, thus Fukui does not disclose "feature extraction network ... the feature extraction network comprises a first neural network."
Examiner's Response: Examiner respectfully disagrees. Fukui disclose “generating, via a feature extraction network, based on an input image, one or more of a first feature map or a second feature map, wherein the feature extraction network comprises a first neural network including a plurality of convolution layers” as recited in amended claim 1. [Fukui, page 2, Fig 2. (a); left col, 3. Attention Branch Network, line 1 – right col, line 7] explicitly recites that the feature extractor receives the input image xi and generates the feature map g(xi), and the feature extractor contains multiple convolution layers and extract feature maps from an input image.
Therefore, arguments regarding claim 1 are not persuasive. Similarly, arguments regarding independent claims 9 and 15 which recite the same features as system claim 1 are not persuasive at least for the same reasons.
Dependent claims
Applicant’s arguments with respect to claims 3-8 and 21 depend from claim 1, 11-14 depend from claim 9, and claims 17-18 and 20 depend from claim 15 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20200143204 A1 (This prior art is pertinent because it discloses generating a saliency map by combining an output of the second neural network and a gradient map)
LORENZO et al. “Hyperspectral Band Selection Using Attention-Based Convolutional Neural Networks”, 2020 (This prior art is pertinent because it discloses combining attention maps which shows feature importance to generate a final attention map)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUN KWON whose telephone number is (571)272-2072. The examiner can normally be reached Monday – Friday 8:00AM – 5:00PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUN KWON/Examiner, Art Unit 2127
/TEWODROS E MENGISTU/Examiner, Art Unit 2127