DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Applicants Amendments filed on July 7, 2026, has been entered and made of record.
Currently pending Claim(s): 1-20
Independent Claim(s): 1, 11, 20
Amended Claim(s): 18, 19
Specification Objections
In view of Applicant’s amendments to the Specification, the previous objections to the Specification are withdrawn.
Claim Objections
In view of Applicant’s amendments to the Claims 18-19, the previous objections to the Claims 18-19 are withdrawn.
Response to Arguments
This office action is responsive to the Applicant’s Arguments/Remarks Made in an Amendment
received on July 7, 2026.
Originally, (in the claim set dated April 15, 2024) Claim 1 was rejected under 35 U.S.C. 103 as being unpatentable over Xiong et al. (US Pub No 20230106716), hereinafter Xiong, in view of Heisler (US Pub No 20220292685), hereinafter Heisler, and further in view of Yu et al. (z. Yu, et al., "Deep Modular Co-Attention Networks for Visual Question Answering," 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 6274-6283), hereinafter Yu.
The Applicant argued (on Remarks, pg. 3), that the cited combination fails to teach or suggest the features of generating the scaled dot product attention matrices based on the word embeddings of the object labels. The Applicant explained (on Remarks, pg. 4) that neither Xiong nor Yu explicitly teach generating the scaled dot product attention matrices based on word embeddings of the object labels. The Applicant then explained that although Hiesler teaches generating word embeddings for object labels, these labels are not used for generating a scaled dot product attention matrix.
The Examiner has found these arguments persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Yu, Huang et al. (Huang, P. et al., “Multi-grained Attention with Object-level Grounding for Visual Question Answering”, ACL 2019, pp. 3595-3600)), hereinafter Huang, and further in view of Sun et al. (CN Pub No 113609355), hereinafter Sun, since Huang and Sun teach the limitation of generating an attention matrix based on the word embeddings of object labels.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-5, 8-15 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (Yu, et al., "Deep Modular Co-Attention Networks for Visual Question Answering," 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 6274-6283), hereinafter Yu, in view of Huang et al. (Huang, P. et al., “Multi-grained Attention with Object-level Grounding for Visual Question Answering”, ACL 2019, pp. 3595-3600), hereinafter Huang, and further in view of Sun et al. (CN Pub No 113609355), hereinafter Sun.
As to Claim 1, Yu teaches a method for controlling an artificial intelligence (Al) device, the method comprising (see pg. 6274, Section 1, “Multimodal learning to bridge vision and language has gained broad interest from both the computer vision and natural language processing communities”):
obtaining an input query, an input image (see pg. 6277, Section 3.2, “the input question and image”, and see Fig. 4 on pg. 4, depicting an input image and question),
bounding boxes for objects detected in the input image (see pg. 6277, Section 4.1, “The input image is represented as a set of regional visual features in a bottom-up manner”, and see Fig. 4, where red bounding boxes are depicted over the input image),
PNG
media_image1.png
387
1100
media_image1.png
Greyscale
Fig. 4 of Yu
and at least one topic label for one or more words in the input query (see pg. 6277, Section 4.1, “The input question is first tokenized into words”, wherein the ‘word’ is the topic label);
generating at least one word embedding for the at least one topic label from the input query, the at least one word embedding being a multi-dimensional vector (see pg. 6277, Section 4.1, “Each word in the question is further transformed into a vector using the 300-D GloVe word embeddings [25] pre trained on a large-scale corpus. This results in a sequence of words of size n×300”);
generating output attention maps corresponding to scaled dot product attention matrices (see pg. 6290, caption under Fig. 7, “Visualizations of the learned attention maps (softmax(qK/√d)”, where it is known to one of ordinary skill in the art that the formula ‘(softmax(qK/√d)’ corresponds to a scaled dot product)
based on the at least one word embedding for the at least one topic label from the input query and each of bounding boxes (see Fig. 7, where the attention is calculated for each object detected in the image);
PNG
media_image2.png
720
1228
media_image2.png
Greyscale
Figure 7 of Yu
combining, via the processor, the output attention maps to generate a final attention map corresponding to the at least one topic label from the input query (see pg. 6290, Fig. 7, where the image comprises multiple combined attention maps);
and executing, via the processor, a function based on the final attention map (see pg. 6290, caption under Fig. 7, “SA(Y) l, SA(X)-l and GA(X,Y)-l denote the question self-attention, image self-attention, and image guided-attention from the l-th layer, respectively. Q, A, P denote the question, answer and prediction respectively”, and see Fig. 7, where the answer ‘3’ is output to the initial user question).
Yu fails to teach obtaining object labels corresponding to the bounding boxes. Yu further fails to teach generating, via the processor, a plurality of word embeddings for the object labels corresponding to the bounding boxes, the plurality of word embeddings being multi-dimensional vectors, and that the scaled dot product is also based on and each of the plurality of word embeddings for the object labels corresponding to the bounding boxes.
However, in an analogous art, Huang teaches a method for controlling question answering system (see pg. 3595, Abstract, “Attention mechanisms are widely used in Visual Question Answering (VQA) to search for visual clues related to the question… this paper proposes a multi-grained attention method”), which comprises
obtaining, an input query (see Fig. 1, where the question is ‘What is the man wearing around his face?’),
an input image (see Section 2.3, pg. 3597, “For the input image in Figure 1”),
bounding boxes for objects detected in the input image, object labels corresponding to the bounding boxes (see Section 2.3, pg. 3597, “For the input image in Figure 1, Faster-RCNN detected objects with labels of “man”, “head””),
generating a plurality of word embeddings for the object labels corresponding to the bounding boxes, the plurality of word embeddings being multi-dimensional vectors (see paragraph [x], “For the k-th object with label ck we encode it into GloVe embedding…
L
G
=
l
1
G
,
…
,
l
k
G
∈
R
D
1
×
K
.., is the GloVe embeddings for the objects labels”),
generating output attention maps corresponding to dot product attention matrices based on the at least one word embedding for the input query and each of the plurality of word embeddings for the object labels corresponding to the bounding boxes, (see pg. 3597, Section 2.3, “Therefore, we compute the WL attention vector, that indicates how much weight we should give to each of the K objects in the image, in terms of the semantic similarity between the category labels of the objects and the words in the question”, and see Formula 1, where a dot product is performed between the word embeddings from the object labels (
L
T
) and the word embedding for the at least on input query (
X
G
), and see caption under Fig. 1, “Figure 1: An example of VQA and the attention maps produced by a state-of-the-art model and our model.”)
Thus, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention to combine the bounding box label embedding taught by Huang with the scaled dot product attention taught by Yu. The motivation for doing so would be to enhance the output attention maps. Huang teaches on pg. 3597, Section 2.3, “In contrast to Anderson et al. (2018) that only use objects’ visual features without the labels, and unlike Wu et al. (2018) that discard the visual features once the labels are generated, we utilize both category labels and the visual features to enhance the fine-grained attention with object level grounding.”
Neither Yu nor Huang explicitly teach a processor. However, in an analogous art Sun teaches a visual questioning answering system comprising a processor (see paragraph [0082], “A computer includes a memory and a processor, and the memory stores a computer program. When the processor executes the computer program, the steps of a video question answering method based on dynamic attention and graph network reasoning are realized”),
which can generate bounding boxes for objects detected in the input image, object labels corresponding to the bounding boxes, and then generate, via the processor, a plurality of word embeddings for the object labels corresponding to the bounding boxes (see paragraph [0097], “The object space feature and object category feature calculation module is used to predict the object tag frame and category label in the video according to the object detection model to obtain the object space feature and object category feature”, and see paragraph [0031], “EL is the word embedding vector representation of the object category label”),
and generate output attention maps corresponding to dot product attention matrices based on the input query and a word embedding for the object label corresponding to the bounding box, (see paragraph [0056], “Use the scaled dot product function to calculate the similarity matrix between the problem feature and the object's joint feature”, wherein the problem feature is a user query, and the object’s joint feature is based on the embedded category label).
Thus, it would have been obvious to combine the processor taught by Sun with the teachings of Yu and Huan. The motivation for doing so would be to implement the method taught by Yu and Huang on a variety of computing devices. Sun teaches in paragraph [0176], “The so-called processor may be a central processing unit, other general-purpose processors, digital signal processors, application specific integrated circuits, ready-made programmable gate arrays or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. . The general-purpose processor may be a microprocessor or the processor may also be any conventional processor or the like”. Thus, it would have been obvious to combine the processor taught by Sun with the method taught by Yu and Huang in order to obtain the invention as claimed in Claim 1.
As to Claim 2, Yu in view of Huang and Sun teaches the function of identifying an object within the input image (see Yu, page 6280, Section 5.3, Figure 7, where bounding boxes are used to identify sheep in the input image).
As to Claim 3, Yu teaches multiplying, via a first linear layer, the at least one word embedding and a first weight matrix to generate a query matrix multiplying, via a second linear layer, and a second weight matrix to generate a key matrix; (see Yu, pg. 6276, Section 3.1 “Given a query q
∈
R
1
×
d
..,, n key-value pairs (packed into a key matrix K
∈
R
1
×
d
and a value matrix V
∈
R
1
×
d
…, and see pg. 3, Formula 3, where each a query, key, and value matrix are multiplied to by first, second, and third weight matrix)
PNG
media_image3.png
59
352
media_image3.png
Greyscale
transposing the key matrix to generate a transposed key matrix (see Yu, pg. 6276 Formula 1, shown below, where
K
T
is a transposed key matrix);
PNG
media_image4.png
96
366
media_image4.png
Greyscale
and multiplying the query matrix and the transposed key matrix to generate a resulting matrix (see Yu, formula shown above, where matrix
Q
and the transposed key matrix are multiplied). Yu fails to teach that the key matrix is obtained from at least one of the plurality of word embeddings.
However, Huang teaches generating, via the processor, a plurality of word embeddings for the object labels corresponding to the bounding boxes, the plurality of word embeddings being multi-dimensional vectors (see pg. 3597, Section 2.3, “For the k-th object with label ck we encode it into GloVe embedding…
L
G
=
l
1
G
,
…
,
l
k
G
∈
R
D
1
×
K
.., is the GloVe embeddings for the objects labels”),
and computing attention using the plurality of word embeddings generating output attention maps corresponding to dot product attention matrices based on the at least one word embedding for the input query and each of the plurality of word embeddings for the object labels corresponding to the bounding boxes, (see pg. 3597, Section 2.3, ““Therefore, we compute the WL attention vector, that indicates how much weight we should give to each of the K objects in the image, in terms of the semantic similarity between the category labels of the objects and the words in the question”, and see Formula 1, where a dot product is performed between the word embeddings from the object labels (
L
T
) and the word embedding for the at least on input query (
X
G
)).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the box label embeddings taught by Huang with the scaled dot product key matrix taught by Yu. The motivation for doing so would be to improve output attention matrices (Huang, pg. 3597, Section 2.3). Thus, it would have been obvious to combine the box label embeddings taught by Huang with the teachings of y Yu and Sun in order to obtain the invention as claimed in Claim 3.
As to Claim 4, Yu in view of Huang and Sun teaches dividing the resulting matrix by a dimension based on the key matrix to generate a scaled matrix (see Yu, pg. 6276, Formula 1, where the resulting matrix
Q
K
T
is divided by dimension
d
k
).
As to Claim 5, Yu in view of Huang and Sun teaches applying softmax to the scaled matrix to generate a normalized matrix (see Xiong, pg. 6276, Formula 1, where a softmax function is applied to the scaled matrix
Q
K
T
d
k
).
As to Claim 8, Yu in view of Huang and Sun teaches wherein the final attention map includes a uniform box overlapping with an object in the input image that corresponds to the at least one topic label from the input query (see Yu, page 6280, Section 5.3, Figure 7, “SA(Y)-l, SA(X)-l and GA(X,Y)-l denote the question self-attention, image self-attention, and question guided-attention from the l-th layer, respectively. Q, A, P denote the question, answer and prediction respectively”, and see how sheep are highlighted in attention map, which corresponds to the word ‘sheep’ from the user query, ‘How many sheep can we see in this picture’).
As to Claim 9, Yu in view of Huang and Sun teaches drawing boxes around the objects detected in the input image to generate the bounding boxes (Yu, see pg. 6277, Section 4.1, “The input image is represented as a set of regional visual features in a bottom-up manner”, and see Fig. 4, where red bounding boxes are depicted over the input image),
As to Claim 10, Yu in view of Huang and Sun teaches determining a characteristic of an object in the input image that corresponds to the at least one topic label from the input query based on the final attention map (see Yu, page 6280, Section 5.3, Figure 7, “SA(Y)-l, SA(X)-l and GA(X,Y)-l denote the question self-attention, image self-attention, and question guided-attention from the l-th layer, respectively. Q, A, P denote the question, answer and prediction respectively”, and see how sheep are highlighted in attention map, which corresponds to the query ‘How many sheep can we see in this picture’, and see the answer ‘3’ output).
As to Claim 11, Yu in view of Huang and Sunt teaches an artificial intelligence (AI) device, the AI device comprising: a memory configured to store attention maps (see Sun, paragraph [0004], “In view of this problem, the present invention proposes a video question answering system, method, computer and storage medium based on dynamic attention and graph network reasoning”);
and a controller (see Sun, paragraph [0082], “A computer includes a memory and a processor”) configured to perform the same steps recited in Claim 1. Therefore, the rejection and rationale are analogous to that of Claim 1.
As to Claim 12, Claim 12 claims the same limitation claimed as Claim 2 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 2.
As to Claim 13, Claim 13 claims the same limitation claimed as Claim 3 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 3.
As to Claim 14, Claim 14 claims the same limitation claimed as Claim 4 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 4.
As to Claim 15, Claim 15 claims the same limitation claimed as Claim 5 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 5.
As to Claim 18, Claim 18 claims the same limitation claimed as Claim 8 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 8.
As to Claim 19, Claim 19 claims the same limitation claimed as Claim 10 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 10.
As to Claim 20, Yu teaches a method for controlling an artificial intelligence (Al) device, the method comprising (see pg. 6274, Section 1, “Multimodal learning to bridge vision and language has gained broad interest from both the computer vision and natural language processing communities”):
obtaining an input query, an input image (see pg. 6277, Section 3.2, “the input question and image”, and see Fig. 4 on pg. 4, depicting an input image and question),
bounding boxes for objects detected in the input image (see pg. 6277, Section 4.1, “The input image is represented as a set of regional visual features in a bottom-up manner”, and see Fig. 4, where red bounding boxes are depicted over the input image),
PNG
media_image1.png
387
1100
media_image1.png
Greyscale
Fig. 1 of Yu
and at least one topic label for one or more words in the input query (see pg. 6277, Section 4.1, “The input question is first tokenized into words”, wherein the ‘word’ is the topic label);
generating at least one word embedding for the at least one topic label from the input query, the at least one word embedding being a multi-dimensional vector (see pg. 6277, Section 4.1, “Each word in the question is further transformed into a vector using the 300-D GloVe word embeddings [25] pre trained on a large-scale corpus. This results in a sequence of words of size n×300”);
generating output attention maps corresponding to product attention matrices (see pg. 6280, caption under Fig. 7, “Visualizations of the learned attention maps (softmax(qK/√d)”, where it is known to one of ordinary skill in the art that the formula ‘(softmax(qK/√d)’ corresponds to a scaled dot product)
based on the at least one word embedding for the at least one topic label from the input query and each of bounding boxes (see Fig. 7, where the attention is calculated for each object detected in the image);
PNG
media_image2.png
720
1228
media_image2.png
Greyscale
Figure 7 of Yu
combining, via the processor, the output attention maps to generate a final attention map corresponding to the at least one topic label from the input query (see pg. 6280, Fig. 7, where the image comprises multiple combined attention maps);
and executing, via the processor, a function based on the final attention map (see pg. 6280, caption under Fig. 7, “SA(Y) l, SA(X)-l and GA(X,Y)-l denote the question self-attention, image self-attention, and image guided-attention from the l-th layer, respectively. Q, A, P denote the question, answer and prediction respectively”, and see Fig. 7, where the answer ‘3’ is output to the initial user question).
Yu fails to teach obtaining object labels corresponding to the bounding boxes. Yu further fails to teach generating, via the processor, a plurality of word embeddings for the object labels corresponding to the bounding boxes, the plurality of word embeddings being multi-dimensional vectors, and that the product attention matrices is also based on and each of the plurality of word embeddings for the object labels corresponding to the bounding boxes.
However, in an analogous art, Huang teaches a method for controlling question answering system (see pg. 3595, Abstract, “Attention mechanisms are widely used in Visual Question Answering (VQA) to search for visual clues related to the question… this paper proposes a multi-grained attention method”), which comprises
obtaining, an input query (see Fig. 1, where the question is ‘What is the man wearing around his face?’),
an input image (see Section 2.3, pg. 3597, “For the input image in Figure 1”),
bounding boxes for objects detected in the input image, object labels corresponding to the bounding boxes (see Section 2.3, pg. 3597, “For the input image in Figure 1, Faster-RCNN detected objects with labels of “man”, “head””),
generating a plurality of word embeddings for the object labels corresponding to the bounding boxes, the plurality of word embeddings being multi-dimensional vectors (see pg. 3597, “For the k-th object with label ck we encode it into GloVe embedding…
L
G
=
l
1
G
,
…
,
l
k
G
∈
R
D
1
×
K
.., is the GloVe embeddings for the objects labels”),
generating output attention maps corresponding to attention matrices based on the at least one word embedding for the input query and each of the plurality of word embeddings for the object labels corresponding to the bounding boxes, (see pg. 3597, Section 2.3, “Therefore, we compute the WL attention vector, that indicates how much weight we should give to each of the K objects in the image, in terms of the semantic similarity between the category labels of the objects and the words in the question”, and see Formula 1, where a dot product is performed between the word embeddings from the object labels (
L
T
) and the word embedding for the at least on input query (
X
G
), and see caption under Fig. 1, “Figure 1: An example of VQA and the attention maps produced by a state-of-the-art model and our model.”)
Thus, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention to combine the bounding box label embedding taught by Huang with the product attention taught by Yu. The motivation for doing so would be to enhance the output attention maps. Huang teaches on pg. 3597, Section 2.3, “In contrast to Anderson et al. (2018) that only use objects’ visual features without the labels, and unlike Wu et al. (2018) that discard the visual features once the labels are generated, we utilize both category labels and the visual features to enhance the fine-grained attention with object level grounding.”
Neither Yu nor Huang explicitly teach a processor. However, in an analogous art Sun teaches a visual questioning answering system comprising a processor (see paragraph [0082], “A computer includes a memory and a processor, and the memory stores a computer program. When the processor executes the computer program, the steps of a video question answering method based on dynamic attention and graph network reasoning are realized”),
which can generate bounding boxes for objects detected in the input image, object labels corresponding to the bounding boxes, and then generate, via the processor, a plurality of word embeddings for the object labels corresponding to the bounding boxes (see paragraph [0097], “The object space feature and object category feature calculation module is used to predict the object tag frame and category label in the video according to the object detection model to obtain the object space feature and object category feature”, and see paragraph [0031], “EL is the word embedding vector representation of the object category label”),
and generate output attention maps corresponding to product attention matrices based on the input query and a word embedding for the object label corresponding to the bounding box, (see paragraph [0056], “Use the scaled dot product function to calculate the similarity matrix between the problem feature and the object's joint feature”, wherein the problem feature is a user query, and the object’s joint feature is based on the embedded category label).
Thus, it would have been obvious to combine the processor taught by Sun with the teachings of Yu and Huan. The motivation for doing so would be to implement the method taught by Yu and Huang on a variety of computing devices (see Sun, paragraph [0176]). Thus, it would have been obvious to combine the processor taught by Sun with the method taught by Yu and Huang in order to obtain the invention as claimed in Claim 20.
Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (Yu, et al., "Deep Modular Co-Attention Networks for Visual Question Answering," 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 6274-6283), hereinafter Yu, in view of Huang et al. (Huang, P. et al., “Multi-grained Attention with Object-level Grounding for Visual Question Answering”, ACL 2019, pp. 3595-3600)), hereinafter Huang, and further in view of Sun et al. (CN Pub No 113609355), hereinafter Sun, and further in view of Bera et al. (A. Bera, Z. Wharton, Y. Liu, N. Bessis and A. Behera, "SR-GNN: Spatial Relation-Aware Graph Neural Network for Fine-Grained Image Categorization," in IEEE Transactions on Image Processing, vol. 31, pp. 6017-6031, 2022), hereinafter Bera.
As to Claim 6, Yu in view of Huang and Sun teaches multiplying the normalized matrix and a value matrix (see Yu, pg. 6277, Formula 1, where the normalized matrix
Q
K
T
d
k
is multiplied by value matrix
V
), Yu further teaches that attention maps can be output through multipoint a normalized matrix and a value matrix (see page 6276, Section 3.1, “Given a query q ∈R1xd, n key-value pairs (packed into a key matrix K ∈ Rn×d and a value matrix V ∈ Rn×d), the attended feature f ∈R1×d is obtained by weighted summation overall values V with respect to the attention learned from q and K”),
the output attention map being one of the output attention maps (see page 6280, Section 5.3, Figure 7, see multiple attention maps),
but fails to explicitly teach that the value matrix is based on an attention map corresponding to object detection.
However, in an analogous art, Bera teaches multiplying a value vector with the dot product of a query and key vector (see page 6021, Section III, Subsection D., “The dot product of Q and K results in the attention weight matrix, which is multiplied with V to produce the desired transformed feature representation”),
wherein the value matrix is based on attention map corresponding to object detection (see page 6021, Section III, Subsection D., “The aim is to generate an attention-focused context vector (i.e., value V ) that enables our model to selectively focus on more relevant regions to generate holistic context information”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the value attention matrix taught by Bera with the teachings of Yu, Huang, and Sun. The motivation for doing so would be to ensure the model focuses on more relevant regions, as taught by Bera in page 6021. Thus, it would have been obvious to combine the teachings of Yu, Huang, Sun, and Bera in order to obtain the invention as claimed in Claim 6.
As to Claim 16, Claim 16 claims the same limitation claimed as Claim 6 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 6.
Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Yu et al. (Yu, et al., "Deep Modular Co-Attention Networks for Visual Question Answering," 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pp. 6274-6283), hereinafter Yu, in view of Huang et al. (Huang, P. et al., “Multi-grained Attention with Object-level Grounding for Visual Question Answering”, ACL 2019, pp. 3595-3600)), hereinafter Huang, and further in view of Sun et al. (CN Pub No 113609355), hereinafter Sun, further in view of Bera et al. (A. Bera, Z. Wharton, Y. Liu, N. Bessis and A. Behera, "SR-GNN: Spatial Relation-Aware Graph Neural Network for Fine-Grained Image Categorization," in IEEE Transactions on Image Processing, vol. 31, pp. 6017-6031, 2022), hereinafter Bera, and further in view of Zhou et al. (US Pub No 20240169733), hereinafter Zhou.
As to Claim 7, Yu in view of Huang, Sun, and Bera fails to teach wherein the attention map for the value matrix is a type of heat map.
Bera teaches that the value matrix may be based on an attention map (see page 6021, Section III, Subsection D.), but fails to explicitly teach that the value matrix is a heat map.
However, in an analogous art, Zhou teaches a heat map where important features are highlighted (see paragraph [0083], “When extracting a target object from a feature map, a target object may be highlighted in response to filtering an image through an non matrix (the value of n may be considered based on elements such as a receptive field and accuracy, for example, may be set to 3*3, 5*5, 7*7, etc.)”),
and that the highlighted features can be used as a value for the value matrix (see paragraph [0115], “First, a linear task may be performed on each of the input Q (e.g., a clip query, an object representation by previous iteration processing, corresponding to the dimension L, C), K (using a video feature as a key feature, corresponding to the dimension THW, C), and V (using a video feature as a value feature”).
Thus, it would have been obvious to combine the feature highlighting taught by Zhou with the teachings of Yu, Huang, Sun, and Bera. The motivation for doing so would be to improve the segmentation of key features from images. Zhou teaches in paragraph [0055], “the accuracy and robustness of a segmentation may be improved to a certain level”). Thus, it would have been obvious to combine the feature highlighting taught by Zhou with the teachings Yu and Huang in order to obtain the invention as claimed in Claim 7.
As to Claim 17, Claim 17 claims the same limitation claimed as Claim 7 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 7.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Liu et al. (Liu, Shikun, et al. “Prismer: A Vision-Language Model with Multi-Task Experts”, 2023, and see corresponding US Pub No 20240265690) teaches a method of answering user questions by using a transformer model. Liu teaches obtaining bounding boxes and corresponding labels, and inputting these labels into a transformer model in order to calculate attention.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOUMYA THOMAS whose telephone number is (571)272-8639. The examiner can normally be reached M-F 8:30-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.T./Examiner, Art Unit 2664
/CHARLOTTE M BAKER/Primary Examiner, Art Unit 2664