DETAILED ACTION
This office action is in response to submission of application on 01/31/2024.
Claims 1-15 are presented for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/31/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-15 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 1:
Step 1: The claim is directed to a system, which falls within the statutory category of a machine/manufacture.
Step 2A Prong 1: The claim is directed to an abstract idea. Specifically, the claim recites:
obtain an attention score matrix based on a dot-product of the obtained query feature map matrix and key feature map matrix; (Abstract idea – mathematical concept. Performing a dot product between two matrices is a mathematical calculation. See MPEP 2106.04(a)(2)(I).)
for each of attention score vectors of the plurality of tokens included in the obtained attention score matrix, set a threshold based on a maximum value from among included attention scores; (Abstract idea – mental process. Setting an attention score threshold for an attention score vector based on a maximum attention score can practically be performed in the human mind or with the aid of pen and paper, for example, by mentally determining the threshold to be equal to the maximum attention score minus a parameter value. The courts have recognized that claims can recite a mental process even if they are claimed as being performed on a computer. See MPEP 2106.04(a)(2)(III).)
bypass a softmax operation for an attention score less than the set threshold for each of the attention score vectors of the plurality of tokens. (Abstract idea – mental process. Bypassing a softmax operation for an attention score less than the threshold can practically be performed in the human mind or with the aid of pen and paper, for example, by mentally comparing the attention score to the threshold, mentally determining that the attention score is less than the threshold, and mentally determining not to perform the softmax operation on that attention score. See MPEP 2106.04(a)(2)(III).)
Step 2A Prong 2: The additional elements recited in the claim do not integrate the abstract idea into a practical application, individually or in combination. Specifically, the claim recites the additional elements:
at least one processor configured to control a process using the artificial intelligence model; and a memory configured to store instructions performed by the at least one processor, wherein, when performing a process of an attention layer included in the artificial intelligence model, the at least one processor is configured to: (This limitation is interpreted as implementation of the disclosed steps in a generic computing environment, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
obtain a query feature map matrix and a key feature map matrix from an input sequence including a plurality of tokens; (Obtaining a query feature map matrix and a key feature map matrix amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Specifically, the claim recites the additional elements:
at least one processor configured to control a process using the artificial intelligence model; and a memory configured to store instructions performed by the at least one processor, wherein, when performing a process of an attention layer included in the artificial intelligence model, the at least one processor is configured to: (This limitation is interpreted as implementation of the disclosed steps in a generic computing environment, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
obtain a query feature map matrix and a key feature map matrix from an input sequence including a plurality of tokens; (Obtaining a query feature map matrix and a key feature map matrix amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).)
Claims 2-5:
Claim 2 recites The computing system of claim 1, wherein a threshold for a specific attention score vector from among the attention score vectors is set based on a maximum value from among attention scores included in the specific attention score vector and a certain parameter. Setting the threshold based on a maximum attention score and a certain parameter can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally determining the threshold to be equal to the maximum attention score minus the parameter value. Therefore, the claim merges with the abstract idea recited in claim 1, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim 3 recites The computing system of claim 2, wherein the threshold is set according to the equation below,
t
h
r
e
s
h
o
l
d
=
l
n
(
β
)
+
m
a
x
(
s
i
)
where β is the parameter, and
m
a
x
(
s
i
)
is the maximum value from among the attention scores included in the specific attention score vector. Setting the threshold according to the recited equation is a mathematical concept. Therefore, the claim merges with the abstract idea recited in claim 2, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim 4 recites The computing system of claim 3, wherein the parameter is set to have a value included in a range of 0.0005 to 0.002. This claim merely specifies the value of the parameter β in the equation of claim 3. Therefore, the claim merges with the abstract idea (mathematical concept) recited in claim 3, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim 5 recites The computing system of claim 1, wherein the at least one processor processes an attention score less than the threshold as 0 based on a threshold set for each of the attention score vectors. Processing an attention score less than the threshold as 0 can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally determining that the attention score is less than the threshold and mentally determining that the attention score should be treated as 0 in downstream operations. Processing the attention score with the processor amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f). Therefore, the claim merges with the abstract idea recited in claim 1, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim 6:
Step 1: The claim is directed to a system, which falls within the statutory category of a machine/manufacture.
Step 2A Prong 1: The claim is directed to an abstract idea. Specifically, the claim recites:
obtain an attention score matrix based on a dot-product of the obtained query feature map matrix and key feature map matrix; (Abstract idea – mathematical concept. Performing a dot product between two matrices is a mathematical calculation. See MPEP 2106.04(a)(2)(I).)
obtain an attention probability matrix through a softmax operation on the obtained attention score matrix; (Abstract idea – mathematical concept. Performing a softmax operation on a matrix is a mathematical calculation. See MPEP 2106.04(a)(2)(I).)
set a threshold based on a length of the input sequence, for each attention probability vector in the attention probability matrix; (Abstract idea – mental process. Setting an attention probability threshold for an attention probability vector based on a length of the input sequence can practically be performed in the human mind or with the aid of pen and paper, for example, by mentally determining the threshold to be equal to a parameter value divided by the sequence length. The courts have recognized that claims can recite a mental process even if they are claimed as being performed on a computer. See MPEP 2106.04(a)(2)(III).)
bypass a dot-product with the value feature map matrix for elements of the attention probability vector less than the set threshold from among respective attention probabilities of the plurality of tokens included in the attention probability matrix. (Abstract idea – mental process. Bypassing a dot-product calculation for an attention probability less than the threshold can practically be performed in the human mind or with the aid of pen and paper, for example, by mentally comparing the attention probability to the threshold, mentally determining that the attention probability is less than the threshold, and mentally determining not to perform the dot-product calculation on that attention probability. See MPEP 2106.04(a)(2)(III).)
Step 2A Prong 2: The additional elements recited in the claim do not integrate the abstract idea into a practical application, individually or in combination. Specifically, the claim recites the additional elements:
at least one processor configured to control a process using the artificial intelligence model; and a memory configured to store instructions performed by the at least one processor, wherein, when performing a process of an attention layer included in the artificial intelligence model, the at least one processor is configured to: (This limitation is interpreted as implementation of the disclosed steps in a generic computing environment, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
obtain a query feature map matrix, a key feature map matrix, and a value feature map matrix from an input sequence including a plurality of tokens; (Obtaining a query feature map matrix, a key feature map matrix, and a value feature map matrix amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g).)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Specifically, the claim recites the additional elements:
at least one processor configured to control a process using the artificial intelligence model; and a memory configured to store instructions performed by the at least one processor, wherein, when performing a process of an attention layer included in the artificial intelligence model, the at least one processor is configured to: (This limitation is interpreted as implementation of the disclosed steps in a generic computing environment, and thus amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f).)
obtain a query feature map matrix, a key feature map matrix, and a value feature map matrix from an input sequence including a plurality of tokens; (Obtaining a query feature map matrix, a key feature map matrix, and a value feature map matrix amounts to adding insignificant extra-solution activity (necessary data gathering) to the judicial exception – see MPEP2106.05(g). Further, the limitation is directed to receiving or transmitting data over a network, which the courts have found to be well-understood, routine, and conventional in the computer arts – see MPEP 2106.05(d).)
Claims 7-10:
Claim 7 recites The computing system of claim 6, wherein the threshold is set based on the length of the input sequence and a certain parameter. Setting the threshold based on the length of the input sequence and a certain parameter can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally determining the threshold to be equal to the parameter value divided by the sequence length. Therefore, the claim merges with the abstract idea recited in claim 6, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim 8 recites The computing system of claim 7, wherein the threshold is set according to the equation below,
t
h
r
e
s
h
o
l
d
=
a
s
e
q
_
l
e
n
where
α
is the parameter, and
s
e
q
_
l
e
n
is the length of the input sequence. Setting the threshold according to the recited equation is a mathematical concept. Therefore, the claim merges with the abstract idea recited in claim 7, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim 9 recites The computing system of claim 8, wherein the parameter is set to have a value included in a range of 0.2 to 0.6. This claim merely specifies the value of the parameter
α
in the equation of claim 8. Therefore, the claim merges with the abstract idea (mathematical concept) recited in claim 8, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim 10 recites The computing system of claim 6, wherein the at least one processor processes a value of an attention probability less than the threshold from among the respective attention probabilities of the plurality of tokens as 0. Processing an attention probability less than the threshold as 0 can practically be performed in the human mind or with the aid of pen and paper (i.e. mental process), for example, by mentally determining that the attention probability is less than the threshold and mentally determining that the attention probability should be treated as 0 in downstream operations. Processing the attention probability with the processor amounts to adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea – see MPEP 2106.05(f). Therefore, the claim merges with the abstract idea recited in claim 6, and does not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claims 11-15 are product claims containing substantially the same elements as system claims 1-5, respectively, and are rejected on the same grounds under 35 U.S.C. 101 as claims 1-5, respectively, mutatis mutandis. The additional components of A non-transitory computer-readable storage medium having recorded thereon instructions for executing an attention-based artificial intelligence model, wherein a processor of the computer is configured to execute the instructions are interpreted as a general-purpose computer and mere instructions to apply the judicial exception on the computer. Therefore, the claims do not recite additional elements that are sufficient to amount to significantly more than the abstract idea.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 5, 11-13 and 15 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by
Ham et al. (hereinafter Ham), “A3: Accelerating Attention Mechanisms in Neural Networks with Approximation” (published 02/22/2020).
Regarding Claim 1,
Ham teaches A computing system that performs a process using an attention-based artificial intelligence model, the computing system comprising: (Pg. 1, Abstract: “The attention mechanism is widely adopted by many state-of-the-art neural networks for computer vision, natural language processing, and machine translation…” Pg. 2, section I: “We design a specialized hardware pipeline for an attention mechanism exploiting parallelism and data path specialization to significantly improve the performance and energy efficiency of the attention mechanism.”)
at least one processor configured to control a process using the artificial intelligence model; and a memory configured to store instructions performed by the at least one processor, (Examiner notes that this limitation is interpreted as a generic computing environment. Pg. 3, section III: “We introduce the base design of A3, a specialized hardware accelerator for attention mechanisms in neural networks which can be integrated to either CPU, GPU, or an existing hardware accelerator.”)
wherein, when performing a process of an attention layer included in the artificial intelligence model, the at least one processor is configured to:
obtain a query feature map matrix and a key feature map matrix from an input sequence including a plurality of tokens; (Pg. 6, section IV.B: “Figure 6 illustrates the basic idea of our algorithm. Given a key matrix and a query vector, this algorithm first replicates query vector across rows to make a replicated query matrix.” Pg. 3, section II.B: “A single vector in the key matrix usually represents an embedding of a word, a sentence, a knowledge, or any other portion of a larger entity.” The algorithm obtains a replicated query matrix (i.e. a query feature map matrix) and a key matrix (i.e. a key feature map matrix) from a larger entity (i.e. input sequence) including a plurality of words, sentences, knowledge, etc. (i.e. tokens).)
obtain an attention score matrix based on a dot-product of the obtained query feature map matrix and key feature map matrix; (Pg. 8, section IV.D: “Once a subset of rows of the key matrix is chosen as candidates, their full scores (i.e., the dot product between the key matrix row and the query) are computed.”)
for each of attention score vectors of the plurality of tokens included in the obtained attention score matrix, set a threshold based on a maximum value from among included attention scores; and (Pg. 8, section IV.D: “To minimize the computation required to calculate the softmax and the weighted sum, it is beneficial to avoid including some of the low-scoring candidates for those steps. One way to perform an approximation on this step is to sort candidates based on their dot product results and then only including top scoring rows for the next steps… we utilize a dynamic post-scoring approximation scheme which decides whether to include a row for the next steps. Basically, a score for a particular row is compared with the top-scoring row’s score and if their difference is larger than threshold
t
, such row is excluded for the next steps.” For a vector of attention scores, a threshold is set based on the top-scoring row’s score (i.e. the maximum value from among included attention scores).)
bypass a softmax operation for an attention score less than the set threshold for each of the attention score vectors of the plurality of tokens. (See the portion of section IV.D cited above. When a row’s attention score differs from the top-scoring row’s score by more than
t
(i.e. the attention score is less than the threshold), it is excluded (i.e. bypassed) for the subsequent steps, which include the softmax operation.)
Regarding Claim 2, Ham teaches The computing system of claim 1, as shown above.
Ham also teaches wherein a threshold for a specific attention score vector from among the attention score vectors is set based on a maximum value from among attention scores included in the specific attention score vector and a certain parameter. (Pg. 8, section IV.D: “[W]e utilize a dynamic post-scoring approximation scheme which decides whether to include a row for the next steps. Basically, a score for a particular row is compared with the top-scoring row’s score and if their difference is larger than threshold
t
, such row is excluded for the next steps… Throughout the paper, we use
T
=
100
*
(
1
/
e
t
)
instead of directly using
t
.” A score is excluded when
m
a
x
_
s
c
o
r
e
-
s
c
o
r
e
>
t
, i.e., when
s
c
o
r
e
<
m
a
x
_
s
c
o
r
e
-
t
, i.e.,
t
h
r
e
s
h
o
l
d
=
m
a
x
_
s
c
o
r
e
-
t
. Given
t
h
r
e
s
h
o
l
d
=
m
a
x
_
s
c
o
r
e
-
t
and
T
=
100
*
(
1
/
e
t
)
, we can rewrite the threshold according to the derivation below:
T
=
100
*
(
1
/
e
t
)
[Wingdings font/0xE0]
t
=
-
ln
T
100
t
h
r
e
s
h
o
l
d
=
m
a
x
_
s
c
o
r
e
-
t
[Wingdings font/0xE0]
t
h
r
e
s
h
o
l
d
=
m
a
x
_
s
c
o
r
e
-
(
-
ln
T
100
)
[Wingdings font/0xE0]
t
h
r
e
s
h
o
l
d
=
ln
T
100
+
m
a
x
_
s
c
o
r
e
.
The threshold is set based on the top-scoring row’s score (represented by
m
a
x
_
s
c
o
r
e
above) (i.e. maximum value from among attention scores included in the attention score vector) and
T
100
(i.e. a certain parameter).)
Regarding Claim 3, Ham teaches The computing system of claim 2, as shown above.
Ham also teaches wherein the threshold is set according to the equation below,
t
h
r
e
s
h
o
l
d
=
l
n
(
β
)
+
m
a
x
(
s
i
)
where
β
is the parameter, and
m
a
x
(
s
i
)
is the maximum value from among the attention scores included in the specific attention score vector. (See the portion of pg. 8, section IV.D cited above in regard to claim 2. The threshold is set according to the equation
t
h
r
e
s
h
o
l
d
=
ln
T
100
+
m
a
x
_
s
c
o
r
e
, where
T
100
corresponds to the claimed parameter
β
, and
m
a
x
_
s
c
o
r
e
corresponds to the claimed maximum value
m
a
x
(
s
i
)
from among attention scores included in the attention score vector.)
Regarding Claim 5, Ham teaches The computing system of claim 1, as shown above.
Ham also teaches wherein the at least one processor processes an attention score less than the threshold as 0 based on a threshold set for each of the attention score vectors. (Pg. 6, section IV.A: “Since score values are transformed to weights with this function [softmax], candidates with relatively small score values get near-zero weight. In addition, these near-zero weights often do not contribute to the accuracy of the model… So for these near-zero weights, it is actually more beneficial to treat them as zeros and avoid including them for softmax computation and the following weighted sum computation.” Attention scores which would be transformed to near-zero weights (i.e. attention scores less than the threshold) are processed as zeros.)
Claims 11-13 and 15 are product claims containing substantially the same elements as system claims 1-3 and 5, respectively. Ham teaches the elements of claims 1-3 and 5, as shown above.
Ham also teaches A non-transitory computer-readable storage medium having recorded thereon instructions for executing an attention-based artificial intelligence model, wherein a processor of the computer is configured to execute the instructions (Examiner notes that this limitation is interpreted as implementation of the disclosed steps in a generic computing environment. Pg. 3, section III: “We introduce the base design of A3, a specialized hardware accelerator for attention mechanisms in neural networks which can be integrated to either CPU, GPU, or an existing hardware accelerator.”)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Ham.
Regarding Claim 4, Ham teaches The computing system of claim 3, as shown above.
While Ham does not appear to explicitly disclose wherein the parameter is set to have a value included in a range of 0.0005 to 0.002, MPEP 2144.05(I) states that “a prima facie case of obviousness exists where the claimed ranges or amounts do not overlap with the prior art but are merely close.” In this case, the threshold parameter
β
is claimed to have a value in the range of 0.0005 to 0.002. As can be seen in Ham’s figure 12 (copied below), Ham’s algorithm is evaluated for
T
∈
{
1
,
2.5
,
5
,
10
,
20
}
, and thus Ham’s parameter
T
100
takes values in the range of 0.01 to 0.2 (Ham, pg. 10). These ranges are sufficiently close, and one of ordinary skill in the art would have been motivated to further optimize Ham’s parameter
T
100
as it is recognized as a configurable, result-effective parameter that can be decreased to improve accuracy: “the lower T indicates more conservative approximation and the higher
T
indicates more aggressive approximation… One of the main strengths of our approach is that
M
and
T
are configurable. By changing
M
and
T
, a user always can select the degree of approximation and choose the tradeoffs between accuracy and performance/energy efficiency.”
PNG
media_image1.png
182
566
media_image1.png
Greyscale
Claim 14 is a product claim containing substantially the same elements as system claim 4. Ham teaches the elements of claim 4, as shown above.
Claims 6-10 are rejected under 35 U.S.C. 103 as being unpatentable over Ham in view of
Yoon et al. (hereinafter Yoon), “Pruning Self-Attention for Zero-Shot Multi-Speaker Text-to-Speech” (published 08/28/2023).
.
Regarding Claim 6,
Ham teaches A computing system that performs a process using an attention-based artificial intelligence model, the computing system comprising: (Pg. 1, Abstract: “The attention mechanism is widely adopted by many state-of-the-art neural networks for computer vision, natural language processing, and machine translation…” Pg. 2, section I: “We design a specialized hardware pipeline for an attention mechanism exploiting parallelism and data path specialization to significantly improve the performance and energy efficiency of the attention mechanism.”)
at least one processor configured to control a process using the artificial intelligence model; and a memory configured to store instructions performed by the at least one processor, (Examiner notes that this limitation is interpreted as a generic computing environment. Pg. 3, section III: “We introduce the base design of A3, a specialized hardware accelerator for attention mechanisms in neural networks which can be integrated to either CPU, GPU, or an existing hardware accelerator.”)
wherein, when performing a process of an attention layer included in the artificial intelligence model, the at least one processor is configured to:
obtain a query feature map matrix, a key feature map matrix, and a value feature map matrix from an input sequence including a plurality of tokens; (Pg. 3, section III.A: “Base A3 takes three inputs — a key matrix (n × d), a value matrix (n × d), and a query vector (d)…” Pg. 6, section IV.B: “Figure 6 illustrates the basic idea of our algorithm. Given a key matrix and a query vector, this algorithm first replicates query vector across rows to make a replicated query matrix.” Pg. 3, section II.B: “A single vector in the key matrix usually represents an embedding of a word, a sentence, a knowledge, or any other portion of a larger entity.” The algorithm obtains a replicated query matrix (i.e. a query feature map matrix), a key matrix (i.e. a key feature map matrix), and a value matrix (i.e. a value feature map matrix) from a larger entity (i.e. input sequence) including a plurality of words, sentences, knowledge, etc. (i.e. tokens).)
obtain an attention score matrix based on a dot-product of the obtained query feature map matrix and key feature map matrix; (Pg. 8, section IV.D: “Once a subset of rows of the key matrix is chosen as candidates, their full scores (i.e., the dot product between the key matrix row and the query) are computed.”))
obtain an attention probability matrix through a softmax operation on the obtained attention score matrix; (Pg. 2, section II.A: “This array is then processed with softmax function…”)
set a threshold [based on a length of the input sequence], for each attention probability vector in the attention probability matrix; and (Pg. 8, section IV.D: “To minimize the computation required to calculate the softmax and the weighted sum, it is beneficial to avoid including some of the low-scoring candidates for those steps. One way to perform an approximation on this step is to sort candidates based on their dot product results and then only including top scoring rows for the next steps… we utilize a dynamic post-scoring approximation scheme which decides whether to include a row for the next steps. Basically, a score for a particular row is compared with the top-scoring row’s score and if their difference is larger than threshold
t
, such row is excluded for the next steps. If a row’s score is smaller than the top row’s score by more than
t
, this means that this row will have a post-softmax weight that is at least
e
t
×
smaller than that of the top-scoring row.” For a row of attention scores, a threshold is set for the corresponding vector of post-softmax attention probabilities.
bypass a dot-product with the value feature map matrix for elements of the attention probability vector less than the set threshold from among respective attention probabilities of the plurality of tokens included in the attention probability matrix. (See the portion of section IV.D cited above. When a row’s post-softmax attention probability differs from the top-scoring row’s probability by more than
e
t
×
(i.e. the attention probability is less than the threshold), it is excluded (i.e. bypassed) for the subsequent steps, which include the weighted sum (i.e. dot product) with the value matrix.)
Ham does not appear to explicitly disclose setting the threshold based on a length of the input sequence
However, Yoon teaches setting the threshold based on a length of the input sequence (Pg. 1, Abstract: “[W]e prune off redundant connections from self-attention layers whose attention weights are below the threshold.” Pg. 3, section 3.1.2: “[W]e first define a hard sparse mask of
h
-th head
S
M
h
a
r
d
h
that inherits the learnable threshold
θ
/
N
… where
θ
is a trainable threshold parameter, and
N
is the length of the sequence used to adjust the threshold value based on variations in input length.” The threshold for pruning attention probabilities is set based on the length of the input sequence
N
.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Ham and Yoon. Ham teaches accelerating an attention layer of a neural network by pruning attention scores/probabilities below a threshold. Yoon teaches pruning attention probabilities below a threshold, where the threshold is dynamically determined based on the length of the input sequence and a trainable parameter. One of ordinary skill would have motivation to combine Ham and Yoon because “the optimal threshold values vary depending on the number of layers, type of generation tasks, and degree of domain mismatch; thus, flexibly setting the threshold is preferable. To this end, we propose a novel differentiable pruning (DP) method with learnable thresholds,” which “adjust[s] the threshold value based on variations in input length” (Yoon, pg. 2-3, section 3.1.2) in order to “flexibly determine the pruning strength for searching optimal degree of generalization” (Yoon, pg. 1, Abstract).
Regarding Claim 7, Ham and Yoon teach The computing system of claim 6, as shown above.
Yoon also teaches wherein the threshold is set based on the length of the input sequence and a certain parameter. (Pg. 3, section 3.1.2: “[W]e first define a hard sparse mask of
h
-th head
S
M
h
a
r
d
h
that inherits the learnable threshold
θ
/
N
… where
θ
is a trainable threshold parameter, and
N
is the length of the sequence used to adjust the threshold value based on variations in input length.” The threshold is set based on the length of the input sequence
N
and a certain parameter
θ
.)
Regarding Claim 8, Ham and Yoon teach The computing system of claim 7, as shown above.
Yoon also teaches wherein the threshold is set according to the equation below,
t
h
r
e
s
h
o
l
d
=
α
/
s
e
q
_
l
e
n
where
α
is the parameter, and
s
e
q
_
l
e
n
is the length of the input sequence. (Pg. 3, section 3.1.2: “[W]e first define a hard sparse mask of
h
-th head
S
M
h
a
r
d
h
that inherits the learnable threshold
θ
/
N
… where
θ
is a trainable threshold parameter, and
N
is the length of the sequence used to adjust the threshold value based on variations in input length.” The threshold is set according to the equation
t
h
r
e
s
h
o
l
d
=
θ
/
N
, where
θ
is the parameter, and
N
is the length of the input sequence.)
Regarding Claim 9, Ham and Yoon teach The computing system of claim 8, as shown above.
While Ham and Yoon do not appear to explicitly disclose wherein the parameter is set to have a value included in a range of 0.2 to 0.6, MPEP 2144.05(I) states that “a prima facie case of obviousness exists where the claimed ranges or amounts do not overlap with the prior art but are merely close.” In this case, the threshold parameter is claimed to have a value in the range of 0.2 to 0.6. As can be seen in Yoon’s table 3 (copied below), Yoon’s learnable threshold parameter
θ
takes values in the range of 0.76 to 5.11 (Yoon, pg. 4). These ranges are sufficiently close, and one of ordinary skill in the art would have been motivated to further optimize Yoon’s parameter
θ
as it is recognized as a learnable, result-effective parameter that can be reduced to decrease loss: “
θ
is a trainable threshold parameter…
L
t
t
s
pulls
θ
down towards 0” in order to “prun[e] off self-attention connections only to the extent that it does not significantly harm the original objective of minimizing
L
t
t
s
” (Yoon, pg. 3, section 3.1.2).
PNG
media_image2.png
106
294
media_image2.png
Greyscale
Regarding Claim 10, Ham and Yoon teach The computing system of claim 6, as shown above.
Ham also teaches wherein the at least one processor processes a value of an attention probability less than the threshold from among the respective attention probabilities of the plurality of tokens as 0. (Pg. 6, section IV.A: “Since score values are transformed to weights with this function, candidates with relatively small score values get near-zero weight. In addition, these near-zero weights often do not contribute to the accuracy of the model… So for these near-zero weights, it is actually more beneficial to treat them as zeros and avoid including them for softmax computation and the following weighted sum computation.” Attention probabilities less than the threshold are processed as zeros.)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN M ROHD whose telephone number is (571)272-6445. The examiner can normally be reached Mon-Thurs 8:00-6:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/B.M.R./Examiner, Art Unit 2147
/ERIC NILSSON/Primary Examiner, Art Unit 2151