DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This Office Action is in response to the communication filed on August 2, 2024.
Claims 1-14 and 16-21 are being considered on the merits.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 23 June 2025 has been considered. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, initialed and dated copies of Applicant's IDS forms 1499 are attached to the instant Office action.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 14 and 16-21 are rejected under 35 U.S.C. 101 because of the following:
Claim 1:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
A system comprising a neural network that is configured to process an input sequence comprising a respective input element at each of a plurality of input positions and to generate a network output representing the input sequence, (Mental Process: Processing an input sequence to generate a network output representation is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting a neural network, nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can evaluate 1 and conceptualize a green dot output entirely in their minds or with the assistance of a pen and paper.)
Step 2A Prong 2: This judicial exception is not integrated into practical application
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
obtaining a block input sequence having a respective block input element at only each of N block input positions, wherein N > 1; and (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
processing the block input sequence to generate a block output sequence having a respective block output element at only each of M block output positions, wherein N > M > 1, the processing comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
applying a learned weight matrix to the block input sequence to generate an intermediate representation comprising, for each block output position, a respective score corresponding to each block input position, and (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
applying the intermediate representation to the block input sequence to generate the block output sequence. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
obtaining a block input sequence having a respective block input element at only each of N block input positions, wherein N > 1; and (Insignificant Extra Solution Activity: Receiving or transmitting data over a network is well-understood, routine, conventional activity – see Berkheimer evidence MPEP § 2106.05(d))
processing the block input sequence to generate a block output sequence having a respective block output element at only each of M block output positions, wherein N > M > 1, the processing comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
applying a learned weight matrix to the block input sequence to generate an intermediate representation comprising, for each block output position, a respective score corresponding to each block input position, and (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
applying the intermediate representation to the block input sequence to generate the block output sequence. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 2:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
The system of claim 1, wherein: the block input sequence is represented as an input matrix
X
∈
R
N
×
D
, wherein each row of the input matrix represents a respective block input element
x
j
∈
R
D
,
j
∈
[
1
,
…
,
N
]
(Mental Process: “Representing” is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components; nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can “represent” an input sequence as an input matrix by crossing out an input sequence converting it to a matrix using pen and paper)
Step 2A Prong 2 and Step 2B: The claim does not include additional elements.
Claim 3:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
The system of claim 1, wherein: the block output sequence is represented as an output matrix
Y
∈
R
M
×
D
, wherein each row of the output matrix represents a respective block output element
y
j
∈
R
D
,
i
∈
[
1
,
…
,
M
]
(Mental Process: “Representing” is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components; nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can “represent” an input sequence as an input matrix by crossing out an input sequence converting it to a matrix using pen and paper)
Step 2A Prong 2 and Step 2B: The claim does not include additional elements.
Claim 4:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
(
X
W
)
T
(Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f)
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f)
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
(
X
W
)
T
(Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f)
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f)
Claim 5:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1 wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
W
T
X
T
(Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1 wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
W
T
X
T
(Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 6:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein: applying the intermediate representation to the block input sequence to generate the block output sequence comprises computing:
Y
=
s
o
f
t
m
a
x
(
S
)
∙
X
((Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
wherein
Y
∈
R
M
×
D
represents the block output sequence,
X
∈
R
N
×
D
represents the block input sequence, and
S
∈
R
M
×
N
represents the intermediate representation. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein: applying the intermediate representation to the block input sequence to generate the block output sequence comprises computing:
Y
=
s
o
f
t
m
a
x
(
S
)
∙
X
((Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
wherein
Y
∈
R
M
×
D
represents the block output sequence,
X
∈
R
N
×
D
represents the block input sequence, and
S
∈
R
M
×
N
represents the intermediate representation. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 7:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein the sequence of network blocks comprises a plurality of merger network blocks, wherein each particular merger network block is configured to generate a block output sequence having a shorter length than the respective block output sequences of any other merger network blocks that precede the particular merger network block in the sequence of network blocks. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein the sequence of network blocks comprises a plurality of merger network blocks, wherein each particular merger network block is configured to generate a block output sequence having a shorter length than the respective block output sequences of any other merger network blocks that precede the particular merger network block in the sequence of network blocks. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 8:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1: wherein the merger network block is configured to obtain a block input sequence comprising a variable number
N
of block input elements and to generate a block output sequence having a predetermined number
M
of block output elements. (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1: wherein the merger network block is configured to obtain a block input sequence comprising a variable number
N
of block input elements and to generate a block output sequence having a predetermined number
M
of block output elements. (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
Claim 9:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein the neural network is configured to perform a first machine learning task, and wherein the neural network has been pre-trained to perform a second machine learning task that is different than the first machine learning task. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein the neural network is configured to perform a first machine learning task, and wherein the neural network has been pre-trained to perform a second machine learning task that is different than the first machine learning task. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 10:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 9, wherein the input sequences corresponding to the first machine learning task have more input elements than the input sequences corresponding to the second machine learning task. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 9, wherein the input sequences corresponding to the first machine learning task have more input elements than the input sequences corresponding to the second machine learning task. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 11:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein: the input sequence represents an input image, and at least some of the input elements represent respective image patches determined from the input image, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
the input sequence represents an input text, and at least some of the input elements represent respective text tokens determined from the input text, or the input sequence represents audio data, and at least some of the input elements represent respective audio tokens determined from the audio data. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein: the input sequence represents an input image, and at least some of the input elements represent respective image patches determined from the input image, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
the input sequence represents an input text, and at least some of the input elements represent respective text tokens determined from the input text, or the input sequence represents audio data, and at least some of the input elements represent respective audio tokens determined from the audio data. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 12:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein the sequence of network blocks further comprises a self-attention network block that is configured to perform operations comprising: obtaining a second block input sequence having a respective second block input element at only each of
P
second block input positions,
P
>
1
; and (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
applying self-attention to the second block input sequence to generate a second block output sequence having a respective second block output element at only each of
P
second block output positions. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein the sequence of network blocks further comprises a self-attention network block that is configured to perform operations comprising: obtaining a second block input sequence having a respective second block input element at only each of
P
second block input positions,
P
>
1
; and (Insignificant Extra Solution Activity: Receiving or transmitting data over a network is well-understood, routine, conventional activity – see Berkheimer evidence MPEP § 2106.05(d))
applying self-attention to the second block input sequence to generate a second block output sequence having a respective second block output element at only each of
P
second block output positions. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 13:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
for each of a plurality of expert subnetworks of the expert network block: determining a subset of the third block input elements; and (Determining a subnet is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. For example, a person can use any arbitrary method for opining on a subset entirely in their minds or with the assistance of a pen and paper.)
generating a third block output sequence having a respective third block output element at only each of Q third block output positions, comprising: for each third block input element, determining a respective third block output element from the sub-outputs generated from the third block input element. (Determining a subnet is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. For example, a person can use any arbitrary method for opining on a subset entirely in their minds or with the assistance of a pen and paper.)
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein the sequence of network blocks further comprises an expert network block that is configured to perform operations comprising: obtaining a third block input sequence having a respective third block input element at only each of
Q
third block input positions,
Q
>
1
; (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
for each third block input element in the subset, processing the third block input element using the expert subnetwork to generate a respective sub-output; and (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein the sequence of network blocks further comprises an expert network block that is configured to perform operations comprising: obtaining a third block input sequence having a respective third block input element at only each of
Q
third block input positions,
Q
>
1
; (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
for each third block input element in the subset, processing the third block input element using the expert subnetwork to generate a respective sub-output; and (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 14:
Step 1: Independent claim 1 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
Step 2A Prong 2: This judicial exception is not integrated into practical application
The system of claim 1, wherein the neural network further comprises a layer normalization neural network layer preceding the merger network block, wherein the layer normalization neural network layer is configured to perform operations comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
obtaining an initial block input sequence having a respective initial block input element at only each of the
N
block input positions; and (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
applying layer normalization to each initial block input element in the initial block input sequence to generate the block input sequence. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
The system of claim 1, wherein the neural network further comprises a layer normalization neural network layer preceding the merger network block, wherein the layer normalization neural network layer is configured to perform operations comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
obtaining an initial block input sequence having a respective initial block input element at only each of the
N
block input positions; and (Insignificant Extra Solution Activity: Receiving or transmitting data over a network is well-understood, routine, conventional activity – see Berkheimer evidence MPEP § 2106.05(d))
applying layer normalization to each initial block input element in the initial block input sequence to generate the block input sequence. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 16:
Step 1: Independent claim 16 recites a non-transitory computer storage media storing instructions and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1:
the processing comprising: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprising, for each block output position, a respective score corresponding to each block input position, and (Evaluating the product of a weight matrix to input is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components; nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can apply a weight matrix for a perceptron into a single input in their mind or with the help of a pen and paper)
Step 2A Prong 2: This judicial exception is not integrated into practical application
One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform implement a neural network that is configured to process an input sequence comprising a respective input element at each of a plurality of input positions and to generate a network output representing the input sequence, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: obtaining a block input sequence having a respective block input element at only each of
N
block input positions, wherein
N
>
1
; and (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
processing the block input sequence to generate a block output sequence having a respective block output element at only each of
M
block output positions, wherein
N
>
M
>
1
, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
applying the intermediate representation to the block input sequence to generate the block output sequence. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform implement a neural network that is configured to process an input sequence comprising a respective input element at each of a plurality of input positions and to generate a network output representing the input sequence, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: obtaining a block input sequence having a respective block input element at only each of
N
block input positions, wherein
N
>
1
; and (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
processing the block input sequence to generate a block output sequence having a respective block output element at only each of
M
block output positions, wherein
N
>
M
>
1
, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
applying the intermediate representation to the block input sequence to generate the block output sequence. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 17:
Step 1: Independent claim 17 recites a method and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1:
applying a learned weight matrix to the block input sequence to generate an intermediate representation comprising, for each block output position, a respective score corresponding to each block input position, and (Evaluating the product of a weight matrix to input is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components; nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can apply a weight matrix for a perceptron into a single input in their mind or with the help of a pen and paper)
Step 2A Prong 2: This judicial exception is not integrated into practical application
A method performed by one or more computers the method comprising: receiving an input sequence comprising a respective input element at each of a plurality of input positions; and (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
processing the input sequence using a neural network to generate a network output representing the input sequence, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
obtaining a block input sequence having a respective block input element at only each of
N
block input positions, wherein
N
>
1
; and (insignificant extra-solution activity to the judicial exception: Receiving or transmitting data over a network – See MPEP § 2106.05(g))
processing the block input sequence to generate a block output sequence having a respective block output element at only each of
M
block output positions, wherein
N
>
M
>
1
, the processing of the block input sequence comprising: (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
applying the intermediate representation to the block input sequence to generate the block output sequence. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 18:
Step 1: Independent claim 17 recites a method and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 1 above. The same rationale applies to this dependent claim.
The method of claim 17, wherein: the block input sequence is represented as an input matrix
X
∈
R
N
×
D
, wherein each row of the input matrix represents a respective block input element
x
j
∈
R
D
,
j
∈
[
1
,
…
,
N
]
(Mental Process: Processing an input sequence to generate a network output representation is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting a neural network, nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can evaluate 1 and conceptualize a green dot output entirely in their minds or with the assistance of a pen and paper.)
Step 2A Prong 2 and Step 2B: The claim does not include additional elements.
Claim 19:
Step 1: Independent claim 17 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 17 above. The same rationale applies to this dependent claim.
The method of claim 17, wherein: the block output sequence is represented as an output matrix
Y
∈
R
M
×
D
, wherein each row of the output matrix represents a respective block output element
y
j
∈
R
D
,
j
∈
[
1
,
…
,
M
]
(Mental Process: Processing an input sequence to generate a network output representation is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting a neural network, nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can “represent” an input sequence as an input matrix by crossing out an input sequence converting it to a matrix using pen and paper.)
Step 2A Prong 2 and Step 2B: The claim does not include additional elements.
Claim 20:
Step 1: Independent claim 17 recites a system and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 17 above. The same rationale applies to this dependent claim.
The method of claim 17, wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
(
X
W
)
T
(Mental Process: Processing an input sequence to generate a network output representation is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting a neural network, nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can compute a representation by evaluating the expression.)
Step 2A Prong 2: This judicial exception is not integrated into practical application
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim 21:
Step 1: Independent claim 17 recites a computer-implemented method and therefore falls under one of the four statutory categories of patent-eligible subject matter.
Step 2A Prong 1: See the rejection of claim 17 above. The same rationale applies to this dependent claim.
The method of claim 17, wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
W
T
X
T
(Mental Process: Processing an input sequence to generate a network output representation is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting a neural network, nothing in this claim element precludes the step from practically being performed in the mind. For example, a person can compute a representation by evaluating the expression.)
Step 2A Prong 2: This judicial exception is not integrated into practical application
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f))
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-14 and 16-21 are rejected under 35 U.S.C. 103 as being unpatentable over Gehring, et. al. (US 2018/0261214 A1) in view of (Le Grand, et. al. (US 2017/0193368 A1) and further in view of (Lee, et. al. (US 2021/0005183 A1).
Claim 1:
A system comprising a neural network that is configured to process an input sequence comprising a respective input element at each of a plurality of input positions and to generate a network output representing the input sequence, (Gehring, para. 0002: “Artificial neural networks have been applied to sequence-to-sequence tasks, in which an input sequence (of potentially unknown length) is mapped to an output sequence. Examples of sequence-to-sequence tasks include machine translation (translating a sequence of words from a source language into a sequence of words from a destination language),” Examiner note Gehring teaches inputting a sequence of elements i.e. words in a sentence and outputting the sequence into a destination language).
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: (Gehring, para. 0030: “The encoder may represent a network made up of one or more encoder blocks 210; similarly, the decoder may represent a network made up of one or more decoder blocks 218 (the term “layers” and “blocks” are sometimes used interchangeably in the context of CNNs).”)
obtaining a block input sequence having a respective block input element at only each of N block input positions, wherein N > 1; and (Gehring, para. 0033: “For a network with a single block and kernel width k, each resulting state hil contains information of k input elements. Stacking several blocks on top of each other increases the number of input elements that are represented in a state. For instance, stacking 6 blocks with k=5 results in an input field of 25 elements, that is, each output depends on 25 input elements.”)
processing the block input sequence to generate a block output sequence having a respective block output element at only each of M block output positions, wherein N > M > 1, the processing comprising: (Le Grand, para. 0026 and fig. 3: “At decision block 206, the computing system 500 can determine whether the current matrix (the input matrix 110 in this example) is larger than the next matrix to be computed (matrix 130 for the second layer 104 in this example). The determination may be based on one or more size-related characteristics of the respective matrices.” Examiner notes Le Grand teaches an input matrix that is larger than an output matrix where Fig. 3 illustrates more than 1 element).
applying a learned weight matrix (Le Grand, para. 0002: “The NN can repeatedly process the input data, and the parameters (e.g., the weight matrices) of the NN can be modified in what amounts to a trial-and-error process until the model produces (or “converges” on) the correct or preferred output.”) to the block input sequence to generate an intermediate representation comprising, (Lee, para. 0008: “In an aspect of the present disclosure, a method for operating a neural network is provided. The method includes receiving an input sequence at an encoder. The method also includes encoding the input sequence to produce hidden representations. Additionally, the method includes calculating attention weights in attention-heads of the neural network based on the hidden representations.” Examiner notes Lee teaches an intermediate representation in the form of a hidden representation) for each block output position, a respective score corresponding to each block input position, and applying the intermediate representation to the block input sequence to generate the block output sequence. (Gehring, para. 0067: “More specifically, FIG. 4 shows examples of heatmaps representing attention scores of various layers of a convolutional neural network based decoder. The attention scores are applied to determine which element of input sequence 202 is most relevant to be operated on next. FIG. 4A-4E show attention scores for decoder layers 1-5, respectively, for the translation of an English sentence to its German equivalent. As can be seen, as shown by the lighter areas of the heatmaps, some layers produce very sharp attention scores, whereas others are more uniform. The lighter areas represent higher attention scores, and indicate which elements of the input sequence 202 are more likely to be processed in the next round. In some embodiments, this may be achieved by summing the attention scores for a given element across multiple decoder blocks, and selecting the element with the highest cumulative score.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand. Le Grand teaches parallelization of artificial neural network processing by conditionally synchronizing, among multiple computer processors, either the input or output of individual operations, and by conditionally using either rows or columns of certain matrices used in the operations; Gehring teaches improvements to neural networks for translation and other sequence-to-sequence tasks. One of ordinary skill would have been motivated to combine the teachings of Gehring into Le Grand in order to achieve better accuracy, capture long-range dependencies, model the syntax structure of a sentence better, and is faster to run on hardware (Gehring, para. 0015).
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand. Lee teaches a method operating a neural network that includes receiving an input sequence at an encoder. One of ordinary skill would have been motivated to combine the teachings of Lee into Le Grand in order to reduce over-fitting and enable larger models to achieve better generalization (Lee, para. 0049).
Claim 2:
The system of claim 1, wherein: the block input sequence is represented as an input matrix
X
∈
R
N
×
D
, wherein each row of the input matrix represents a respective block input element
x
j
∈
R
D
,
j
∈
[
1
,
…
,
N
]
(Gehring, para. 0067, fig. 4E: “FIG. 4A-4E show attention scores for decoder layers 1-5, respectively, for the translation of an English sentence to its German equivalent.” Examiner notes Fig 4E shows English sentence input where each element i.e. word is associated with a row of the matrix).
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand as set forth above with respect to claim 1.
Claim 3:
The system of claim 1, wherein: the block output sequence is represented as an output matrix
Y
∈
R
M
×
D
, wherein each row of the output matrix represents a respective block output element
y
j
∈
R
D
,
i
∈
[
1
,
…
,
M
]
(Lee, para. 0058: “Multiplication units 418a-i receive the output of encoder 402 (e.g., hidden representation h[t]) and the attention weights αi[t], and performs a multiplication operation to compute a context vector ci. Accordingly, each attention-head 404a-i may output a context vector ci, which may be calculated by the weighted sum as follows: [equation omitted]”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand, as modified, as set forth above with respect to claim 1.
Claim 4:
The system of claim 1, wherein: applying a learned weight matrix (Le Grand, para. 0002: “The NN can repeatedly process the input data, and the parameters (e.g., the weight matrices) of the NN can be modified in what amounts to a trial-and-error process until the model produces (or “converges” on) the correct or preferred output.”) to the block input sequence to generate an intermediate representation comprises computing:
S
=
(
X
W
)
T
(Lee, para. 0063: “As such, the keyword spotting network 400 may directly compute the context vector ci from the encoder output (hidden representation h[t]) by multiplying the attention weights αi[t].”)
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Le Grand, para. 0029 and fig. 3: “FIG. 3 illustrates an example of multiplying input matrix 110 by weight matrix 120 to generate matrix 130 according to the conditional parallel processing described above.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand, as modified, as set forth above with respect to claim 1.
Claim 5:
The system of any one of claim 1 wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
W
T
X
T
(Le Grand, para. 0030: “Processor 502 multiplies column subset 302 from input matrix 110 by row subset 312 from weight matrix 120 to generate intermediate matrix 322.”)
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Le Grand, para. 0029 and fig. 3: “FIG. 3 illustrates an example of multiplying input matrix 110 by weight matrix 120 to generate matrix 130 according to the conditional parallel processing described above.”)
Claim 6:
The system of claim 1, wherein: applying the intermediate representation to the block input sequence to generate the block output sequence comprises computing:
Y
=
s
o
f
t
m
a
x
(
S
)
∙
X
(Gehring, para. 0039-0040: “Linear mappings are added to project between the embedding size f and the hidden size 2d. A transform is applied to w when feeding it to the encoder network, to the encoder output zju 212, to the final layer of the decoder just before the softmax hL, and to all decoder layers hl before computing attention scores as discussed below” Examiner notes Gehring teaches softmax to decode)
wherein
Y
∈
R
M
×
D
represents the block output sequence,
X
∈
R
N
×
D
represents the block input sequence, and
S
∈
R
M
×
N
represents the intermediate representation. (Le Grand, para. 0029 and fig. 3: “FIG. 3 illustrates an example of multiplying input matrix 110 by weight matrix 120 to generate matrix 130 according to the conditional parallel processing described above.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand as set forth above with respect to claim 1.
Claim 7:
The system of claim 1, wherein the sequence of network blocks comprises a plurality of merger network blocks, (Le Grand, para. 0006: “FIG. 3 is a block diagram illustrating parallel processing of a matrix multiplication operation in which an output matrix is smaller than an input matrix.” Examiner note Le Grand teaches processing a plurality of matrices with an output smaller than an input matrix)
wherein each particular merger network block is configured to generate a block output sequence having a shorter length than the respective block output sequences of any other merger network blocks that precede the particular merger network block in the sequence of network blocks. (Le Grand, para. 0006: “FIG. 3 is a block diagram illustrating parallel processing of a matrix multiplication operation in which an output matrix is smaller than an input matrix.” Examiner note Le Grand teaches processing a plurality of matrices with an output smaller than the preceding input matrix. Examiner notes a “network block” is interpreted as a layer of a neural network, as set forth in the specification).
Claim 8:
The system of claim 1: wherein the merger network block is configured to obtain a block input sequence comprising a variable number
N
of block input elements and to generate a block output sequence having a predetermined number
M
of block output elements. (Le Grand para. 0016 and 0018: “ FIG. 1 illustrates an example NN 100 in which conditional parallel processing may be implemented. As shown, the NN 100 has a first layer 102 with a plurality of nodes, a second layer 104 with a plurality of nodes, and a third layer 106 with a plurality of nodes.” “As shown, NN 100 is a fully-connected NN.”)
Claim 9:
The system of claim 1, wherein the neural network is configured to perform a first machine learning task, and wherein the neural network has been pre-trained to perform a second machine learning task that is different than the first machine learning task. (Le Grand, para. 0010: “Such efficiently parallelized artificial neural networks may be used in a variety of machine learning applications and other systems, including but not limited to: product recommendation generation, automatic speech recognition, facial recognition, handwriting recognition, and image recognition.” Examiner notes Le Grand teaches parallel neural networks which teaches at least a first and second task running parallel)
Claim 10:
The system of claim 9, wherein the input sequences corresponding to the first machine learning task have more input elements than the input sequences corresponding to the second machine learning task. (Le Grand, para. 0029: “FIG. 3 illustrates an example of multiplying input matrix 110 by weight matrix 120 to generate matrix 130 according to the conditional parallel processing described above. As shown, input matrix 110 includes five rows of input vectors, and four columns of data elements corresponding to the four nodes of the first layer 102 of the NN 100. Weight matrix 120 includes two rows of weights corresponding to the two nodes of the second layer 104, and four columns of weights corresponding to the four columns of the input matrix 110. Matrix 130 includes five rows for data computed from the five rows of input vectors, and two columns of data elements corresponding to the two nodes of the second layer 104”)
Claim 11:
The system of claim 1, wherein: the input sequence represents an input image, and at least some of the input elements represent respective image patches determined from the input image, (Gehring, para. 0026: “Other types of logic may be provided for other types of inputs 104 (e.g., optical character recognition logic for converting input handwriting or typing, image analysis logic for converting input photographs, etc.).” Examiner note Gehring teaches image input in the form of characters or handwriting where the characters are an input element representing patches of the entire handwritten input).
the input sequence represents an input text, and at least some of the input elements represent respective text tokens determined from the input text, or (Gehring, para. 0023 and para 0059: “In the case of a translation task, the input 104 may be in the form of text in a source language, such as text input from a keyboard via a web browser or application” Examiner notes para. 0059 teaches tokens of data for mini-batch processing)
the input sequence represents audio data, and at least some of the input elements represent respective audio tokens determined from the audio data. (Gehring, para. 0025 and para. 0059: “Accordingly, to handle multiple different types of inputs 104, logic may be provided for converting the input 104 into text. For example, FIG. 1 depicts automatic speech recognition (ASR) logic 106 that is configured to convert input audio in the source language into text in the source language” Examiner notes para. 0059 teaches tokens of data for mini-batch processing)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand as set forth above with respect to claim 1.
Claim 12:
The system of claim 1, wherein the sequence of network blocks further comprises a self-attention network block that is configured to perform operations comprising: obtaining a second block input sequence having a respective second block input element at only each of
P
second block input positions,
P
>
1
; and (Le Grand, para. 0029: “As shown, input matrix 110 includes five rows of input vectors, and four columns of data elements corresponding to the four nodes of the first layer 102 of the NN 100. Weight matrix 120 includes two rows of weights corresponding to the two nodes of the second layer 104, and four columns of weights corresponding to the four columns of the input matrix 110.” Examiner notes Le Grand teaches four nodes in a first layer followed by two nodes in a second layer i.e. a second block having input positions only at two particular nodes).
applying self-attention to the second block input sequence to generate a second block output sequence having a respective second block output element at only each of
P
second block output positions. (Lee, para. 0007 and fig. 1: “Multi-head attention based systems are proposed to alleviate this issue. For example, the proposed multi-head attention based systems replace convolutional neural networks (CNNs) and recurrent neural networks (RNNs) by self-attention networks for speech tasks.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand, as modified, as set forth above with respect to claim 1.
Claim 13:
The system of claim 1, wherein the sequence of network blocks further comprises an expert network block that is configured to perform operations comprising: obtaining a third block input sequence having a respective third block input element at only each of
Q
third block input positions,
Q
>
1
; (Le Grand, para. 0037: “Weight matrix 140 includes four rows of weights corresponding to the four nodes of the third layer 106, and two columns of weights corresponding to the two nodes of the second layer 104. Matrix 150 includes five rows for data computed from the five rows of input vectors, and four columns of data elements corresponding to the four nodes of the third layer 106.” Examiner notes Le Grand teaches four nodes in a first layer followed by two nodes in a second layer followed by four nodes in a third layer i.e. a third block having input positions at four nodes).
for each of a plurality of expert subnetworks of the expert network block: determining a subset of the third block input elements; and for each third block input element in the subset, processing the third block input element using the expert subnetwork to generate a respective sub-output; and (Le Grand, para. 0037: “ Weight matrix 140 includes four rows of weights corresponding to the four nodes of the third layer 106, and two columns of weights corresponding to the two nodes of the second layer 104. Matrix 150 includes five rows for data computed from the five rows of input vectors, and four columns of data elements corresponding to the four nodes of the third layer 106.” Examiner notes Le Grand teaches processing matrix 150 using the third layer subnetwork to receive an output)
generating a third block output sequence having a respective third block output element at only each of Q third block output positions, comprising: for each third block input element, determining a respective third block output element from the sub-outputs generated from the third block input element. (Le Grand, para. 0039: “The output matrix of the NN can then be accessed by a downstream process. For example, a function may be performed on the columns of the output matrix stored on each individual processor to generate a scalar output value. The scalar output value may then be used by a downstream process, thereby avoiding the need to perform a final “reduction” or “gather” operation to gather the individual columns from the corresponding processors into a single matrix that is transmitted or stored.” Examiner note Le Grand para. 0037 above describes a third block)
Claim 14:
The system of claim 1, wherein the neural network further comprises a layer normalization neural network layer preceding the merger network block, wherein the layer normalization neural network layer is configured to perform operations comprising: (Gehring, para. 0053: “Normalizing activations in the network when adding the output of different layers, e.g. residual connections, requires a careful weight initialization. For initialization, the goal is to maintain variances of activations throughout the forward and backward passes”)
obtaining an initial block input sequence having a respective initial block input element at only each of the
N
block input positions; and (Gehring, para. 0033: “For a network with a single block and kernel width k, each resulting state hil contains information of k input elements. Stacking several blocks on top of each other increases the number of input elements that are represented in a state. For instance, stacking 6 blocks with k=5 results in an input field of 25 elements, that is, each output depends on 25 input elements.”)
applying layer normalization to each initial block input element in the initial block input sequence to generate the block input sequence. (Gehring, para. 0053: “Normalizing activations in the network when adding the output of different layers, e.g. residual connections, requires a careful weight initialization. For initialization, the goal is to maintain variances of activations throughout the forward and backward passes…This ensures that the variance of a normally distributed input is retained.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand as set forth above with respect to claim 1.
Claim 16:
One or more non-transitory computer storage media storing instructions that when executed (Le Grand, para. 0043: “In addition to the components shown in FIG. 5 and otherwise described herein, a computing system 500 may include various other components, such as one or more network interfaces (e.g., network interface cards), one or more computer readable medium drives (e.g., high density disks, solid state drives, flash drives, and/or other persistent non-transitory computer-readable media), an input/output device interface (e.g. an IO interface in communication with one or more microphones or display screens), and one or more computer readable memories (e.g., random access memory and/or other volatile non-transitory computer-readable media).”) by one or more computers cause the one more computers to perform implement a neural network that is configured to process an input sequence comprising a respective input element at each of a plurality of input positions and to generate a network output representing the input sequence, (Gehring, para. 0002: “Artificial neural networks have been applied to sequence-to-sequence tasks, in which an input sequence (of potentially unknown length) is mapped to an output sequence. Examples of sequence-to-sequence tasks include machine translation (translating a sequence of words from a source language into a sequence of words from a destination language),” Examiner note Gehring teaches inputting a sequence of elements i.e. words in a sentence and outputting the sequence into a destination language).
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: (Gehring, para. 0030: “The encoder may represent a network made up of one or more encoder blocks 210; similarly, the decoder may represent a network made up of one or more decoder blocks 218 (the term “layers” and “blocks” are sometimes used interchangeably in the context of CNNs).”)
obtaining a block input sequence having a respective block input element at only each of
N
block input positions, wherein
N
>
1
; and (Gehring, para. 0033: “For a network with a single block and kernel width k, each resulting state hil contains information of k input elements. Stacking several blocks on top of each other increases the number of input elements that are represented in a state. For instance, stacking 6 blocks with k=5 results in an input field of 25 elements, that is, each output depends on 25 input elements.”)
processing the block input sequence to generate a block output sequence having a respective block output element at only each of
M
block output positions, wherein
N
>
M
>
1
, the processing comprising: (Le Grand, para. 0026 and fig. 3: “At decision block 206, the computing system 500 can determine whether the current matrix (the input matrix 110 in this example) is larger than the next matrix to be computed (matrix 130 for the second layer 104 in this example). The determination may be based on one or more size-related characteristics of the respective matrices.” Examiner notes Le Grand teaches an input matrix that is larger than an output matrix where Fig. 3 illustrates more than 1 element).
applying a learned weight matrix (Le Grand, para. 0002: “The NN can repeatedly process the input data, and the parameters (e.g., the weight matrices) of the NN can be modified in what amounts to a trial-and-error process until the model produces (or “converges” on) the correct or preferred output.”) to the block input sequence to generate an intermediate representation comprising, (Lee, para. 0008: “In an aspect of the present disclosure, a method for operating a neural network is provided. The method includes receiving an input sequence at an encoder. The method also includes encoding the input sequence to produce hidden representations. Additionally, the method includes calculating attention weights in attention-heads of the neural network based on the hidden representations.” Examiner notes Lee teaches an intermediate representation in the form of a hidden representation) for each block output position, a respective score corresponding to each block input position, and applying the intermediate representation to the block input sequence to generate the block output sequence. (Gehring, para. 0067: “More specifically, FIG. 4 shows examples of heatmaps representing attention scores of various layers of a convolutional neural network based decoder. The attention scores are applied to determine which element of input sequence 202 is most relevant to be operated on next. FIG. 4A-4E show attention scores for decoder layers 1-5, respectively, for the translation of an English sentence to its German equivalent. As can be seen, as shown by the lighter areas of the heatmaps, some layers produce very sharp attention scores, whereas others are more uniform. The lighter areas represent higher attention scores, and indicate which elements of the input sequence 202 are more likely to be processed in the next round. In some embodiments, this may be achieved by summing the attention scores for a given element across multiple decoder blocks, and selecting the element with the highest cumulative score.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand as set forth above with respect to claim 1.
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand, as modified, as set forth above with respect to claim 1.
Claim 17:
A method performed by one or more computers the method comprising: receiving an input sequence comprising a respective input element at each of a plurality of input positions; and processing the input sequence using a neural network to generate a network output representing the input sequence, (Gehring, para. 0002: “Artificial neural networks have been applied to sequence-to-sequence tasks, in which an input sequence (of potentially unknown length) is mapped to an output sequence. Examples of sequence-to-sequence tasks include machine translation (translating a sequence of words from a source language into a sequence of words from a destination language),” Examiner note Gehring teaches inputting a sequence of elements i.e. words in a sentence and outputting the sequence into a destination language).
the neural network comprising a sequence of one or more network blocks, the sequence comprising a merger network block configured to perform operations comprising: (Gehring, para. 0030: “The encoder may represent a network made up of one or more encoder blocks 210; similarly, the decoder may represent a network made up of one or more decoder blocks 218 (the term “layers” and “blocks” are sometimes used interchangeably in the context of CNNs).”)
obtaining a block input sequence having a respective block input element at only each of
N
block input positions, wherein
N
>
1
; and (Gehring, para. 0033: “For a network with a single block and kernel width k, each resulting state hil contains information of k input elements. Stacking several blocks on top of each other increases the number of input elements that are represented in a state. For instance, stacking 6 blocks with k=5 results in an input field of 25 elements, that is, each output depends on 25 input elements.”)
processing the block input sequence to generate a block output sequence having a respective block output element at only each of
M
block output positions, wherein
N
>
M
>
1
, the processing of the block input sequence comprising: (Le Grand, para. 0026 and fig. 3: “At decision block 206, the computing system 500 can determine whether the current matrix (the input matrix 110 in this example) is larger than the next matrix to be computed (matrix 130 for the second layer 104 in this example). The determination may be based on one or more size-related characteristics of the respective matrices.” Examiner notes Le Grand teaches an input matrix that is larger than an output matrix where Fig. 3 illustrates more than 1 element).
applying a learned weight matrix (Le Grand, para. 0002: “The NN can repeatedly process the input data, and the parameters (e.g., the weight matrices) of the NN can be modified in what amounts to a trial-and-error process until the model produces (or “converges” on) the correct or preferred output.”) to the block input sequence to generate an intermediate representation comprising, (Lee, para. 0008: “In an aspect of the present disclosure, a method for operating a neural network is provided. The method includes receiving an input sequence at an encoder. The method also includes encoding the input sequence to produce hidden representations. Additionally, the method includes calculating attention weights in attention-heads of the neural network based on the hidden representations.” Examiner notes Lee teaches an intermediate representation in the form of a hidden representation) for each block output position, a respective score corresponding to each block input position, and applying the intermediate representation to the block input sequence to generate the block output sequence. (Gehring, para. 0067: “More specifically, FIG. 4 shows examples of heatmaps representing attention scores of various layers of a convolutional neural network based decoder. The attention scores are applied to determine which element of input sequence 202 is most relevant to be operated on next. FIG. 4A-4E show attention scores for decoder layers 1-5, respectively, for the translation of an English sentence to its German equivalent. As can be seen, as shown by the lighter areas of the heatmaps, some layers produce very sharp attention scores, whereas others are more uniform. The lighter areas represent higher attention scores, and indicate which elements of the input sequence 202 are more likely to be processed in the next round. In some embodiments, this may be achieved by summing the attention scores for a given element across multiple decoder blocks, and selecting the element with the highest cumulative score.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand as set forth above with respect to claim 1.
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand, as modified, as set forth above with respect to claim 1.
Claim 18:
The method of claim 17, wherein: the block input sequence is represented as an input matrix
X
∈
R
N
×
D
, wherein each row of the input matrix represents a respective block input element
x
j
∈
R
D
,
j
∈
[
1
,
…
,
N
]
(Gehring, para. 0067, fig. 4E: “FIG. 4A-4E show attention scores for decoder layers 1-5, respectively, for the translation of an English sentence to its German equivalent.” Examiner notes Fig 4E shows English sentence input where each element i.e. word is associated with a row of the matrix).
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gehring into Le Grand as set forth above with respect to claim 1.
Claim 19:
The method of claim 17, wherein: the block output sequence is represented as an output matrix
Y
∈
R
M
×
D
, wherein each row of the output matrix represents a respective block output element
y
j
∈
R
D
,
j
∈
[
1
,
…
,
M
]
(Lee, para. 0058: “Multiplication units 418a-i receive the output of encoder 402 (e.g., hidden representation h[t]) and the attention weights αi[t], and performs a multiplication operation to compute a context vector ci. Accordingly, each attention-head 404a-i may output a context vector ci, which may be calculated by the weighted sum as follows: [equation omitted]”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand, as modified, as set forth above with respect to claim 1.
Claim 20:
The method of claim 17, wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing:
S
=
(
X
W
)
T
(Lee, para. 0063: “As such, the keyword spotting network 400 may directly compute the context vector ci from the encoder output (hidden representation h[t]) by multiplying the attention weights αi[t].”)
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Le Grand, para. 0029 and fig. 3: “FIG. 3 illustrates an example of multiplying input matrix 110 by weight matrix 120 to generate matrix 130 according to the conditional parallel processing described above.”)
It would have obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Lee into Le Grand, as modified, as set forth above with respect to claim 1.
Claim 21:
The method of claim 17, wherein: applying a learned weight matrix to the block input sequence to generate an intermediate representation comprises computing: Filed Herewith
S
=
W
T
X
T
(Le Grand, para. 0030: “Processor 502 multiplies column subset 302 from input matrix 110 by row subset 312 from weight matrix 120 to generate intermediate matrix 322.”)
wherein
X
∈
R
N
×
D
represents the block input sequence,
W
∈
R
D
×
M
represents the learned weight matrix,
S
∈
R
M
×
N
represents the intermediate representation, and each element
s
i
,
j
of
S
represents the score corresponding to the
i
t
h
block output position and the
j
t
h
block input position. (Le Grand, para. 0029 and fig. 3: “FIG. 3 illustrates an example of multiplying input matrix 110 by weight matrix 120 to generate matrix 130 according to the conditional parallel processing described above.”)
Search Notes
best results from IP.com using application number as main concept with priority date filters and additional keywords including self-attention and weight matrix
Prior Art
Hargil et. al. (US 2020/0356836 A1) teaches classifying information using a fully-connected layer of a convolutional neural network. A method for classifying information using a fully-connected layer of a convolutional neural network includes calculating a first partial output for a first block of elements by performing a dot product operation using a first row of elements of the first block of elements and a first weight block, where the first row of elements of the first block of elements corresponds to a first batch of elements.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sally T. Ley whose telephone number is (571)272-3406. The examiner can normally be reached Monday - Thursday, 10:00am - 6:00pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/STL/Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147