DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on May 25, 2026. Claims 1-7, 10-16, and 19-20 are pending and have been examined. Hence, this action has been made FINAL.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Response to Arguments
The reply filed on May 25, 2026 has been entered. Applicant’s arguments with respect to claim 1 have been considered but are not persuasive.
With respect to the applicant’s arguments to claim objections, Applicant has amended each of the independent claims and asserts that “informalities of these claims are corrected as suggested. Applicant respectfully requests the withdrawal of the objection.” The examiner agrees with these assertions, and has thus withdrawn the claim objections.
With respect to the applicant’s arguments to claim rejections under 35 U.S.C. § 112, Applicant has amended each of the independent claims and requests that “In view of the above amendments, Applicant respectfully requests the withdrawal of the rejection.” The examiner agrees that the claim amendments properly addressed the indefiniteness of the claims, and has thus withdrawn each of the claim rejections under 35 U.S.C. § 112.
With respect to the applicant’s arguments to claim rejections under 35 U.S.C § 101, Applicant has amended each of the independent claims and asserts that “Under the Alice/Mayo framework, all steps of amended claim 1, … are performed by a processor and cannot be performed in the human mind. Moreover, specialized computer tools (e.g., word2vec, reinforcement learning) are required, and these tools do not fall within the scope of human mental activities (Step 2A prong 1).” The examiner respectfully disagrees with these assertions. The additional elements of a “processor”, “word2vec”, and “reinforcement learning” are recited at a high level of generality (¶ [00144], [0072], and [0073]-[0074], respectively) and merely equate to “apply it” or otherwise merely uses a generic computer as a tool to perform an abstract which are not indicative of integration into a practical application as per MPEP 2106.05(f), as described further below with respect to claim rejections under 35 USC 101.
Applicant further asserts that “Manually designing prompt templates or tag words for different tasks is a laborious matter. The method as recited claim 1 reduces the candidate tag word search space …, automatically constructs effective prompt templates, simultaneously reduces the difference between different prompt templates, and improves the accuracy of downstream tasks. Thus, even if claim 1 involves an abstract idea, it improves the existing prompt-based fine-tuning method and integrates the abstract idea into a practical application (Step 2A prong 2). Therefore, claim 1 of this application integrates the recited exception into a practical application of that exception, rendering it patent eligible.” The examiner respectfully disagrees with these assertions. As per MPEP 2106.05(a), “it is important to keep in mind that an improvement in the abstract idea itself (e.g. a recited fundamental economic concept) is not an improvement in technology. For example, in Trading Technologies Int’l v. IBG, 921 F.3d 1084, 1093-94, 2019 USPQ2d 138290 (Fed. Cir. 2019), the court determined that the claimed user interface simply provided a trader with more information to facilitate market trades, which improved the business process of market trading but did not improve computers or technology. ... To show that the involvement of a computer assists in improving the technology, the claims must recite the details regarding how a computer aids the method, the extent to which the computer aids the method, or the significance of a computer to the performance of the method. Merely adding generic computer components to perform the method is not sufficient. Thus, the claim must include more than mere instructions to perform the method on a generic component or machinery to qualify as an improvement to an existing technology.” Accordingly, constructing a candidate tag word search space, constructing effective prompt templates, computing the difference between different prompt templates, and performing downstream tasks are all considered abstract ideas, since a human could perform each of the listed tasks. For example, a human could brainstorm a candidate tag word search space, create prompt templates, calculate the cosine difference between different prompt templates, and perform downstream tasks such as text classification. Mere improvements in these abstract ideas cannot be equated to improvements in computer technology itself. Further, the recitation of computer components to perform these tasks cannot constitute an improvement to the computer technology itself without explicitly explaining how such computer components are improved.
Applicant further asserts that “The use of the small-sample fine-tuning method for a pretrained language model reduces memory requirements and system complexity, especially preventing small-sample overfitting. Moreover, the present application adopts a reinforcement learning process to search for optimal tag words and templates, thereby solving the problem that general algorithms easily fall into a local optimum.” The examiner respectfully disagrees with these assertions. The mere assertion that the claimed invention reduces memory requirements and system complexity is considered a conclusory statement; that is, the applicant asserts improvements to computer technology without further explaining how such improvements are realized. The recitation of “preventing small-sample overfitting” and “solving the problem that general algorithms easily fall into a local optimum” are stated in a conclusory manner, with no apparent justification for how these improvements are actually realized. See MPEP 2106.04(d)(1), “if the specification explicitly sets forth an improvement but in a conclusory manner (i.e., a bare assertion of an improvement without the detail necessary to be apparent to a person of ordinary skill in the art), the examiner should not determine the claim improves technology.” As amended, there is no language in the independent claims that would prevent a human from performing these steps, as addressed in further detail below with respect to claim rejections under 35 USC § 101.
With respect to the applicant’s arguments to claim rejections under 35 U.S.C § 103, Applicant has amended each of the independent claims and asserts that “Gao does not disclose or teach determining a near-synonym set for each category in the training set via a cosine similarity, let alone using the intersection of the near-synonym set and the conditional probability set as candidate tag words.” The examiner respectfully disagrees with these assertions. Pg. 3820, Section 5.1, Paragraph 1 of Gao et al. states, “for each class
c
∈
Υ
, we construct a pruned set
V
c
⊂
V
of the top
k
vocabulary words based on their conditional likelihood … To further narrow down the search space, we find the top
n
assignments over the pruned space that maximize zero-shot accuracy on
D
t
r
a
i
n
(both
n
and
k
are hyper-parameters, see Appendix C.2). Then we fine-tune all top
n
assignments, and re-rank to find the best one using
D
d
e
v
.” Pg. 3828, Appendix C.2, Paragraph 1 further states, “we observe that filtering
V
c
using conditional likelihood alone is still noisy, thus we set k = 1000, and then re-rank
V
c
by the nearest neighbors of the original manual label words and take the top 30 per class. We set
n
to 100 in all experiments.” Accordingly, re-ranking a vocabulary set
V
c
by the nearest neighbors of the original manual label words and taking the top 30 per class (i.e. category) is considered analogous to determining a near-synonym set for each category in a training set via a cosine similarity. Further, the step of “[fine-tuning] all top
n
assignments” involves the cosine intersection
c
o
s
(
e
(
x
i
n
)
;
e
(
x
)
)
as described in Section 6.2 of Gao et al. Thus, the subsequent step of “[fine-tuning] all top
n
assignments, and [re-ranking] to find the best one using
D
d
e
v
” in Section 5.1 of Gao et al. is considered analogous to using the intersection of a near-synonym set and the conditional probability set to determine a candidate tag word under each category, (e.g. “the best [assignment],” found by re-ranking using
D
d
e
v
).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-7, 10-16, and 19-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. All of the claims are method claims (1-7, 10-16), apparatus/machine claims (19-20) or manufacture claim under (Step 1), but under Step 2A all of these claims recite abstract ideas and specifically mental processes. These mental processes are more particularly recited in claims 1, 19, and 20 as:
forming an input sample according to a fixed template…
constructing, by a processor, a candidate tag word set and a candidate prompt template set…
searching, by a processor, an optimal tag word corresponding to the input sample from the candidate tag word set, and a prompt template…
outputting, by a processor, a mapping relationship of the optimal tag word and an optimal prompt template format corresponding to the prompt template…
automatically selecting, by the processor, an optimal candidate tag word…
automatically selecting, by the processor, a candidate prompt template…
initializing a vocabulary by the processor…
vectorizing, by the processor, each word in the vocabulary using a word2vec method…
determining a near-synonym set for each category in the training set…
selecting, by the processor, a word in the vocabulary that maximizes a conditional probability…
determining, by the processor, a candidate tag word under each category…
integrating, by the processor, candidate tag words under various categories…
determining an assignment mode which maximizes an accuracy rate of the training set as the optimal candidate tag word…
Under Step 2A Prong One, claims 1, 19, and 20 are directed to an abstract idea and specifically a mental process. As detailed above, the steps of forming, constructing, searching, outputting, selecting, vectorizing, determining, integrating, etc. may be practically performed in the human mind with the use of a physical aid such as a pen and paper. For example, a human researcher could receive a training dataset with sentence/label tuples, format each input using a fixed template, randomly sample the formatted tuples to obtain a set of candidate labels, create a set of candidate templates for each input by formatting the input sentence according to various placeholder templates, and task a second human to find a mapping that indicates an optimal label and optimal template from the sets of candidate labels and candidate templates respectively based on their personal experience. The human researcher, in order to obtain a set of candidate labels, could receive a vocabulary of label words, convert each word into a vector, calculate the cosine similarity between the vocabulary and a set of original manual label words, select the top 30 most similar vocabulary words to the set of original manual label words as a near-synonym set, calculate the conditional probability for each word in the vocabulary of label words, take the top k words with the most conditional probability, select the candidate label with the highest similarity between the conditional probability set and the near-synonym set for each label, and then utilize the set of chosen label words for downstream classification tasks due to them maximizing the accuracy rate of a training set.
Under Step 2A Prong Two, this judicial exception is not integrated into a practical application because claims 2-16, 21, and 22 do not recite additional elements that integrate the exception into a practical application. In particular, claims 1, 19, and 20 recite the additional elements of a non-volatile readable storage medium (¶ [00143]), a processor (¶ [00144]), word2vec (¶ [0072]), and a pre-trained language model (¶ [00103]). These additional elements are recited at a high level of generality and merely equate to “apply it” or otherwise merely uses a generic computer as a tool to perform an abstract which are not indicative of integration into a practical application as per MPEP 2106.05(f). Further, claims 1, 19, and 20 recite the additional element of “inputting…” which amounts to insignificant extra-solution activities which are not indicative of integration into a practical application as per MPEP 2106.05(g). Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Under Step 2B, the claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, s discussed above with respect to the integration of the abstract idea into a practical application, the additional elements of using a computer is noted as a general computer {non-volatile readable storage medium (¶ [00143]); processor (¶ [00144])}. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Further, the additional limitations in the claim noted above are directed towards insignificant extra-solution activities. The claim is not patent eligible.
With respect to claims 2, the claim relates to splitting a dataset into validation, training, and test datasets. This relates to a human researcher splitting a dataset manually by randomly sampling a fixed number of tuples from a larger set to form each individual set. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 3, the claim relates to data in a dataset having ID, sentence, and label attributes. This relates to a human researcher receiving a dataset including a dataset name, sentence type information, and label information. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 4, the claim relates to forming the input sample by measuring cosine distance and random sampling. This relates to a human researcher picking tuples for the training or validation set by defining a test set comprising a select number of tuples, measuring a cosine similarity for each remaining tuple in the dataset to the tuples included in the testing set, and then random sampling the remaining tuples in order to form the training or validation sets. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 5 and 7, the claim relates to initializing and converting prompt templates. This relates to a human researcher formatting input data using placeholder templates. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 6, the claim relates to encoding and calculating cosine similarities. This relates to a human researcher following the step-wise procedure of Sentence BERT to create sentence embeddings for each input tuple, and then using the sentence embeddings to calculate a cosine similarity between all sentences in a dataset. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 10, the claim relates to using an equation to determine conditional probability. This limitation is directed towards an abstract idea and more specifically, a mathematical concept which cannot constitute patentable subject matter as per MPEP 2106.04(a)(2) Section I. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 11, the claim relates to further defining the process of automatically selecting candidate templates. This relates to a human researcher using the above mental process of claim 9 to determine an optimal candidate tag word, converting training sentences into input sequences using various placeholder templates, and then performing beam search to obtain a candidate prompt template among the various input sequences. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 12 and 13, the claim relates to further defining the process of searching. This relates to a human researcher randomly sampling three candidate label for each category and then concatenating the candidate label set with a template set in order to obtain a search space list with a masked token. The human researcher could then hand this search list to the second human, who could then rely on their experience with language in order to replace the masked token with the most appropriate label. The second human could also indicate which template of the template set they were most confident about predicting. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 14 and 15, the claim relates to using reinforcement learning to perform learning. This relates to a second human receiving a search space list from the human researcher and predicting a label for a given input. The human researcher could then calculate a loss value based measuring the distance between the predicted label and the test label. The human researcher could then inform the second human of this loss value, and the second human can re-align their mental processes to better account for that loss. The second human’s re-aligned mental processes may select a different or better selection direction than the previous prediction process based on the loss. The additional element of a “language model" is recited at a high level of generality (¶ [0096]) and merely equate to “apply it” or otherwise merely uses a generic computer as a tool to perform an abstract which are not indicative of integration into a practical application as per MPEP 2106.05(f). No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
With respect to claims 16, the claim relates to averaging and normalizing conditional probabilities in order to calculate a correction matrix. This limitation is directed towards an abstract idea and more specifically, a mathematical concept which cannot constitute patentable subject matter as per MPEP 2106.04(a)(2) Section I. No additional limitations are present. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Claim 19 is drawn to a “signal” per se as recited in the preamble and as such is non-statutory subject matter. On paragraph [00143] of the as filed Specification, the term “non-volatile readable storage medium" is not defined as to what the scope of the term is meant to encompass. Hence, one of ordinary skilled in the art can interpret such term to include transitory signals and non-transitory signals. It does not appear that a claim reciting a signal encoded with functional descriptive material falls within any of the categories of patentable subject matter set forth in § 101. First, a claimed signal is clearly not a "process" under § 101 because it is not a series of steps. The other three § 101 classes of machine, compositions of matter and manufactures "relate to structural entities and can be grouped as 'product' claims in order to contrast them with process claims." 1 D. Chisum, Patents § 1.02 (1994).
The Applicant's Specification presents a broad definition as to what the “non-volatile readable storage medium” covers and is being made to include transitory and non-transitory signals. The Applicant's as filed Specification in paragraph [00143], refers to the “storage medium”. Hence, it appears that the claims appear to be drawn towards transitory signals, which is not subject matter eligible. In order to overcome the present rejection, the Applicant is advised to amend the claims by using the following terminology: "non-transitory machine readable storage medium." Such example terminology has been also found in the Official Gazette 1351 OG 212.
For all of the above reasons, taken alone or in combination, claims 1-7, 10-16, and 19-20 recite a non-statutory mental process.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-7, 10-15 and 19-20 are rejected under 35 U.S.C. 103 as obvious over "Making Pre-trained Language Models Better Few-shot Learners" (Gao et al.) in view of US Patent Publication 20190370219 A1 (Efstathiou et al.) in view of US Patent Publication 20200409948 A1 (Corvinelli et al.).
Claim 1
Regarding claim 1, Gao et al. disclose a small sample fine-tuning method for a pre-trained language model, comprising:
inputting a data set (Gao et al. pg. 3818, Section 3, Paragraph 2, "We conduct a systematic study across 8 single-sentence and 7 sentence-pair English tasks, including 8 tasks from the GLUE benchmark (Wang et al., 2019), SNLI (Bowman et al., 2015), and 6 other popular sentence classification tasks (SST-5, MR, CR, MPQA, Subj, TREC). All of the dataset details are provided in Appendix B.") to the pre-trained language model (Gao et al. pg. 3816, Section 1, Paragraph 2, "In this work, we study a more practical scenario in which we only assume access to a moderately-sized language model such as BERT (Devlin et al., 2019) or RoBERTa (Liu et al., 2019), and a small number of examples (i.e., a few-shot setting), which we can use to fine-tune the weights of the language model." BERT or RoBERTa is considered analogous to a pre-trained language model), and forming an input sample according to a fixed template (Gao et al. pg. 3819, Section 4, Paragraph 2, "we can formulate a binary sentiment classification task using a prompt with input
x
1
(e.g., “No reason to watch it .”) as:
x
p
r
o
m
p
t
=
C
L
S
x
1
I
t
w
a
s
M
A
S
K
.
[
S
E
P
]
"), wherein the data set comprises a training set, a validation set, and a test set (Gao et al. pg. 3818, Section 3, Paragraph 1, "we only assume
K
training examples per class for the task’s training set
D
t
r
a
i
n
, such that the total number of examples is
K
t
o
t
=
K
×
|
Y
|
, and
D
t
r
a
i
n
=
x
i
n
i
,
y
i
i
=
1
K
t
o
t
. Our goal is then to develop task-agnostic learning strategies that generalize well to an unseen test set
x
i
n
t
e
s
t
,
y
t
e
s
t
~
D
t
e
s
t
. For model selection and hyper-parameter tuning, we assume a development set
D
d
e
v
, of the same size as the few-shot training set"
D
t
r
a
i
n
is considered analogous to a training set.
D
t
e
s
t
is considered analogous to a testing set.
D
d
e
v
is considered analogous to a validation set.);
constructing, by a processor (Gao et al. pg. 3816, Section 1, Paragraph 2, "This setting is appealing as (1) such models can be trained on typical research hardware" Typical research hardware is construed to include a processor), a candidate tag word set (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "for each class
c
∈
Y
, we construct a pruned set
V
c
∈
V
of the top
k
vocabulary words based on their conditional likelihood using the initial ℒ.") and a candidate prompt template set (Gao et al. pg. 3820, Section 5.2, Paragraph 1, "we study how to generate a diverse set of templates {𝒯} automatically from a fixed set of label words ℳ(𝒴).");
searching, by the processor, an optimal tag word corresponding to the input sample from the candidate tag word set (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "To further narrow down the search space, we find the top
n
assignments over the pruned space that maximize zero-shot accuracy on
D
t
r
a
i
n
(both
n
and
k
are hyper-parameters, see Appendix C.2). Then we fine-tune all top
n
assignments, and re-rank to find the best one using
D
d
e
v
." Finding the best label assignment is considered analogous to searching an optimal tag word), and a prompt template corresponding to the input sample from the candidate prompt template set (Gao et al. pg. 3821, Section 5.2, Paragraph 4, "we use a wide beam width (e.g., 100) to cheaply obtain a large set of diverse templates.") [by means of reinforcement learning]; and
outputting, by the processor, a mapping relationship of the optimal tag word (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "We first study how to construct a label word mapping ℳ that maximizes accuracy on
D
d
e
v
after fine-tuning, given a fixed template 𝒯.") and an optimal prompt template format corresponding to the prompt template (Gao et al. pg. 3821, Section 5.2, Paragraph 4, "We then fine-tune each generated template on
D
t
r
a
i
n
and use
D
d
e
v
to either pick the single template with the best performance (Table 3), or the top k templates to use as an ensemble (Table 4).") that are used by the pre-trained language model to perform downstream task (Gao et al. pg. 3818, Section 3, Paragraph 2, "We conduct a systematic study across 8 single-sentence and 7 sentence-pair English tasks, including 8 tasks from the GLUE benchmark (Wang et al., 2019), SNLI (Bowman et al., 2015), and 6 other popular sentence classification tasks (SST-5, MR, CR, MPQA, Subj, TREC)." English tasks are considered analogous to downstream tasks);
wherein the constructing, by the processor, the candidate tag word set and the candidate prompt template set comprises:
automatically selecting, by the processor, an optimal candidate tag word (Gao et al. pg. 3820, Section 5, Paragraph 1, "We now explore principled ways of automating the search process for label words (§5.1) and templates (§5.2)."); and
automatically selecting, by the processor, a candidate prompt template (Gao et al. pg. 3820, Section 5, Paragraph 1, "We now explore principled ways of automating the search process for label words (§5.1) and templates (§5.2).");
wherein the automatically selecting, by the processor, the candidate tag word comprises:
initializing a vocabulary by the processor (Gao et al. pg. 3819, Section 4.1, Paragraph 1, "Let
M
:
Y
→
V
be a mapping from the task label space
Y
to individual words in the vocabulary 𝒱 of ℒ.");
vectorizing, by the processor, each word in the vocabulary using a [word2vec] method (Gao et al. pg. 3822, Section 6.2, Paragraph 1, "we use a pre-trained SBERT (Reimers and Gurevych, 2019) model to obtain embeddings for all input sentences (for sentence-pair tasks, we use the concatenation of the two sentences)." pg. 3828, Appendix C.2, Paragraph 1, "For TREC, we ... re-rank
V
c
by the nearest neighbors of the original manual label words and take the top 30 per class." Re-ranking vocabulary words
V
c
by nearest neighbors to the original manual label words implies that
V
c
must be vectorized, since nearest neighbor algorithms require that data must be represented in a vector space for computing distance metrics), and determining a near-synonym set for each category in the training set via a cosine similarity (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "To further narrow down the search space, we find the top
n
assignments over the pruned space that maximize zero-shot accuracy on
D
t
r
a
i
n
(both
n
and
k
are hyper-parameters, see Appendix C.2)." pg. 3828, Appendix C.2, Paragraph 1, "For TREC, we observe that filtering
V
c
using conditional likelihood alone is still noisy, thus we set
k
=
1000
, and then re-rank
V
c
by the nearest neighbors of the original manual label words and take the top 30 per class." Taking the top 30 vocabulary words in a label
c
that are nearest neighbors to the original manual label words is considered analogous to determining a near-synonym set for each category);
for each category in the training set, selecting, by the processor, a word in the vocabulary that maximizes a conditional probability, and a conditional probability set comprising the word (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "let
D
t
r
a
i
n
c
⊂
D
t
r
a
i
n
be the subset of all examples of class
c
. We take
V
c
as
T
o
p
-
k
v
∈
V
∑
x
i
n
∈
D
t
r
a
i
n
c
l
o
g
P
L
M
A
S
K
=
v
|
T
(
x
i
n
)
, where
P
L
denotes the output probability distribution of ℒ"
P
L
is considered analogous to a conditional probability. Thus,
V
c
is considered analogous to selecting a word in the vocabulary for each category that maximizes a conditional probability. The Top-K computation of the above equation is considered analogous to a conditional probability set), by the pre-trained model that is not fine-tuned (Gao et al. pg. 3818, Section 3, Paragraph 1, "In this work, we assume access to a pre-trained language model ℒ that we wish to fine-tune on a task
D
with a label space 𝒴." pg. 5, Section 5.1, Paragraph 1, "for each class
c
∈
Y
, we construct a pruned set
V
c
∈
V
of the top
k
vocabulary words based on their conditional likelihood using the initial ℒ.");
determining, by the processor, a candidate tag word under each category (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "Then we fine-tune all top
n
assignments, and re-rank to find the best one using
D
d
e
v
." The highest ranked label assignment selected from vocabulary of labels
c
is considered analogous to a candidate tag word under each category) as a maximum value of an intersection (Gao et al. pg. 3822, Section 6.2, Paragraph 1, "we devise a simple strategy in which we only sample examples that are semantically close to
x
i
n
. ... For each query
x
i
n
and each label
c
∈
Y
, we sort all training instances with the label
x
∈
D
t
r
a
i
n
c
by their similarity score to the query
c
o
s
(
e
(
x
i
n
)
;
e
(
x
)
)
" Section 5.1 describes fine-tuning assignments, which can be seen to to utilize cosine distance as described in the above citation) of the near-synonym set and the conditional probability set (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "Then we fine-tune all top
n
assignments, and re-rank to find the best one using
D
d
e
v
." Re-ranking the fine-tuned top
n
assignements and selecting the "best one" is considered analogous to using a maximum value of an intersection between a near-synonym set (e.g. "top
n
assignments") and a conditioanl probability set (e.g. "
T
o
p
-
k
v
∈
V
∑
x
i
n
∈
D
t
r
a
i
n
c
l
o
g
P
L
M
A
S
K
=
v
|
T
(
x
i
n
)
(Equation 3)")); and
integrating, by the processor, candidate tag words under various categories (Gao et al. pg. 3819, Section 4.1, Paragraph 1, "Let
M
:
Y
→
V
be a mapping from the task label space
Y
to individual words in the vocabulary 𝒱 of ℒ." It is to be understood that Section 5.1 was directed towards the construction of
M
. Thus, proceeding with a classificaiton task using the mapping found using techniques described in section 5.1 is considered analogous to integrating the candidate tag words under various categories, the various categories being label space
Y
), and determining an assignment mode which maximizes an accuracy rate of the training set as the optimal candidate tag word (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "We first study how to construct a label word mapping
M
that maximizes accuracy on
D
d
e
v
after fine-tuning, given a fixed template
T
." See Table 2, which illustrates an example of determining an assignment mode which maximizes an accuracy rate (the label words "great/terrible" maximizes the accuracy metric in the right column of Table 2) of the training set as the optimal candidate tag word (e.g. "great/terrible"). Choosing to utilize the label words "great/terrible" for downstream tasks is considered analogous to determining an assignment mode).
Gao et al. do not explicitly disclose all of reinforcement learning.
However, Efstathiou et al. disclose inputting a data set (Efstathiou et al. ¶ [0117], "FIG. 3 shows a method of labeling unlabeled data according to an embodiment. The method starts 30 with the retrieval of a set of labeled data and a set of unlabeled data.") [to the pre-trained language model, and forming an input sample according to a fixed template, wherein the data set comprises a training set, a validation set, and a test set];
constructing, by a processor (Efstathiou et al. ¶ [0169], "Execution of the classification controller software 107 by the processor 101 will cause embodiments as described herein to be implemented."), a candidate tag word set (Efstathiou et al. ¶ [0118], "Labels are then assigned to the unlabeled data 34 based on the confidence score output by the classifier. The classifier may be a multi-output classifier (a classifier that classifies data into one of a plurality of classes)." A set of labels/classes are considered analogous to a candidate tag word set) [and a candidate prompt template set];
searching, by the processor, an optimal tag word corresponding to the input sample from the candidate tag word set (Efstathiou et al. ¶ [0153], "The training system is therefore able to learn, via reinforcement learning, the best actions for retraining the classifier. Each action may include generating new instances of labeled data and retraining the classifier on these new instances. The finally trained policy can then be used to train a classifier to improve its performance"), [and a prompt template corresponding to the input sample from the candidate prompt template set] by means of reinforcement learning (Efstathiou et al. ¶ [0146], "FIG. 7 shows a flow chart for training a reinforcement learning system to train a classifier according to embodiments described herein."); and
outputting, by the processor, a mapping relationship of the optimal tag word (Efstathiou et al. ¶ [0153], "Each action may include generating new instances of labeled data and retraining the classifier on these new instances.") [and an optimal prompt template format corresponding to the prompt template that are used by the pre-trained language model to perform downstream task]….
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify Gao et al.’s small sample fine-tuning method to incorporate Efstathiou et al.’s reinforcement learning.
The suggestion/motivation for doing so would have been that, “By utilizing the reinforcement learning methods described herein, a system can be trained to improve the classification performance of a classifier without requiring additional manually labeled data,” as noted by the Efstathiou et al. disclosure in paragraph [0062].
Gao et al. in view of Efstathiou et al. do not explicitly disclose all of a word2vec method.
However, Corvinelli et al. disclose a word2vec method (Corvinelli et al. ¶ [0019], "FIG. 2 depicts a SQL embedding layer 200 in accordance with at least one embodiment of the present invention. ... In at least one embodiment, SQL embedding layer is configured to utilize a Word2vec model to create word embeddings.").
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify Gao et al.’s small sample fine-tuning method to include Corvinelli et al.’s word2vec embedding method because such a modification is the result of simple substitution of one known element for another producing a predictable result. More specifically, Gao et al.’s SBERT encoding and Corvinelli et al.’s word2vec encoding perform the same general and predictable function, the predictable function being generating vectors from datasets. Since each individual element and its function are shown in the prior art, albeit shown in separate references, the difference between the claimed subject matter and the prior art rests not on any individual element or function but in the very combination itself - that is in the substitution of Gao et al.’s SBERT encoding by replacing it with Corvinelli et al.’s word2vec encoding. Thus, the simple substitution of one known element for another producing a predictable result renders the claim obvious.
Claim 2
Regarding claim 2, the rejection of claim 1 is incorporated.
Gao et al. further disclose wherein:
the training set is configured to sample randomly to form the input sample (Gao et al. pg. 3818, Section 3, Paragraph 1, "we measure average performance across 5 different randomly sampled
D
t
r
a
i
n
and
D
d
e
v
splits."); and
the validation set is configured to calculate the cosine similarity (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "we fine-tune all top
n
assignments, and re-rank to find the best one using
D
d
e
v
." pg. 3822, Section 6.2, Paragraph 1, "For each query
x
i
n
and each label
c
∈
Y
, we sort all training instances with the label
x
∈
D
t
r
a
i
n
c
by their similarity score to the query
c
o
s
(
e
(
x
i
n
)
;
e
(
x
)
)
").
Claim 3
Regarding claim 3, the rejection of claim 1 is incorporated.
Gao et al. further disclose forming data in the data set according to ID attributes, sentence attributes, and label attributes wherein, the ID attributes are configured to represent IDs of the data, the sentence attributes are configured to represent contents of the data, and the label attributes are configured to represent tag words of the data (Gao et al. pg. 3829, Table B.1 describes each of the datasets used for training and evaluation. Column "Dataset" is considered analogous to ID attributes. Column "Category" is considered analogous to sentence attributes. Column "Labels (classification tasks)" is considered analogous to label attributes.).
Claim 4
Regarding claim 4, the rejection of claim 1 is incorporated.
Gao et al. further disclose acquiring input content (Gao et al. pg. 3819, Section 4, Paragraph 2, "we can formulate a binary sentiment classification task using a prompt with input
x
1
(e.g., “No reason to watch it .”)");
representing the input content in the fixed template (Gao et al. pg. 3819, Section 4, Paragraph 2, "we can formulate a binary sentiment classification task using a prompt with input
x
1
(e.g., “No reason to watch it .”) as:
x
p
r
o
m
p
t
=
C
L
S
x
1
I
t
w
a
s
M
A
S
K
.
[
S
E
P
]
");
calculating the cosine similarity between the input content and all samples in a training set (Gao et al. pg. 3822, Section 6.2, Paragraph 1, "For each query
x
i
n
and each label
c
∈
Y
, we sort all training instances with the label
x
∈
D
t
r
a
i
n
c
by their similarity score to the query
c
o
s
(
e
(
x
i
n
)
;
e
(
x
)
)
, and only sample from the top r = 50% instances for each class to use as demonstrations."); and
random sampling from a preset percentage of training set samples to obtain the input sample (Gao et al. pg. 3822, Section 6.2, Paragraph 1, "For each query
x
i
n
and each label
c
∈
Y
, we sort all training instances with the label
x
∈
D
t
r
a
i
n
c
by their similarity score to the query
c
o
s
(
e
(
x
i
n
)
;
e
(
x
)
)
, and only sample from the top r = 50% instances for each class to use as demonstrations." pg. 3821, Section 6.1, Paragraph 1, "at each training step, we randomly sample one9 example
(
x
i
n
c
,
y
c
)
∈
D
t
r
a
i
n
from each class").
Claim 5
Regarding claim 5, the rejection of claim 4 is incorporated.
Gao et al. further disclose initializing a prompt template format (Gao et al. pg. 3819, Section 4, Paragraph 2, "we can formulate a binary sentiment classification task using a prompt with input
x
1
(e.g., “No reason to watch it .”) as:
x
p
r
o
m
p
t
=
C
L
S
x
1
I
t
w
a
s
M
A
S
K
.
[
S
E
P
]
"); and
representing the input content in the initialized prompt template format (Gao et al. pg. 3819, Section 4, Paragraph 2, "we can formulate a binary sentiment classification task using a prompt with input
x
1
(e.g., “No reason to watch it .”) as:
x
p
r
o
m
p
t
=
C
L
S
x
1
I
t
w
a
s
M
A
S
K
.
[
S
E
P
]
").
Claim 6
Regarding claim 6, the rejection of claim 4 is incorporated.
Gao et al. further disclose encoding the input content using an SBERT method (Gao et al. pg. 3822, Section 6.2, Paragraph 1, "we use a pre-trained SBERT (Reimers and Gurevych, 2019) model to obtain embeddings for all input sentences"); and
calculating, for each input content in a validation set, the cosine similarity to all samples in the training set respectively (Gao et al. pg. 3822, Section 6.2, Paragraph 1, "For each query
x
i
n
and each label
c
∈
Y
, we sort all training instances with the label
x
∈
D
t
r
a
i
n
c
by their similarity score to the query
c
o
s
(
e
(
x
i
n
)
;
e
(
x
)
)
, and only sample from the top r = 50% instances for each class to use as demonstrations.").
Claim 7
Regarding claim 7, the rejection of claim 3 is incorporated.
Gao et al. further disclose converting the input sample to a prompts input (Gao et al. pg. 3819, Section 4, Paragraph 2, "we can formulate a binary sentiment classification task using a prompt with input
x
1
(e.g., “No reason to watch it .”) as:
x
p
r
o
m
p
t
=
C
L
S
x
1
I
t
w
a
s
M
A
S
K
.
[
S
E
P
]
").
Claim 10
Regarding claim 10, the rejection of claim 1 is incorporated.
Gao et al. further disclose determining the conditional probability set through a formula:
T
o
p
-
k
v
∈
V
∑
x
i
n
∈
D
t
r
a
i
n
c
l
o
g
P
L
M
A
S
K
=
v
|
T
(
x
i
n
)
wherein Topk is a word with a maximum conditional probability;
V
is an initialization vocabulary;
L
is the pre-trained model that is not fine-tuned; c is each category in the training set;
P
L
represents an output probability distribution based on the model
L
; and T(
X
i
n
) is an input sample (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "for each class
c
∈
Y
, we construct a pruned set
V
c
∈
V
of the top
k
vocabulary words based on their conditional likelihood using the initial ℒ. That is, let
D
t
r
a
i
n
c
⊂
D
t
r
a
i
n
be the subset of all examples of class
c
. We take
V
c
as
T
o
p
-
k
v
∈
V
∑
x
i
n
∈
D
t
r
a
i
n
c
l
o
g
P
L
M
A
S
K
=
v
|
T
(
x
i
n
)
, where
P
L
denotes the output probability distribution of ℒ").
Claim 11
Regarding claim 11, the rejection of claim 1 is incorporated.
Gao et al. further disclose determining the optimal candidate tag word (Gao et al. pg. 3820, Section 5.1, Paragraph 1, "We first study how to construct a label word mapping ℳ that maximizes accuracy on
D
d
e
v
after fine-tuning, given a fixed template 𝒯.");
generating an initial prompt template by filling a placeholder (Gao et al. pg. 3821, Section 5.2, Paragraph 2-3, "Given an input example
(
x
i
n
;
y
)
∈
D
t
r
a
i
n
, we consider the following simple conversions, denoted as
T
g
(
x
i
n
;
y
)
, for formulating the T5 model inputs: [see mappings following paragraph]. As shown in Figure 2, we rely on the T5 model to fill in the placeholders."); wherein the initial prompt template is configured to maximize an output probability in the training set (Gao et al. pg. 3821, Section 5.2, Paragraph 3, "When decoding, our goal here is to find an output that can work well for all examples in
D
t
r
a
i
n
, i.e., the output template 𝒯 that maximizes
∑
(
x
i
n
;
y
)
∈
D
t
r
a
i
n
l
o
g
P
T
5
(
T
|
T
g
(
x
i
n
;
y
)
)
, where
P
T
5
denotes the output probability distribution of T5."); and
decoding the initial prompt template using a bundle search algorithm to obtain the candidate prompt template (Gao et al. pg. 3821, Section 5.2, Paragraph 4, "We use beam search to decode multiple template candidates. Concretely, we use a wide beam width (e.g., 100) to cheaply obtain a large set of diverse templates. We then fine-tune each generated template on
D
t
r
a
i
n
and use
D
d
e
v
to either pick the single template with the best performance (Table 3), or the top k templates to use as an ensemble (Table 4)." Beam search is considered analogous to a bundle search algorithm).
Claim 12
Regarding claim 12, the rejection of claim 11 is incorporated.
Gao et al. further disclose determining a preset number of candidate tag word set for each category (Gao et al. pg. 3821, Section 6.1, Paragraph 1, "at each training step, we randomly sample one9 example
(
x
i
n
c
,
y
c
)
∈
D
t
r
a
i
n
from each class" pg. 6, Column 2, Footnote 9, "We also explored sampling multiple examples per class, but did not observe any improvements." Experimenting with different sampling numbers and settling on a standard value (e.g. “one”) is considered analogous to determining a preset number);
combining the candidate tag word set with a template set corresponding to the candidate prompt template to obtain a search space list (Gao et al. pg. 3821, Section 6.1, Paragraph 1, "at each training step, we randomly sample one9 example
(
x
i
n
c
,
y
c
)
∈
D
t
r
a
i
n
from each class, convert it into
T
(
x
i
n
c
)
in with [MASK] replaced by
M
(
y
(
c
)
)
—we denote this as
T
~
(
x
i
n
c
,
y
c
)
—and then concatenate them with
x
i
n
:" See Figure 1(c), which illustrates the combined search space list); and
by means of the search space list, determining an optimal tag word corresponding to the input sample from the candidate tag word set, and a prompt template corresponding to the input sample from the candidate prompt template set (Gao et al. pg. 3820, Section 5, Paragraph 1, "We now explore principled ways of automating the search process for label words (§5.1) and templates (§5.2). Our goals are to ... find more optimal settings than those that we manually choose.").
Claim 13
Regarding claim 13, the rejection of claim 12 is incorporated.
Gao et al. further disclose by combining the candidate tag word set with a template set corresponding to the candidate prompt template, obtaining the search space list (Gao et al. pg. 3821, Section 6.1, Paragraph 1, "at each training step, we randomly sample one9 example
(
x
i
n
c
,
y
c
)
∈
D
t
r
a
i
n
from each class, convert it into
T
(
x
i
n
c
)
in with [MASK] replaced by
M
(
y
(
c
)
)
—we denote this as
T
~
(
x
i
n
c
,
y
c
)
—and then concatenate them with
x
i
n
:" See Figure 1(c), which illustrates the combined search space list) to determine the optimal assignment mode of the candidate tag word and the candidate prompt template in the finetuning process (Gao et al. pg. 3821, Section 5.2, Paragraph 4, "We then fine-tune each generated template on
D
t
r
a
i
n
and use
D
d
e
v
to either pick the single template with the best performance (Table 3), or the top k templates to use as an ensemble (Table 4).").
Claim 14
Regarding claim 14, the rejection of claim 1 is incorporated.
Gao et al. further disclose determining the optimal tag word and the prompt template (Gao et al. pg. 3820, Section 5, Paragraph 1, "We now explore principled ways of automating the search process for label words (§5.1) and templates (§5.2).") [by key factors in reinforcement learning, wherein the key factors comprise agent, environment, action, status, and reward].
Efstathiou et al. further disclose key factors in reinforcement learning, wherein the key factors comprise agent, environment, action, status, and reward (Efstathiou et al. ¶ [0063], "FIG. 1 shows an example of a reinforcement learning process. This shows a single episode of reinforcement learning by a single agent. The agent is on an observed state
s
t
at a particular time point
t
. The environment in reinforcement learning is typically considered to be a Markov Decision Process (MDP)" ¶ [0031], "Training the agent may comprise selecting and storing in the policy the actions that provide the highest value. The value of each action may be based on a reward for that action.").
Claim 15
Regarding claim 15, the rejection of claim 14 is incorporated.
Gao et al. further disclose inputting text into the model (Gao et al. pg. 3819, Section 4, Paragraph 2, "we can formulate a binary sentiment classification task using a prompt with input
x
1
(e.g., “No reason to watch it .”) as:
x
p
r
o
m
p
t
=
C
L
S
x
1
I
t
w
a
s
M
A
S
K
.
[
S
E
P
]
") to obtain an output result (Gao et al. pg. 3819, Section 4.1, Paragraph 1, "we can treat our task as an MLM, and model the probability of predicting class
y
∈
Y
as:
p
y
x
i
n
=
p
M
A
S
K
=
M
(
y
)
x
i
n
=
e
x
p
(
w
M
(
y
)
∙
h
[
M
A
S
K
]
)
∑
y
'
∈
Y
e
x
p
(
w
M
(
y
)
∙
h
[
M
A
S
K
]
)
'
, where
h
[
M
A
S
K
]
is the hidden vector of
[
M
A
S
K
]
and
w
M
(
y
)
denotes the pre-softmax vector corresponding to
v
∈
V
."); the model comprising a language model environment (Gao et al. pg. 3818, Section 3, Paragraph 1, "In this work, we assume access to a pre-trained language model ℒ that we wish to fine-tune on a task
D
with a label space 𝒴.");
calculating a loss of the output result and a current tag word (Gao et al. pg. 3819, Section 4.1, Paragraph 1, "When supervised examples
(
x
i
n
,
y
)
are available, ℒ can be fine-tuned to minimize the cross-entropy loss.");
[feeding back the loss as the reward to the agent;] and
determining, [by the agent,] selection directions of subsequent templates and tag words [according to the reward] until the optimal tag word and the prompt template are determined (Gao et al. pg. 3820, Section 5, Paragraph 1, "We now explore principled ways of automating the search process for label words (§5.1) and templates (§5.2). Our goals are to ... find more optimal settings than those that we manually choose.").
Efstathiou et al. further disclose inputting text into the model (Efstathiou et al. ¶ [0113]-[0114], "AlSynth is a classification method that aims to assign labels to unlabeled data based on an initial labeled training set of data.") to obtain an output result (Efstathiou et al. ¶ [0063]-[0064], "The agent is on an observed state
s
t
at a particular time point
t
. ... The agent determines an action
a
t
to be performed in response to the state
s
t
based on a policy for the agent."); the model comprising a language model environment (Efstathiou et al. ¶ [0126]-[0127], "ChopSynth operates in the same scenario as AlSynth with the same goal—to synthesize labeled data from unlabeled data based on labeled train data. ... ChopSynth is described with reference to the classification of words from text data.");
calculating a loss of the output result and a current tag word (Efstathiou et al. ¶ [0065], "By applying the action
a
t
, the agent traverses to a new state
s
t
+
1
in the next time point
t
+
1
. It then receives a reward
r
t
from the environment at the state
s
t
+
1
.");
feeding back the loss as the reward to the agent (Efstathiou et al. ¶ [0065], "By applying the action
a
t
, the agent traverses to a new state
s
t
+
1
in the next time point
t
+
1
. It then receives a reward
r
t
from the environment at the state
s
t
+
1
."); and
determining, by the agent, selection directions of subsequent [templates and] tag words according to the reward (Efstathiou et al. ¶ [0068]-[0069], "The most appropriate action for a given state may be determined using a gradient ascent method based on the reward values. This involves locating values, connected to particular actions, that maximise the reward function's result. In this way, after the end of the learning process, the agent has successfully learnt how to traverse to the most desirable state through a series of other states, by selecting the highest in value actions (i.e. those of the optimal policy).") until the optimal tag word [and the prompt template] are determined (Efstathiou et al. ¶ [0104], "Each reinforcement learning agent works on an individual class within a multi-class classification problem." Classifying data into multiple classes using an optimal policy trained overtime is considered analogous to selecting directions of tag word classification using a reward).
Claim 19
Regarding claim 19, Gao et al. disclose a non-transitory readable storage medium having stored thereon a computer program that, when executed by a processor, implements the steps of the method according to claim 1 (Gao et al. pg. 3816, Section 1, Paragraph 2, "In this work, we study a more practical scenario in which we only assume access to a moderately sized language model such as BERT (Devlin et al., 2019) or RoBERTa (Liu et al., 2019).... This setting is appealing as (1) such models can be trained on typical research hardware").
The remaining limitations of claim 19 are identical to that of claim 1 and therefore are rejected for similar reasons as described above.
Claim 20
Regarding claim 19, Gao et al. disclose an electronic device, comprising a memory having stored thereon a computer program, and a processor that implements the steps of the method according claim 1 when calling the computer program in the memory (Gao et al. pg. 3816, Section 1, Paragraph 2, "In this work, we study a more practical scenario in which we only assume access to a moderately sized language model such as BERT (Devlin et al., 2019) or RoBERTa (Liu et al., 2019).... This setting is appealing as (1) such models can be trained on typical research hardware").
The remaining limitations of claim 20 are identical to that of claim 1 and therefore are rejected for similar reasons as described above.
Claim 16 is rejected under 35 U.S.C. 103 as obvious over Gao et al. in view of Efstathiou et al. in view of Corvinelli et al. as applied to claim 1, and further in view of "Calibrate Before Use: Improving Few-Shot Performance of Language Models" (Zhao et al.).
Claim 16
Regarding claim 16, the rejection of claim 1 is incorporated. Gao et al. in view of Efstathiou et al. in view of Corvinelli et al. disclose all the elements of the claimed invention as stated above.
Gao et al. further disclose [when an input is textless,] averaging an output tag word corresponding probability (Gao et al. pg. 3828, Appendix C.3, Paragraph 1, "When using demonstrations, we sample 16 different sets of demonstrations for each input and average the predicted log probability for each class during inference.") [and then normalizing to obtain a normalized probability p_cf ; and calculating a correction matrix according to the formula].
Efstathiou et al. further disclose textless input (Efstathiou et al. ¶ [0127], "the ChopSynth method uses frequent (but not the overly common and therefore, meaningless) words to identify sequences of words in the unlabeled instances. If applied to other data types such as image data, common features in the labeled set are used to identify equivalent features in the unlabeled set to generate new samples.").
Gao et al. in view of Efstathiou et al. in view of Corvinelli et al. do not explicitly disclose all of a normalization or calculation of a correction matrix.
However, Zhao et al. disclose averaging an output tag word corresponding probability (Zhao et al. pg. 5, Section 5, Paragraph 4, "In all our experiments, we average the probabilities from three content-free inputs: “N/A”, “[MASK]”, and the empty string.") and then normalizing to obtain a normalized probability p_cf (Zhao et al. pg. 5, Section 5, Paragraph 1, "For classification tasks,
p
^
is the set of probabilities that are associated with each label name, renormalized to one."); and calculating a correction matrix according to the formula
[
d
i
a
g
p
c
f
]
-
1
(Zhao et al. pg. 5, Section 5, Paragraph 3, "We first obtain
p
^
for the content-free input, denoted
p
^
c
f
. We then set
W
=
d
i
a
g
(
p
^
c
f
)
-
1
").
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify Gao et al. in view of Efstathiou et al. in view of Corvinelli et al. to incorporate Zhao et al.’s normalization and correction.
The suggestion/motivation for doing so would have been that, “LMs are biased towards outputting answers that are (1) frequent in the prompt (majority label bias), (2) towards the end of the prompt (recency bias), and (3) common in the pre-training data (common token bias) … we look to correct this [bias] by “calibrating” the model’s output probabilities. A common technique for adjusting output probabilities is to apply an affine transformation… where a weight matrix
W
and a bias vector
b
are applied to the original probabilities
p
^
to get the new probabilities,” as noted by the Zhao et al. disclosure in pg. 4, Section 4, Paragraph 1, and pg. 5, Section 5, Paragraph 1.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB B VOGT whose telephone number is (571)272-7028. The examiner can normally be reached Monday - Friday, 11am - 8pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PARAS D SHAH can be reached at (571)270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACOB B VOGT/Examiner, Art Unit 2653
/JESSE S PULLIAS/Primary Examiner, Art Unit 2655 08/13/26