DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 11/27/2024 was filed in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
1. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claims 1 and 11, “A system” and “A method” are recited, which are each directed to one of the four statutory categories of invention (machine, process; Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea.
The following limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts:
display a visualization of the stance, the sentiment, and the sarcasm: a person can write down visuals/text describing stance/sentiment/sarcasm
preprocess…an input text from the text dataset based on the plurality of user-defined parameters to obtain a preprocessed text batch: a person can preprocess text according to parameters from a user (e.g. splitting up text according to user specified lengths)
encode and tokenize…the preprocessed text batch to obtain a multi-task dataset having a tokenized input text: a person writes down a written representation of the text as tokens
adjust…a plurality of task weights of the multi-task model with a weighting scheme: adjusting weights according to a weighting scheme amounts to mathematical calculations
determine the stance, the sentiment, and the sarcasm based on…the plurality of task weights: a person can use the task weights information to reach decisions about the stance, sentiment, and sarcasm in a piece of text
Claims 1 and 11 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The additional limitations are “A system…comprising: a graphical processing unit having a memory; an input device configured to receive a plurality of user-defined parameters and connected to the graphical processing unit; and a display device configured to display…and connected to the graphical processing unit and the memory” (claim 1), “preprocess…by input layers” (claims 1 and 11), “encode and tokenize…by shared layers” (claims 1 and 11), “train, by task specific layers, a multi-task model having a stance head, a sentiment head, and a sarcasm head with the multi-task dataset” (claims 1 and 11), “adjust, by the task-specific layers” (claims 1 and 11), and “determine the stance, the sentiment, and the sarcasm based on the multi-task model” (claims 1 and 11). These limitations are recited at a high level of generality and amount to mere instructions to implement the judicial exception using a generic computer, which even when viewed in combination with the claim as a whole, do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract ideas. Therefore, claims 1 and 11 are directed to abstract ideas.
Claims 1 and 11 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which even when viewed in combination with the claim as a whole, do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 1 and 11 are not patent eligible.
Regarding claims 2-10 and 12-20, “The system” and “The method” are recited, which are each directed to one of the four statutory categories of invention (machine, process; Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite further mental processes or mathematical concepts which fall into the category of abstract idea.
The following limitations, under their broadest reasonable interpretation, recite further mental processes or mathematical concepts:
Claims 2 and 12:
transform the tokenized input text into a plurality of representations including a token embedding, a segment embedding, and a position embedding; generate a unified representation by adding the plurality of representations; and tune the unified representation…: forming representations via adding together token, segment, and position embeddings and tuning amounts to mathematical calculations
Claim 2 contains the additional limitation “tune…with a pre-trained language model”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 3 and 13:
Claims 3 and 13 recite the additional limitation of “wherein the multi-task model is selected from the group consisting of a parallel multi-task model and a sequential multi-task model”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 4 and 14:
wherein the weighting scheme is selected from the group consisting of a static weighting sum, a hierarchical weighting, and an uncertainty weighting: weighting via static, hierarchical, and uncertainty weighting schemes amounts to mathematical calculations
Claims 5 and 15:
Claims 5 and 15 contain the additional limitation “where the stance head is a primary head, and the sentiment head and the sarcasm head are auxiliary heads”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 6 and 16:
wherein the visualization includes an attention visualization configured to provide a plurality of information of the text dataset: a person can present information using a written report to a user regarding attention
Claims 7 and 17:
wherein the plurality of information includes attention weights, relevance level, and a prominence level: a person can write down information relating to attention weights, relevances, and prominences to show a user
Claims 8 and 18:
Claims 8 and 18 contain the limitation “wherein the multi-task model is a multi-target sequential multi-task learning model with hierarchical weighting (SMTL-HW)”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 9 and 19:
Claims 9 and 19 contain the limitation “wherein the plurality of user-defined parameters includes a maximum sequence length, a feature dimension, a batch size, a dropout rate, a patience parameter, a number of epochs, and a learning rate”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 10 and 20:
Claims 10 and 20 contain the limitation “wherein the maximum sequence length is 128 tokens, the feature dimension is 786, the batch size is 32, the dropout rate is 0.1, the patience parameter is 5, the number of epochs is 20, and the learning rate is 2e^-5”, which amounts to mere instructions to implement the judicial exception using a generic computer.
Claims 2-10 and 12-20 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which even viewed in combination with the claims as a whole, do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract ideas. Therefore, claims 2-10 and 12-20 are directed to abstract ideas.
Claims 2-10 and 12-20 do not contain any additional limitations which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which even when viewed in combination with the claims as a whole, do not amount to significantly more than the judicial exception as they do not provide an inventive concept. Therefore, claims 2-10 and 12-20 are not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
2. Claims 1, 3-7, 11, and 13-17 are rejected under 35 U.S.C. 103 as being unpatentable over Fu et al. (NPL Incorporate opinion-towards for stance detection, hereinafter Fu) in view of Ghosh et al. (NPL “Laughing at you or with you”: The Role of Sarcasm in Shaping the Disagreement Space, hereinafter Ghosh) and further in view of NPL Document “tweet-preprocessor 0.6.0”, hereinafter NPL1.
Regarding claim 1, Fu discloses A system for simultaneously predicting a stance, sentiment, … from a text dataset, comprising (see Fig. 2; the use of the disclosed machine learning model architecture, along with other computer based machine learning models (e.g. BERT, Table 4) and machine learning text datasets (section 4.1), to obtain experimental results (Tables 4-9) inherently reads on a generic system): …a visualization of the stance, the sentiment…(Figs. 3-4)…encode and tokenize, by shared layers, the…text batch to obtain a multi-task dataset having a tokenized input text (pg. 3, section 3.1 “Input layer. The training set in our benchmark data SemEval-2016 Task 6A can be denotes as S…where each sample in S has a tweet s, claim c, stance label lstance (Against (-1), None (0) and Favor (1)), opinion-towards label lot (Other (-1), No one (0), and Target (1)), and sentiment label lsent (Negative (-1), Neither (0) and Positive (1))…We leverage a word embedding matrix X…to map each word into a continuous dense vector, where v is the vocabulary size, d is the dimension of dense vectors. Correspondingly, the tweet and the claim in S can then be denoted as x^s…and x^c…respectively, based on their word individual representations…”); train, by task-specific layers, a multi-task model having a stance head, a sentiment head, and a … head with the multi-task dataset (pg. 3, section 3.1 “Input layer. The training set in our benchmark data SemEval-2016 Task 6A can be denotes as S…; task specific layers utilized by model (see Fig. 2, “GAT”, “BiLSTMsem”, “BiLSTMot”, “BiLSTMsent”); training is performed using training data to minizimise multi-task loss function (see pg. 5, “Training” and Eqs. 12-13); the multi-task layer has a stance head (softmax to generate p_stance), sentiment head (softmax to generate p_sentiment), and a third head (softmax to generate p_ot)); adjust, by the task-specific layers, a plurality of task weights of the multi-task model with a weighting scheme (static weighting sum used for weighting multiple tasks: (see Eq 13, hyperparameters lambda 1-3)); determine the stance, the sentiment, … based on the multi-task model and the plurality of task weights (model trained with static weight summed tasks is used to determine stance, sentiment, and a third task (opinion-towards classification) (see outputs of Fig. 2)).
Fu does not specifically disclose:
a graphical processing unit having a memory; an input device configured to receive a plurality of user-defined parameters and connected to the graphical processing unit; and a display device configured to display a visualization… wherein the memory includes program instructions configured to:
[train, by task-specific layers, a multi-task modal having…] a sarcasm head [with the multi-task dataset] and
[determine…] the sarcasm [based on the multi-task model and the plurality of task weights].
Ghosh teaches a graphical processing unit having a memory (pg. 13 “BERT based models…Each training epoch took between 08:46 ~ 9 minutes over a K-80 GPU with 48GB vRAM…”); an input device configured to receive a plurality of user-defined parameters and connected to the graphical processing unit (various hyperparameters (e.g. user-defined parameters) are recited (e.g. pg. 13, batch size, dropout value, epochs, hidden state vector sizes, learning rates) which are used to further train the model using a computer (e.g. GPU), which inherently reads on an input device configured to receive user-defined parameters and connected to the GPU); and a display device configured to display a visualization… (pg. 7, 2nd Col. “We display the heat maps of the attention weights for a pair of prior and current turns (LSTM-based model) (Figure 3) whereas for BERT we display word-to-word attentions (Figures 4, 5, 6, 7, and 8) using visualization tools (Vig, 2019; Yang and Zhang, 2018)…”; the use of visualization tools to generate attention based displays based on executions of the taught multi-task model inherently reads on a display device to display ) wherein the memory includes program instructions configured to (pg. 2 reference is made to source code with reference to GitHub repository: “We make the code from our experiments publicly available…”); [train, by task-specific layers, a multi-task modal having…] a sarcasm head [with the multi-task dataset] (pg. 3, section 3 “Our training and test data are collected from the Internet Argument Corpus (IAC) (Walker et al., 2012a). This corpus consists of posts from conversations in online forums on a range of controversial political and social topics such as Evolution, Abortion, Gun Control, and Gay Marriage (Abbott et al., 2011, 2016). …were annotated using Mechanical Turk for argumentative relations (agree/disagree/none) and other characteristics such as sarcasm/non-sarcasm, respect/insult, nice/nastiness. Median Cohen’s κ is 0.5 across all topics.”; Fig, 2 “Classification Head Sarcasm Detection”; pg. 5 section 4.3 “Multitask Learning with BERT…Here, we pass the BERT output embeddings to two classification heads- one for each task (i.e., detection of argumentative relation and sarcasm), and the relevant gold labels are passed to them.”) and [determine…] the sarcasm [based on the multi-task model and the plurality of task weights] (pg. 5, Eq. 1 shows dynamic weighting of task-specific losses for training; trained model is used to determine sarcasm: Fig. 1, “Sarcasm” output from “Dense + SoftMax” layer).
Fu and Ghosh are considered to be analogous to the claimed invention as
they both are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu to incorporate the teachings of Ghosh in order to specifically train a model having a sarcasm head, and to determine the sarcasm back on the multi-task model and plurality of task weights. Doing so would be beneficial, as sarcasm-related features improve the accuracy of identifying argumentative relation/stance identification (Ghosh, pg. 9, conclusion).
Fu in view of Ghosh the text dataset (Fu, section 4.1) does not specifically disclose preprocess, by input layers, an input text …based on the plurality of user-defined parameters to obtain a preprocessed text batch.
NPL1 teaches preprocess, by input layers, an input text …based on the plurality of user-defined parameters to obtain a preprocessed text batch (pg. 1, section “Fully customizable…Preprocessor will go through all of the options by default unless you specify some options…”; see pg. 2, “Available options” for preprocessing specific user input (URL, hashtag symbols, emojis, etc.) according to specific user parameters (p.OPT.URL, p.OPT.MENTION, etc.)).
Fu, Ghosh, and NPL1 are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu in view of Ghosh to incorporate the teachings of NPL1 in order to preprocess by input layers an input text from the text dataset based on a plurality of user-defined parameters to obtain a preprocessed text batch. Doing so would be beneficial, as this would clean user-specified non-text from the input text, enabling raw data to be processed further by natural language processing models.
Regarding claim 3, Fu in view of Ghosh and NPL1 discloses wherein the multi-task model is selected from the group consisting of a parallel multi-task model and a sequential multi-task model (Fu, parallel multi-task model, see Fig. 2, task-specific layers in parallel with each other).
Regarding claim 4, Fu in view of Ghosh and NPL1 discloses wherein the weighting scheme is selected from the group consisting of a static weighting sum, a hierarchical weighting, and an uncertainty weighting (Fu, static weighting sum, see Eq. 13, hyperparameters lambda 1-3).
Regarding claim 5, Fu in view of Ghosh and NPL1 discloses wherein the stance head is a primary head (Fu, Fig. 2, “Main Task Stance Detection”), and the sentiment head and the sarcasm head are auxillary heads (Fu teach auxiliary head for sentiment: Fig. 2, “Auxiliary Task Sentiment Classification”; Ghosh teaches auxiliary head for sarcasm: pg. 7, 1st para. “However, adding more data for the auxiliary task (i.e., sarcasm detection)…”, see Fig. 2 “Classification Head Sarcasm Detection”).
Fu, Ghosh, and NPL1 are considered to be analogous to the claimed invention as they both are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Ghosh in order to have an auxiliary sarcasm head. Doing so would be beneficial, as sarcasm-related features improve the accuracy of identifying argumentative relation/stance identification (pg. 9, conclusion).
Regarding claim 6, Fu in view of Ghosh and NPL1 discloses wherein the visualization includes an attention visualization configured to provide a plurality of information of the text dataset (Fu, see Figs. 3-4, visualizations of attention weights for examples 1 and 2).
Regarding claim 7, Fu in view of Ghosh and NPL1 discloses wherein the plurality of information includes attention weights (Fu, attention weights represented via color intensity, with a higher color indicating higher attention (Figs. 3-4); see also section 4.6.2), a relevance level (Fu, visualization indicates how relevant (how intense the color is) of each word for each particular task (stance/opinion towards/sentiment)), and a prominence level (Fu, visualization indicates how prominent a particular token stands out compared to others for the particular task (e.g. Fig. 3, words “scandal” and “her” are more prominent than words “i”, and “would” for the stance detection task)).
Regarding claim 11, claim 11 is a method claim with limitations similar to those in system claim 1, and is thus rejected under similar rationale.
Regarding claim 13, claim 13 is rejected for analogous reasons to claim 3.
Regarding claim 14, claim 14 is rejected for analogous reasons to claim 4.
Regarding claim 15, claim 15 is rejected for analogous reasons to claim 5.
Regarding claim 16, claim 16 is rejected for analogous reasons to claim 6.
Regarding claim 17, claim 17 is rejected for analogous reasons to claim 7.
3. Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Fu in view of Ghosh and NPL1, and further in view of Devlin et al. (NPL BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, hereinafter Devlin).
Regarding claim 2, Fu in view of Ghosh and NPL1 does not specifically disclose transform the tokenized input text into a plurality of representations including a token embedding, a segment embedding, and a position embedding; generate a unified representation by adding the plurality of representations; and tune the unified representation with a pre-trained language model.
Devlin teaches transform the tokenized input text into a plurality of representations including a token embedding, a segment embedding, and a position embedding (pg. 4, 3rd para. “For a given token, its input representation is constructed by summing the corresponding token, segment, and position embeddings…”; Fig. 2); generate a unified representation by adding the plurality of representations (pg. 4, 3rd para. “For a given token, its input representation is constructed by summing the corresponding token, segment, and position embeddings…”; Fig. 2); and tune the unified representation with a pre-trained language model (pg. 4, 3rd para. “For a given token, its input representation is constructed by summing the corresponding token, segment, and position embeddings…”; pg. 14, section A.5 “Illustrations of Fine-tuning on Different Tasks…In the figure, E represents the input embedding”; Fig. 4, input embeddings are fed into pre-trained language model (BERT)).
Fu, Ghosh, NPL1, and Devlin are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu in view of Ghosh and NPL1 to incorporate the teachings of Devlin in order to transform the tokenized input text into a plurality of representations including a token embedding, a segment embedding, and a position embedding; generate a unified representation by adding the plurality of representations; and tune the unified representation with a pre-trained language model. Doing so would be beneficial, as the pre-trained BERT model is a language representation model which can be fine-tuned to create SOTA models for a wide variety of natural language processing tasks (Devlin, Abstract).
Regarding claim 12, claim 12 is rejected for analogous reasons to claim 2.
4. Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Fu in view of Ghosh and NPL1, and further in view of Song et al. (NPL Hierarchical Multi-task Learning for Organization Evaluation of Argumentative Student Essays, hereinafter Song).
Regarding claim 8, Fu in view of Ghosh and NPL1 discloses wherein the multi-task model (Fu, Fig. 2) is a multi-target (multi-target models predict multiple target variables for a single set of input features; the disclosed model in Fu, Fig. 2 outputs multiple targets (a stance, an opinion-towards classification, and a sentiment classification), and is thus multi-target)…multi-task learning model (Fu, Fig. 2, disclosed model is trained to perform one main task and two auxiliary tasks, thus making it multi-task)…
Fu in view of Ghosh and NPL1 does not specifically disclose [wherein the multi-task model is a multi-target] sequential [multi-task learning model] with hierarchical weighting…
Song teaches a multi-task learning model (Fig. 2, model with three tasks) which is sequential (Fig. 2, “Sentence function representations”, used for task 1, is also passed to Bi-LSTM for task 2; paragraph and sentence function representations are both passed to organization evaluation component block for task 3) and with hierarchical weighting (pg. 5, section 4.5, see loss functions (6) and (7) such that the weight of for SFI and PFI are adjusted relative to OE to form a hierarchy of losses).
Fu, Ghosh, NPL1, and Song are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu in view of Ghosh and NPL1 to incorporate the teachings of Song in order to make the multi-task learning model be sequential and have hierarchical weighting. Doing so would be beneficial, as the taught sequential model structure exploits hierarchical supervisions from different levels, allowing for lower level tasks to provide information for higher level tasks (Song, pg. 2, 1st para.).
Regarding claim 18, claim 18 is rejected for analogous reasons to claim 8.
5. Claims 9-10 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Fu in view of Ghosh and NPL1, and further in view of Devlin and Hussein & Shareef (NPL An Empirical Study on the Correlation between Early Stopping Patience and Epochs in Deep Learning, hereinafter Hussein).
Regarding claim 9, Fu in view of Ghosh and NPL1 discloses wherein the plurality of user-defined parameters includes…a feature dimension (Fu, pg. 5, section 4.1, last para. “Word vectors are initialized using GloVe [30] with dimension 300 of BERT [31] with dimension 768…”), a batch size (Ghosh, pg. 13, section A.1, “Dual LSTM and Multi-task Learning experiment:…Particularly we experimented with different mini-batch size (e.g., 8, 16, 32)…”), a dropout rate (Fu, pg. 5, section 4.1, last para. “The experimental results are presented based on the average F1 scores over five runes…The size of hidden units in LSTM is 100 and the dropout rate is 0.3 for dense layers…”)…a number of epochs (Ghosh, pg. 13, section A.1, “Dual LSTM and Multi-task Learning experiment:…Particularly we experimented with different …number of epochs (e.g., 40, 50)”), and a learning rate (Fu, pg. 5, section 4.1, last para. “The learning rate is 0.001 in the GloVe-based model and 0.00001 in the BERT-based model…”).
Fu, Ghosh, and NPL1 are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Ghosh in order to have the plurality of user-defined parameters include a batch size and a number of epochs. Doing so would be beneficial, as a batch size parameter would allow for the model to train on smaller chunks (batches) instead of training on the whole dataset at once which are often too large to process at once, while the epoch parameter allows for multiple passes of the training data to the model, preventing underfitting.
Fu in view of Ghosh and NPL1 does not specifically disclose wherein the plurality of user-defined parameters includes a maximum sequence length.
Devlin teaches wherein the plurality of user-defined parameters includes a maximum sequence length (pg. 13, section A.2 “To generate each training input sequence, we sample two spans of text from the corpus, which we refer to as “sentences”…They are sampled such that the combined length is <= 512 tokens…” pg. 13 2nd Col. “To speed up pretraining in our experiments, we pre-train the model with sequence lgnth of 128 for 90% of the steps. Then, we train the rest 10% of the steps of sequence of 512 to learn the positional embeddings…”).
Fu, Ghosh, NPL1, and Devlin are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu in view of Ghosh and NPL1 to incorporate the teachings of Devlin in order to have the plurality of user-defined parameters include a maximum sequence length. Doing so would be beneficial, as this would enable the use of the taught language model (BERT) to be used, which is a language representation model which can be fine-tuned to create SOTA models for a wide variety of natural language processing tasks (Devlin, Abstract).
Fu in view of Ghosh, NPL1, and Devlin does not specifically disclose wherein the plurality of user-defined parameters includes…a patience parameter.
Hussein teaches wherein the plurality of user-defined parameters includes…a patience parameter (pg. 4, section 4, table contains results of different experiments using different patience values (2, 3, 4, 5 10, 15, 20)).
Fu, Ghosh, NPL1, Devlin, and Hussein are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu in view of Ghosh, NPL1, and Devlin to incorporate the teachings of Hussein in order to have the plurality of user-defined parameters include a patience parameter. Doing so would be beneficial, as patience parameters allow for the use of early stopping, which is a common technique for preventing model overfitting and model memorization of the training data (Hussein, pgs. 1-2, Introduction).
Regarding claim 10, Fu in view of Ghosh and NPL1 discloses …the feature dimension is 786 (Fu, pg. 5, section 4.1, last para. “Word vectors are initialized using GloVe [30] with dimension 300 of BERT [31] with dimension 768…”), the batch size is 32 (Ghosh, pg. 13, section A.1, “Dual LSTM and Multi-task Learning experiment:…Particularly we experimented with different mini-batch size (e.g., 8, 16, 32)…”)…
Fu, Ghosh, and NPL1 are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Ghosh in order to have the plurality of user-defined parameters include a batch size of 32. Doing so would be beneficial, as a batch size parameter would allow for the model to train on smaller chunks (batches) instead of training on the whole dataset at once which are often too large to process at once.
Fu in view of Ghosh and NPL1 does not specifically disclose wherein the maximum sequence length is 128 tokens…the dropout rate is 0.1…and the learning rate is 2e-5.
Devlin teaches wherein the maximum sequence length is 128 tokens (pg. 13 2nd Col. “To speed up pretraining in our experiments, we pre-train the model with sequence length of 128 for 90% of the steps. …”)…the dropout rate is 0.1 (pg. 13, section A.3 “The droupout probability was always kept at 0.1”)…and the learning rate is 2e-5 (pg. 14 “Learning rate (Adam): 5e-5, 3e-5, 2e-5…”).
Fu, Ghosh, NPL1, and Devlin are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu in view of Ghosh and NPL1 to incorporate the teachings of Devlin in order to have maximum sequence length be 128, the dropout rate be 0.1 and the learning rate be 2e-5. Doing so would be beneficial, as the taught maximum sequence length would enable the use of the taught language model (BERT) to be used, which is a language representation model which can be fine-tuned to create SOTA models for a wide variety of natural language processing tasks, while speeding up pretraining (Devlin, Abstract). Further, utilizing the dropout rate would be beneficial, as this rate would provide regularization to prevent overfitting without removing essential features (NPL Das, How Dropout Layers in Neural Networks Work Like Random Forests: A brief overview, pg. 4, section “Practical Considerations: Fine-Tuning Dropout rates”). Further, the taught learning rate would help ensure that BERT would overcome catastrophic forgetting during training (NPL Sun et al., How to Fine-Tune BERT for Text Classification, section 5.3.3).
Fu in view of Ghosh, NPL1, and Devlin does not specifically disclose wherein the patience parameter is 5, the number of epochs is 20.
Hussein teaches the patience parameter is 5 (pg. 3, section 3, 2nd para. “We trained the model for a fixed number of 20 epochs and used early stopping with varying patience values, with patience values of 2, 3, 4, 5, 10, 15, and 20…”), the number of epochs is 20 (pg. 3, section 3, 2nd para. “We trained the model for a fixed number of 20 epochs…”).
Fu, Ghosh, NPL1, Devlin, and Hussein are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Fu in view of Ghosh, NPL1, and Devlin to incorporate the teachings of Hussein in order to have patience parameter be 5 and the number of epochs be 20. Doing so would be beneficial, as this configuration would lead to high validation accuracies (Hussein, Table 4 and section 4, 2nd para.).
Regarding claim 19, claim 19 is rejected for analogous reasons to claim 9.
Regarding claim 20, claim 20 is rejected for analogous reasons to claim 10.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Kulla et al. (US 2025/0384213 A1): text preprocessing operation, followed by sentiment classification (Fig. 1)
Roy et al. (US 2025/0252261 A1): multi-task learning for natural language processing tasks, text tokenizations, embeddings, shared layers, and task specific layers (Fig. 2)
Wang et al. (US 2024/0153495 A1): multi-task learning for ASR and auxiliary tasks (Abstract)
Le et al. (US 2023/0186072 A1): providing explanations for attention-based models, relevance scores providing explanation of token’s contributions towards or against the outcome (Abstract)
Hosseini-Asl & Liu (US 2022/0366145 A1): multi-task generative language model for sentiment analysis (Abstract)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CODY DOUGLAS HUTCHESON/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659