DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is in responsive to communication(s): original application filed on 02/14/2024, said application claims a priority filing date of 02/16/202. Claims pending. Claims 1, 7, and 13 are independent.
Specification
The disclosure is objected to because of the following informalities:
Table 2 shown in ¶ [0075] is different to Table 3 shown in ¶ [0077]; however, the description of Table 2 in ¶ [0074] ("Table 2: F1 comparisons in fields of law, healthcare, and document writing") is the same as the description of Table 3 in ¶ [0076] ("Table 3: F1 comparisons in fields of law, healthcare and document writing"). Clarification between the description of Table 2 and the description of Table 3 is required.
Appropriate correction is required.
Claim Objections
Claims 3-4, 9-10, and 15-16 are objected to because of the following informalities:
In Claim 3, lines 3-4; Claim 9, lines 3-4; and Claim 15, lines 3-4, , "… according to whether an error in the source occurs for the first time for the training model …" appears to be "… according to whether an error in the source occurs for a first time for the training model …";
in Claim 4, lines 13-15; Claim 10, lines 13-15; and Claim 16, lines 13-15, "… I-F1 is a harmonic mean F1 … on those errors having been seen in training" appears to be "… I-F1 is a harmonic mean F1 … on those errors having been seen in the training".
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 7, and 13 recite the limitation "… obtaining a sample label comprising a source and a target … calculating a precision and a recall according to the source, the target, and the prediction result of the sample label; calculating an average precision and an average recall for the precisions and the recalls of a plurality of sample labels …" in lines 2-10, 7-15, and 4-12 respectively, which rendering these claims indefinite because (1) it is unclear whether ".
Claims 2-6, 8-13, and 9-18 are rejected for fully incorporating the deficiency of their respective base claim.
Claims 4, 10, and 16 recite the limitation "… averaging precisions and recalls of the plurality of sample labels in the first class and the second class, respectively, to obtain an average precision and an average recall of the first-class sample labels and an average precision and an average recall of the second-class sample labels; and calculating E-F1 according to the average precision and the average recall of the first-class sample labels, and calculating I-F1 according to the average precision and the average recall of the second-class sample labels, wherein E-F1 is a harmonic mean F1 of the precisions and the recalls of the training model on those errors having not been seen in training, and I-F1 is a harmonic mean F1 of the precisions and the recalls of the training model on those errors having been seen in training" in lines 3-15, which rendering these claims indefinite because ".
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 6-8, 12-14, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over LI et al. (CN 111523306 A, pub. date: 08/11/2020), hereinafter LI in view of He et al., ("EA-MLM: Error-aware Masked Language Modeling for Grammatical Error Correction", 2021 International Conference on Asian Language Processing (IALP), Dec. 11-13, 2021, pp.363-368), hereinafter He.
Independent Claims 1, 7, and 13
LI discloses a text error correction method (LI, ¶¶ [0006] and [0029]-[0043] with FIG. 2: a text correction method comprising: Step S21: obtain sentence information, wherein the sentence information includes at least one of the following feature information: text feature information, pinyin feature information, and stroke feature information; step S23: Use a text correction model to process the sentence information, wherein a neural network model is trained based on the training corpus to obtain the text correction model; Step S25: determine the error correction result of the sentence based on the processing result of the text error correction model), comprising:
obtaining a sample label comprising a source and a target, and confusion characters to the source in the sample label to obtain the source with the confusion characters (LI, ¶¶ [0037]-[0040]: the training corpus mentioned above can be obtained from a confusion set, which is a collection of similar words; similar words refer to words that sound similar or have similar strokes; a confusion set of Chinese characters can be obtained by changing the pinyin of the Chinese characters, and the confusion set includes homophones; alternatively, a confusion set of Chinese characters can be obtained by changing the strokes of the Chinese characters, and the confusion set includes characters with similar shapes; the training corpus can be obtained from historical dialogue records; after obtaining the correct statement from the dialogue record, the incorrect statement corresponding to the correct statement is obtained by transforming the correct statement; each erroneous statement and a correct statement can form a training corpus; ¶¶ [0091]-[0094] and [0098]: obtaining training corpus based on a confusion set, wherein the training corpus includes correct text and erroneous text corresponding to the correct text obtained from the confusion set; obtaining sentence information of the erroneous text; since the confusion set includes a collection of characters with similar shapes or sounds, after determining the correct text, the incorrect text of the correct text can be obtained based on the confusion set of each character in the correct text, thus obtaining at least one set of training corpus; obtaining test corpus; the test corpus can be obtained through manual annotation; e.g., the incorrect text and the correct text corresponding to the incorrect text can be manually annotated to obtain the test corpus of <incorrect text, correct text>; the test corpus can also be obtained from the training corpus; e.g., the training corpus can be sampled to generate test data such as <"What is your Weixin number", "What is your WeChat number">, which is used for the evaluation of the trained model; ¶¶ [0107]-[0117]: obtaining training corpus based on the confusion set includes: obtaining correct text; obtaining a confusion set corresponding to at least one character in the text, wherein the confusion set includes characters whose edit distance from the at least one character's pinyin and/or strokes is less than a preset value; replacing at least one character with other characters in the confusion set that are different from at least one character to obtain incorrect text; the training corpus includes <correct text, incorrect text>; the correct text is the preset correct text, and the incorrect text is the text obtained by replacing one or more characters in the correct text with other characters in its confusion set; after obtaining the correct text, obtain the misspelled words corresponding to one or more characters in the correct text from the preset confusion set, and replace one or more characters in the correct text with the misspelled words from the confusion set, thereby obtaining one or more incorrect texts corresponding to the correct text, which can then constitute at least one set of training corpus; the edit distance mentioned above is used to represent the number of operations required to change one pinyin to another, or one stroke to another; the smaller the edit distance between two characters, the more similar the two characters are; performing at least one of the following processing on the pinyin or strokes of at least one character to obtain a confusion set corresponding to at least one character: adding, deleting, changing, and swapping; there are four editing methods: adding, deleting, changing, and swapping. By performing any of these operations on the pinyin or strokes of the text, the corresponding confusion set can be obtained; uses the confusion set of correct text in terms of pinyin and strokes to construct the training corpus, which greatly expands the training corpus for error correction; it also has the sequence-to-sequence fitting ability of neural networks, which can effectively solve the error correction of homophones, near-homophones and near-similar characters; ¶¶ [0120]-[0133] with S41-S47 in FIG. 4: S41, Obtain training corpus; the corresponding training data can be obtained according to the scenario required for error correction; the training data is in the format of a text sequence; e.g., for keyword and product error correction, it can be obtained from product search keyword logs or product databases; for dialogue and Q&A error correction, it can be obtained from user dialogue and Q&A history logs; for document proofreading error correction, it can be obtained from document databases; S42, extract the pinyin sequence and stroke sequence; S43, construct an embedding dictionary for pinyin and stroke order; S44, based on edit distance, construct confusion sets for similar pinyin and strokes respectively; S45 converts sequences of text, pinyin, and strokes into the model input format; S46, construct training pseudo-corpus; S47, Build test data; high-quality test data can be generated based on user history logs, using manually labeled data or by sampling and labeling pseudo-corpora, such as <“What is your WeChat ID?”, “What is your WeChat ID?”>, for evaluating the trained model);
inputting the source with the confusion characters into a training model to obtain a prediction result (LI, ¶¶ [0036]-[0040] with S23 in FIG. 2: Step S23: Use a text correction model to process the sentence information, wherein a neural network model is trained based on the training corpus to obtain the text correction model; the text correction model is a neural network model obtained through training; ¶¶ [0091], [0095], and [0102]-[0106]: training an initial neural network model based on the sentence information of the erroneous text and the correct text to obtain the text correction model; erroneous texts from the test corpus can be input into the training results, the training results can be used to correct the erroneous texts; when the sentence information of the erroneous text includes textual feature information, pinyin feature information, and stroke feature information of the erroneous text, a neural network model is trained based on the sentence information of the erroneous text and the correct text to obtain a text correction model; this includes: concatenating the textual feature information, pinyin feature information, and stroke feature information of the erroneous text; inputting the concatenation result into the encoder of the neural network model for encoding; inputting the encoding result into the decoder of the neural network model, wherein the decoder includes an attention mechanism for the encoder; obtaining the error between the decoding result of the decoder and the correct text corresponding to the erroneous text; adjusting the parameters of the neural network model according to the error until the error meets a preset condition, and determining the neural network model whose error meets the preset condition as the text correction model; the neural network model is trained by backpropagation (BP); in each iteration, the weights and biases in the neural network model are updated in a predetermined manner so that the output of the neural network model is closer to the expectation, thereby obtaining a text correction model; trains the text correction model by fusing multiple feature information of the sentences to be corrected, thereby realizing the training of a multi-granularity fusion neural network model: it can be trained using a Seq2Seq-based NMT (Neural Machine Translation) neural machine translation model, with training corpus as the training input; during model training, the model training is complete when the metric on the validation set (which can be the F-value) stops decreasing; the aforementioned neural network is an end-to-end model training method, which can directly train the model with training corpus and obtain high-quality error correction results through loss function optimization; ¶¶ [0134]-[0137] with S46-S49 in FIG. 4: S48, multi-granularity fusion neural network training; The Seq2Seq NMT (Neural Machine Translation) neural machine translation model can be used for training; the training input is the training pseudo-corpus constructed in step S46; the model can use an Encoder-Decoder architecture, with the Encoder and Decoder each using a multi-layer Bi-LSTM model, and an Attention mechanism for the Encoder added to the Decoder; during model training, the model training is complete when the metrics on the validation set stop decreasing; S49, offline model training; the effectiveness of the text correction model can be evaluated based on the test set constructed by S47);
calculating a precision and a recall according to the source, the target, and the prediction result of the sample label; calculating an average precision and an average recall for the precisions and the recalls of a plurality of sample labels, and calculating a harmonic mean F1 of the precisions and the recalls according to the average precision and the average recall; and adjusting the training model according to the harmonic mean F1 and taking the adjusted training model as a text error correction model (LI, ¶¶ [0094]-[0100]: verifying the evaluation parameters of the trained text correction model through the test corpus, wherein the evaluation parameters include one or more of the following: accuracy, recall, and harmonic mean; if the evaluation parameters are higher than a preset parameter threshold, then the trained text correction model is allowed to correct sentences; the accuracy, recall, and harmonic mean of the training results can be determined based on the error correction results of the training results; the aforementioned precision is used to represent the ratio of the test data with accurate test results to the test data of the input text error correction model, recall is used to represent the ratio of the test data with accurate test results to all test data, and the harmonic mean is used to represent the mean of precision and recall; the aforementioned test corpus is used to test the accuracy of the trained model; if the evaluation parameters of the trained model exceed the preset threshold, it indicates that the model has been successfully trained and can be used as a text correction model; if the evaluation parameters of the trained model do not exceed the preset threshold, it indicates that the accuracy of the model is low; even if it is used as a text correction model, it is difficult to obtain accurate correction results; therefore, further training is needed to further correct the network parameters and improve the accuracy of the model; the evaluation parameters mentioned above include one or more of precision, recall, and harmonic mean; if only one is included, for example, the evaluation parameter is precision, then the evaluation parameter threshold is the precision threshold; the precision is compared with the precision threshold; if the precision is greater than the precision threshold, the training result is determined to be a text correction model, which can be used to correct sentences; if the evaluation parameters include accuracy, recall, and harmonic mean, then weights can be set for accuracy, recall, and harmonic mean, and a weighted average can be calculated based on the weights of accuracy, recall, and harmonic mean from the training results; the resulting weighted average is the evaluation parameter; then, the evaluation parameters of the training results are compared with the evaluation parameter threshold; if the evaluation parameter is less than or equal to the evaluation parameter threshold, the training results need to be trained again; ¶¶ [0137]-[0141] with S410-S411 and S46 FIG. 4: the evaluation method is to calculate the matching degree between the model's predicted output and the test set results, and calculate the accuracy, recall and F-score at the character or word level according to the granularity of the training data used by the model; S410 determines whether the test result is higher than the target; if the test result is higher than the target, proceed to step S411; otherwise, proceed to step S46; different levels of accuracy can be determined according to the needs of different businesses, and the model can be judged to meet the requirements based on the determined accuracy; if the model performs better than the target on the test set, it means that the trained model can be used; S411 produces a text error correction model; save the model with the test results labeled, and use the model prediction module for online error correction later).
LI further discloses an electronic device (LI, ¶ [0023] with 10 in FIG. 1: a computer terminal 10 (or mobile device 10)), comprising: a memory (LI, ¶ [0023] with 104 in FIG. 1: a memory 104for storing data); and a processor (LI, ¶¶ [0023]-[0024] with 102 in FIG. 1:one or more processors102 (shown as 102a, 102b, …, 102n in the figure) (processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) coupled to the memory, the memory having therein stored instructions which, when executed by the processor, cause the electronic device to perform a text error correction method described above (LI, ¶¶ [0023]-[0025]: a hardware block diagram of a computer terminal (or mobile device) for implementing a text error correction method; the memory 104 can be used to store software programs and modules of application software, such as the program instructions/data storage device corresponding to the text error correction method; the processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, to implement the above-mentioned application vulnerability detection method).
LI fails to explicitly disclose randomly adding a mask to the source in the sample label to obtain the source with the mask; and inputting the source with the mask into a training model.
He teaches a system and a method relating to text error correction (He, 1st paragraph of Section I in Page 363), wherein randomly adding a mask to the source in the sample label to obtain the source with the mask; and inputting the source with the mask into a training model (He, Abstract of Page 363: in recent years, BERT has been used in the task of grammatical error correction (GEC) and achieved good performance; however, few previous studies have investigated the incorporation of real grammatical errors into BERT for the GEC task; we argue that the distribution of GEC data (containing several types of errors) is different from the distribution of BERT pre-training data (usually error-free); to fill this gap, we extend masked language modeling and propose a novel error-aware masked language modeling strategy (EA-MLM) to fine-tune BERT so that the representation distribution of the pretrained BERT is better adapted to the GEC task; Section I with FIG. 1 of Page 363-364: Grammatical Error Correction (GEC) is the task of automatically correcting different kinds of errors in text, such as spelling, grammar, and word choice errors; GEC is typically formulated as a sentence correction task, which takes a potentially erroneous sentence as input and converts it into its correct form; we believe that the data distribution of the GEC task is very different from that of the pre-trained BERT, because the text used in the GEC task contain several types of errors, such as grammatical and spelling errors; therefore, it is critical to fine-tune the pre-trained BERT with GEC data to make it more suitable for the GEC task; in existing works, Kaneko et al. make an attempt with similar ideas; they introduce GEC knowledge to BERT by finetuning it through masked language modeling and error detection with annotated GEC data; their method enables BERT to learn the knowledge about whether a sentence contains grammatical errors; however, an erroneous sentence usually contains several types of errors; these fine-grained errors can be leveraged to further improve the adaptability of BERT to the GEC task; inspired by this, we propose a novel error-aware masked language modeling strategy (EA-MLM) to fine-tune BERT, which enables BERT to be aware of the grammatical errors in text and enhances its ability of error detection, error type recognition and even error correction; in our proposed EA-MLM strategy, we extend the masked language modeling (MLM) strategy and design several well tailored error-aware masking schemes to better transfer the general language knowledge of the pre-trained BERT to the GEC task; as shown in Fig. 1, we list the sentences masked by the vanilla MLM and our proposed EA-MLM respectively; compared with MLM, the sentence masked by EA-MLM is more likely to be made by English learners and is more beneficial for BERT to learn to be error-aware; the main contributions of our work are as follow: (a) we propose a novel error-aware masked language modeling strategy (EA-MLM) to fine-tune BERT for the GEC task; to the best of our knowledge, this is the first study to leverage the fine-grained errors for BERT fine-tuning and use it for the GEC model; (b) we conduct extensive experiments and compare our EA-MLM strategy with the vanilla MLM strategy as well as some strong GEC models; Section II with FIGS. 2-3 of Pages 364-365: the overall architecture of our method consists of two stages: the BERT fine-tuning stage and the correction model training stage, which is shown in Fig. 2; in the first stage, we fine-tune the pre-trained BERT using our proposed EA-MLM which contains several error-aware masking schemes and obtain BERTfine-tuned, which is supposed to be aware of errors with its ability of error detection enhanced; in the second stage, we fuse the output representations of BERTfine-tuned into Transformer using the BERT-fused mechanism for correction model training; specifically, assume that the source sentence is X; first, BERTfine-tuned receives X and produces its BERT representations XB; then, the BERT-fused Transformer encodes X, additionally incorporates XB with a BERT-attention module in each layer, and outputs XE; the BERT-fused Transformer decoder receives XE and additionally incorporates XB with BERT-attention modules, producing the final correct sentence
Y
^
; in order to make BERTfine-tuned be aware of more types of errors, we propose an error-aware masked language modeling (EA-MLM) finetuning strategy, which includes several error-aware masking schemes; assume that the input sentence of EA-MLM is X = (x1, x2, …, xm), where xi is the ith word; and the set of the error-aware masking schemes is F = { f1, f2, …, fn}, where fj is the jth masking scheme; for each word xi to be masked, a masking scheme fj is randomly chosen from the set of the error-aware masking schemes F; the chosen scheme fj is then applied to the word xi and produces the masked version
x
^
i
, as in Equation (1); the masked version
x
^
i
may be different in length with xi due to BERT tokenization; in this case we discard the current masked version
x
^
i
and randomly re-select a masking scheme to corrupt xi; after randomly masking some words in X, we obtain the masked sentence
Y
^
; BERT receives the masked
Y
^
and is asked to predict the masked words; by predicting the masked words, BERT learns to recover the wrong (masked) sentences into their correct forms, enabling it to be more error-aware; we give each masking scheme the same probability to be chosen rather than assigning a specific probability to each scheme because not all schemes are applicable to a certain word; the comparison shows that our proposed EA-MLM can generate masked words that are more similar to real errors; we carefully develop 6 error-aware masking schemes to mask words in the input sentence for BERT fine-tuning; Real Error Masking: some of the GEC corpora, such as the FCE corpus and the Cambridge English Write & Improve dataset, contain high-quality of error annotations made by humans; we first extract all of the man-made grammatical error patterns from the error annotation files (e.g. M2 files) and represent each pattern in the form of {worderror : wordcorrect} key-value pair; then, we exchange the key and the value of each pair and obtain {wordcorrect : worderror} which will be used for the real error masking scheme; the intuition is that we replace the correct word with an erroneous word; receiving a word xi, we randomly look for a pair whose key is xi, and use the value of the pair as the masked version
x
^
i
; Synonym Masking: the English learners tend to mistake words for their synonyms; we propose the synonym masking scheme; receiving a word xi to be masked, the scheme replaces it with one of its synonyms as the masked result
x
^
i
; Inflection Masking: inflection is one of the features of the English language, including noun declension such as "cat" declining to "cats", and verb conjugation such as "be" conjugating to "is"; inflection-related errors are frequently made by the English learners; we take the inflection phenomenon into account and propose the inflection masking scheme; the input word xi of the scheme will be replaced by one of its inflection variants; Functional Word Masking: functional words such as prepositions, articles and pronouns perform certain functions in a sentence, and they are easy to be misused by the English learners; we believe that replacing a functional word with another one of the same functional type can mimic the confusion behavior of the English learners; therefore, we propose the functional word masking scheme; Case Masking: case errors are common errors that are frequently made by the English learners, most of which are concern about proper nouns such as personal names, country names and so on; we propose the case masking scheme to mimic the case misuse errors; Misspell Masking: misspelling errors account for a huge part of the grammatical errors in writings of the English learners and even native speakers; people may mistake a word with another one that looks similar to it, or simply write a word in a wrong way; we propose the misspell masking scheme and mimic such behavior by replacing the input word xi with another one that looks similar to it; Section III-IV with FIG. 4 of Pages 365-367: for the first stage, we fine-tune the BERT-base-cased model provided by Huggingface transformers; for the error-aware masking schemes, we utilize WordNet to obtain the synonyms of a word, word_forms to obtain all inflection forms of a word, and the Python standard library difflib to generate the similar words of each word; for the second stage, we use bert-nmt [5] as the BERT-fused architecture; we use the gec-pseudodata pre-trained model to initialize the weights of the transformer correction model; for EA-MLM BERT Fine-tuning, we fine-tune the BERT-base-cased model for 3 epochs with the AdamW optimizer; the batch size is set to 32, the max sentence length is 128, and the learning rate is set to 4e-5; for correction model training, we use the Transformer (big) architecture; we train the model for 30 epochs with the Adam optimizer; the maximum length of tokens is set to 4096, the learning rate is 3e-5; we use Label smoothed cross-entropy as the loss function with the smoothing factor being 0.1; we set the dropout as 3e-1; and the gradient clipping is set to 0.1; in inference, the beam size is set to 5; we use the MaxMatch (M2) scorer for evaluating the FCE-test and CoNLL-2014 test sets, the ERRANT evaluation metric for the W&I+L-test set, and report the precision, recall and F0.5 score; we use the GLEU metric for reporting the GLEU score on the JFLEG test set; we train 3 variants with different random seeds and take the average result for report; we observe that the precision of the BERT-fused EA-MLM model is the highest, indicating that EA-MLM indeed strengthens the ability of the model to detect errors, recognize error types and correct errors).
LI and He are analogous art because they are from the same field of endeavor, a system and a method relating to text error correctio. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of He to LI. Motivation for doing so would improve adaptation of the pre-trained BERT to the Grammatical Error Correction.
Claims 2, 8, and 14
LI in view of He discloses all the elements as stated in Claims 1, 7, 13 and further discloses wherein the harmonic mean F1=2*(precision*recall)/(precision + recall) (LI, ¶¶ [0094]-[0100]: verifying the evaluation parameters of the trained text correction model through the test corpus, wherein the evaluation parameters include one or more of the following: accuracy, recall, and harmonic mean; if the evaluation parameters are higher than a preset parameter threshold, then the trained text correction model is allowed to correct sentences; the accuracy, recall, and harmonic mean of the training results can be determined based on the error correction results of the training results; the aforementioned precision is used to represent the ratio of the test data with accurate test results to the test data of the input text error correction model, recall is used to represent the ratio of the test data with accurate test results to all test data, and the harmonic mean is used to represent the mean of precision and recall (see also in "en.wikipedia.org/wiki/F-score" that "the traditional F-measure or balanced F-score (F1 score) is the harmonic mean of precision and recall:
PNG
media_image1.png
82
648
media_image1.png
Greyscale
With precision = TP / (TP + FP) and recall = TP / (TP + FN), it follows that the numerator of F1 is the sum of their numerators and the denominator of F1 is the sum of their denominators"); the aforementioned test corpus is used to test the accuracy of the trained model; if the evaluation parameters of the trained model exceed the preset threshold, it indicates that the model has been successfully trained and can be used as a text correction model; if the evaluation parameters of the trained model do not exceed the preset threshold, it indicates that the accuracy of the model is low; even if it is used as a text correction model, it is difficult to obtain accurate correction results; therefore, further training is needed to further correct the network parameters and improve the accuracy of the model; the evaluation parameters mentioned above include one or more of precision, recall, and harmonic mean; if only one is included, for example, the evaluation parameter is precision, then the evaluation parameter threshold is the precision threshold; the precision is compared with the precision threshold; if the precision is greater than the precision threshold, the training result is determined to be a text correction model, which can be used to correct sentences; if the evaluation parameters include accuracy, recall, and harmonic mean, then weights can be set for accuracy, recall, and harmonic mean, and a weighted average can be calculated based on the weights of accuracy, recall, and harmonic mean from the training results; the resulting weighted average is the evaluation parameter; then, the evaluation parameters of the training results are compared with the evaluation parameter threshold; if the evaluation parameter is less than or equal to the evaluation parameter threshold, the training results need to be trained again; ¶¶ [0137]-[0141] with S410-S411 and S46 FIG. 4: the evaluation method is to calculate the matching degree between the model's predicted output and the test set results, and calculate the accuracy, recall and F-score at the character or word level according to the granularity of the training data used by the model; S410 determines whether the test result is higher than the target; if the test result is higher than the target, proceed to step S411; otherwise, proceed to step S46; different levels of accuracy can be determined according to the needs of different businesses, and the model can be judged to meet the requirements based on the determined accuracy; if the model performs better than the target on the test set, it means that the trained model can be used; S411 produces a text error correction model; save the model with the test results labeled, and use the model prediction module for online error correction later).
Claims 6, 12, and 18
LI in view of He discloses all the elements as stated in Claims 1, 7, 13 and further discloses inputting a source to be corrected into the text error correction model to obtain a corrected target (LI, ¶¶ [0059]-[0062] and [0081]-[0100] with S36 in FIG. 3: the processing result of the text correction model includes: a predicted correction result and a confidence level corresponding to the predicted correction result; determining the correction result of a statement based on the processing result of the text correction model includes: obtaining a confidence threshold; if the confidence level of the predicted correction result is greater than the confidence threshold, the predicted correction result is determined as the correction result of the statement to be corrected; if the confidence level of the predicted correction result is less than or equal to the confidence threshold, the statement to be corrected is determined as the correction result; since the predicted error correction results output by the text error correction model are not completely accurate, it is necessary to judge the accuracy of the predicted error correction results after the text error correction model outputs the predicted error correction results; the above scheme judges its accuracy by predicting the confidence level of the error correction result; the confidence threshold mentioned above is used to determine whether the prediction results output by the current text correction model are reliable; the text correction model outputs the predicted correction result and the corresponding confidence level, and judges whether the predicted correction result is usable based on the confidence level; specifically, the confidence level mentioned above can be in the range (0, 1), which is used to represent the credibility of the prediction and error correction results; the confidence threshold can be set to 0.9; if the confidence level of the text correction model's prediction of the correction result for a sentence to be corrected is 0.97, which is greater than the preset confidence threshold of 0.9, then the prediction correction result is considered reliable and can be used as the correction result of the sentence; S36 is predicted using a neural network text correction model; if the confidence of the prediction and correction result is greater than the confidence threshold, the prediction and correction result is considered reliable; if the prediction and error correction result is deemed unreliable, the text entered by the user is considered correct, and the original text is returned; S310 returns the error correction result and confidence level; S311 outputs the error correction results and confidence scores to the downstream task model; the aforementioned test corpus is used to test the accuracy of the trained model; if the evaluation parameters of the trained model exceed the preset threshold, it indicates that the model has been successfully trained and can be used as a text correction model; the precision is compared with the precision threshold; if the precision is greater than the precision threshold, the training result is determined to be a text correction model, which can be used to correct sentences; ¶ [0141]: use the model prediction module for online error correction later).
Claims 3, 5, 9, 11, 15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over LI in view of He as applied to Claims 1, 7, 13 respectively above, and further in view of HUI et al. (WO 2021/189851 A1, pub. date: 09/30/2021), hereinafter HUI.
Claims 3, 9, and 15
LI in view of He discloses all the elements as stated in Claims 1, 7, 13 and except failing to explicitly disclose further discloses according to whether an error in the source occurs for the first time for the training model, classifying the plurality of sample labels into two classes, a first class being sample labels where the error in the source occurs for the first time for the training model, and a second class being sample labels where the error in the source does not occur for the first time for the training model.
HUI teaches a system and a method relating to text error corrections (HUI, Abstract), according to whether an error in the source occurs for the first time for the training model, classifying the plurality of sample labels into two classes, a first class being sample labels where the error in the source occurs for the first time for the training model, and a second class being sample labels where the error in the source does not occur for the first time for the training model (HUI, ¶¶ [0057]-[0069] with FIG. 2: before step S10, the method further includes: Step A1: obtain labeled training data, which includes sentences without erroneous words, sentences with erroneous words, and the correct sentences corresponding to the sentences with erroneous words; Step A2: based on the labeled training data, perform FINE-TUNE fine-tuning on the BERT-based pre-trained language model to obtain a BERT-based masked language model; the BERT-based masked language model is obtained by fine-tuning the parameters of a BERT-based pre-trained language model using labeled training data, where the labeled training data is text data related to the business scenario, and different business scenarios may have different labeled training data; furthermore, Step A2 described above includes: the sentences in the labeled training data that do not contain erroneous words are masked according to a preset BERT masking method to obtain the first mask data, and the predicted word of the masked word is set as the word before the mask (i.e., errors introduced into training data for the first time in the training data via masking as the first training data); the sentences in the labeled training data that contain erroneous words are masked in the original words to obtain the second mask data, and the predicted word of the masked word is set as the corresponding correct word (i.e., errors has already been seen in the training data before masking as the second training data); based on the first mask data, the second mask data, and their respective predicted words, the BERT-based pre-trained language model is fine-tuned to obtain a BERT-based masked language model; the labeled training data include sentences without erroneous words, which can be used as the first training data; the first training data are masked according to a preset BERT masking method, where the preset BERT masking method refers to masking a preset proportion/ratio of words in the first training data to obtain the first mask data; the first mask data is also associated with the corresponding correct word, i.e., the predicted word; the predicted word of the first mask data is itself; the specific masking method is as follows: 80% of the preset proportion/ratio of words in the first training data are masked with [MASK] so that the model can predict the masked words in the text through context before and after, and learn cloze tests (i.e., learn how to fill the missing words); 10% of the preset proportion/ratio of words in the first training data are masked with random words so that the model can learn how to correct incorrect words; and 10% of the preset proportion/ratio of words in the first training data are preserved in their original form so that the model can learn to detect whether a word is incorrect (i.e., based on the ratio of 8: 1: 1 replacing the as [mask], random other words, the original word among the preset proportion/ratio of words to be masked); the preset ratio is less than or equal to 20%, for example, it can be selected as 10%, 15%, or 20%; the labeled training data also includes sentences containing erroneous words, which can be used as the second training data; the erroneous words in the second training data are masked by preserving the original words to obtain the second mask data; the second mask data is also associated with the corresponding correct words, i.e., the predicted words; after obtaining the first mask data, the second mask data, and their corresponding predicted words, these data are input into a BERT-based pre-trained language model to train the pre-trained language model, thus obtaining a BERT-based masked language model; furthermore, to further prevent overfitting, some correct words in the second training data can also be masked to obtain third mask data; the third mask data is also associated with the corresponding predicted word, i.e., the word itself, where the proportion/ratio of masking the correct words in the second training data can be the same as the proportion/ratio of masking the incorrect words in the second training data; correspondingly, after obtaining the first mask data, the second mask data, the third mask data, and their respective predicted words, these data are input into the BERT-based pre-trained language model to train the pre-trained language model, thus obtaining the BERT-based masked language model; this embodiment uses a pre-trained language model that has already been pre-trained using a large number of normal samples; only a small amount of business-related training data is needed to fine-tune the pre-trained language model to obtain a BERT-based masked language model, thereby avoiding the overfitting problem caused by insufficient parallel corpora for Chinese text correction in existing technologies)
LI in view of He, and HUI are analogous art because they are from the same field of endeavor, a system and a method relating to text error corrections. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of HUI to LI in view of He. Motivation for doing so would prevent the overfitting problem caused by insufficient parallel corpora (HUI, ¶ [0069]).
Claims 5, 11, and 17
LI in view of He discloses all the elements as stated in Claims 1, 7, 13 respectively and further discloses wherein (He, Section II with FIGS. 2-3 of Pages 364-365: the overall architecture of our method consists of two stages: the BERT fine-tuning stage and the correction model training stage, which is shown in Fig. 2; in the first stage, we fine-tune the pre-trained BERT using our proposed EA-MLM which contains several error-aware masking schemes and obtain BERTfine-tuned, which is supposed to be aware of errors with its ability of error detection enhanced; in the second stage, we fuse the output representations of BERTfine-tuned into Transformer using the BERT-fused mechanism for correction model training; specifically, assume that the source sentence is X; first, BERTfine-tuned receives X and produces its BERT representations XB; then, the BERT-fused Transformer encodes X, additionally incorporates XB with a BERT-attention module in each layer, and outputs XE; the BERT-fused Transformer decoder receives XE and additionally incorporates XB with BERT-attention modules, producing the final correct sentence
Y
^
; in order to make BERTfine-tuned be aware of more types of errors, we propose an error-aware masked language modeling (EA-MLM) finetuning strategy, which includes several error-aware masking schemes; assume that the input sentence of EA-MLM is X = (x1, x2, …, xm), where xi is the ith word; and the set of the error-aware masking schemes is F = { f1, f2, …, fn}, where fj is the jth masking scheme; for each word xi to be masked, a masking scheme fj is randomly chosen from the set of the error-aware masking schemes F; the chosen scheme fj is then applied to the word xi and produces the masked version
x
^
i
, as in Equation (1); the masked version
x
^
i
may be different in length with xi due to BERT tokenization; in this case we discard the current masked version
x
^
i
and randomly re-select a masking scheme to corrupt xi; after randomly masking some words in X, we obtain the masked sentence
Y
^
; BERT receives the masked
Y
^
and is asked to predict the masked words; by predicting the masked words, BERT learns to recover the wrong (masked) sentences into their correct forms, enabling it to be more error-aware; we give each masking scheme the same probability to be chosen rather than assigning a specific probability to each scheme because not all schemes are applicable to a certain word; the comparison shows that our proposed EA-MLM can generate masked words that are more similar to real errors; we carefully develop 6 error-aware masking schemes to mask words in the input sentence for BERT fine-tuning; Real Error Masking: some of the GEC corpora, such as the FCE corpus and the Cambridge English Write & Improve dataset, contain high-quality of error annotations made by humans; we first extract all of the man-made grammatical error patterns from the error annotation files (e.g. M2 files) and represent each pattern in the form of {worderror : wordcorrect} key-value pair; then, we exchange the key and the value of each pair and obtain {wordcorrect : worderror} which will be used for the real error masking scheme; the intuition is that we replace the correct word with an erroneous word; receiving a word xi, we randomly look for a pair whose key is xi, and use the value of the pair as the masked version
x
^
i
; Synonym Masking: the English learners tend to mistake words for their synonyms; we propose the synonym masking scheme; receiving a word xi to be masked, the scheme replaces it with one of its synonyms as the masked result
x
^
i
; Inflection Masking: inflection is one of the features of the English language, including noun declension such as "cat" declining to "cats", and verb conjugation such as "be" conjugating to "is"; inflection-related errors are frequently made by the English learners; we take the inflection phenomenon into account and propose the inflection masking scheme; the input word xi of the scheme will be replaced by one of its inflection variants; Functional Word Masking: functional words such as prepositions, articles and pronouns perform certain functions in a sentence, and they are easy to be misused by the English learners; we believe that replacing a functional word with another one of the same functional type can mimic the confusion behavior of the English learners; therefore, we propose the functional word masking scheme; Case Masking: case errors are common errors that are frequently made by the English learners, most of which are concern about proper nouns such as personal names, country names and so on; we propose the case masking scheme to mimic the case misuse errors; Misspell Masking: misspelling errors account for a huge part of the grammatical errors in writings of the English learners and even native speakers; people may mistake a word with another one that looks similar to it, or simply write a word in a wrong way; we propose the misspell masking scheme and mimic such behavior by replacing the input word xi with another one that looks similar to it).
LI in view of He fails to explicitly disclose wherein a ratio of randomly adding the mask to the source is 0.1 to 0.3.
HUI teaches a system and a method relating to text error corrections (HUI, Abstract), wherein a ratio of randomly adding the mask to the source is 0.1 to 0.3 (HUI, ¶¶ [0057]-[0069] with FIG. 2: before step S10, the method further includes: Step A1: obtain labeled training data, which includes sentences without erroneous words, sentences with erroneous words, and the correct sentences corresponding to the sentences with erroneous words; Step A2: based on the labeled training data, perform FINE-TUNE fine-tuning on the BERT-based pre-trained language model to obtain a BERT-based masked language model; the BERT-based masked language model is obtained by fine-tuning the parameters of a BERT-based pre-trained language model using labeled training data, where the labeled training data is text data related to the business scenario, and different business scenarios may have different labeled training data; furthermore, Step A2 described above includes: the sentences in the labeled training data that do not contain erroneous words are masked according to a preset BERT masking method to obtain the first mask data, and the predicted word of the masked word is set as the word before the mask; the sentences in the labeled training data that contain erroneous words are masked in the original words to obtain the second mask data, and the predicted word of the masked word is set as the corresponding correct word; based on the first mask data, the second mask data, and their respective predicted words, the BERT-based pre-trained language model is fine-tuned to obtain a BERT-based masked language model; the labeled training data include sentences without erroneous words, which can be used as the first training data; the first training data are masked according to a preset BERT masking method, where the preset BERT masking method refers to masking a preset proportion/ratio of words in the first training data to obtain the first mask data; the first mask data is also associated with the corresponding correct word, i.e., the predicted word; the predicted word of the first mask data is itself; the specific masking method is as follows: 80% of the preset proportion/ratio of words in the first training data are masked with [MASK] so that the model can predict the masked words in the text through context before and after, and learn cloze tests (i.e., learn how to fill the missing words); 10% of the preset proportion/ratio of words in the first training data are masked with random words so that the model can learn how to correct incorrect words; and 10% of the preset proportion/ratio of words in the first training data are preserved in their original form so that the model can learn to detect whether a word is incorrect (i.e., based on the ratio of 8: 1: 1 replacing the as [mask], random other words, the original word among the preset proportion/ratio of words to be masked); the preset ratio is less than or equal to 20%, for example, it can be selected as 10%, 15%, or 20%; the labeled training data also includes sentences containing erroneous words, which can be used as the second training data; the erroneous words in the second training data are masked by preserving the original words to obtain the second mask data; the second mask data is also associated with the corresponding correct words, i.e., the predicted words; after obtaining the first mask data, the second mask data, and their corresponding predicted words, these data are input into a BERT-based pre-trained language model to train the pre-trained language model, thus obtaining a BERT-based masked language model; furthermore, to further prevent overfitting, some correct words in the second training data can also be masked to obtain third mask data; the third mask data is also associated with the corresponding predicted word, i.e., the word itself, where the proportion/ratio of masking the correct words in the second training data can be the same as the proportion/ratio of masking the incorrect words in the second training data; correspondingly, after obtaining the first mask data, the second mask data, the third mask data, and their respective predicted words, these data are input into the BERT-based pre-trained language model to train the pre-trained language model, thus obtaining the BERT-based masked language model; this embodiment uses a pre-trained language model that has already been pre-trained using a large number of normal samples; only a small amount of business-related training data is needed to fine-tune the pre-trained language model to obtain a BERT-based masked language model, thereby avoiding the overfitting problem caused by insufficient parallel corpora for Chinese text correction in existing technologies)
LI in view of He, and HUI are analogous art because they are from the same field of endeavor, a system and a method relating to text error corrections. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of HUI to LI in view of He. Motivation for doing so would prevent the overfitting problem caused by insufficient parallel corpora (HUI,.
Allowable Subject Matter
Claims 4, 10, and 16 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Claims 4, 10, and 16
LI in view of He and HUI discloses all the elements as stated in Claims 3, 9, 15.
Dahlmeier et al. (US 2013/0325442 A1, pub. Date: 12/05/2013) discloses in Abstract and ¶¶ [0015]-[] that (1) systems and methods for automated text correction through analysis according to a single text correction model, wherein the single text correction model may be generated through analysis of both a corpus of learner text and a corpus of non-learner text; (2) Grammatical error correction (GEC) has also been recognized as an interesting and commercially attractive problem in natural language processing (NLP); (3) despite the growing interest, research has been hindered by the lack of a large annotated corpus of learner text that is available for research purposes; (3) as a result, the standard approach to GEC has been to train an off-the-shelf classifier to re-predict words in non-learner text; (4) learning GEC models directly from annotated learner corpora is not well explored, as are methods that combine learner and non-learner text; (4) the de facto standard approach to GEC is to build a statistical model that can choose the most likely correction from a confusion set of possible correction choices; (5) the way the confusion set is defined depends on the type of error; (6) work in context-sensitive spelling error correction has traditionally focused on confusion sets with similar spelling (e.g., { dessert, desert}) or similar pronunciation (e.g., {there, their}); (7) in other words, the words in a confusion set are deemed confusable because of orthographic or phonetic similarity; (8) other work in GEC has defined the confusion sets based on syntactic similarity, for example all English articles or the most frequent English prepositions form a confusion set; (9) receiving a natural language text input, the text input comprising a grammatical error in which a portion of the input text comprises a class from a set of classes; (10) generating a plurality of selection tasks from a corpus of non-learner text that is assumed to be free of grammatical errors, wherein for each selection task a classifier re-predicts a class used in the non-learner text; (11) generating a plurality of correction tasks from a corpus of learner text, wherein for each correction task a classifier proposes a class used in the learner text; (12) additionally, training a grammar correction model using a set of binary classification problems that include the plurality of selection tasks and the plurality of correction tasks; (13) using the trained grammar correction model to predict a class for the text input from the set of possible classes; (14) the non-learner text and the learner text have a different feature space, the feature space of the learner text including the word used by a writer; (15) training the grammar correction model may include minimizing a loss function on the training data; (16) training the grammar correction model may also include identifying a plurality of linear classifiers through analysis of the non-learner text; (17) the linear classifiers further comprise a weight factor included in a matrix of weight factors; (18) training the grammar correction model further comprises performing a Singular Value Decomposition (SYD) on the matrix of weight factors; (19) training the grammar correction model may also include identifying a combined weight value that represents a first weight value element identified through the analysis of the non-learner text and a second weight value component that is identified by analyzing a learner text by minimizing an empirical risk function; (20) correcting semantic collocation errors; (21) identifying one or more translation candidates in response to analysis of a corpus of parallel-language text; (22) determining a feature associated with each translation candidate; (23) generating a set of one or more weight values from a corpus of learner text; (24) calculating a score for each of the one or more translation candidates in response; (25) identifying one or more translation candidates may include selecting a parallel corpus of text from a database of parallel texts, each parallel text comprising text of a first language and corresponding text of a second language, segmenting the text of the first language; (26) tokenizing the text of the second language, automatically aligning words in the first text with words in the second text, extracting phrases from the aligned words in the first text and in the second text, and calculating a probability of a paraphrase match associated with one or more phrases in the first text and one or more phrases in the second text; (26) collocation corrections with features derived from spelling edit distance; and (27) generating a phrase table having collocation corrections with features derived from a homophone dictionary; generating a phrase table having collocation corrections with features derived from synonym dictionary; additionally, generating a phrase table having collocation corrections with features derived from native language-induced paraphrases.
WANG et al. (CN 114781358 A, pub. Date:07/22/2022) discloses in Abstract that (1) performing pronunciation similar masking and font similar masking on a text in a training corpus, constructing a first training sample by using the masked training corpus, importing the first training sample into a pre-training language model, outputting a first pre-training text error correction result, adjusting the first training sample, and generating a second training sample; (2) performing iterative training on the pre-training language model by using the second training sample to obtain a text error correction model, and finally importing the text to be corrected into the text error correction model, and outputting a text error correction result; and (3) the pronunciation information and the font information are introduced when the error correction model is trained, the training sample with rich noise is constructed through pronunciation similar masking and font similar masking, and the error correction model is further trained through a reinforcement learning technology, so that the model can better recognize spelling errors, and the generalization performance is higher. WANG further discloses in ¶¶ [0004]-[0051] that (1) A text correction method based on reinforcement learning includes: (a) collect training data and perform pronunciation similarity masking and character shape similarity masking on the text in the training data according to the preset text masking ratio; (b) the first training sample is constructed using the masked training corpus; (c) the first training sample is transformed into a vector to obtain the first sample embedding vector; (d) the first sample embedding vector is imported into the pre-trained language model, and the first pretrained text error correction result is output; (e) the first training sample is adjusted based on the first pre-trained text correction result to generate the second training sample; (f) the pre-trained language model is iteratively trained using the second training sample to obtain the text correction model; and (g) receive text correction instructions, obtain the text to be corrected, import the text to be corrected into the text correction model, and output the text correction result; (2) collecting training data and performing pronunciation similarity masking and character shape similarity masking on the text in the training data according to a preset text masking ratio include: (a) the target text is masked in the training corpus fragments according to the preset text masking ratio; (b) identify texts with similar pronunciation and similar glyphs that correspond to the target masked text within a pre-defined text obfuscation set; and (c) based on texts with similar pronunciation and similar glyphs, target text is masked using both pronunciation similarity and glyph similarity; (3) constructing the first training sample using the masked training corpus specifically includes: (a) the training corpus segments that complete pronunciation similarity masking and those that complete character shape similarity masking are combined to form the first training sample; (b) obtain random text from a preset vocabulary list, and use the random text to randomly mask the target masked text; and (c) the training data segments that have undergone pronunciation similarity masking, character shape similarity masking, random text masking, and no text masking are combined to form the first training sample; (4) adjusting the first training samples based on the first pre-trained text correction results to generate the second training samples specifically includes: (a) the action value score of each training corpus segment in the first training sample is calculated based on the error correction results of the first pre-trained text; and (b) the proportion of each training corpus segment in the first training sample is adjusted based on the action value score of each training corpus segment to obtain the second training sample; (5) calculating the action value score of each training corpus segment in the first training sample based on the error correction results of the first pre-trained text specifically includes: (a) the F1 score of the pre-trained language model is determined based on the error correction results of the first pre-trained text; (b) obtain the training overhead of the pre-trained language model; and (c) the action value score of each training corpus segment in the first training sample is calculated based on the F1 score of the pre-trained language model and the training cost; (6) adjusting the proportion of each training corpus segment in the first training sample based on the action value score of each training corpus segment to obtain the second training sample specifically includes: (a) determine the positive or negative value of the action value score for each training corpus segment; (b) when the action value score is positive, the proportion of the training corpus segment corresponding to the action value score is increased; and (c) when the action value score is negative, the proportion of the training corpus segment corresponding to the action value score is reduced; (7) iteratively training the pre-trained language model using the second training samples to obtain the text correction model specifically include: (a) the second training sample is imported into the pre-trained language model to obtain the second pretrained text correction result; (b) the action value score of each training corpus segment in the second training sample is calculated based on the error correction results of the second pre-trained text; (c) the action value scores of each training corpus segment in the second training sample are summed to obtain the final action value score; and (d) the pre-trained language model is iteratively trained based on the final action value score to obtain the text correction model; (8) the first sample embedding vector includes a text embedding vector, a positional embedding vector, a pronunciation embedding vector, and a glyph embedding vector; (9) transforming the first training sample to obtain the first sample embedding vector specifically includes: (a) feature extraction is performed on the first training sample to obtain text features, position features, pronunciation features, and glyph features; and (b) vector transformations are performed on text features, positional features, pronunciation features, and glyph features respectively to obtain text embedding vectors, positional embedding vectors, pronunciation embedding vectors, and glyph embedding vectors; and (11) a text correction device based on reinforcement learning, comprising: (a) the sample construction module is used to construct the first training sample using the masked training corpus; (b) the vector transformation module is used to transform the first training sample into a vector to obtain the first sample embedding vector; (c) the pre-training module is used to import the embedding vector of the first sample into the pre-trained language model and output the error correction result of the first pre-trained text; (d) the sample adjustment module is used to adjust the first training sample based on the first pre-trained text correction result to generate the second training sample; (e) the iterative training module is used to iteratively train the pre-trained language model using the second training samples to obtain the text correction model; and (f) the text correction module is used to receive text correction instructions, obtain the text to be corrected, import the text to be corrected into the text correction model, and output the text correction result. WANG further discloses in ¶¶ [0119]-[0124] that (1) obtain the training overhead of the pre-trained language model; (2) the action value score of each training corpus segment in the first training sample is calculated based on the F1 score of the pre-trained language model and the training cost; (3) the F1 score, also known as the balanced F score, is defined as the harmonic mean of precision and recall, which can be seen as a harmonic average of the model's precision and recall, with a maximum value of 1 and a minimum value of 0, and is used to evaluate the analytical performance of the binary classification model; (4) specifically, the first pre-trained text correction result determines the F1 score of the pre-trained language model, obtains the training overhead of the pre-trained language model, and calculates the action value score for each training corpus segment in the first training sample based on the F1 score and training overhead; (5) the action value score is calculated using the following formula in ¶ [0123]; (6) in the formula, Si is the action value score of the i-th pre-training, F1i is the F1 score of the i-th pre-training, and TCosti is the training cost of the i-th pre-training; (7) it should be noted that the action value score can be positive or negative; (8) if the F1 score obtained after one round of training is higher than that after the previous round, then the training is beneficial to model fitting, and its action value score is F1/TCosti; and (9) if the F1 score is lower than that after the previous round, then the training is not beneficial to model fitting, and its action value score is -F1/TCost.
Bravo-Candel et al. ("Automatic Correction of Real-Word Errors in Spanish Clinical Texts", Sensors 2021, 21, 2893, April 21, 2021, pp. 1-27) disclose in Section 3 with FIG. 1 and Table 3 of Pages 6-15 that (1) propose the use of seq2seq for Neural Machine Translation; (2) erroneous texts must be tokenized in sentences, which are the input to the model; (3) the model then maps the erroneous sentences to the correct sentences, provided that it learnt context from previous examples; (4) additionally, pretrained embeddings containing semantic features have been implemented in the model; (5) the overall process and stages adopted in this research are depicted in Figure 1; (6) first, two corpora are collected to train and test the system; (7) for that, different datasets are gathered and integrated; (8) next, a number of text preprocessing steps are executed in order to remove unwanted characters and symbols; (9) then, two datasets for each corpus are generated by using two complementary error generation strategies; (10) the implemented seq2seq model is then trained with the generated datasets by considering different settings including the use of GloVe and Word2Vec pretrained embeddings; (11) finally, a post-processing stage is required to deal with some issues of the seq2seq model when correcting sentences; (12) Seq2seq [15,26] is based on an encoder–decoder system that translates input sentences into different ones; (13) difference in length between the input and the output is achieved by an intermediate vector of fixed-size; (14) this characteristic is desirable in language translation, where the word length varies from one language to another; (15) however, the output length may be the same as the input when correcting sentences with errors. In order to extract features from the sequences, the encoder-decoder system consists of two connected RNNs; (16) first, the encoder RNN maps the input sentence into an intermediate vector; (17) second, the decoder RNN generates the output sentence from the intermediate vector; (18) two corpora have been used to train and test the proposed system for real-word error correction: the Wikicorpus and the medicine corpus; (19) the Wikicorpus is the largest one with more than 611 million words extracted from Wikipedia articles in Spanish, whereas the medicine corpus is a collection of approximately 5750 clinical cases and 2 million words derived from three different corpora; (20) the Wikicorpus served as a first test on general language due to the large amount of data required for the correction systems to work; (21) therefore, once the system performed satisfactorily on this data, it was then tested on the created medical corpus, a much smaller source although indispensable for the purpose of this work; (23) a sufficient amount of data are needed to train the seq2seq model, however, such amount of text is not available when it comes to the task of real-word error correction; (24) as a solution, errors were introduced in clean text by using rules as it is done in [36]; (25) errors were generated straightforwardly in the absence of text annotations and syntactic parsers; e.g., some corpora contain part-of-speech information about words that could be useful to generate grammatical errors; (26) on the other hand, syntactic parsers might be used to introduce the errors in the sentences; (27) among the errors generated, six types were distinguished: (a) grammatical genre, (b) number, (c) grammatical genre-number, (d) homophones, (e) Hunspell-generated and (f) subject-verb concordance; (28) the rules were designed to identify specific substrings in the target word and replace them to generate a new word; (29) no syntactic or morphological information was provided to the rules; (30) therefore, sentence words were iterated by the rules and the first match was used to introduce an error in the sentence; (31) the new words generated by the rules were searched in a dictionary to determine whether they were real words; (32) the dictionary was compiled from the training corpus; (33) all the errors generated differed as much as three edit operations from the correct word, in other words, three or less characters were deleted or added to the word to obtain the error; (34) Table 3 shows an example of an erroneous sentence generated for each type; (35) rules were created to generate erroneous sentences; (36) Seq2seq input data was organized in one dataset containing the source or erroneous sentences and a second dataset containing the target or correct sentences; and (37) similar to [11], different datasets were compiled according to two strategies: (1) many sentences with different errors were generated from one correct sentence; and (2) only one erroneous sentence was generated from one correct sentence. Bravo-Candel further discloses in Section 4 of Pages 15-22 that (1) the seq2seq model was evaluated with the medicine corpus and the Wikicorpus; (2) both sets were also combined in a third set in order to evaluate the model performance in mixed corpora; (3) Recall (R) (Equation (5)), precision (P) (Equation (6)) and F0.5 (Equation (7)) were obtained in order to evaluate the seq2seq models; (4) these are the most common measures for the task of real-word error correction; (5) for each set of sentences, errors were generated by following the two aforementioned strategies to compile two training datasets for each corpus; (6) Seq2seq models were also evaluated according to each error type; (7) the same evaluation dataset was used for both datasets; (8) Table 8 shows the evaluation of 18 Seq2seq models on all error types; (9) the same models were evaluated on each error type: (a) grammatical genre, (b) number, (c) grammatical genre-number, (d) homophones and (e) Hunspell-generated and (d) subject-verb concordance errors; (10) models trained on the Wikicorpus were expected initially to perform better than those trained on medicine sentences, due to the larger amount of data; (11) Wikicorpus sentences varied in length from 5 to 60 tokens; (12) Moreover, article topics were diverse, which resulted in a very large vocabulary and unknown words in the evaluation set, whereas known words were used in different contexts; (13) at last, Wikipedia is open source and widely contributed, so that they may contain syntactic and grammatical errors; (14) on the contrary, clinical text vocabulary and context was reduced to one single domain and sentences were written in a simple style; and (14) this could answer why results improved for models trained on the medicine corpus even if the dataset was smaller.
However, closest arts of records, as discussed above, singly or in combination do not teach or suggest at least following features "averaging precisions and recalls of the plurality of sample labels in the first class and the second class, respectively, to obtain an average precision and an average recall of the first-class sample labels and an average precision and an average recall of the second-class sample labels; and calculating E-F1 according to the average precision and the average recall of the first-class sample labels, and calculating I-F1 according to the average precision and the average recall of the second-class sample labels, wherein E-F1 is a harmonic mean F1 of the precisions and the recalls of the training model on those errors having not been seen in training, and I-F1 is a harmonic mean F1 of the precisions and the recalls of the training model on those errors having been seen in training" when combining with all other limitations of these claims as a whole..
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HWEI-MIN LU/Primary Examiner, Art Unit 2142