Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claims 1-20 are pending. Claims 1, 7 and 14 are independent. Claims 2-3 and 7 and 14 are amended.
This Application was published as U.S. 2024/0062570.
Apparent priority: 19 August 2022.
Claims 1-6 are allowed.
As to the remainder of the Claims 7-20, applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection that if presented were necessitate by the amendments to the Claims.
This action is Final.
Response to Amendments and Arguments
Rejection of Claims 2-3 is withdrawn in view of the amendments to these Claims.
Claim 1 was previously allowed:
1. A computer-implemented method for detecting Unicode injection in text, the method comprising:
training a language model to determine if text data conforms with human writing habits using a corpus of negative samples comprising text written by humans and a corpus of positive samples formed by randomly inserting Unicode characters into said corpus of negative samples, wherein said human writing habits comprise text written by humans obtained from publicly accessible online sources;
training said language model to recognize a region of text with Unicode characters using an entity recognition method comprising a bidirectional encoder representations from transformers (BERT) model and a conditional random fields (CRF) model to learn features of text via said BERT model and obtain sequence-level tag information via a CRF layer to identify text fragments that do not conform with said human writing habits; and
receiving text data by said language model to determine if said text data is suspect for containing Unicode characters by identifying one or more regions within said received text that do not conform with said human writing habits represented in said negative samples.
Claims 7 and 14:
Claims 7 and 14 are amended as follows and the added language is taught by the previously cited Malkiel. No specific arguments were presented and the Response appears to have relied on the amendments.
Brown as shown in the modified rejection teaches sanitizing the inputs by removing or modifying the homographs so that a subsequent model training is not performed with bad data. Determining a difference between vectors of text was mapped to Malkiel which also teaches the added “threshold” to the Claims.
7. A computer-implemented method for detecting Unicode injection in text, the method comprising:
identifying original data to be used in a text training set for a natural language processing task;
recording image data from a copy of original data;
recording a first set of text data from said original data;
performing optical character recognition on said recorded image data of the original data to generate a first set of text data and a second set of text data;
generating a first feature vector for said first set of text data;
generating a second feature vector for said second set of text data;
based on a difference between the first feature vector and the second feature vector exceeding a defined threshold, identifying the original data as suspect for containing Unicode characters; and
modifying said text training set by removing said first set of text data in response to identifying said first set of text data as being suspect for containing Unicode characters.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 7-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 7 is amended as follows and Claim 14 is amended similarly. The amendments have created an antecedent basis issue. The dependent Claims inherit the indefiniteness.
7. A computer-implemented method for detecting Unicode injection in text, the method comprising:
identifying original data to be used in a text training set for a natural language processing task;
recording image data from a copy of original data;
recording a first set of text data from said original data;
performing optical character recognition on said recorded image data of the original data to generate a first set of text data and a second set of text data;
generating a first feature vector for said first set of text data;
generating a second feature vector for said second set of text data;
based on a difference between the first feature vector and the second feature vector exceeding a defined threshold, identifying the original data as suspect for containing Unicode characters; and
modifying said text training set by removing said first set of text data in response to identifying said first set of text data as being suspect for containing Unicode characters.
There are two “a first set of text data” in the Claim and each is generated differently which means that they should not both be the same.
Applicant has referenced the following paragraphs [0015]-[0016] and [0030]-[0031] as support for the amendments:
PNG
media_image1.png
318
748
media_image1.png
Greyscale
PNG
media_image2.png
196
694
media_image2.png
Greyscale
PNG
media_image3.png
248
716
media_image3.png
Greyscale
PNG
media_image4.png
640
704
media_image4.png
Greyscale
Of the above cited paragraphs only [0030] is on point and this is what it says:
[0030] The embodiments of the present disclosure provide a means for detecting Unicode injection in text, such as text (e.g., text training set) used by natural language processing tasks, by training a language model to detect text (e.g., text fragment) that does not conform with human writing habits which indicates a Unicode injection (e.g., direct Unicode injection) as discussed further below. Alternatively, Unicode injection in text (e.g., indirect Unicode injection) may be detected by performing optical character recognition on the recorded image data of the original text data to generate a first set of text data as well as recording the original text data as a second set of text data. A feature extraction network of pre-trained natural language processing task(s) may be utilized to generate feature vectors using the first and second sets of text data. For example, a pre-trained natural language processing task (e.g., semantic matching) may generate a first feature vector for the first set of text data and a second feature vector for the second set of text data. If the difference between the measurements of such feature vectors exceeds a threshold value, then the original text data may be identified as being suspect for containing Unicode characters. A more detailed description of these and other features will be provided below.
As such the limitation of “performing optical character recognition on said recorded image data of the original data to generate a first set of text data and a second set of text data;” has no support in the Specification as filed and cannot be explained or disambiguated by referring to the Disclosure.
The Disclosure supports:
recording a first set of text data from said original data;
performing optical character recognition on said recorded image data of the original data to generate
Thus, for applying art, the Claim is interpreted as shown with the modification above.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 7-20 are r rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
The material added to Claims 7 and 14 by the most recent amendments does not find support in the Specification and Drawings as filed and is thus considered new matter. The remaining Claims inherit the new matter.
Claims 7 and 14 are amended to include:
“performing optical character recognition on said recorded image data of the original data to generate a first set of text data and a second set of text data;”
Applicant has referenced the following paragraphs [0015]-[0016] and [0030]-[0031] as support for the amendments:
PNG
media_image1.png
318
748
media_image1.png
Greyscale
PNG
media_image2.png
196
694
media_image2.png
Greyscale
PNG
media_image3.png
248
716
media_image3.png
Greyscale
PNG
media_image4.png
640
704
media_image4.png
Greyscale
Of the above cited paragraphs only [0030] is on point and this is what it says:
[0030] The embodiments of the present disclosure provide a means for detecting Unicode injection in text, such as text (e.g., text training set) used by natural language processing tasks, by training a language model to detect text (e.g., text fragment) that does not conform with human writing habits which indicates a Unicode injection (e.g., direct Unicode injection) as discussed further below. Alternatively, Unicode injection in text (e.g., indirect Unicode injection) may be detected by performing optical character recognition on the recorded image data of the original text data to generate a first set of text data as well as recording the original text data as a second set of text data. A feature extraction network of pre-trained natural language processing task(s) may be utilized to generate feature vectors using the first and second sets of text data. For example, a pre-trained natural language processing task (e.g., semantic matching) may generate a first feature vector for the first set of text data and a second feature vector for the second set of text data. If the difference between the measurements of such feature vectors exceeds a threshold value, then the original text data may be identified as being suspect for containing Unicode characters. A more detailed description of these and other features will be provided below.
As such the limitation of “performing optical character recognition on said recorded image data of the original data to generate a first set of text data and a second set of text data;” has no support in the Specification as filed. The Specification provides for conducting OCR to “generate a first set of text data” and “recording the original text data as a second set of text data” such that there are two sets of text data one of which is generated by OCR and other is a recording of the original set of text data. The OCR does not generate two sets of data as the current Claim language states. The added language is thus New Matter. This is in addition to the fact that it has created an antecedent basis problem and indefiniteness which cannot be resolved.
The Disclosure supports:
recording a first set of text data from said original data;
performing optical character recognition on said recorded image data of the original data to generate
Thus, for applying art, the Claim is interpreted as shown with the modification above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 7-20 are rejected under 35 U.S.C. 103 as being unpatentable over Brown (U.S. 10943067) in view of Malkiel (U.S. 11836175).
Regarding Claim 7, Brown teaches:
7. A computer-implemented method for detecting Unicode injection in text, the method comprising: [Brown is directed to detecting homographs which are created by injecting Unicode into text. “Disclosed are various embodiments for detecting homograph attacks using text recognition. A first string of untrusted text is received. A second string is determined corresponding to what the first string of untrusted text appears to be in a particular language. The second string is determined to differ from the first string of untrusted text….” Abstract. “Although typical Latin character fonts may include homographs, the potential for homographs is greatly magnified in considering extended Unicode, with thousands of different characters for over a hundred different writing systems. For example, the lower case Greek letter omega may resemble the lower case Latin “w,” and the Latin character “a” may look the same as the Cyrillic character “a.”…” 2:5-11. “… Even without external dependencies, homographs can be used to hide malicious behavior in plain sight. For instance, in a random on-call selector, a Unicode space may be utilized to generate a parsing error on a particular name, ensuring that the name will never be randomly selected.” 10:1-12.]
identifying original data to be used in a text training set for a natural language processing task; [Brown teaches that is directed to “using text recognition to defeat homograph attacks.” Text recognition is a natural language processing task.]
recording image data from a copy of said original data; [Brown in Figure 1 shows capture/recording of an image of the URL region 106a. “According to the present disclosure, an image 109 of the address bar 106b may be captured, and an optical character recognition (OCR) process performed.” 3:20-22. Figure 2, 218: “Although the homograph recognition engine 218 may include an image generator 224, it is understood that in some embodiments, the homograph recognition engine 218 may receive previously generated images, which may be screen captures or portions of screen captures.” 4:33-38.]
recording a first set of text data from said original data; [Brown, Figure 1, receiving the URL at 106a which is the “original data” of the Claim. Figure 3, “receive first string of untrusted text 303.” “Beginning with box 303, the homograph recognition engine 218 receives a first string of untrusted text. For example, the homograph recognition engine 218 may execute as a service, and the client computing device 206 (FIG. 2) may have submitted the untrusted text for homograph detection as a service. Alternatively, the homograph recognition engine 218 may execute on the client computing device 206 and intercept the untrusted text before or as it is being rendered for display by a client application 266 (FIG. 2).” 6:61-7:3.]
performing optical character recognition on said recorded image data of the original data to generate a first set of text data and a second set of text data; [Brown, Figure 1, “OCR” being performed on the captured image of 109. “According to the present disclosure, an image 109 of the address bar 106b may be captured, and an optical character recognition (OCR) process performed.” 3:20-23. Figure 2, 227: “The text recognition application 227 is executed to perform an optical character recognition (OCR) process on an input image. Specifically, the text recognition application 227 examines the image and determines which characters or glyphs are present in the image….” 4:39-44.]
generating a first feature vector for said first set of text data;
generating a second feature vector for said second set of text data; and
based on a difference [Brown, Figure 3, 318: “Strings differ?”] between the first feature vector and the second feature vector exceeding a defined threshold, identifying the original data as suspect for containing Unicode characters; and [Brown, Figure 3, 315. “In box 315, the homograph recognition engine 218 compares the first string of untrusted text that was received with the second string determined from the text recognition application 227. In box 318, the homograph recognition engine 218 determines whether the strings differ. …” 7:48-59.]
modifying said text training set by removing said first set of text data in response to identifying said first set of text data as being suspect for containing Unicode characters. [Brown, Figure 5, 503 to 515 remove the infected/suspect samples. “In box 515, the data processing application 221 may process the result string. As an example, the data processing application 221 may use the result string to train a machine learning model, thus defeating homograph attacks meant to poison machine learning training data sets. ….” 9:47-69. One of the goal cited by Brown includes sanitizing before retraining the model: “(4) sanitizing data corpuses to remove homographic data, thereby ensuring the integrity of the data corpus for use in machine learning, search engines, and so forth, and (5) improve computing efficiencies by, for example, reducing network and processing demands if a user realizes that he or she is on the wrong site and needs to navigate to another site, avoiding training a machine learning model with bad data and having to retrain the model, and so on. The additional elements in the user interfaces enable users to more efficiently locate, and navigate to, what they are looking for and allow the user to navigate to the associated pages with fewer clicks, taps, or other interactions.” which implies removing the bad samples. 2:55-67. Brown in several places teaches Sanitizing the data before it can be used for training which directly and expressly teaches fixing/modifying the Unicode injected text portions: “The data processing applications 221 are executed to perform a data processing function with respect to a data corpus. For example, a data processing application 221 may be training a machine learning model, indexing network pages for a search engine, performing plagiarism detection, sanitizing source code repositories, or performing other functions. The homograph recognition engine 218 may be used to sanitize the data corpus to remove homographic strings or to replace them with the strings that they appear to be before the data corpus is processed.” 4:60-5:2.]
PNG
media_image5.png
526
426
media_image5.png
Greyscale
Brown does not include the details of determining the similarity and thus does not teach the comparison of feature vectors which is a well-known method in the art. Additionally, the added “threshold” is not mentioned in Brown while the use of a threshold is implied in most difference evaluations.
Malkiel teaches:
generating a first feature vector for said first set of text data; [Malkiel, Figure 2, “obtain a first feature vector representative of the search query 204.”]
generating a second feature vector for said second set of text data; and [Malkiel, Figure 2, “… each of a plurality of second feature vectors … representative of a … respective first text-based content item … 206.” Figure 1, “vector repository 112.” “Vector repository 112 is configured to store a plurality of feature vectors 120, each being representative of a summary of a respective content item stored in content item repository 120….” 5:42-45.]
based on a difference between the first feature vector and the second feature vector exceeding a defined threshold, identifying the original data as suspect for containing Unicode characters; and [Malkiel, Figure 2, 206, “At step 206, a respective semantic similarity score between the first feature vector and each of a plurality of second feature vectors generated by a transformer-based machine learning model is determined….” 7: 11-15. Figure 1, “similarity determiner 114.” “Semantic search techniques via focused summarizations are described. For example, a search query is received for a text-based content item in a data set comprising a plurality of text-based content items. A first feature vector representative of the search query is obtained. A respective semantic similarity score is determined between the first feature vector and each of a plurality of second feature vectors. Each of the second feature vectors is representative of a machine-generated summarization of a respective text-based content item…. The search result comprises a subset of the plurality of text-based content items associated with a respective second feature vector having a semantic similarity score that has a predetermined relationship with a predetermined threshold value.” Abstract. “… The search result comprises a subset of the plurality of text-based content items associated with a respective second feature vector having a semantic similarity score that has a predetermined relationship with a predetermined threshold value….” 6:10-20. “At step 208, a search result comprising a subset of the first plurality of text-based content items associated with a respective second feature vector having a semantic similarity score that has a predetermined relationship with a predetermined threshold value is provided. For example, with reference to FIG. 1, similarity determiner 114 provides search result 126 comprising a subset of the first plurality of text-based content items of content item repository 110 that are associated with a respective second feature vector 122 having a semantic similarity score that has a predetermined relationship with a predetermined threshold value.” 7:33-43.]
Brown and Malkiel pertain processing text including Unicode characters and it would have been obvious to combine the feature vector comparison of Malkiel which is very commonly used method of comparison of pieces of text with the system of Brown which leaves out the details of implementation of the comparison process. This combination falls under simple substitution of one known element for another to obtain predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396.
Regarding Claim 8, Brown teaches:
8. The method as recited in claim 7 further comprising:
copying said original data that is to be processed by said natural language processing task; and [Brown, Figure 1, the “capture” is a form of copying.]
recording said image data from said copy of original data that is to be processed by said natural language processing task. [Brown, Figure 1, the “capture” copies and records for a subsequent OCR.]
Regarding Claim 9, Brown does not teach generating feature vectors.
Malkiel teaches:
9. The method as recited in claim 7 further comprising:
generating said first and second feature vectors by said natural language processing task. [Malkiel: Figure 1, the “feature vectors 118” are the product of the “transformer based machine learning model 106 which extracts words and keywords as features and is thus performing an NLP: “To generate a feature vector 118, a representation of each of the plurality of multi-word fragments may be provided to transformer-based machine learning model 106. The representation may be a feature vector representative of the multi-word fragment. For instance, features (e.g., words) may be extracted from the multi-word fragment and such features may be utilized to generate the feature vector. The feature vectors may take any form, such as a numerical, visual and/or textual representation, or may comprise any other suitable form and may be generated using various techniques, such as, but not limited to, keyword featurization, semantic-based featurization, bag-of-words featurization, and/or n-gram-TF-IDF (term frequency-inverse document frequency) featurization.” 4:64 to 5:10.]
Rationale for combination as provided for Claim 7 considering that Malkiel was cited for teaching the generation of the feature vectors.
Regarding Claim 10, Brown teaches:
10. The method as recited in claim 9, wherein said natural language processing task comprises one of the following:
text classification, entity recognition, machine reading comprehension, semantic matching and machine translation. [Brown performs entity recognition by its “domain name list 233” as shown in Figure 4 and also teaches text classification as a homograph or not in Figure 3.]
Regarding Claim 11, Brown teaches:
11. The method as recited in claim 7 further comprising:
identifying said first set of text data as being suspect for containing Unicode characters in response to a difference between measurements of said first and second features exceeding a threshold value. [Brown in Figure 1 and Figure 3, 315 teaches that the difference between two strings indicates if there has been a homograph attack or the text is normal. “Strings differ? 318” to Yes. “However, if the strings differ, a homograph attack may be occurring, and the homograph recognition engine 218 implements one or more actions in box 321. For example, the homograph recognition engine 218 may cause an alert to be generated in a user interface 269 rendered by a client computing device 206. …” 7:60 to 8:10.]
Brown does not teach comparison of vector features and a threshold.
Malkiel teaches:
identifying said first set of text data as being suspect for containing Unicode characters in response to a difference between measurements of said first and second features exceeding a threshold value. [Malkiel, Figure 1, “similarity determiner 114”: “… For example, similarity determiner 114 may determine whether a respective semantic similarity score determined between feature vectors 122 and feature vector 124 has a predetermined relationship with a predetermined threshold value. For instance, similarity determiner 114 may determine whether a respective semantic similarity score reaches and/or exceeds a predetermined threshold value. Content item(s) corresponding to feature vector(s) 122 being associated with semantic similarity scores having the predetermined relationship with the predetermined threshold may be returned to application 128 via one or more search results 126….” 6:10-21.]
Rationale for combination as provided for Claim 7 considering that Malkiel was cited for teaching the generation of the feature vectors and comparison of the vectors to determine similarity. The use of thresholds is common in the art.
Regarding Claim 12, Brown teaches:
12. The method as recited in claim 7 further comprising:
identifying said first set of text data as being normal in response to a difference between measurements of said first and second features not exceeding a threshold value. [Brown in Figure 1 and Figure 3, 315 teaches that the difference between two strings indicates if there has been a homograph attack or the text is normal. 318: “Strings differ?” to NO. “In box 315, the homograph recognition engine 218 compares the first string of untrusted text that was received with the second string determined from the text recognition application 227. In box 318, the homograph recognition engine 218 determines whether the strings differ. … If the strings are the same, no homograph attack is detected, and the operation of the homograph recognition engine 218 ends.” 7:48-59.]
Brown does not teach comparison of vector features and a threshold.
Malkiel as applied to Claim 11 above, teaches the use of thresholds in the comparison and obviously if exceeding the threshold indicates similarity then not reaching the threshold would indicate absence of similarity. Rationale as provided for Claim 11.
Regarding Claim 13, Brown teaches:
13. The method as recited in claim 7, further comprising:
providing said first set of text data to a user for further handling in response to identifying said first set of text data as being suspect for containing Unicode characters. [Brown, Figure 5, the infected/suspect samples are shown to an administrator: “Beginning with box 503, the data processing application 221 receives a string of untrusted text from an untrusted data corpus 239 (FIG. 2). In box 509, the data processing application 221 submits the string to the homograph recognition engine 218 (FIG. 2) for processing. This processing may result in removal of homographs from other scripts or languages, thus normalizing or sanitizing the string. Further, this processing may result in the removal of unused characters, zero-width spaces, or extraneous characters that merge into a single characters. In particular, this processing may remove steganographic data encoded into the non-rendering characters and zero-width spaces. The removal of such steganographic data may prevent data exfiltration from the computing environment 203. In some cases, if such steganographic data is detected (e.g., an abnormal quantity or pattern of non-rendering characters or zero-width spaces), the occurrence may be flagged for further review by a system administrator. In box 512, the data processing application 221 obtains a result string from the homograph recognition engine 218.” 9:27-47.]
Claim 14 is a computer program product system claim with limitations corresponding to the limitations of method Claim 7 and is rejected under similar rationale.
Claim 15 is a computer program product system claim with limitations corresponding to the limitations of method Claim 8 and is rejected under similar rationale.
Claim 16 is a computer program product system claim with limitations corresponding to the limitations of method Claim 9 and is rejected under similar rationale.
Claim 17 is a computer program product system claim with limitations corresponding to the limitations of method Claim 10 and is rejected under similar rationale.
Claim 18 is a computer program product system claim with limitations corresponding to the limitations of method Claim 11 and is rejected under similar rationale.
Claim 19 is a computer program product system claim with limitations corresponding to the limitations of method Claim 12 and is rejected under similar rationale.
Claim 20 is a computer program product system claim with limitations corresponding to the limitations of method Claim 13 and is rejected under similar rationale.
Allowable Subject Matter
Claims 1-6 are allowed.
The following is an examiner’s statement of reasons for allowance: In view of each of the particular limitations of the independent Claims when considered in the order established by the Claim language and in the context of the language of the independent Claims when each Claim is considered as a whole, the independent Claims of this Application were not found in the prior art that was viewed.
In particular, while use of positive and negative samples for training a model including a large language model is known, and while the use of BERT+CRF is also known, as evidenced by the art cited below, the use of positive and negative samples to train a BERT model and a CRF model to identify phishing by injection of Unicode which creates language that in the language of the Claim “does not conform to human writing habits” (meaning it has Unicode injected for the purpose of phishing attacks) when considered in the context of the independent Claims as a whole and considering each and every limitation of these claims was not found in the prior art.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Close Art of Record
In addition to the art applied during the prosecution, note the following:
Positive and Negative Samples for Training
Lee (US 20200193153):
[0047] In some embodiments, the search engine is implemented using and comprises a machine learning computer program. The machine learning program may be trained based on a training set of positive and negative examples. The positive examples may include pairs of text segments, comprising a patent claim text segment and an associated stored text segment from a prior art document that is identified as similar. The negative examples may include pairs of text segments, comprising a patent claim text segment with an associated stored text segment from a prior art document that is identified as not being similar. The positive and negative examples may be generated by humans or through automated techniques.
[0051] In one embodiment, the output of the search engine may take the form of an office action, such as in the format of office actions generated by the U.S. Patent & Trademark Office or by the patent offices of foreign countries or the International Bureau (TB) of the World Intellectual Property Organization (WIPO). Office actions may include reasons for rejection, such as for lack of novelty or obviousness. Office actions are currently written by human patent examiners.
[0140] Language model 544 may be trained to generate word embeddings using large amounts of training text. The training text may be representative of inputs to text similarity model 505. The training text may be tokenized using tokenizer 542 to generate a sequence of training tokens. A word embedding algorithm may be trained on the sequence of training tokens to generate word embeddings for each token. The word embedding algorithm may comprise skip-gram, continuous bag of words (CBOW), Word2Vec, GloVe, ELMo, or other algorithms.
Rarick (US 20130103695):
[0054] FIGS. 5 and 6 illustrate exemplary sentence pairs that demonstrate the difference between legitimate OOV tokens and tokens that result from machine translation. FIG. 5 shows an example of a toy model number that is likely OOV for both English and Japanese ("GAT-X105"), but would not be indicative of machine translation (e.g., this sentence pair in fact seems to be human-translated). In contrast, FIG. 6 Error! Reference source not found. includes two likely OOV words on the Japanese side: one is probably the result of a word left untranslated by a machine translation system ("templatized"), and one appears to refer to an identifier in a piece of code ("back_insert_iterator"). It would not be surprising to see OOV tokens such as the token that refers to the identifier in a piece of code in a human-written Japanese sentence that discusses a piece of code.
Starkie (US 20080126078):
[0014] As described in Gold, E. M. [1967] Language identification in the limit, in Information and Control, 10(5):447-474, 1967 ("Gold"), it was demonstrated in 1967 that the grammars used to model natural languages at that time could be learnt deterministically from examples sentences generated by that grammar, but that it was possible for a language to be learnt from both examples sentences generated from that grammar, referred to as positive examples, and examples of bad sentences that are not generated from that grammar, referred to as negative examples.
[0066] The present invention also provides a process for inferring a grammar from a plurality of positive and negative example sentences and a starting grammar, including:
[0070] The present invention also provides a process for inferring a grammar in the limit from a plurality of positive and negative example sentences, including: [0071] identifying in the limit a grammar from only the positive example sentences using machine learning; and [0072] generating, on the basis of said grammar and said plurality of positive and negative example sentences, an output grammar that can generate all of the positive example sentences but cannot generate any of the negative example sentences.
[0073] The present invention also provides a process for inferring a grammar from a plurality of positive and negative example sentences, including: [0074] generating in the limit a class of grammar from only the positive example sentences; and [0075] removing recursion from the grammar.
[0096] As shown in FIG. 1, a grammatical inference system includes a merger or merging component 102, a splitter or splitting component 106, and one or more converters 108. Optionally, the grammatical inference system may also include an unfolder or unfolding component 104. The grammatical inference system executes a grammatical inference process, as shown in FIG. 2, that infers a final or inferred grammar 216 from a set of positive example sentences 206 and a set of negative example sentences 222. The inferred grammar 216 can then be converted into a form suitable for use with a spoken or text-based dialog system, as described below.
[0098] The grammatical inference process takes as its input a set of positive example sentences 206 which represent a subset of the sentences that the final grammar 216 should be able to generate….
BERT/CRF:
Nguyen (US 20230096939):
[0058] The classification task and the sequence tagging task are jointly trained using the BERT-based model. Specifically, the classification task is trained using the BERT-based model and the softmax layer; while the sequence tagging task is trained using the BERT-based model and the CRF layer. The following describes the training of the classification task and the sequence tagging task.
Hosseini-Asl (US 20220366145):
[0009] Another approach is based on a conditional random field (CRF) combined with BERT for aspect term extraction and term polarity prediction. The two modules are employed for improving aspect term extraction and term polarity prediction of the BERT model. First, a parallel approach is used which combines predictions for aspect term and polarity from the last four layers of BERT in a parallel way. Moreover, a hierarchical aggregation module is also examined, where predictions of previous layers of BERT are fed into the next layer.
Wang (US 20220318506):
[0116] Exemplarily, the BERT sub-layer performs word segmentation on an input object text with a tokenizer to obtain a sequence X after word segmentation, which is then encoded into a token embedding matrix W.sub.t and a position embedding matrix W.sub.p, and then the two matrixes (or vectors) are added to form an embedding expression vector h.sub.0 that obtains an output text semantic expression vector h.sub.L through an L-layer transformer, where the corresponding process is represented as: …
[0229] A known extraction model having a “BERT+CRF (conditional random field)” structure in some related arts is taken as a comparative example extraction model, called BERT+CRF.
Huang (US 20210312230):
[0044] By fine-tuning the BERT model with training samples, better word embedding representation can be obtained. The process of training the BI-LSTM model and the CRF layer is also a process of fine-tuning the BERT model. Compared with a static word vector used in a traditional sequence labeling method, more context information and position information is introduced during the extraction model training process, and the semantic vector is dynamically generated according to the context of the word, and has richer context semantics, and can better solve the problem of enhancement ambiguity. Performing boundary correction on the first enhanced text output by the extraction model enables the model to better fit the actual scenario, further enhance the accuracy of the extraction model, and provides support for the automatic style enhancement of the text in the landing page field.
Chen (US 20240037339):
[0040] NER research has a long history and recent approaches using Neural Network models like BiLSTM-CNN-CRF and contextual embeddings such as BERT and FLAIR have improved the NER performance in the general domain to the human-level. However, the NER performance for specific domains is still moderate due to the challenges of limited annotations and dealing with complicated domain-specific contexts.
[0054] Decoding Layer: Finally, a Conditional Random Field (CRF) layer is used to decode the enriched embeddings F=[f1, f2, . . . , fn] into a sequence of labels y={y1, y2, . . . , yn}. Wherein the sequence of labels are an enriched entity predictions of the input. In the training phrase, then optimize the whole model by minimizing the negative log-likelihood loss with respect to golden labels.
Zhou (U.S. 20230367966):
[0070] With continued reference to FIG. 3, the method 200 continues with enabling the user to compare performances of a plurality of different algorithms for training the natural language understanding model (block 240). Particularly, in response to a corresponding user selection via the graphical user interface, the processor 110 operates the display 130 to display a graphic user interface enables the user to selectively train and test different algorithms for the natural language understanding model. A wide variety of natural language understanding algorithm can be supported by the development platform, such as Joint BERT (with CRF), Joint BERT (without CRF), Stack Propagation+BERT, Bi-Model, Attention Encoder-Decoder, or InterNLU. To these ends, in response to corresponding user selections via the graphical user interface, the processor 110 trains a plurality of instances of the natural language understanding model(s) 122 using a plurality of different algorithms. Additionally, the processor 110 determines a performance of each of the plurality of instances of the natural language understanding model(s) 122.
Removing Samples from Training Data:
Lee (U.S. 20200193153):
[0154] The text parsed from the comparison documents may be modified to remove irregularities or improve training performance. For example, irregularities such as non-ASCII characters may be removed or replaced with placeholder characters and common errors from OCR such as confusion between lower-case letter ‘L’ and the number ‘1’ may be corrected. In another example, in order to optimize training performance, a spell-checking algorithm may be applied, or text which is too short or too long for optimal training may be removed, expanded, or trimmed.
Rarrick (U.S. 20130103695):
[0015] FIG. 10 is a flow diagram that illustrates an exemplary methodology for detecting and removing machine translated content from a set of document pairs used to train a machine translation engine.
[0037] The system 200 further includes a training component 206 that trains a machine translation engine 208 using the filtered remainder of the document pairs and without using the subset of the document pairs detected as being generated through machine translation. Hence, web-scraped parallel corpora can be cleaned for use in training statistical machine translation systems (e.g., the machine translation engine 208, etc.). The web-scraped parallel corpora can be cleaned by detecting (e.g., with the feature extraction component 104 and the classification component 110) and removing (e.g., with the filter component 204) machine translated content. Accordingly, since the machine translated content included in the web-scraped parallel corpora can be removed, the training component 206 can train the machine translation engine 208 on human translated content without training the machine translation engine 208 on machine translated content. Training the machine translation engine 208 using machine translated content can introduce errors and detrimentally impact performance of the machine translation engine 208. Hence, exclusion of the machine translated content from the web-scraped parallel corpora by the filter component 204 can improve the quality of the machine translation engine 208 over time as trained by the training component 206. Moreover, it is contemplated that a different document can be translated with the machine translation engine 208 as trained.
[0104] … At 1010, a machine translation engine can be trained using the filtered set of the document pairs and without using document pairs removed from the filtered set of the document pairs detected as being generated through machine translation.
Liu (U.S. 8380488):
“… Generating the language model can include, for each document in the training corpus, segmenting sequences of characters in the document into a plurality of model features; and removing from the plurality of model features every model feature comprising characters not among the plurality of characters used in writing the language of the document. …” 2:29-67.
Starkie (U.S. 20080126078):
[0283] Next, all of the negative examples are parsed using the grammar and a chart parser. If there is a sentence in the set of negative examples 222 that cannot be parsed using the grammar, or can be parsed using the grammar but is assigned a different set of attributes, then it is deleted from the local copy of the set of negative examples, but is not removed from the file (or similar medium) from which the training examples were obtained.
[0284] At this point, the process has a set of training sentences that are to be removed from the grammar, together with a set of sentences that are required to be generated from the grammar.
Datta (U.S. 20070288230):
[0043] The process 200 also removes any languages associated with a variant that have a frequency that does not exceed a pre-defined absolute threshold (step 230). The absolute threshold is pre-determined and specified on a per-language basis. This threshold is used to remove variants that are likely to be misspellings or mistakes in the training corpus. For a language that is well represented in the training corpus, a large threshold (e.g., 40 for English) will generally omit obscure misspellings. The threshold for a small language which is not well represented would be set lower (e.g., 10) to preserve legitimate but rare words. The threshold can be turned off (or set to 0) for languages which are poorly-represented in the corpus.
Previously Cited:
Starbuck (U.S. 20040260776) Figure 8. “Train filter using machine learning algorithm 850” which identifies spam from non-spam. Non-spam conforms with human writing habits.
Chhichhia (U.S. 20160137891): Figure 5, “feature extraction module 510.” “[0083] The feature extraction module 510 extracts features from documents in the content catalog database 402….” “[0092] After extracting the feature vectors from the content entity samples ….” “classification module 550.” The extracted text is in Unicode and the feature vectors will contain Unicode characters. “[0058] Text is extracted from the pages of the original document tagged as having text. The text extraction may be done at the individual character level, together with markers separating words, lines, and paragraphs. The extracted text characters and glyphs are represented by the Unicode character mapping determined for each….” “[0060] The output of text extraction 302, therefore, a dataset referenced by the page number, comprising the characters and glyphs in a Unicode character mapping with associated location information and embedded fonts used in the original document.” “[0084] … In one embodiment, the entity relationship analysis module 520 determines a similarity between the content entities based on the feature vectors received from the feature extraction module 510. The entity relationship analysis module uses the document features to determine a similarity between each content entity and each of the other content entities. For example, the entity relationship analysis module 520 builds a classifier to classify documents of one content entity into each of the other content entities based on the documents' features….”
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARIBA SIRJANI whose telephone number is (571)270-1499. The examiner can normally be reached 9 to 5, M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Fariba Sirjani/
Primary Examiner, Art Unit 2659