DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/19/2025 and 05/16/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Priority/Benefit
Acknowledgment is made of applicant’s claim for priority under 35 U.S.C. 119 (a)-(d). The certified copy of Republic of Korean Application KR10-2024-0064347 filed on May 17, 2024 has been received on 07/07/2025.
Claim Objections
Claim 8 is objected to because of the following informalities: typographical error. Claim 8 recites a limitation “a computer which is hardware is recorded” in the following clause: “A non-transitory computer-readable recording medium on which a computer program executed by a computer which is hardware is recorded”. This limitation is not grammatically correct.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-8 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Applying the subject matter eligibility test, as outlined in MPEP 2106:
Step 1: Statutory Category
The claims fall within a statutory category. Claims 7-8 are considered “machines” based claims and claims 1-6 are considered “processes”. Both machines and processes are members of the statutory categories. Thus, the analysis moves towards step 2A, prong one of the subject matter eligibility test.
Step 2A, Prong One: Judicial Exception
The claims recite a judicial exception, specifically an abstract idea. For example, claim 1 and 7-8 recite acquiring data (e.g. raw dark web data from a database), analyzing the data to extract first dark web data (e.g. by preprocessing the raw dark web data) and pretraining a bidirectional encoder representations from transformers (BERT)-based language model (e.g. using the first dark web data) and fine-tuning the pretrained BERT-based language model (e.g. using second dark web data). Such processes are akin to a mental process or methods of organizing human activity, which have been recognized as abstract ideas. Thus, the analysis moves towards step 2A, prong two.
Step 2A, Prong Two: Integration into a Practical Application
The claims do not integrate the abstract idea into a practical application. The additional elements, such as a processor, memory and performing a task for cybersecurity using the fine-tuned BERT-based language model do not impose any meaningful limits of on the abstract idea. In Recentive Analytics, Inc. v. Fox Corp., 2023-2437 (Fed. Cir. Apr. 18, 2025), the Federal Circuit held that applying generic machine learning techniques to a specific field without improving the underlying technology does not constitute a practical application. The court emphasized that claims must delineate how the machine learning technology achieves a technological improvement. Thus, the analysis moves towards step 2B.
Step 2B: Inventive concept
Finally, the claims do not recite an inventive concept that transforms the abstract idea into a patent-eligible application. The use of machine learning in a generic manner, without specifying a novel algorithm or unique training methodologies, fails to add significantly more to the abstract idea. As noted in Recentive, merely applying existing machine learning models to new data environments, without disclosing improvements to the models themselves, is insufficient for patent eligibility.
To summarize, there is no actual improvement to the machine learning model disclosed in the claim. In Recentive, claims that “do no more than claim the application of generic machine learning to a new data environment, without disclosing improvements to the learning model” were held ineligible. Here, the specification does nothing more than say use a generic trained machine learning model to decide how to perform a task for cybersercurity. There is no disclosure of a novel network architecture, no unusual training regimen, no non-routine feature extraction, and no explanation of why the machine learning model is not just an abstract idea of analyzing data.
Because none of these recitations describe a fundamental improvement to computer technology itself (e.g., a new mitigation method or machine learning algorithm), the claim is “directed to” an abstract idea.
Claims 2-4 and 6 merely add details to dark web data analysis that do not alter the outcome of the analysis above.
Claim 5 further describes the task for cybersecurity and outputting the result of performing the task. They do not alter the outcome of the analysis above.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-8 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Specifically, while the claim recites “pretraining a bidirectional encoder representations from transformers (BERT)-based language model using the first dark web data;fine-tuning the pretrained BERT-based language model using second dark web data”, the specification lacks a detailed description of the machine learning model or Artificial Intelligence, including its architecture, algorithm, training methodology, data inputs, or how it operates to achieve the claimed result. Courts have in the past (see MPEP2161.01) found that generic claim language in the original disclosure does not satisfy the written description requirement if it fails to support the scope of the genus claimed. Ariad, 598 F.3d at 1349-50, 94 USPQ2d at 1171 ("[A]n adequate written description of a claimed genus requires more than a generic statement of an invention’s boundaries.").
Dependent claims 2-6 are also rejected for inheriting the deficiencies of the independent claims from which they depend on.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-8 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Regarding claims 1 and 7-8, Limitations, "… second dark web data " is Indefinite because the claim does not provide objective boundaries or metrics by which a person having an ordinary skill in the art can determine or what constitutes as a “second dark web data”, not it is clear whether the second dark web data differs from the first dark web data.
Claims 1 and 7-8 are further rejected because the claims recite limitations involving machine learning, such as “pretraining a bidirectional encoder representations from transformers (BERT)-based language model using the first dark web data;fine-tuning the pretrained BERT-based language model using second dark web data”, without providing sufficient detail to inform, with reasonable certainty, those skilled in the art about the scope of the invention. Specifically, the claim language lacks clarity regarding the algorithm of BERT-based language model used, or how the model achieves the claimed determination.
As established in Nautilus, Inc. v. Biosig Instruments, Inc., 572 U.S. 898, 901, 910, 110 USPQ2d 1688, 1693 (2014), a claim is indefinite if, when read in light of the specification and the prosecution history, it fails to inform, with reasonable certainty, those skilled in the art about the scope of the invention. Additionally, MPEP § 2173.02 emphasizes that claims must be clear and precise to delineate the metes and bounds of the subject matter to be protected.
Dependent claims 2-13, 25-29 are also rejected for inheriting the deficiencies of the independent claims from which they depend on.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 5, 7 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (CN 117201055, hereinafter Liu) in view of Rao et al. (US 20240354556, hereinafter Rao).
Regarding claim 1: Liu teaches: A method of performing a task for cybersecurity on the basis of a dark web which is performed by a device, the method comprising:
acquiring raw dark web data from a database (Liu - [Page 4, Line 36]: S1, collecting text information in a normal webpage and a webpage hidden chain to form a data set, [Page 4, Line 54-55]: S1.3, setting a quantity threshold of dark chain samples in the dark chain data set, stopping acquisition of the data set when the quantity of the dark chain samples exceeds the quantity threshold);
acquiring first dark web data by preprocessing the raw dark web data (Liu - [Page 4, Line 38]: s2, extracting text information in the webpage to be detected);
pretraining a bidirectional encoder representations from transformers (BERT)-based language model using the first dark web data (Liu - [Page 4, Line 38-39]: extracting feature vectors from the text information by using the Bert model, and classifying the feature vectors by using three decision algorithms to obtain a decision set of the webpage to be detected);
fine-tuning the pretrained BERT-based language model using second dark web data (Liu - [Page 4, Line 48-50]: s5, judging whether to manually recheck the abnormal text in the abnormal text database according to the frequency of a certain type of abnormal text in the abnormal text database or whether the total number recorded in the abnormal text database exceeds a corresponding threshold value, and updating a data set and retraining the Bert model after the recheck);
However, Liu doesn’t explicitly teach, but Rao discloses: performing a task for cybersecurity using the fine-tuned BERT-based language model (Rao - [0085]: the generator model 310 is incrementally trained to generate sequences that resemble plausible sequences resembling actual sequences).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Liu with Rao so that the trained model can provide an output of plausible sequences resembling actual sequences. The modification would have allowed the system to output a sequence resembling the actual one.
Regarding claim 5: Liu as modified teaches: further comprising, when the task is threat keyword inference, masking one or more elements in the raw dark web data (Rao - [0081]: The training module 230 then generates a “corrupt” sequence by replacing original tokens for one or more masked positions in the training sequence with the generated sample (i.e., output token) from the generator model 310),
wherein the performing of the task for cybersecurity using the fine-tuned BERT-based language model comprises outputting a possibility value of at least one element corresponding to a masked position using the fine-tuned BERT-based language model (Rao - [0085]: The training module 230 obtains one or more terms from the loss function, and backpropagates parameters of the discriminator 320 model and the generator model 310 to reduce the loss function. In this manner, the generator model 310 is incrementally trained to generate sequences that resemble plausible sequences resembling actual sequences).
The reason to combine is in the same rational as claim 1.
Regarding claim 7: Claim is directed to apparatus/device claim and do not teach or further define over the limitations recited in claim 1. Therefore, claim 7 is also rejected for similar reasons set forth in claim 1.
Regarding claim 8: Claim is directed to computer readable medium claim and do not teach or further define over the limitations recited in claim 1. Therefore, claim 8 is also rejected for similar reasons set forth in claim 1.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (CN 117201055, hereinafter Liu) in view of Rao et al. (US 20240354556, hereinafter Rao) and Sofer et al. (US 2021/0056219).
Regarding claim 2: Liu as modified teaches wherein the acquiring of the first dark web data by preprocessing the raw dark web data comprises:
acquiring a dark web text dataset from the raw dark web data (s1.1, collecting URL addresses of a large number of websites);
balancing the dark web text dataset on the basis of categories (S1.2, classifying, labeling and storing each URL address in the collected URL address set to form a data set);
However, Liu as modified doesn’t explicitly teaches but Sofer discloses removing duplicate data of the dark web text dataset using a text similarity algorithm (Sofer - [0048]: a token-based similarity algorithm may be applied to them (or to their prefix-less/suffix-less versions) to detect duplicates and remove the redundancy).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Liu and Rao with Sofer so that duplicated data is removed from the data set. The modification would have allowed the system to be more efficient.
Claims 3-4 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (CN 117201055, hereinafter Liu) in view of Rao et al. (US 20240354556, hereinafter Rao) and GOVARDHAN et al. (CN 110413908).
Regarding claim 3: Liu as modified teaches collecting dark web data, classifying/labeling the data and training the BERT language model. However, Liu as modified doesn’t explicitly teaches but GOVARDHAN discloses ransomware leak sites from the raw dark web data (GOVARDHAN - [Page 4, Line 26-28]: the output analyzing module 224 uses machine learning techniques to determine the access classification of URL. access classification may include, but are not limited to, a phishing classification, adult classification, tissue specific constraint classification, malicious software classification, ransomware classification).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Liu and Rao with GOVARDHAN so that ransomware leak sites from the raw dark web data is used to train the language model. The modification would have allowed the system to be used for ransomware leak sites.
Regarding claim 4: Liu as modified teaches collecting dark web data, classifying/labeling the data and training the BERT language model. However, Liu as modified doesn’t explicitly teaches but GOVARDHAN discloses threat threads ffrom the raw dark web data (GOVARDHAN - [Page 4, Line 26-28]: the output analyzing module 224 uses machine learning techniques to determine the access classification of URL. access classification may include, but are not limited to, a phishing classification, adult classification, tissue specific constraint classification, malicious software classification, ransomware classification).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Liu and Rao with GOVARDHAN so that threat threads from the raw dark web data is used to train the language model. The modification would have allowed the system to be used for threat threads.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (CN 117201055, hereinafter Liu) in view of Rao et al. (US 20240354556, hereinafter Rao) and Yuan et al. (US 2025/0292074, hereinafter Yuan).
Regarding claim 6: Liu as modified teaches further comprising, when the task is threat keyword inference, masking one or more elements in the raw dark web data (Rao - [0081]: The training module 230 then generates a “corrupt” sequence by replacing original tokens for one or more masked positions in the training sequence with the generated sample (i.e., output token) from the generator model 310),
However, Liu as modified doesn’t explicitly teaches but Yuan some of the nonlinguistic elements from which linguistic meaning is inferable are included among targets of masking (Yuan - [0045]: Its Value Vector basically stores a vector close to the Embedding of the word “Beijing”. When transformer's input is “the capital of China is [Mask]”, the node detects this knowledge pattern from the input layer, so it generates a larger response output. It is assumed that other neurons in the key layer do not respond to this input, and the corresponding nodes in the value layer will receive the word embedding corresponding to the Value of “Beijing”, and further numerical amplification is performed through the RELU function. Therefore, the output corresponding to the Mask position will naturally output the word “Beijing”).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method of Liu and Rao with Yuan so that the actual word is outputted in place of masking position. The modification would have allowed the system to be more usable.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Uziel. et al. US 20250307550 - receiving a source dataset comprising a plurality of textual data instances and corresponding labels in two or more classes; training a machine learning classifier on the source dataset; performing inference by the trained machine learning classifier over a subset of the data instances in the source dataset, to extract a hidden representation for each of said data instances in said subset; applying a trained multilayer perceptron (MLP) network to the extracted hidden representations, to generate a set of corresponding soft prompts; and feeding the generated set of soft-prompts as prompts for a trained language model, to tune the trained language model to reconstruct the data instances in the subset.
AL-QURISHI US 9076021 B2- Shared Language model 620 (SLM): A pre-trained Language model (such as BERT, GPT, DeBERTa, ROBERTa, ELECTRA) is used as the base model, which is responsible for learning contextualized representations of the input text 602. This base model has several layers of transformers 614 and a final output layer 616 that produces hidden states H for each token in the input sequence.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MENG LI whose telephone number is (571)272-8729. The examiner can normally be reached M-F 8:30-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexander Lagor can be reached on (571) 270-5143. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MENG LI/
Primary Examiner, Art Unit 2437