DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
The following is a Non-Final Office Action in response to applicant’s filing on 11/04/2024. Claims 1-20 are pending.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 11/04/2024. The submissions are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Examiner Note
The term “one or more computer-readable storage media” recited in claim 10 has been interpreted to cover only a set of computer-executable instructions stored in a non-transitory and physical computer readable storage media in view of paragraphs [0021] of the specification which states “A computer program product embodiment (“CPP embodiment” or “CPP”) is a term used in the present disclosure to describe any set of one, or more, storage media (also called “mediums”) collectively included in a set of one, or more, storage devices that collectively include machine readable code corresponding to instructions and/or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device that can retain and store instructions for use by a computer processor. Without limitation, the computer readable storage medium may be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these mediums include: diskette, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), static random-access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device (such as punch cards or pits/lands formed in a major surface of a disc) or any suitable combination of the foregoing”.
Specification
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
The abstract of the disclosure is objected to because it is drafted in a manner that tracks the claim language and does not provide a concise summary of the technical disclosure. Specifically, the abstract reproduces claim terminology, including “adding an incoming first feature vector to an outer window. In response to a determination that the first feature vector is a qualifying feature vector, the first feature vector is added into a voting window. In response to a determination that the outer window is full, a relatively oldest feature vector is removed from the outer window”, rather than describing the invention in narrative form. Correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 9 and 18 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claim 9 recites “wherein the first feature vector is determined to be a qualifying feature vector in response to a determination that write operations are detected in a period covering the first feature vector”. The non-provisional specification fails to provide written description support for the claim limitation of “a period covering the first feature vector” (i.e., a qualifying feature vector may be a feature vector having write operations that are detected in a period covering the feature vector. Accordingly, in one or more of such approaches, a determination may be made as to whether the first feature vector is a qualifying feature vector (see decision 208), and more specifically, whether write operations are detected in a period covering the first feature vector, see paragraph [0048]). The specification fails to describe what constitutes the claimed “period”, how the boundaries of the period are established, and how a first feature vector is determined to be “covered” by the period. Accordingly, the originally filed disclosure does not provide adequate written description support for the claimed limitation.
The same reasons apply to dependent claim 18.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 is rejected as being indefinite. Claim 1 recites “adding an incoming first feature vector to an outer window; in response to a determination that the first feature vector is a qualifying feature vector, adding the first feature vector into a voting window;
in response to a determination that the outer window is full, removing a relatively oldest feature vector from the outer window; and using feature vectors in the voting window to infer ransomware activity”.
Claim 1 fails to clearly define the relationship between these windows and how their interaction contributes to inferring ransomware activity. In particular, claim 1 recites removing feature vectors from the outer window when full, while ransomware inference is performed using feature vectors in the voting window. However, the claim does not specify whether feature vectors removed from outer window remain in the voting window, whether the voting window is bounded by or derived from the outer window, or how management of the outer window affects the inference process. Accordingly, the scope of the relationship between the outer window, voting window and ransomware inference is unclear.
The term “relatively” in claim 1 is a relative term which renders the claim indefinite. The term “relatively” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The term “relatively” introduces ambiguity because it does not specify the “oldest” criterion used to determine a time frame for a feature vector such as a time defined of a feature vector in to the time window. Therefore, because the claim fails to define how relatively oldest feature vector is identified, the metes and bounds of the claim are unclear.
The term “relatively” in claim 2 is a relative term which renders the claim indefinite. The term “relatively” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The term “relatively” introduces ambiguity because it does not specify criterion used to determine a time frame for a feature vector such as a time defined of a feature vector in to the time window. Therefore, because the claim fails to define how relatively oldest feature vector is identified, the metes and bounds of the claim are unclear.
The term “relatively” in claim 3 is a relative term which renders the claim indefinite. The term “relatively” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The term “relatively” introduces ambiguity because it does not specify criterion used to determine a time frame for a feature vector such as a time defined of a feature vector in to the time window. Therefore, because the claim fails to define how relatively oldest feature vector is identified, the metes and bounds of the claim are unclear.
The same reasons apply to independent claims 10 and 19, and to the dependent claims by virtue of dependency to their independent claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1- 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ransomware Detection with Machine Learning in Storage Systems by Gagulic in view of Gao et al. (US 2016/0063357 A1), hereinafter Gao.
Regarding claim 1, the combination of Gagulic in view of Gao teaches a method comprising:
adding an incoming first feature vector to an outer window (Gagulic, Page. 19, In their research from 2019, their best-performing model reached an accuracy of 98 %. They collected input/output (IO) traces of a storage system where two ransomware (WannaCry and TeslaCrypt) and one benign workload (Zip) were run separately by using Wayback Visor, a hypervisor that is located between the hardware and the operating system. They started tracing immediately after the workload was started and continued tracing for 30 seconds. From these traces, five features were computed 4.1: These features are computed over different window sizes Twindow of one,five,ten,15,20 and 25 seconds with shifting the windows one second until the end of the trace time is reached);
in response to a determination that the first feature vector is a qualifying feature vector (Gao, Para. 0009, selecting the one or more non-support vector positive samples based on a distance between each of the one or more non-support vector positive samples and a decision boundary),
adding the first feature vector into a voting window; in response to a determination that the outer window is full (Gao, Para. 0053, assume a cache capacity of 20,000 (e.g., 20 k) samples, including a buffer of 1,000 (e.g., 1 k) samples. After a training procedure (e.g., after learning alpha values), one or more non-support vector positive samples 422 may be pruned in order to avoid exceeding 19,000 samples stored in the cache (or to avoid diminishing the buffer below 1,000). In other words, one or more non-support vector positive samples 422 may be pruned (e.g., pruned second) if necessary to avoid exceeding a sample number threshold),
removing a relatively oldest feature vector from the outer window (Gao, Para. 0056, the one or more non-support vector positive samples 422 may be ordered for pruning based on the age, where the one or more non-support vector positive samples 422 with greater ages (e.g., older samples) are pruned before one or more non-support vector positive samples 422 with lesser ages); and
using feature vectors in the voting window to infer ransomware activity (Gagulic, Page. 55, Feature I is a method of horizontally stacking all the other features original + A-C + E-H computed over windows. Thus, the impact of including the time aspect in the feature vector using horizontal stacking is evaluated in this experiment. Therefore, different feature vectors were extracted using the Feature Extractor explained in Subsection 4.3.5 from the self collected storage access patterns from six different ransomware and one benign workload as explained in Subsection 4.3.1.).
Regarding claim 2, the combination of Gagulic in view of Gao teaches the method of claim 1, further comprising:
in response to the determination that the outer window is full (Gao, Para. 0053, assume a cache capacity of 20,000 (e.g., 20 k) samples, including a buffer of 1,000 (e.g., 1 k) samples. After a training procedure (e.g., after learning alpha values), one or more non-support vector positive samples 422 may be pruned in order to avoid exceeding 19,000 samples stored in the cache (or to avoid diminishing the buffer below 1,000). In other words, one or more non-support vector positive samples 422 may be pruned (e.g., pruned second) if necessary to avoid exceeding a sample number threshold),
determining whether the relatively oldest feature vector is present in the voting window (Gao, Para. 0056, the one or more non-support vector positive samples 422 may be ordered for pruning based on the age, where the one or more non-support vector positive samples 422 with greater ages (e.g., older samples) are pruned before one or more non-support vector positive samples 422 with lesser ages); and
in response to a determination that the relatively oldest feature vector is present in the voting window (Gao, Para. 0053, assume a cache capacity of 20,000 (e.g., 20 k) samples, including a buffer of 1,000 (e.g., 1 k) samples. After a training procedure (e.g., after learning alpha values), one or more non-support vector positive samples 422 may be pruned in order to avoid exceeding 19,000 samples stored in the cache (or to avoid diminishing the buffer below 1,000). In other words, one or more non-support vector positive samples 422 may be pruned (e.g., pruned second) if necessary to avoid exceeding a sample number threshold),
removing the relatively oldest feature vector from the voting window (Gao, Para. 0056, the one or more non-support vector positive samples 422 may be ordered for pruning based on the age, where the one or more non-support vector positive samples 422 with greater ages (e.g., older samples) are pruned before one or more non-support vector positive samples 422 with lesser ages).
Regarding claim 3, the combination of Gagulic in view of Gao teaches the method of claim 1, wherein the relatively oldest feature vector is a second feature vector (Gao, Para. 0056, the non-support vector positive samples may then be ordered (e.g., sorted, indexed, etc.) to indicate an order (e.g., increasing order, decreasing order, etc.) of ages).
Regarding claim 4, the combination of Gagulic in view of Gao teaches the method of claim 1, wherein the first feature vector details feature information about operations performed within a storage system (Gagulic, Page. 12, the feature extraction and inference part can be executed directly in the storage system stack by using a computational storage architecture).
Regarding claim 5, the combination of Gagulic in view of Gao teaches the method of claim 4, wherein the feature information is selected from the group consisting of:
read transfer size (Gagulic, Page. 56, the read transfer size), write transfer size (Gagulic, Page. 56, write transfer size), an entropy of writes (Gagulic, Page. 56, write operations Entropy_mean_wr), an application tag, a logical block address (LBA) of a write operation (Gagulic, Page. 29, Table 4.4), and an LBA of a read operation (Gagulic, Page. 29, Table 4.4). Regarding claim 6,
the method of claim 4, the combination of Gagulic in view of Gao teaches further comprising: determining a response for mitigating a ransomware attack associated with the ransomware activity (Gagulic, Page. 44, to avoid information leakage between the train and the test set, windows 19 to 23 are thus removed from the train set and windows 14 to 18 remain in the test set); and
causing the response to be performed (Gagulic, Page. 14, provides specific behavior usable in incident response and forensic analysis).
Regarding claim 7, the combination of Gagulic in view of Gao teaches the method of claim 6, wherein the using feature vectors in the voting window to infer the ransomware activity comprises:
determining whether a majority of inferred votes on the feature vectors in the voting window exceed a dynamically adjustable threshold (Gagulic, Page. 22, Figure 4.2 shows the confusion matrix of seven ransomware and five benign software on Windows 7 with F1 Scores in 26, 12, and two classes evaluated on a Random Forest model with Twindow = ten seconds and Td = 90 seconds and using half of the files of the RanSAP dataset for training and half of the files for testing. It shows high F1 Scores of 96.2% for binary classification, 85.8% for multi-classification in 12 classes, and significant lower F1 Scores of 58% in 26 classes. In conclusion, ransomware can effectively be distinguished from benign software using Hirano et al.’s proposed model within the given conditions); and in response to a determination that the majority of the inferred votes on the feature vectors in the voting window exceed the dynamically adjustable threshold (Gagulic, Page. 43, The extract() function automatically appends the endings _ws{window_size}_wo{window_offset} to each file),
determining that the storage system is targeted by a ransomware attack (Gagulic, Page. 22, in conclusion, ransomware can effectively be distinguished from benign software using Hirano et al.’s proposed model within the given conditions).
Regarding claim 8, the combination of Gagulic in view of Gao teaches the method of claim 6, further comprising:
training an artificial intelligence (AI) engine to use training feature vectors in a training voting window to infer simulated ransomware activity (Gagulic, Page. 63, to test our models capability of detecting ransomware behavior in more realistic noisy environments, mixed workload traces used for testing have been collected for Lockbit (+Benign), BlackBasta (+Benign) and WannaCry (+Benign). The bash script (see Subsection 4.3.3) is executed just before the ransomware is run. As depicted in Table 5.7, all parameter of the bash script are set to 1 (random pauses, random traversing and random read only) to effectively simulate typical user behavior);
determining an accuracy of the AI engine (Gagulic, Page. 60); and
in response to a determination that the AI engine exceeds a predetermined threshold of accuracy (Gagulic, Page. 60, the models were trained and evaluated using 5-fold cross validation method described in Subsection4.3.5 and the average of evaluation F1-Scores are shown in Figure 5.12),
causing the AI engine to use the feature vectors in the voting window to infer the ransomware activity and determine the response for mitigating the ransomware attack associated with the ransomware activity (Gagulic, Page. 28, the established backdoor allows the monitoring of network traffic, which gives insights into the recovery approaches of the infected device. Further, this ransomware accelerates its spreading phase by using multithreading, which emphasizes the importance of quick detection and efficient countermeasures. Similar to the other selected ransomware, Conti employs a double extortion scheme).
Regarding claim 9, the combination of Gagulic in view of Gao teaches the method of claim 1, further comprising:
determining whether the first feature vector is a qualifying feature vector (Gagulic, Pages. 12-13, feature extraction and inference part can be executed directly in the storage system stack by using a computational storage architecture. And thus, (3) all these tasks can be executed without a significant impact on the host IO traffic. Nevertheless, interesting insights have also been presented from the research on detecting ransomware on the storage system level. Hirano and Kobayashi [10] were successful in distinguishing ransomware from benign workload by training three different machine learning models only with information gathered from storage systems, such as write and read throughput, the variance of logical block addresses (LBAs) accessed and the Shannon entropy of written sectors.),
wherein the first feature vector is determined to be a qualifying feature vector in response to a determination that write operations are detected in a period covering the first feature vector (Gagulic, Page. 13, Hirano and Kobayashi [10] were successful in distinguishing ransomware from benign workload by training three different machine learning models only with information gathered from storage systems, such as write and read throughput, the variance of logical block addresses (LBAs) accessed and the Shannon entropy of written sectors).
Regarding claim 10, the claim is interpreted and rejected for the same rational set forth
in claim 1.
Regarding claim 11, the claim is interpreted and rejected for the same rational set forth
in claim 2.
Regarding claim 12, the claim is interpreted and rejected for the same rational set forth
in claim 3.
Regarding claim 13, the claim is interpreted and rejected for the same rational set forth
in claim 4.
Regarding claim 14, the claim is interpreted and rejected for the same rational set forth
in claim 5.
Regarding claim 15, the claim is interpreted and rejected for the same rational set forth
in claim 6.
Regarding claim 16, the claim is interpreted and rejected for the same rational set forth
in claim 7.
Regarding claim 17, the claim is interpreted and rejected for the same rational set forth
in claim 8.
Regarding claim 18, the claim is interpreted and rejected for the same rational set forth
in claim 9.
Regarding claim 19, the claim is interpreted and rejected for the same rational set forth
in claim 1 and 10
Regarding claim 20, the claim is interpreted and rejected for the same rational set forth
in claim 2 and 11.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO-892.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GITA FARAMARZI whose telephone number is (571)272-0248. The examiner can normally be reached Monday- Friday 9:00 am- 6:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jorge L. Ortiz-Criado can be reached at (571)272-7624. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GITA FARAMARZI/Examiner, Art Unit 2496
/JORGE L ORTIZ CRIADO/Supervisory Patent Examiner, Art Unit 2496