Prosecution Insights
Last updated: October 02, 2026
Application No. 18/624,424

DATA LABELING USING A PREVALENCE-DRIVEN ARTIFICIAL INTELLIGENCE MODEL

Non-Final OA §101§103
Filed
Apr 02, 2024
Examiner
NIU, JIAHE
Art Unit
Tech Center
Assignee
CrowdStrike Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
6 currently pending
Career history
5
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Step 1 analysis for all claims: In the instant case, claims 1-7 are directed to a process, claims 8-14 are directed to manufacture and claims 15-20 are directed to a machine. Thus, each of the claims falls within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter). Claim 1: Step 2A, Prong 1 analysis: The claim(s) recite(s) in part: producing … a confidence level based on the hash; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses analyzing the hashing equation and determining on a scale of how clean a file is compared to it being malware. associating a label to the sample file based on the confidence level to produce a labeled sample file; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses using the level of how clean or malicious the file is to categorize the file. Step 2A, Prong 2 analysis: The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of: receiving a hash that corresponds to a sample file; which amounts to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. providing the hash to an artificial intelligence (AI) model which amounts to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. wherein the AI model is trained to utilize prevalence data corresponding to the hash to predict whether the sample file comprises malware; which are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)) by a processing device using the AI model; which are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements of: receiving a hash that corresponds to a sample file; This limitation is directed to receiving input at an interface on a computing device, wherein the input comprises a dataset, an analysis for the dataset, and an output medium which amounts to extra-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). providing the hash providing the hash to an artificial intelligence (AI) model, This limitation is directed to receiving input at an interface on a computing device, wherein the input comprises a dataset, an analysis for the dataset, and an output medium which amounts to extra-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). wherein the AI model is trained to utilize prevalence data corresponding to the hash to predict whether the sample file comprises malware; The machine learning model is recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training of the machine learning model is recited at a high-level of generality with no detail of the training process such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) by a processing device using the AI model; The computer elements are recited at a high level of generality and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. The training of the machine learning model is recited at a high-level of generality with no detail of the training process such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 2: Step 2A, Prong 1 analysis: The claim(s) recite(s) in part: analyzing content information corresponding to the sample file against a label rule, wherein the label rule determines whether the sample file comprises the malware; this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses determining whether of not a sample file has malware using information about the content information. comparing the confidence level to a threshold; and; this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses comparing two values. in response to the label rule determining that the sample file comprises the malware, and that the confidence level at least meets the threshold, flagging the sample file for further analysis; this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses categorizing the file so we can continue analyzing, if the file’s content is determined to malware and the confidence is above a threshold. Step 2A, Prong 2 analysis: There are no additional elements that individually or in combination integrate the judicial exception into a practical application. Step 2B analysis: There are no additional elements individually or in combination that amount to significantly more than the judicial exception. Claim 3: Step 2A, Prong 1 analysis: The claim(s) recite(s) in part: in response to the label rule determining that the sample file comprises the malware, and that the confidence level is below the threshold, using a dirty label to associate to the sample file.; this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses after deciding that the file needs further analysis from above, giving a label of dirty to the file. Step 2A, Prong 2 analysis: There are no additional elements that individually or in combination integrate the judicial exception into a practical application. Step 2B analysis: There are no additional elements individually or in combination that amount to significantly more than the judicial exception. Claim 4: Step 2A, Prong 1 analysis: The claim(s) recite(s) in part: in response to the label rule determining that the sample file is clean from the malware, and that the confidence level at least meets the threshold, using a clean label to associate to the sample file.; this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses after deciding that the file does not need further analysis from above, giving a label of clean to the file. Step 2A, Prong 2 analysis: There are no additional elements that individually or in combination integrate the judicial exception into a practical application. Step 2B analysis: There are no additional elements individually or in combination that amount to significantly more than the judicial exception. Claim 5: Step 2A, Prong 1 analysis: The claim(s) recite(s) in part: generating, based on the hash, a feature vector utilizing prevalence metadata, wherein the prevalence metadata comprises incidence information about the sample file; and; this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses using the hashing equation to come up with a vector of numbers to describe the metadata. utilizing … the feature vector during the producing of the confidence level.; this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses using the vector of numbers that describe the metadata to decide how clean a file is. Step 2A, Prong 2 analysis: AI model was addressed in the rejection of Claim 1. There are no additional elements that individually or in combination integrate the judicial exception into a practical application. Step 2B analysis: AI model was addressed in the rejection of Claim 1. There are no additional elements individually or in combination that amount to significantly more than the judicial exception. Claim 6: Step 2A, Prong 2 analysis: The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of: initiating a training of an AI-driven malware detector using the labeled sample file to reduce an amount of false positives of the malware by the AI-driven malware detector.; which are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)) Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements of: initiating a training of an AI-driven malware detector using the labeled sample file to reduce an amount of false positives of the malware by the AI-driven malware detector.; Training is recited at a high-level of generality with no detail of the training process such and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 7: Step 2A, Prong 1 analysis: The claim(s) recite(s) in part: generating a subsequent hash from the subsequent file; this limitation encompasses a Mathematical Concept (mathematical relationships, mathematical formulas or equations or mathematical calculations) producing … a subsequent confidence level based on the subsequent hash; and; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses analyzing the next hashing equation and determining on a scale of how clean a file is compared to it being malware. determining whether the subsequent file comprises the malware based on the subsequent confidence level; As drafted and under its broadest reasonable interpretation, this limitation covers performance of the limitation in the mind (including an observation, evaluation, judgment, opinion) or with the aid of pencil and paper but for the recitation of generic computer components. For example, this limitation encompasses using the level of how clean or malicious the next file is to categorize it if it actually has malware. Step 2A, Prong 2 analysis: The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of: receiving a subsequent file that is marked as comprising the malware; which amounts to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. providing the subsequent hash to the AI model; which amounts to extra-solution activity of gathering data for use in the claimed process. As described in MPEP 2106.05(g), limitations that amount to merely adding insignificant extra-solution activity to a judicial exception do not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application. by the processing device using the AI model; which are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements of: receiving a subsequent file that is marked as comprising the malware; This limitation is directed to receiving input at an interface on a computing device, wherein the input comprises a dataset, an analysis for the dataset, and an output medium which amounts to extra-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). providing the subsequent hash to the AI model; This limitation is directed to receiving input at an interface on a computing device, wherein the input comprises a dataset, an analysis for the dataset, and an output medium which amounts to extra-solution activity of gathering data for use in the claimed process. The courts have found limitations directed to obtaining information electronically, recited at a high level of generality, to be well-understood, routine, and conventional (see MPEP 2106.05(d)(II), “receiving or transmitting data over a network”, "electronic record keeping," and "storing and retrieving information in memory"). by the processing device using the AI model; Computer elements are recited at a high-level of generality such and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 8: Claim 8 recites substantially similar limitations for claim 1 with an additional mental process and additional elements. and is therefore rejected on the same basis. Step 2A, Prong 1 analysis: The claim(s) recite(s) in part: generate a hash from a sample file; This limitation encompasses a Mathematical Concept (mathematical relationships, mathematical formulas or equations or mathematical calculations) Step 2A, Prong 2 analysis: The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of: a memory to store instructions that, when executed by the processing device, cause the processing device to; which are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the exception using a generic computer component (See MPEP 2106.05(f)) Accordingly, at Step 2A, prong two, the additional elements individually or in combination do not integrate the judicial exception into a practical application. Step 2B analysis: In accordance with Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more that the judicial exception. As discussed above, the additional elements of: a memory to store instructions that, when executed by the processing device, cause the processing device to Computer elements are recited at a high-level of generality such and amounts to no more than adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea (See MPEP 2106.05(f)) Accordingly, at Step 2B the additional elements individually or in combination do not amount to significantly more than the judicial exception. Claim 9: Claim 9 recites substantially similar limitations for claim 2 and is therefore rejected on the same basis. Claim 10: Claim 10 recites substantially similar limitations for claim 3 and is therefore rejected on the same basis. Claim 11: Claim 11 recites substantially similar limitations for claim 4 and is therefore rejected on the same basis. Claim 12: Claim 12 recites substantially similar limitations for claim 5 and is therefore rejected on the same basis. Claim 13: Claim 13 recites substantially similar limitations for claim 6 and is therefore rejected on the same basis. Claim 14: Claim 14 recites substantially similar limitations for claim 7 and is therefore rejected on the same basis. Claim 15: Claim 15 recites substantially similar limitations for claim 8 and is therefore rejected on the same basis. Claim 16: Claim 16 recites substantially similar limitations for claim 2 and is therefore rejected on the same basis. Claim 17: Claim 17 recites substantially similar limitations for claim 3 and is therefore rejected on the same basis. Claim 18: Claim 18 recites substantially similar limitations for claim 4 and is therefore rejected on the same basis. Claim 19: Claim 19 recites substantially similar limitations for claim 5 and is therefore rejected on the same basis. Claim 20: Claim 20 recites substantially similar limitations for claim 7 and is therefore rejected on the same basis. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C 103 as being unpatentable over Oliver et al. ( US 11182481 B1, hereinafter Oliver) in view of Briliauskas (US 20240232349 A1, hereinafter Briliauskas) in further view of Friedrichs et a. (US 20140165203 A1, hereinafter Friedrichs). Claim 1: Oliver teaches: A method comprising: Oliver [col 9 lines 42-43] Methods and systems for evaluating files for cyber threats have been disclosed receiving a hash that corresponds to a sample file; Oliver [col 4 lines 5-6] In the example of FIG. 1, the front end system 210 receives a query (see arrow 201) for a target file 218. [col 4 lines 16-18] In other embodiments, the query includes the target locality sensitive hash but not the file 218, in which case the front end system 210 receives the file 218 from some other source. EN: target file reads on sample file; the front end receives a query that has the hash providing the hash to an artificial intelligence (AI) model, Oliver [col 2 lines 31-34] In the example of FIG. 1, the front end system 210 includes a machine learning model 211 for detecting malicious files, e.g., viruses, worms, advanced persistent threats, Trojans, and other cyber threats. EN: as cited above the query which includes a hash is provided to the front end; this passage denotes that the front end includes a machine learning model; thus the hash is provided to the ai model; machine learning model reads on ai model wherein the AI model is trained to utilize … data corresponding to the hash to predict whether the sample file comprises malware; Oliver [col 2 lines 41-44] Various features of the sample files may be used in the training of the machine learning model 211, including header data, file size, opcodes used, sections, etc. Nechita [0015] The prevalence data includes, for example, metadata pertaining to the occurrence or frequency of file types, file names, file properties, or a combination thereof. In some embodiments, the prevalence data is collected from various customer systems. EN: specifications of this patent application show that prevalence data reads on metadata by a processing device using the AI model Oliver [col 9 lines 13-21] Referring now to FIG. 6, there is shown a logical diagram of a computer system 100 that may be employed with embodiments of the present invention. The computer system 100 may be employed as the front end system 210, the backend system 216, or other computer described herein. The computer system 100 may have fewer or more components to meet the needs of a particular cybersecurity application. The computer system 100 may include one or more processors 101. EN: ML model is part of the front end; front end reads on computer system which includes processors associating a label to the sample file … to produce a labeled sample file. Oliver [col 4 lines 19 - 30] In response to the query, the front end system 210 uses the machine learning model 211 to evaluate the file 218. More particularly, the front end system 210 may input the file 218 to the machine learning model 211, which classifies the file 218. Depending on implementation details of the machine learning model 211, features may be extracted from the file 218, the contents of the file 218 loaded in memory, or some other form of the file 218 and input to the machine learning model 211 for classification. In one embodiment, the machine learning model 211 gives a positive result when the file 218 is classified as malicious, and a negative result when the file 218 is classified as normal (i.e., not malicious). EN: the machine learning model classifies the file as normal or malicious which reads on labeling Oliver does not explicitly teach a confidence level being produced based on the hash, however Briliauskas teaches: producing … a confidence level based on the hash; and Briliauskas [0044] Locality-Sensitive Hashing Scan in a Vantage-Point Tree Structure. At the client device(s) 104a, 104b, the model 108 (shown as 108′) of the locality-sensitive hashing operation with the vantage-point tree data structure (also referred to as a VPT hash classification model 108) can be employed to predict or provide a likelihood or confidence value or score (i) whether a target code 119 (e.g., operating system files, application files, emails, browser data, API calls, etc. stored in memory 116 of the device) is malicious, or non-malicious, based on the fuzzy hash space to the nearest neighbor to known malicious files or code and (ii) whether the target code 119 is non-malicious, or malicious, based on the distance to nearest neighbor known clean files or codes. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the method of evaluating files for cyber threats using hashing and machine learning of Oliver with the method of classifying malicious and non-malicious files using hashing along with confidence scores of Briliauskas in order to see how reliable the prediction is and whether or not further analysis is needed. A high confidence determination would save computational resources as no further analysis is necessary. However, with low confidence determination allows the system can continue to analyze the file in order to make a more reliable decision. Briliauskas [0038] Various implementations of the present disclosure disclose methods for performing malware detection/classification operations that improve the efficiency and/or reliability of these steps/operations. [0046] The threshold operator 124 determines whether the output value 121 is in a range 126a associated with high confidence that the target code 119 is a malicious file, a range 126b associated with high confidence that the target code 119 is a non-malicious file, or a range 126c associated with low confidence of either (shown as “low confidence” or unknown). That is, the fuzzy hash space of the target code 119 indicates it is being searched against a set of malicious or non-malicious files that appears to be different from those used in the training data set 112, 114. The combination of Oliver and Briliauskas does not explicitly teach a confidence level being produced based on the hash, however Friedrichs teaches: prevalence data Friedrichs [0064] According to another aspect of the present invention is an intelligent filtering component. This component examines metadata gathered on a plurality of files from a plurality of devices on which these files reside and identifies a subset of these files that are suitable candidates for rescanning. This component can use numerous characteristics for determining whether a file is suitable for rescanning. … The characteristics used to determine whether a file is a suitable candidate can include, but are not limited to, the following. [0069] A fifth consideration is whether the file's prevalence among users exceeds a pre-defined threshold (for example, the file is on known to be on more than 50 systems). EN: this passage describes the meta data that is used; one part if the metadata that is gathered is the prevalence data Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the method of evaluating files for cyber threats using hashing and machine learning of Oliver and Briliauskas with the method of classifying malicious and non-malicious files using hashing along with the use of prevalence data for detecting malicious software of Friedrichs in order to provide more useful information for whether or not a file is malicious. Friedrichs [0069] This information is useful in the context of identifying malware since it is known in the art that most malware has low prevalence (e.g., is either unique or is on a handful of systems). If a file was previously marked malicious, and now is seen to have high prevalence, then it is possible a mistake was made earlier and the file is now clean. More so, even if the file is still believed to be malicious, there may be additional intelligence regarding that file that is now available (such as the malware family it belongs to, the category of malware to which it can be classified, or the actions associated with this type of malware). Along similar lines, if a file was marked clean earlier, and it's prevalence is higher, it is helpful to recheck whether it is still believed to be clean since we expect that as a file's prevalence increases, so too does the likelihood that there is useful intelligence regarding that file). Claim 2: The combination of Oliver, Briliauskas, and Friedrichs teaches all of the limitations of claim 1 as cited above and Oliver further teaches: analyzing content information corresponding to the sample file against a label rule, wherein the label rule determines whether the sample file comprises the malware; Oliver [col 4 lines 19 - 30] Depending on implementation details of the machine learning model 211, features may be extracted from the file 218, the contents of the file 218 loaded in memory, or some other form of the file 218 and input to the machine learning model 211 for classification. EN: content information is input into a ML model to label the file; ML model reads on label rule in response to the label rule determining that the sample file comprises the malware … flagging the sample file for further analysis. Oliver [col 5 lines 2 - 6] Otherwise, when the machine learning model 211 classifies the file 218 as malicious but response actions for the file 218 have been disabled, the front end system 210 declares the file 218 to be normal and returns a negative result in response to the query. EN: ML model (which reads on label rule as cited above) determines file comprises malware and returns a negative result [col 5 lines 32 - 38] More particularly, when an entry in the query log 215 indicates that a query for a particular file is returned a negative result, the backend system 216 may retrieve the particular file from the query log 215 or other source. The backend system 216 may reevaluate the particular file for cyber threats using cybersecurity evaluation procedures that are more extensive than the machine learning model 211. EN: files that have a negative result are retrieved for further analysis Oliver does not teach a confidence level, but as cited above Briliauskas teaches a confidence level. Briliauskas further teaches: comparing the confidence level to a threshold; and Briliauskas [0046] The output 121 of the VPT hash classification model 108″ is provided to a threshold operator 124. The threshold operator 124 determines whether the output value 121 is in a range 126a associated with high confidence that the target code 119 is a malicious file, a range 126b associated with high confidence that the target code 119 is a non-malicious file, or a range 126c associated with low confidence of either (shown as “low confidence” or unknown). Claim 3: The combination of Oliver, Briliauskas, and Friedrichs teaches all of the limitations of claim 2 as cited above and Oliver further teaches: in response to the label rule determining that the sample file comprises the malware … using a dirty label to associate to the sample file. Oliver [col 4 lines 19 - 30] In response to the query, the front end system 210 uses the machine learning model 211 to evaluate the file 218. More particularly, the front end system 210 may input the file 218 to the machine learning model 211, which classifies the file 218. Depending on implementation details of the machine learning model 211, features may be extracted from the file 218, the contents of the file 218 loaded in memory, or some other form of the file 218 and input to the machine learning model 211 for classification. In one embodiment, the machine learning model 211 gives a positive result when the file 218 is classified as malicious, and a negative result when the file 218 is classified as normal (i.e., not malicious). EN: classified as malicious reads on labeling as dirty with malware Oliver does not explicitly teach a confidence level below a threshold, however Briliauskas further teaches a range that the confidence level would be below a threshold when a file is labeled malicious: and that the confidence level is below the threshold, Briliauskas [0046] The output 121 of the VPT hash classification model 108″ is provided to a threshold operator 124. The threshold operator 124 determines whether the output value 121 is in a range 126a associated with high confidence that the target code 119 is a malicious file, a range 126b associated with high confidence that the target code 119 is a non-malicious file, or a range 126c associated with low confidence of either (shown as “low confidence” or unknown). EN: low confidence reads on confidence below a threshold as the model is does not have a high confidence of whether it is malicious Claim 4: The combination of Oliver, Briliauskas, and Friedrichs teaches all of the limitations of claim 2 as cited above, and Oliver further teaches: in response to the label rule determining that the sample file is clean from the malware … using a clean label to associate to the sample file. Oliver [col 4 lines 19 - 30] In response to the query, the front end system 210 uses the machine learning model 211 to evaluate the file 218. More particularly, the front end system 210 may input the file 218 to the machine learning model 211, which classifies the file 218. Depending on implementation details of the machine learning model 211, features may be extracted from the file 218, the contents of the file 218 loaded in memory, or some other form of the file 218 and input to the machine learning model 211 for classification. In one embodiment, the machine learning model 211 gives a positive result when the file 218 is classified as malicious, and a negative result when the file 218 is classified as normal (i.e., not malicious). EN: classified as normal reads on labeling as clean from malware Oliver does not explicitly teach labeling when the confidence label is in a specific range, however Briliauskas further teaches a range that a confidence level is above a threshold when a file is labeled normal: and that the confidence level at least meets the threshold, Briliauskas [0046] The output 121 of the VPT hash classification model 108″ is provided to a threshold operator 124. The threshold operator 124 determines whether the output value 121 is in a range 126a associated with high confidence that the target code 119 is a malicious file, a range 126b associated with high confidence that the target code 119 is a non-malicious file, or a range 126c associated with low confidence of either (shown as “low confidence” or unknown). EN: high confidence that a file is non-malicious reads on the confidence level of a sample file being clean from malware being at least at a threshold Claim 5: The combination of Oliver, Briliauskas, and Friedrichs teaches all of the limitations of claim 1 as cited above, and Oliver further teaches: generating, based on the hash, a feature vector Oliver [col 2 lines 58 -61] Generally speaking, a locality sensitive hashing algorithm may extract many very small features (e.g., 3 bytes) of a file and put the features into a histogram, which is encoded to generate the digest of the file. EN: the encoded digest reads on a vector made using small features about the file The combination of Oliver and Briliauskas does not distinctly teach using prevalence metadata, however, Friedrichs teaches: feature vector utilizing prevalence metadata, wherein the prevalence metadata comprises incidence information about the sample file; and [0064] According to another aspect of the present invention is an intelligent filtering component. This component examines metadata gathered on a plurality of files from a plurality of devices on which these files reside and identifies a subset of these files that are suitable candidates for rescanning. This component can use numerous characteristics for determining whether a file is suitable for rescanning. … The characteristics used to determine whether a file is a suitable candidate can include, but are not limited to, the following. [0069] A fifth consideration is whether the file's prevalence among users exceeds a pre-defined threshold (for example, the file is on known to be on more than 50 systems). [0077] The prevalence of the file either as determined from the community of users who are running a specific piece of client software or through third-party intelligence (note that for these purposes, knowing the exact prevalence is not strictly necessary--for example, it may be sufficient to know whether the file was never seen before, whether it was seen but only on a small number of machines, or whether it was seen on a large number of machine); EN: this reads on incidence information [0074] additional ancillary data can be computed from the meta-data; e.g., the overall prevalence of the file, the frequency with which the file was seen during given time window, the number of malicious files associated with a given user (both overall and within a specific window), etc [0061] This metadata can include one or more cryptographic hash values, one or more fuzzy hash values, one or more machine learning feature vectors, or some similar pieces of information that can be used in the art for identifying malware. The meta-data can also include behavioral characteristics and broader system characteristics (that can be encoded within machine learning feature vectors). EN: this passages reads on metadata being ended within machine learning feature vectors; utilizing, by the AI model, the feature vector during the producing of the confidence level. Friedrichs [0074] Third, a disposition is determined based on the information provided (the disposition can be determined through a separate module or through any technique known in the art, including, but not limited to: checking whitelists/blacklists for the presence of fingerprints or fuzzy fingerprints; applying a machine learning classifier to the feature vectors; using any characteristics of the system on which the file recently came, such as its recent infection history; or using any aggregate information gathered about the file such as its patterns across a plurality of users and its prevalence). The disposition can be good, bad, or unknown. Note that from the perspective of a client system, in some instances an unknown file might be allowed to continue remaining on the system (i.e., it will be treated in a manner similar to that of a good file), whereas in other instances (such as a system that contains sensitive data or is in a sensitive location, an unknown can be blocked (i.e., treated in a manner similar to that of a malicious file). Further, the disposition can include a confidence value (or both the disposition and the confidence value can be encoded in a single number; for example, a number between 0 and 1 where 0 means good and 1 means bad, and numbers closer to one are more likely to be bad, in which case 0.85 would mean an approximately 85% chance the file is malicious). EN: machine learning classifier reads on AI model; this reads on using feature vectors for the calculation of confidence Claim 6: The combination of Oliver, Briliauskas, and Friedrichs teaches all of the limitations of claim 1 as cited above including an AI model and Oliver further teaches: initiating a training of an AI-driven malware detector using the labeled sample file to reduce an amount of false positives of the malware by the AI-driven malware detector. Oliver [col 8 lines 36-44] As can be appreciated, false positive and false negative classifications of a machine learning model are more properly addressed by retraining the machine learning model with the corresponding false positive/false negative files. However, retraining the machine learning model takes time. Embodiments of the present invention advantageously address false positive and false negative classifications of a machine learning model, until such time that the machine learning model can be retrained. EN: retraining the ML model using the false positives reads on reducing the amount of false positives Claim 7: The combination of Oliver, Briliauskas, and Friedrichs teaches all of the limitations of claim 1 as cited and Oliver further teaches: providing the … hash to the AI model; Oliver [col 2 lines 31-34] In the example of FIG. 1, the front end system 210 includes a machine learning model 211 for detecting malicious files, e.g., viruses, worms, advanced persistent threats, Trojans, and other cyber threats. EN: as cited above the query which includes a hash is provided to the front end; this passage denotes that the front end includes a machine learning model; thus the hash is provided to the ai model; machine learning model reads on ai model by the processing device using the AI model, Oliver [col 9 lines 13-21] Referring now to FIG. 6, there is shown a logical diagram of a computer system 100 that may be employed with embodiments of the present invention. The computer system 100 may be employed as the front end system 210, the backend system 216, or other computer described herein. The computer system 100 may have fewer or more components to meet the needs of a particular cybersecurity application. The computer system 100 may include one or more processors 101. EN: ML model is part of the front end; front end reads on computer system which includes processors determining whether the … file comprises the malware ... Oliver [col 4 lines 19 - 30] In response to the query, the front end system 210 uses the machine learning model 211 to evaluate the file 218. More particularly, the front end system 210 may input the file 218 to the machine learning model 211, which classifies the file 218. Depending on implementation details of the machine learning model 211, features may be extracted from the file 218, the contents of the file 218 loaded in memory, or some other form of the file 218 and input to the machine learning model 211 for classification. In one embodiment, the machine learning model 211 gives a positive result when the file 218 is classified as malicious, and a negative result when the file 218 is classified as normal (i.e., not malicious). EN: the machine learning model classifies the file as normal or malicious which reads on determining Oliver does not explicitly teach a confidence level being produced based on the hash, however Briliauskas teaches: producing a … confidence level based on the … hash; and Briliauskas [0044] Locality-Sensitive Hashing Scan in a Vantage-Point Tree Structure. At the client device(s) 104a, 104b, the model 108 (shown as 108′) of the locality-sensitive hashing operation with the vantage-point tree data structure (also referred to as a VPT hash classification model 108) can be employed to predict or provide a likelihood or confidence value or score (i) whether a target code 119 (e.g., operating system files, application files, emails, browser data, API calls, etc. stored in memory 116 of the device) is malicious, or non-malicious, based on the fuzzy hash space to the nearest neighbor to known malicious files or code and (ii) whether the target code 119 is non-malicious, or malicious, based on the distance to nearest neighbor known clean files or codes. determining whether the … file comprises the malware based on the … confidence level. Briliauskas [0046] The output 121 of the VPT hash classification model 108″ is provided to a threshold operator 124. The threshold operator 124 determines whether the output value 121 is in a range 126a associated with high confidence that the target code 119 is a malicious file, a range 126b associated with high confidence that the target code 119 is a non-malicious file, or a range 126c associated with low confidence of either (shown as “low confidence” or unknown). EN: if output value 121 is in range 126a then the target file is determined with high confidence that the file is malicious; this reads on determining if file comprises hardware based on confidence level The combination of Oliver and Briliauskas does not teach a subsequent file, however Friedrichs further teaches: receiving a subsequent file that is marked as comprising the malware; Friedrichs [0042] In addition to an initial scan, a file may be periodically rescanned if either its initial disposition was inconclusive, or the confidence associated with the disposition was below a certain threshold, or there is reason to believe that the initial disposition was incorrect, or even as a general safety mechanism to safeguard against the possibility that the initial disposition was incorrect, or to see if there is additional information available about that file that might be of interest to an end user or system administrator (e.g., a file may have been determined malicious through some generic means, but now more information is known such as the malicious software family it belongs to, or in what category of malware it can be placed, or what types of actions are associated with this type of malware). File rescanning typically involves cross referencing large file sample collections against the latest threat intelligence. EN: rescanned and processed again reads on receiving a subsequent filed that is already marked; this passage describes a malicious file that was marked malicious and is rescanned in order to classify it further generating a subsequent hash from the subsequent file; Friedrichs [0060] According to one aspect of the present invention, a system is provided for intelligently rescanning files. The system includes a client and server component, which are capable of communicating with each other either directly or indirectly. [0061] … a client-side metadata extraction component is provided that can execute the following steps: First, it identifies Files of interest. … Second, the meta-data extraction module extracts meta-data from this new file. This metadata can include one or more cryptographic hash values, one or more fuzzy hash values, one or more machine learning feature vectors, or some similar pieces of information that can be used in the art for identifying malware. The meta-data can also include behavioral characteristics and broader system characteristics (that can be encoded within machine learning feature vectors). EN: after a file is chosen to get rescanned, the extraction module will extract a hash value Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the method of evaluating files for cyber threats using hashing and machine learning of Oliver and Briliauskas with the method of classifying malicious and non-malicious files using hashing along with the method of rescanning to detect malicious software of Friedrichs in order to improve reliability of the malware classification by using the new information the system receives. Friedrichs [0018] According to another aspect of the present invention is a server-side rescanning module that re-examines files and file metadata from the intelligent filtering module, and updates an intelligence database accordingly. This module can also be used to identify endpoints on which a discrepancy exists and inform those endpoints about this discrepancy. Claim 8: Claim 8 recites substantially similar limitations for claim 1 other than the generation of a hash and the memory and is therefore rejected on the same basis. Oliver further teaches: a memory to store instructions that, when executed by the processing device, cause the processing device to: Oliver [col 9 lines 32-41] The computer system 100 is a particular machine as programmed with one or more software modules 110, comprising instructions stored non-transitory on the main memory 108 for execution by the processor 101 to cause the computer system 100 to perform corresponding programmed steps. An article of manufacture may be embodied as computer-readable storage medium including instructions that when executed by the processor 101 cause the computer system 100 to be operable to perform the functions of the one or more software modules 110. generate a hash from a sample file; Oliver [col 6 lines 46-48] The file evaluation interface 273 may generate a target locality hash of the target file or receive the target locality hash as part of the query. Claim 9: Claim 9 recites substantially similar limitations for claim 2 and is therefore rejected on the same basis. Claim 10: Claim 10 recites substantially similar limitations for claim 3 and is therefore rejected on the same basis. Claim 11: Claim 11 recites substantially similar limitations for claim 4 and is therefore rejected on the same basis. Claim 12: Claim 12 recites substantially similar limitations for claim 5 and is therefore rejected on the same basis. Claim 13: Claim 13 recites substantially similar limitations for claim 6 and is therefore rejected on the same basis. Claim 14: Claim 14 recites substantially similar limitations for claim 7 and is therefore rejected on the same basis. Claim 15: Claim 15 recites substantially similar limitations for claim 8 and is therefore rejected on the same basis. Claim 16: Claim 16 recites substantially similar limitations for claim 2 and is therefore rejected on the same basis. Claim 17: Claim 17 recites substantially similar limitations for claim 3 and is therefore rejected on the same basis. Claim 18: Claim 18 recites substantially similar limitations for claim 4 and is therefore rejected on the same basis. Claim 19: Claim 19 recites substantially similar limitations for claim 5 and is therefore rejected on the same basis. Claim 20: Claim 20 recites substantially similar limitations for claim 7 and is therefore rejected on the same basis. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIAHE NIU whose telephone number is (571)270-0152. The examiner can normally be reached 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Omar Fernandez Rivas can be reached at (571) 272-2589. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JIAHE NIU/Examiner, Art Unit 2128 /OMAR F FERNANDEZ RIVAS/Supervisory Patent Examiner, Art Unit 2128
Read full office action

Prosecution Timeline

Apr 02, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month