Prosecution Insights
Last updated: October 01, 2026
Application No. 18/756,540

TECHNIQUES FOR DETECTING FILE SIMILARITY

Final Rejection §103
Filed
Jun 27, 2024
Examiner
ADAMS, CHARLES D
Art Unit
2152
Tech Center
2100 — Computer Architecture & Software
Assignee
CrowdStrike Inc.
OA Round
4 (Final)
45%
Grant Probability
Moderate
5-6
OA Rounds
2y 8m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 45% of resolved cases
45%
Career Allowance Rate
194 granted / 432 resolved
-10.1% vs TC avg
Strong +44% interview lift
Without
With
+43.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 11m
Avg Prosecution
23 currently pending
Career history
462
Total Applications
across all art units

Statute-Specific Performance

§101
21.6%
-18.4% vs TC avg
§103
56.0%
+16.0% vs TC avg
§102
11.6%
-28.4% vs TC avg
§112
8.6%
-31.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 432 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 8-12, and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al. (US Pre-Grant Publication 2021/0294840), in view of Park et al. (US Pre-Grant Publication 2024/0160890), and further in view of Larkin et al. (US Pre-Grant Publication 2023/0057414). As to claim 1, Lee teaches a method comprising: … providing to the model, a set of files, wherein the ML model is configured to generate, based on the set of files, a feature vector database comprising a set of feature vectors, wherein each of the set of feature vectors corresponds to a particular file of the set of files and wherein the set of feature vectors is grouped based on … characteristics (see Lee paragraphs [0024]-[0025]. A set of music files may be used to train a neural network. Feature vectors of the music files are generated. Each of the feature vectors is generated from a music file. The feature vectors may be grouped based on characteristics into subspaces corresponding to musical attributes); in response to receiving a query file to be compared to the set of files, processing the query file using the ML model to generate a query feature vector (see Lee paragraphs [0026]-[0027]. The user may supply a query music file to the system. The system will generate a feature vector of the query music file); and querying, by a processing device, the feature vector database using the query feature vector to identify one or more of the set of files that are similar to the query file (see Lee paragraphs [0027]-[0028]. Based on the feature vector, the system will search for music file similar to the query music file. As noted in paragraph [0028], the searching compares the query music file vector to the stored feature vectors). Lee does not explicitly teach: Training, over a plurality of steps, a machine learning model to group files based on a hierarchy of characteristics, wherein at each of the plurality of steps, the ML model is trained to group files iteratively, and wherein at each progressive iteration the ML model learns to group files based on a characteristic from the hierarchy of characteristics that is progressively lower on the hierarchy of characteristics; wherein the set of feature vectors is grouped based on a hierarchy of characteristics; Park teaches: Training, over a plurality of steps, a machine learning model to group files based on … characteristics, wherein at each of the plurality of steps, the ML model is trained to group files iteratively, and wherein at each progressive iteration the ML model learns to group files based on a particular characteristic from the … characteristics (see paragraph [0123]. Park teaches to cluster (or group) nodes (or files) based on characteristics); It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Lee by the teachings of Park, because both references are directed towards training data. Park merely adds to Lee an additional method to train the data, which will help to identify groups of entities within an entity network based on relationships between the entities. The learning process of Park will help to categorize data in a more accurate manner (see Park paragraph [0032]). Larkin teaches: Training, over a plurality of steps, a machine learning model to [match] files based on a hierarchy of characteristics, wherein at each of the plurality of steps, the ML model is trained to group files iteratively (see Larkin paragraphs [0043] and [0094]. Larkin relies upon a hierarchy of match models, such that each iteration is associated with a different matching model of a hierarchical string matching machine learning framework), and wherein at each progressive iteration the ML model learns to group files based on a characteristic from the hierarchy of characteristics that is progressively lower on the hierarchy of characteristics (see Larkin paragraphs [0043] and [0094]-[0096]. A sequence of progressively “lower” match models in a hierarchy is used with each iteration. It is noted that Applicant does not define “lower” nor provide any details regarding the nature of the hierarchy of characteristics); wherein the set of feature vectors is grouped based on the hierarchy of characteristics (see paragraphs [0043] and [0094]-[0096] for a hierarchy of characteristics. It is noted that Lee paragraphs [0024]-[0025] teach to calculate feature vectors based on characteristics. Larkin is simply relied upon to show wherein those characteristics used by a machine learning model may be based on a hierarchy); It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Lee by the teachings of Larkin, because both references are directed towards searching for files using feature vectors. Larkin merely adds to Lee an additional method to recognizing matching of data, notably based on a hierarchy of matching characteristics. The process of Larkin will help to improve predictive accuracy of string-based machine learning models (see Larkin paragraph [0021]). As to claim 2, Lee as modified by Park teaches the method of claim 1, wherein the ML model is trained using training data comprising a plurality of training data batches, wherein each of the plurality of training data batches comprises a set of training files with a label for each characteristic in the hierarchy of characteristics (see Lee paragraph [0025]. The machine learning model is trained using a plurality of training data. Each of the model subspaces is trained using a different label for each characteristic. Larkin teaches the existence of a hierarchy of characteristics, see paragraphs [0043] and [0094]-[0096]) and wherein training the ML model comprises: at each of the plurality of steps: grouping, using the ML model, a respective training data batch iteratively based on the hierarchy of characteristics to generate an output for each iteration (see Park paragraph [0061] for training based on different clusters. See Park paragraph [0074], which indicates how there are multiple iterations) ; and for each iteration: analyzing the output with a hierarchical contrastive learning (HCL) loss function to determine a loss value (see Park paragraph [0061]. A contrastive learning loss function is used to identify hierarchical community loss); and adjusting one or more weights of the ML model based at least in part on the loss value (see Park paragraph [0075]. Loss is measured and weights are adjusted in the model). As to claim 3, Lee as modified by Park teaches the method of claim 2, wherein training the ML model further comprises: for each iteration: analyzing the output with a focal loss function to determine a second loss value (see Park paragraph [0061]. Two samples may be used to calculate a total loss value); and adding the loss value and the second loss value to generate a total loss value, wherein the one or more weights of the ML model are adjusted based on the total loss value (see Park paragraph [0061]. The two samples are added to generate a community loss. As noted in paragraph [0075], node weights may be adjusted to minimize a loss). As to claim 4, Lee as modified teaches the method of method of claim 1, wherein querying the feature vector database using the query feature vector comprises: using a nearest neighbors algorithm to identify from the feature vector database, one or more of the set of feature vectors that are similar to the query feature vector (see Lee paragraph [0160]). As to claim 5, Lee as modified teaches the method of method of claim 4, further comprising: for each of the identified one or more feature vectors, retrieving a file from the set of files corresponding to the identified feature vector to obtain the one or more of the set of files that are similar to the query file set (see Lee paragraphs [0028]-[0029] and [0160]); and providing the one or more of the set of files that are similar to the query file as a result set (see Lee paragraphs [0028]-[0029] and [0160]). As to claims 8 and 15, see the rejection of claim 1. As to claims 9 and 16, see the rejection of claim 2. As to claims 10 and 17, see the rejection of claim 3. As to claims 11 and 18, see the rejection of claim 4. As to claims 12 and 19, see the rejection of claim 5. Claims 6, 13, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al. (US Pre-Grant Publication 2021/0294840), in view of Park et al. (US Pre-Grant Publication 2024/0160890), in view of Larkin et al. (US Pre-Grant Publication 2023/0057414), and further in view of Srivastava et al. (US Pre-Grant Patent 8,561,193). As to claim 6, Lee as modified teaches the method of method of claim 1. Lee does not teach wherein the hierarchy of characteristics comprises: threat type, malware family, subtype, compiler, packer, and library. Srivastava teaches wherein the hierarchy of characteristics comprises: threat type, malware family, subtype, compiler, packer, and library (see Srivastava 5:10-37. Srivastava teaches wherein each of these characteristics may be extracted and recorded as part of a file. It is noted that Lee extracts characteristics from a file to use when creating feature vectors. Larkin teaches the creation of a hierarchy of features. Srivastava simply shows wherein such features may be related to malware attributes, including those claimed. It is additionally note that the specific data types of the hierarchy do not appear to functionally change the invention, and that no order for the data types is claimed). It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Lee by the teachings of Srivastava, because both references are directed towards extracting data. Srivastava merely adds to Lee an additional type of data entity that may be categorized and searched for using the system of Lee. This will make the search system of Lee be able to respond to additional types of user requests. As to claims 13 and 20, see the rejection of claim 6. Claims 7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al. (US Pre-Grant Publication 2021/0294840), in view of Park et al. (US Pre-Grant Publication 2024/0160890), in view of Larkin et al. (US Pre-Grant Publication 2023/0057414), and further in view of Alme et al. (US Pre-Grant Patent 8,561,193). As to claim 7, Lee as modified teaches the method of method of claim 1, Lee does not teach wherein each of the set of files and the query file are portable executable files Alme teaches wherein each of the set of files and the query file are portable executable files (see Alme paragraphs [0033], [0035] and [0064]. Malicious executable files may be used for training. Additionally, executable files are searched when received). It would have been obvious to one of ordinary skill in the art before the earliest filing date of the invention to have modified Lee by the teachings of Alme, because both references are directed towards extracting data and searching vectors of data. Alme merely adds to Lee an additional type of data entity that may be categorized and searched for using the system of Lee. This will make the search system of Lee be able to respond to additional types of user requests. As to claim 14, see the rejection of claim 7. Response to Arguments Applicant's arguments filed 13 July 2026 have been fully considered but they are not persuasive. Applicant argues that “As can be seen from the foregoing discussion, Larkin discloses generating a database mapping prediction iteratively using a hierarchy of match models, where each matching iteration is associated with a particular match model from the hierarchy of match models. This is not the same as training the same ML model over a series of steps, "wherein at each of the plurality of steps, the ML model is trained to group files iteratively, and wherein at each progressive iteration the ML model learns to group files based on a characteristic from the hierarchy of characteristics that is progressively lower on the hierarchy of characteristics," as recited in claim 1. Using multiple different match models to generate a database mapping prediction is completely different from iteratively training a single model "wherein at each progressive iteration the ML model learns to group files based on a characteristic from the hierarchy of characteristics that is progressively lower on the hierarchy of characteristics," as recited in claim 1. Thus, there is nothing in Larkin that would suggest combining its disclose with the disclosure of Park and Lee to achieve the above-recited feature of claim 1.” In response to this argument, it is noted that the cited iterative steps of Larkin occur in the context of a “machine learning framework.” Within this context, Larkin describes the steps of Larkin “[introduce] innovative machine learning techniques that enable using trained models in combination with rule-based techniques in a manner that is configured to improve training speed and training efficiency of string-based machine learning models.” Thus, Larkin does suggest the use of the cited techniques within a training system. Additionally, Applicant is reminded that Larkin is cited as part of a combination with both Park and Lee. Both Park and Lee clearly describe a training of machine learning models in the cited portions, with Park describing iterative training steps based on characteristics (see Park paragraph [0123]). Because Larkin indicates that the steps of Larkin occur in the context of a “machine learning framework” that “enable[s] using trained models in combination with rule-based techniques in a manner that is configured to improve training speed and training efficiency of string-based machine learning models” (see Larkin paragraph [0021]) and because the combination of references describes an iterative training of a model in Park and a training of a model in Lee, one of ordinary skill in the art would find the claimed invention obvious. It is noted that nothing in the claim language excludes additional machine learning models from being a part of the claimed machine learning model’s training. Applicant is reminded that unclaimed features from the specification, such as a prohibition on different match models, receives no patentable weight until claimed. Applicant argues that “This is further evidenced by the fact that the cited portions of Larkin do not even disclose a training process. Indeed, the cited portions of Larkin disclose "generating a database mapping prediction for an input string, utilizing the hierarchical string matching machine learning framework" that comprises already trained models. There is nothing in the cited portions of Larkin that discuss the training of a machine learning model. Therefore, there is no teaching or suggestion to combine Larkin, Park and Lee to achieve the above-recited feature of claim 1. It follows that the combination of Larkin, Park and Lee do not disclose or suggest the above-recited feature of claim 1.” In response to this argument, it is noted that Larkin describes at paragraph [0021] how the techniques in Larkin are “machine learning techniques that enable using trained models in combination with rule-based techniques to improve training speed and training efficiency of string-based machine learning models.” Cited paragraphs [0043] and [0094] of Larkin are both techniques that use trained models in combination with rule-based techniques in a machine learning framework. Because Larkin describes how such techniques “improve training speed and training efficiency of string-based machine learning models,” it would be obvious to one of ordinary skill in the art before the earliest filing date of the invention to use the cited techniques of Larkin for “a training process.” In addition to this, it is noted that Larkin is relied upon in combination with Lee and Park. Park describes the use of the use of an iterative process by which a system learns to group nodes (see Park paragraph [0123]). Lee explicitly describes training a neural network (see Lee paragraph [0025]). Thus, cited references each individually and as a combined whole are directed towards an iterative training process. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHARLES D ADAMS whose telephone number is (571)272-3938. The examiner can normally be reached M-F, 9-5:30 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aleksandr Kerzhner can be reached at 5712701760. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CHARLES D ADAMS/ Primary Examiner, Art Unit 2165
Read full office action

Prosecution Timeline

Show 3 earlier events
Nov 05, 2025
Final Rejection mailed — §103
Jan 29, 2026
Applicant Interview (Telephonic)
Jan 30, 2026
Examiner Interview Summary
Feb 05, 2026
Request for Continued Examination
Feb 17, 2026
Response after Non-Final Action
Mar 11, 2026
Non-Final Rejection mailed — §103
Jul 13, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748985
DEVICE AND COMPUTER IMPLEMENTED METHOD FOR ADDING A QUANTITY FACT TO A KNOWLEDGE BASE
3y 7m to grant Granted Sep 29, 2026
Patent 12730822
CLOUD DATA CONSOLIDATION AND PROCESSING SYSTEM
4y 9m to grant Granted Sep 08, 2026
Patent 12717766
SYSTEMS AND METHODS FOR DATABASE ORIENTATION TRANSFORMATION
3y 6m to grant Granted Aug 25, 2026
Patent 12717840
DYNAMIC SEARCH INPUT SELECTION
3y 5m to grant Granted Aug 25, 2026
Patent 12717803
MONITORING AND ALERTING PLATFORM FOR EXTRACT, TRANSFORM, AND LOAD JOBS
1y 11m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
45%
Grant Probability
89%
With Interview (+43.8%)
4y 11m (~2y 8m remaining)
Median Time to Grant
High
PTA Risk
Based on 432 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month