Prosecution Insights
Last updated: October 01, 2026
Application No. 18/773,123

IMAGE ANALYSIS USING A MULTIMODAL LARGE LANGUAGE MODEL

Non-Final OA §102§103
Filed
Jul 15, 2024
Examiner
LE, VU
Art Unit
2668
Tech Center
2600 — Communications
Assignee
CrowdStrike Inc.
OA Round
1 (Non-Final)
54%
Grant Probability
Moderate
1-2
OA Rounds
9m
Est. Remaining
64%
With Interview

Examiner Intelligence

Grants 54% of resolved cases
54%
Career Allowance Rate
24 granted / 44 resolved
-7.5% vs TC avg
Moderate +9% lift
Without
With
+9.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
10 currently pending
Career history
55
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
54.1%
+14.1% vs TC avg
§102
26.0%
-14.0% vs TC avg
§112
8.1%
-31.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 44 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Election/Restrictions Applicant’s election without traverse of Group 1 (claims 6-20) in the reply filed on 6/29/2026 is acknowledged. Claims 1-5 withdrawn from further consideration pursuant to 37 CFR 1.142(b) as being drawn to a nonelected Group 2, there being no allowable generic or linking claim. Election was made without traverse in the reply filed on 6/29/2026. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 6, 8, 10, 13, 16-17 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US20250147957A1 (VERKRUYSE et al) Regarding claim 6, VERKRUYSE et al discloses the following: One or more non-transitory computer-readable media storing instructions executable by one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations (pars. 0032-33; also Fig. 8 and par. 0158 for non-transitory computer readable media and CPU) “[0033] Cross-modal data enrichment refers to the process of integrating and enhancing data from different modalities (e.g., text, images, audio, video, or structured data) by dynamically linking entities and relationships across these diverse sources. Existing solutions are often siloed and static, handling only specific data types (e.g., text-only or image-only) or lacking the ability to provide continuous automated updates, leading to incomplete or fragment analysis. In contrast, the implementations herein may leverage advanced techniques in data integration, NLP, ML, and AI to create a unified and enriched dataset that provides deeper insights and more comprehensive understanding of the dataset. Cross-modal data enrichment may comprise identifying and extracting entities (e.g., people, places, organizations) from different data modalities, connecting these entities across different data sources to create a cohesive knowledge graph, identifying, extracting, and continuously updating relationships between entities within and across different data modalities, and/or enhancing the dataset by adding contextual information and insights derived from the integrated data.” comprising: inputting, into a multimodal large language model (m-LLM), first data associated with one of: a data stream, a byte slice, or a byte array, the first data including: first image data representing a first set of metrics associated with a computing device over a first time period, and second image data representing a second set of metrics associated with the computing device over the first time period (Fig. 2, 202, 208A-C: “first time period” is defined by time series data loader 208C for example; “first/second image data” representing “first/second sets of metrics” are defined by unstructured data loader 208A in conjunction with 208C based on data sources 216 e.g., image data input to graph database core 202 based on confidence scores and data quality metrics as discussed in par. 0114 for example); determining, by the m-LLM, a context between a first metric of the first set of metrics and a second metric of the second set of metrics based at least in part on comparing the first metric and the second metric to a metric threshold (pars. 0035-37; 0091); “[0035] In some implementations, another advantage of the system is multi-scale temporal-semantic reasoning, seamlessly integrating time series patterns with semantic knowledge. This approach seamlessly integrates time series patterns with semantic knowledge, enabling a comprehensive and dynamic understanding of data. As noted above, traditional systems often rely on static schemas and manual updates, which can be time-consuming and prone to errors. In contrast, multi-scale temporal-semantic reasoning dynamically evolves the data schema based on incoming data and user interactions.” “[0091] … The real-time processing engine enables immediate insights and alerts. For example, continuous evaluation of predefined conditions or thresholds may trigger alerts or actions when met. Furthermore, real-time updating of dashboards and visualizations of the UI 212 connected to the graph database 204 may be performed. Immediate propagation of detected anomalies or significant pattern changes to relevant parts of the graph representation enable rapid response to evolving situations.” determining, by the m-LLM, semantic information describing a function or a meaning of the first image data or the second image data (pars. 0035-37); “[0036] Additionally, semantic knowledge integration incorporates contextual information, such as meanings, relationships, and contexts, into the analysis. As such, the system may understand the significance and implications of the temporal patterns. For instance, integrating medical research articles and treatment guidelines with patient health data provides a richer context for interpreting health trends. Multi-scale reasoning combines insights from different time scales and semantic contexts to provide a holistic understanding of the data. This is particularly useful in complex scenarios where short-term fluctuations need to be understood in the context of long-term trends and broader semantic knowledge. Additionally, the system may be configured for dynamic adaptation, continuously updating, and refining the analysis as new data and semantic information become available. This ensures that the insights remain current and relevant, adapting to new developments and user interactions. [0037] In some implementations, this approach offers a comprehensive analysis by combining temporal patterns with semantic knowledge, leading to deeper and more nuanced insights than existing solutions can provide. The integration of contextual information enhances the accuracy and relevance of the insights, making those insights more actionable.” and storing the context and the semantic information as stored data in a storage device for access by the computing device at a later time (par. 0047), “[0047] This integration enables the system to leverage the semantic graph database 204 for efficient storage and retrieval of complex, interconnected data, representation of domain-specific ontologies and knowledge structures, facilitation of inferential reasoning capabilities, and support for flexible schema evolution to accommodate dynamic data landscapes.” the computing device configured to determine presence of a malicious event in third data based at least in part on the stored data (par. 0091). “[0091] … The real-time processing engine enables immediate insights and alerts. For example, continuous evaluation of predefined conditions or thresholds may trigger alerts or actions when met. Furthermore, real-time updating of dashboards and visualizations of the UI 212 connected to the graph database 204 may be performed. Immediate propagation of detected anomalies or significant pattern changes to relevant parts of the graph representation enable rapid response to evolving situations.” Regarding claim 8, the rejection of claim 6 is incorporated herein. VERKRUYSE et al further teach: wherein the first image data represents a first graph and the second image data represents a second graph, and the operations further comprising: detecting text in one of: the first graph or the second graph, the text associated with an axis, a title, or a label of the first graph or the second graph, wherein determining the context is further based at least in part on the text (Figs. 1-22; par. 0025). [0025] FIG. 1 illustrates an example graph representation according to some implementations herein. In some implementations, the graph representation 100 comprises a plurality of nodes 102 and a plurality of edges 104 between nodes 102. The graph representation comprises a structured representation of information that captures relationships, represented by edges 104, between entities, represented by nodes 102, in a way that is both human-readable and machine-interpretable. … Both nodes 102 and edges 104 may also comprise properties, which are additional pieces of information that provide more context about the entity or relationship, such as names, addresses, industries, dates, and/or positions, among others. In some implementations, labels may also be used to categorize nodes 102 and edges 104 in the graph representation 100. A node 102 or edge 104 can have one or more labels that help in identifying the type of entity or relationship.” Regarding claim 10, The one or more non-transitory computer-readable media of claim 6, the operations further comprising: transmitting the stored data to the computing device; and causing the computing device to determine presence of the malicious event in the third data based at least in part on accessing the stored data from the storage device (The rejection of claim 1 is applicable here). Regarding claim 13, The one or more non-transitory computer-readable media of claim 6, wherein: the first data represents first computer-readable instructions associated with a first operating system or first data format, the second data represents second computer-readable instructions associated with a second operating system or second data format, and determining the context or the semantic information is performed independent of requiring input from a user (Fig. 2, Data Source(s) 216 depicts various data formats independent of user input; see also pars. 0050, 0082). Regarding claim 16, The one or more non-transitory computer-readable media of claim 6, wherein the first data is received from one of: an event-based message queue, a service, or a third-party queue (pars. 0052, 0105, 0112, 0152). Regarding claim 17, A computer-implemented method comprising: inputting, into a multimodal large language model (m-LLM), first data associated with one of: a data stream, a byte slice, or a byte array, the first data including: first image data representing a first set of metrics associated with a computing device over a first time period, and second image data representing a second set of metrics associated with the computing device over the first time period; determining, by the m-LLM, a context between a first metric of the first set of metrics and a second metric of the second set of metrics based at least in part on comparing the first metric and the second metric to a metric threshold; determining, by the m-LLM, semantic information describing a function or a meaning of the first image data or the second image data; and storing the context and the semantic information as stored data in a storage device for access by the computing device at a later time, the computing device configured to determine presence of a malicious event in third data based at least in part on the stored data (Claim 17 is a method claim wherein the scope corresponds to that of independent claim 6 and thus, the rejection of claim 6 is fully incorporated herein). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over US20260017972A1 (ABERLE) in view of US 20240330446 A1 (BULUT et al). Regarding claim 6, ABERLE discloses the following: One or more non-transitory computer-readable media storing instructions executable by one or more processors, wherein the instructions, when executed, cause the one or more processors to perform operations (Claims 18-19) “18. A non-tangible computer readable medium that stores instructions, the instructions when executed by a processor programs the processor to: access an input query comprising an input to search for images in an image database; obtain a text description to search based on the input; generate an input vector based on the text description; compare the input vector against a plurality of vectors in the image database, each vector from among the plurality of vectors in the image database being based on a text description of a corresponding image in the image database; and identify one or more images in the image database based on the comparison, wherein each of the one or more images has a corresponding text description that is semantically similar to the text description. 19. The non-tangible computer readable medium of claim 18, wherein the input comprises an image input, and wherein the instructions when executed further programs the processor to: determine the text description based on the input image; and generating an input vector based on the description of the input image.” comprising: inputting, into a multimodal large language model (m-LLM) (Par. 0007; Fig. 9, 904), first data associated with one of: a data stream, a byte slice, or a byte array, the first data including: first image data representing a first set of metrics associated with a computing device over a first time period, and second image data representing a second set of metrics associated with the computing device over the first time period (Fig. 5, “primary/secondary” objects with associated metrics input to mLLM; see also pars. 0029-0031) ; determining, by the m-LLM, a context between a first metric of the first set of metrics and a second metric of the second set of metrics based at least in part on comparing the first metric and the second metric to a metric threshold (pars. 0007, claims 3 & 14); determining, by the m-LLM, semantic information describing a function or a meaning of the first image data or the second image data (pars. 0008, 0023); and storing the context and the semantic information as stored data in a storage device for access by the computing device at a later time (Fig. 1, 125), Aberle does not teach the limitation as indicated in the strike-through above. However, Bulut et al teach said limitations (par. 0003; Figs. 2A-2C). “[0003] Systems and methods for improving the performance and energy efficiency of machine learning systems that generate security specific machine learning models, generate security related information using the security specific machine learning models, and/or detect security related anomalies are provided. One example of a security specific machine learning model is a security specific large language model. A security specific large language model may be trained and deployed to generate and output semantically related security information. For example, the security specific large language model may be used to determine whether a particular security log that stores log lines regarding security events (e.g., failed logins, password changes, failed authentication requests, and file deletions) for a networked computing environment includes a log line associated with a malicious security event. The security specific large language model may be pretrained with a security specific dataset that was generated using similarity deduplication and long line handling, and with security specific objectives, such as next log line prediction based on host, system, application, and cyber attackers' behavior. Further, a security specific similarity dataset may be generated to align the security specific large language model to capture similarity between different security events. The security specific large language model may be fine-tuned using the security specific similarity dataset and then stored within a datastore.” At the time of effective filing, it would have been obvious to incorporate the teaching of Bulut et al into Aberle to derive at claim 6 for the technical benefits of reduced energy consumption, computing and storage costs, network/system improvements, and data security (Bulut et al, par. 0004). Allowable Subject Matter Claims 7, 9, 11-12, 14-15, 18-20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: the most relevant prior arts are applied in the aforementioned rejections. However, neither of them, alone or in combination, teach the further limitations as recited in the identified dependent claims. Contact Info Any inquiry concerning this communication or earlier communications from the examiner should be directed to VU LE whose telephone number is (571)272-7332. The examiner can normally be reached M-F 8:00 - 17:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. Vu Le can be reached at 2-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /VU LE/Supervisory Patent Examiner, Art Unit 2668
Read full office action

Prosecution Timeline

Jul 15, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738058
VIDEO PROCESSING
3y 5m to grant Granted Sep 15, 2026
Patent 12731433
TARGET TRACKING METHOD AND APPARATUS BASED ON PATH
4y 0m to grant Granted Sep 08, 2026
Patent 12725266
MEDICAL IMAGE PROCESSING APPARATUS, MEDICAL IMAGE PROCESSING METHOD, AND PROGRAM
3y 0m to grant Granted Sep 01, 2026
Patent 12725404
METHOD AND SYSTEM FOR DETECTING RAILWAY FOREIGN OBJECT INTRUSION, DEVICE AND MEDIUM
1y 11m to grant Granted Sep 01, 2026
Patent 12718164
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND IMAGE PROCESSING PROGRAM
4y 0m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
54%
Grant Probability
64%
With Interview (+9.0%)
2y 11m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 44 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month