Prosecution Insights
Last updated: October 02, 2026
Application No. 18/963,450

CORRELATING STRUCTURED AND UNSTRUCTURED DOMAIN-SPECIFIC DATA

Non-Final OA §103§112
Filed
Nov 27, 2024
Examiner
AHMED, ZAIN JIM
Art Unit
2432
Tech Center
2400 — Computer Networks
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-58.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
3 currently pending
Career history
6
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The claims recite an "entity extraction agent" (claims 1, 4, 8, 9, 15, 16), a "filter agent" (claims 2, 3, 7, 10, 17, 18), a "GAI filter agent" (claim 2), a "validation agent" (claims 4, 9, 16), and an "attack techniques discovery agent" (claims 5, 12). These limitations pair the word "agent" with functional language. The claims do not use the word "means," so there is a rebuttable presumption that 35 U.S.C. 112(f) does not apply. The presumption is not rebutted here. In the context of these claims and the specification, "agent" names a known class of software components - a model or program given a defined role such as extraction, filtering, validation, or discovery - and therefore connotes sufficiently definite structure to a person of ordinary skill in the art. Accordingly, these limitations are not being interpreted under 35 U.S.C. 112(f). It is noted that claim 2 expressly recites a "GAI filter agent," while the remaining claims recite "agent" without the "GAI" modifier. The applicant thus specifies a generative artificial intelligence implementation where one is intended. Under the broadest reasonable interpretation, the unmodified "agent" limitations are not limited to generative artificial intelligence models. Claims 1, 7, and 14 each conclude with the clause "wherein the knowledge graph is used for augmenting a generative artificial intelligence (GAI) query." Nothing recited in the claims performs the augmenting; no step or instruction uses the knowledge graph to augment any query. The clause is a statement of the intended use of the knowledge graph. Had the applicant intended to claim the use, a positive step or instruction using the knowledge graph would have been recited. The examiner finds that this clause is not required by the claims and does not further limit them. If the applicant believes the clause is required, or amends the claims to require it, the clause is preemptively addressed arguendo in the rejections below, and is also addressed under 35 U.S.C. 112(b) below because its effect on claim scope cannot be determined. Claims 6, 13, and 20 recite storing the knowledge graph "in a database accessible by a GAI application, the GAI application being operable to augment user-submitted queries to a GAI model using the knowledge graph." The positive act recited is storing the knowledge graph in a database. The remaining recitations describe capabilities of a GAI application with which the claimed invention merely interacts; every database could be accessible by an application, and what a separate application is operable to do does not further limit the claimed system, method, or medium. The examiner finds that these recitations do not further limit the claims beyond the storing of the knowledge graph in a database. They are nonetheless preemptively addressed arguendo in the rejections below. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor regards as the invention. Claims 1, 7, and 14 recite "wherein the knowledge graph is used for augmenting a generative artificial intelligence (GAI) query." As set forth in the Claim Interpretation section above, no step or instruction recited in the claims performs the augmenting, and the clause reads as a statement of intended use. It cannot be determined whether the clause limits the claims, and the scope of the claims therefore cannot be determined. Claims 2-6, 8-13, and 15-20 are rejected by virtue of their dependence from claims 1, 7, and 14, respectively. For purposes of examination, the clause is interpreted as a positive step or instruction of using the knowledge graph to augment a generative artificial intelligence query, and is addressed arguendo in the rejections below. Claim 1 recites "wherein: the first entity data is generated using a first entity extraction agent... the second entity data is generated using a second entity extraction agent." Claim 1 is a system claim reciting a memory storing instructions configured to cause a processor to perform operations; no data is actually generated within the scope of the claim. The recitation that the entity data "is generated" describes, in the past/passive tense, data that is never generated by the claim, and the metes and bounds of the limitation therefore cannot be determined. Claims 2-6 are rejected by virtue of their dependence from claim 1. For purposes of examination, the limitation is interpreted as reciting that the instructions are configured to cause the processor to generate the first entity data using a first entity extraction agent and to generate the second entity data using a second entity extraction agent. Claim 15 recites that "the entity data is generated by first and second entity extraction agents," with each entity extraction agent "instructed to identify," from the input data, entities of its respective entity type. Claim 15 depends from claim 14, a computer storage medium claim in which no data is generated; the recitation that the entity data "is generated" describes data that is never generated within the scope of the claim. Further, nothing recited in the claim instructs the agents; the claim recites no actor performing the instructing. The metes and bounds of the limitation therefore cannot be determined. Claim 16 is rejected by virtue of its dependence from claim 15. For purposes of examination, the limitation is interpreted as reciting that the instructions further cause the processor to generate the entity data using first and second entity extraction agents, the first caused to identify, from the input data, entities of the first entity type and the second caused to identify, from the input data, entities of a second entity type different from the first. Claim 10 recites "the first and second relationship types." There is insufficient antecedent basis for this limitation in the claim. Claim 7, from which claim 10 depends, does not recite any relationship types. For purposes of examination, the limitation is interpreted as referring to relationship types of the relationships recited in claim 7. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The references relied upon in the following rejections are: Jeong et al. (US-20220197923-A1), hereinafter “Jeong.” Cheng et al., “CTINEXUS: Leveraging Optimized LLM In-Context Learning for Constructing Cybersecurity Knowledge Graphs Under Data Scarcity”, October 28, 2024, hereinafter “Cheng.” Mihindukulasooriya et al. (US-20210109995-A1), hereinafter “Mihin.” Zhang et al. (CN-112131882-A), hereinafter “Zhang.” Zhu et al. (CN-117914516-A), hereinafter “Zhu.” Huang et al. (CN-117875327-A), hereinafter “Huang.” Claims 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng Regarding claim 14, the claim recites "[a] computer storage medium embodying instructions that, upon execution by a processor, cause the processor to" perform operations. Jeong provides for this limitation because Jeong discloses that its apparatus may be implemented in a computer system including a computer-readable recording medium, one or more processors, and memory ([0115]-[0116]). The claim further recites "receive input data comprising unstructured data related to a knowledge domain." Jeong provides for this limitation because Jeong discloses a collection engine that collects cyber threat information ([0050]), and the collected cyber threat information may be unstructured data, including reports written in unstructured natural language such as cyber threat analysis reports, malware analysis reports, and vulnerability analysis reports, as well as news, blogs, and tweets ([0054]). The knowledge domain is cyber threat intelligence ([0002]). The claim further recites "generate entity data regarding a first entity of a first entity type and a second entity of a second entity type, the first and second entity types being defined by a domain-specific ontology for the knowledge domain." Jeong provides for generating entity data regarding entities of defined types because Jeong discloses a named-entity recognition model that matches each word in an input sentence with the most suitable label and collects the labels for each piece of metadata ([0087], [0089], FIG. 6), where the metadata types for the cyber threat domain include Threat_Actor, Victim_Target, IP_Attack, Domain_Attack, Email_Address, CVE_Numbers, Malware, Attack_Vector, and Attack_Tool ([0065], Table 4). Jeong further discloses defining twelve entities and six relationships for the cyber threat domain through integration and selection of the extracted metadata ([0097]), the entities including Attack_Objective, Victim_Location, Victim_Target, IP, Domain, Email, CVE, Threat_Actor, Malware, Attack_Vector, and Attack_Tool ([0098]). While Jeong uses a domain-specific ontology, because Jeong describes the cybersecurity domain and defines the entity types for that domain with formal definitions (Table 4; [0065]) and defines a closed set of relationship types for the domain ([0097], [0099]), Jeong does not explicitly disclose that the first and second entity types are "defined by a domain-specific ontology." However, Cheng provides for the entity types being defined by a domain-specific ontology because Cheng discloses that construction of a cybersecurity knowledge graph typically follows an ontology, which specifies the entity types and their allowed relations (§2.1), and Cheng applies a cybersecurity ontology (MALOnt) that defines the entity types and relation types for the domain (§4.1). The claim further recites "generate relationship data regarding relationships between the first entity and the second entity using relationship definitions from the domain-specific ontology." Jeong provides for this limitation because Jeong discloses a defined set of six relationships for the domain - Include, Use, Relate, Attack, Target, and Exploit ([0099]) - and discloses defining triples for the relationships between the extracted entities according to those definitions, such as a triple for the relationship between an attack nation and a victim nation and a triple for a tool used for an attack ([0101]). Jeong defines the component entities and the relationship using <head, relation, tail> ([0102]), for example "Attack_Nation, Attack(exploit), Victim_Nation" (Table 6). The claim further recites "construct a knowledge graph including adding nodes corresponding to the entities and edges corresponding to the relationships." Jeong provides for this limitation because Jeong discloses constructing a cyber threat knowledge graph by converting the defined triples into an RDF dataset ([0096], [0100]) that is stored in a graph database in which the data is connected and schematized ([0047]). As per the 112(b) indefiniteness/Claim Interpretation section above, the following limitation is being addressed arguendo. The claim further recites "wherein the knowledge graph is used for augmenting a generative artificial intelligence (GAI) query." Jeong does not provide for the knowledge graph being used for augmenting a GAI query. Cheng provides for this limitation because Cheng discloses that a question-answering system can be developed upon the constructed cybersecurity knowledge graph using a large language model's retrieval-augmented generation, to provide grounded answers to threat-related questions (§6). The proposed modification is to add the retrieval-augmented question answering suggested by Cheng to the cyber threat knowledge graph of Jeong, such that a query submitted to a generative artificial intelligence model is augmented with the entities and relationships of Jeong's knowledge graph before the model generates its answer. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to make this modification. Cheng expressly suggests it: a question-answering system can be developed upon the constructed cybersecurity knowledge graph using retrieval-augmented generation to provide grounded answers to threat-relations questions (§6). One of ordinary skill would have been motivated to augment the query with the knowledge graph because a generative model alone generates its answer from its training data, which may not contain, and may be inconsistent with, the collected threat information organized in Jeong's graph; augmenting the query with the graph's entities and relationships causes the answer to be generated from the collected information itself - a grounded answer, exactly as Cheng suggests (§6). Regarding claim 20, the claim depends from claim 14 and further recites that the instructions "store the knowledge graph in a database accessible by a GAI application, the GAI application being operable to augment user-submitted queries to a GAI model using the knowledge graph." As set forth in the Claim Interpretation section above, the positive act recited is storing the knowledge graph in a database; the recitations regarding the GAI application describe capabilities of an application with which the claimed medium interacts and are addressed arguendo. Jeong provides for storing the knowledge graph in a database because Jeong discloses that the data is automatically recognized and stored in a graph database ([0047]). Arguendo, Cheng provides for the database being accessible by a GAI application operable to augment user-submitted queries because Cheng's question-answering system is developed upon the constructed knowledge graph and answers threat-related questions submitted to it using retrieval-augmented generation (§6). The proposed modification and the motivation to combine are the same as set forth above with respect to claim 14. Claims 1, 6, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng, and further in view of Zhang. Regarding claim 1, the claim recites a system comprising "a processor; and a memory storing instructions." Jeong provides for this limitation because Jeong discloses an apparatus that includes memory in which at least one program is recorded and a processor for executing the program ([0021]). The claim further recites receiving input data comprising unstructured data related to a knowledge domain, generating relationship data regarding relationships between the entities, constructing a knowledge graph by adding nodes corresponding to the entities and edges corresponding to the relationships, and the knowledge graph being used for augmenting a GAI query. These limitations are substantially similar to the corresponding limitations of claim 14 and are provided for by Jeong and Cheng for the same reasons set forth above with respect to claim 14. As per the 112(b) indefiniteness/Claim Interpretation section above, the following limitation is being addressed arguendo, interpreted as reciting that the instructions are configured to cause the processor to generate the first entity data using a first entity extraction agent and to generate the second entity data using a second entity extraction agent. The claim further recites that "the first entity data is generated using a first entity extraction agent, the first entity data regarding entities of a first entity type, the second entity data is generated using a second entity extraction agent, the second entity data regarding entities of a second entity type, wherein the first and second entity types are defined in a domain-specific ontology." Jeong and Cheng do not provide for the first entity data being generated using a first entity extraction agent and the second entity data being generated using a second entity extraction agent. Zhang provides for this limitation because Zhang discloses identifying entities of different types using different identification modes: entities of a preset category are identified from the text data according to a preset regular expression, and the entities of the other types defined by the network security knowledge ontology are identified using a trained entity identification model (Zhang, claims 1-2; step S12). Each identification mode is a software component given the role of identifying the entities of its assigned entity types, and the entity types are defined by the network security knowledge ontology (Zhang, claim 1). The proposed modification is to divide the entity extraction of the Jeong/Cheng system between two separate extraction components according to entity type, as Zhang discloses: the entities of preset, pattern-shaped categories identified by a regular-expression component, and the remaining ontology-defined entity types identified by a trained entity identification model (Zhang, claims 1-2; step S12). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to make this modification. One of ordinary skill would have been motivated to do so because identifying the entities according to their different types, using a different identification mode for each, makes the identification of the entities "more accurate and comprehensive" (Zhang, claims 1-2): a pattern-shaped entity such as an IP address or a CVE number is identified exactly by its pattern, while an entity without a fixed pattern, such as the name of a threat actor or a malware, is identified by the trained model, so each entity type is identified by the mode suited to it. It is further noted that, in the combination, the entity extraction is performed with the model-based extractors of Cheng, which extract entities with a generative model (GPT-4) whose instruction incorporates the ontology (Cheng, §§3, 4.1). Implementing Zhang's type-based division with Cheng's model-based extractors yields a first and a second generative extraction agent, each caused to identify the entities of its assigned entity types. The combination therefore provides for the claimed first and second entity extraction agents under either reading of "agent" (see Claim Interpretation above). Regarding claim 6, the claim depends from claim 1 and recites limitations substantially similar to those of claim 20. Claim 6 is rejected for the same reasons set forth above with respect to claims 1 and 20. Regarding claim 15, the claim depends from claim 14. As per the 112(b) indefiniteness/Claim Interpretation section above, the limitations of claim 15 are being addressed arguendo, interpreted as reciting that the instructions further cause the processor to generate the entity data using first and second entity extraction agents, the first caused to identify, from the input data, entities of the first entity type and the second caused to identify, from the input data, entities of a second, different entity type. Zhang provides for these limitations for the same reasons set forth above with respect to claim 1: Zhang's regular expression mode identifies the entities of the preset categories, and Zhang's trained entity identification model identifies the entities of the other, different ontology-defined types (Zhang, claims 1-2; step S12), so each component identifies entities of types different from the other's. The proposed modification and the motivation to combine are the same as set forth above with respect to claim 1. Claims 7, 10, 11, 13, 17, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng, and further in view of Mihin. Regarding claim 7, the claim recites a computerized method comprising receiving input data related to a knowledge domain comprising unstructured data; generating entity data regarding a first entity of a first entity type and a second entity of a second entity type, the entity types being defined by a domain-specific ontology; and constructing a knowledge graph including adding nodes corresponding to the entities and edges corresponding to the relationships, wherein the knowledge graph is used for augmenting a GAI query. These limitations are substantially similar to the corresponding limitations of claim 14 and are provided for by Jeong and Cheng for the same reasons set forth above with respect to claim 14. The claim further recites "generating initial relationship data regarding relationships between the first and second entities by including in the initial relationship data every relationship permitted for the first and second entities by the domain-specific ontology" and "generating, by a filter agent, relationship data including evaluating a level of consistency between the relationships in the initial relationship data and the input data, the relationship data including only those relationships in the initial relationship data that are consistent with the input data." Jeong and Cheng do not provide for generating initial relationship data that includes every relationship permitted by the ontology and then filtering that initial relationship data down, by a filter agent, to only the relationships consistent with the input data. Mihin provides for these limitations. Mihin discloses that conventional systems treat the external knowledge graph as infallible and simply retrieve a relationship between nodes that match the extracted terms, without checking whether the relationship applies or is relevant to the terms as used in the input corpus, so that the retrieved set can contain relationships that are true but irrelevant ([0001], [0006]). Mihin discloses retrieving from the knowledge graph the relationships between the identified nodes to generate an intermediate domain taxonomy that contains such spurious relationships ([0029]). Mihin further discloses analyzing the corpus semantics, as captured in context-based embeddings, to determine whether the extracted terms are used and distributed throughout the corpus in a manner consistent with the existence of the relationship retrieved from the knowledge graph, and filtering out as irrelevant any relationship whose similarity value falls below a threshold, leaving a refined domain taxonomy that lacks the spurious relationships ([0006], [0030], [0033]). Mihin illustrates the filtering with a concrete example in which a retrieved relationship that is technically true is nonetheless removed because the corpus does not use the terms consistently with it ([0035]-[0036]). In the combination, the relationships "permitted by the domain-specific ontology" are Jeong's defined relationship set - Include, Use, Relate, Attack, Target, and Exploit ([0099]) - a small closed set, so including every permitted relationship for a pair of entities and then filtering for consistency follows directly from applying Mihin's generate-then-filter approach to Jeong's definitions. The filter agent is Mihin's filtration component, a software component given the filtering role (see Claim Interpretation above). The proposed modification is to generate the relationship data of the Jeong/Cheng system by first including, for the extracted entities, the relationships permitted by Jeong's defined relationship set ([0099]), and then filtering that initial set with Mihin's consistency-based filtration so that only the relationships consistent with the input data are retained ([0006], [0030], [0033]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to make this modification. One of ordinary skill would have been motivated to do so because a relationship drawn from a defined external set can be true in the abstract yet unsupported by the particular input data (Mihin, [0001], [0006]); Jeong defines its triples after heuristic analysis ([0101]) and verifies them only through ontology visualization analysis ([0105]); applying Mihin's automated consistency filtering to the relationship data keeps relationships that the input data does not support out of the knowledge graph, so that the constructed graph reflects only the collected information. Although Mihin builds a domain taxonomy rather than a security knowledge graph, Mihin is reasonably pertinent to the problem faced by the inventors: constructing a domain-specific knowledge structure from a text corpus with the help of an external knowledge source, and removing the relationships that the corpus does not support. Regarding claim 10, the claim depends from claim 7 and, as best understood in view of the rejection under 35 U.S.C. 112(b) above, further recites that relationship definitions for the relationship types are derived from the domain-specific ontology, that each relationship definition includes a source entity type, a relationship type, and a target entity type, and that the relationship definitions are used in the generating of the relationship data and by the filter agent. Jeong provides for relationship definitions including a source entity type, a relationship type, and a target entity type because Jeong defines the component entities and the relationship using <head, relation, tail> ([0102]), with typed heads and tails, for example "Attack_Nation, Attack(exploit), Victim_Nation" (Table 6), and the definitions come from Jeong's defined entity and relationship sets ([0097] - [0099]). Mihin provides for the definitions being used by the filter agent because Mihin discloses selecting the embedding used for the consistency evaluation based on the type of the relationship retrieved ([0032]), so the relationship definition informs the filtering. The motivation to combine is the same as set forth above with respect to claim 7. Regarding claim 11, the claim depends from claim 7 and further recites that the knowledge domain is cyber threat intelligence and the input data is a threat intelligence report. Jeong provides for this limitation because Jeong's collected unstructured data is cyber threat information including cyber threat analysis reports ([0002], [0054]). Regarding claim 13, the claim depends from claim 7 and recites limitations substantially similar to those of claim 20. Claim 13 is rejected for the same reasons set forth above with respect to claims 7 and 20. Regarding claim 17, the claim depends from claim 14 and further recites generating initial relationship data by including "information regarding relationships permitted by the domain-specific ontology for the entities" and evaluating, using a filter agent, a level of consistency to the input data, including in the relationship data only the relationships that are consistent with the input data. These limitations are narrower forms of the limitations addressed above with respect to claim 7 (claim 17 does not require every permitted relationship), and Jeong, Cheng, and Mihin provide for them for the same reasons set forth above with respect to claim 7. The motivation to combine is the same as set forth above with respect to claim 7. Regarding claim 18, the claim depends from claim 17 and further recites deriving, from the domain-specific ontology, a relationship definition including a source entity type, the relationship type, and a target entity type, the definition being used in the generating of the initial relationship data and by the filter agent. Jeong provides for deriving the relationship definition because Jeong's twelve entities and six relationships are defined through integration and selection of the extracted metadata ([0097]) and the triples are defined after heuristic analysis according to those definitions ([0101]-[0102], Table 6). Mihin provides for the definition being used by the filter agent ([0032]), as set forth above with respect to claim 10. The motivation to combine is the same as set forth above with respect to claim 7. Claims 2, 3, and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng, Zhang, and Mihin. Regarding claim 2, the claim depends from claim 1 and further recites generating initial relationship data by including "information regarding every possible relationship permitted by the domain-specific ontology for the entities in the entity data," evaluating, "using a GAI filter agent," a level of consistency between each of the relationships and the input data, and including in the relationship data only the relationships that are consistent with the input data. The generate-every-permitted-relationship and consistency-filtering limitations are provided for by Jeong and Mihin for the same reasons set forth above with respect to claim 7. As to the filter agent being a "GAI filter agent", Mihin's filtration component performs the consistency evaluation but Mihin does not provide for the filter agent being a generative artificial intelligence model. Cheng provides for implementing the checking role with a generative model because Cheng discloses performing its knowledge extraction and relation prediction with a generative model (GPT-4) using in-context learning (§3, §4.1) and expressly suggests using stronger large language models for verification (§6). The proposed modification is to implement Mihin's filtration role in the combination with a large language model, as Cheng discloses for its own extraction and relation prediction. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to make this modification because the combined system already performs its extraction and inference with such models, and Cheng identifies model-based verification as the way to catch erroneous outputs (§6). The remaining proposed modifications and motivations are the same as set forth above with respect to claims 1 and 7. Regarding claim 3, the claim depends from claim 2 and recites limitations substantially similar to those of claim 18. Claim 3 is rejected for the same reasons set forth above with respect to claims 2 and 18. Regarding claim 8, the claim depends from claim 7 and further recites that a first entity extraction agent is instructed to identify, from the input data, entities of the first entity type and a second entity extraction agent is instructed to identify, from the input data, entities of a second entity type that is different from the first entity type, the entity data comprising the entities identified by both. Zhang provides for this limitation for the same reasons set forth above with respect to claim 1: Zhang's regular-expression mode identifies the entities of the preset categories and Zhang's trained model identifies the entities of the other, different ontology-defined types, and the identified entities together form the extracted entity set (Zhang, claims 1-2; step S12). The motivation to combine Zhang is the same as set forth above with respect to claim 1, and the motivation to combine Mihin is the same as set forth above with respect to claim 7. Claims 4 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng, Zhang, and Huang. Regarding claim 4, the claim depends from claim 1 and further recites that the first and second entity extraction agents generate initial entity data, that the initial entity data is provided to corresponding first and second validation agents, that the validation agents "perform natural language processing operations to validate that the entities of the initial entity data are consistent with the input data," that the validation agents remove from the initial entity data the entities that are inconsistent with the input data, and that the entity data is an aggregation of the initial entity data that was validated and not removed. Jeong, Cheng, and Zhang do not provide for validation agents that validate the extracted entities against the input data and remove the inconsistent entities. Huang provides for these limitations. Huang discloses detecting entities using a named entity identification model - a natural language processing operation - in both the model-generated content and the source key information, and calculating an entity difference set that identifies the entities of the generated content that are not present in the source information (Huang, abstract; description). Huang further discloses secondarily judging those candidate false entities against a local knowledge base to confirm which are truly false (Huang, abstract; description). Huang thereby validates entities produced by a model against the source information and separates the unsupported entities from the supported ones, so that only the supported entities are treated as real. Huang does not provide for the entity extraction pipeline of the combination; Huang is relied upon only for its entity-level verification technique, which in the combination is applied to the entities extracted by the system of Jeong, Cheng, and Zhang. The proposed modification is to provide the initial entity data generated by each extraction component of the Jeong/Cheng/Zhang system to a corresponding validation component that performs Huang's entity-level verification, removing from the initial entity data the entities not supported by the input data before the knowledge graph is constructed. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to make this modification. Cheng expressly identifies hallucinated extractions as a known problem of model-based extraction and suggests hallucination detection and model-based verification as the solution (§6); Huang is exactly such a hallucination-detection mechanism, operating at the entity level. One of ordinary skill would have been motivated to make the modification because a model-based extractor can output entities that do not appear in the input data; verifying each extracted entity against the input data and removing the unsupported ones ensures that the entity data - and the knowledge graph constructed from it - contains only entities the input data actually supports. Regarding claim 16, the claim depends from claim 15 and recites validating, using first and second validation agents corresponding to the extraction agents, that the entities of the initial entity data are "supported by or relevant to the input data," dropping the entities that are not, the entity data comprising only the validated entities that are not dropped. These limitations are substantially similar to those of claim 4 (without the express natural-language-processing requirement) and are provided for by Huang for the same reasons set forth above with respect to claim 4, with the limitations of claim 15 addressed arguendo as set forth above. The motivations to combine are the same as set forth above with respect to claims 1 and 4. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng, Zhang, Mihin, and Huang. Regarding claim 9, the claim depends from claim 8 and recites limitations substantially similar to those of claim 16. Claim 9 is rejected for the same reasons set forth above with respect to claims 7, 8, and 16. In this combination, Jeong's pipeline collects the threat reports and defines the entities, relationships, and knowledge graph; Zhang's type-based division supplies the per-type extraction components; Huang's entity-level verification supplies the validation components that drop unsupported entities; Mihin's consistency-based filtration supplies the relationship filtering of claim 7; and Cheng supplies the use of the resulting graph for retrieval-augmented question answering and the express suggestion to verify model outputs. The motivation to combine each reference are the same as set forth above with respect to claims 14, 1, 7, and 4, respectively. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng, Zhang, and Zhu. Regarding claim 5, the claim depends from claim 1 and further recites that the knowledge domain is cyber threat intelligence and the input data is a threat intelligence report, which Jeong provides for as set forth above with respect to claim 11. The claim further recites summarizing, "using an attack techniques discovery agent, the threat intelligence report into a plurality of attack steps," identifying, from the plurality of attack steps, an attack technique, and the relationship data further including information regarding relationships between entities and the attack technique. Jeong provides for identifying attack techniques and relating entities to them, because Jeong extracts an Attack_Vector metadata type listing attack methods according to industry-standard categories, including the MITRE categories ([0065], Table 4), and defines triples relating the entities to the attack elements, such as a tool used for an attack ([0101], Table 6). Jeong, Cheng, and Zhang do not provide for summarizing the threat intelligence report into a plurality of attack steps and identifying the attack technique from those steps. Zhu provides for these limitations. Zhu discloses filtering the non-attack-related sentences out of a cyber threat intelligence report using a pre-trained redundancy removing model, extracting the entities and their interaction relations from the filtered report, and organizing the attack into four attack stages abstracted from the ATT&CK tactics - an initial invasion stage, a malicious code executing stage, a breaking-through-access-control stage, and a leakage and damage stage - where the action specification emphasizes the causal relationship between the attack stages (Zhu, claims 1-2; abstract; description). Zhu further discloses that the target of each stage can be achieved using a plurality of techniques, each stage modeling a plurality of malicious event flows (Zhu, description). Zhu thereby summarizes the report into a plurality of causally ordered attack steps and identifies the techniques associated with those steps. The attack techniques discovery agent is Zhu's set of trained models given the summarizing and identifying role (see Claim Interpretation above). The proposed modification is to add, to the report processing of the Jeong/Cheng/Zhang system, Zhu's organization of the report's attack content into causally ordered attack stages - filtering the non-attack sentences with the redundancy removing model and mapping the attack into the stages abstracted from the ATT&CK tactics - and to identify the attack technique from those stages, relating the entities to the identified technique in the relationship data as Jeong already does for its extracted attack methods (Table 4; Table 6). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to make this modification. One of ordinary skill would have been motivated to do so because Jeong extracts attack-method labels directly from the text without representing the order in which the attack occurred; organizing the report into causally ordered stages and identifying the techniques from those stages causes the knowledge graph to represent not only which attack methods were used but the ordered process by which the attack occurred, improving the representation of the attack in the graph exactly as Zhu describes. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng, Mihin, and Zhu. Regarding claim 12, the claim depends from claim 11 and further recites summarizing, by an attack technique discovery agent, the threat intelligence report into a plurality of attack steps, identifying an attack technique from the plurality of attack steps, and the relationship data further including information regarding a relationship between the first entity and the attack technique. These limitations are substantially similar to those of claim 5 and are provided for by Zhu and Jeong for the same reasons set forth above with respect to claim 5. The relationship between the first entity and the attack technique is provided for by Jeong's triples relating entities to the attack elements ([0101]-[0102], Table 6). The motivations to combine are the same as set forth above with respect to claims 7 and 5. Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Jeong in view of Cheng and Zhu. Regarding claim 19, the claim depends from claim 14 and further recites that the knowledge domain is cyber threat intelligence, that the input data is a threat intelligence report, and that the instructions summarize the threat intelligence report into a plurality of attack steps and identify, from the plurality of attack steps, an attack technique, the relationship data further including information regarding a relationship between the first entity and the attack technique. Claim 19 does not recite an agent. These limitations are provided for by Jeong and Zhu for the same reasons set forth above with respect to claims 11, 5, and 12. The motivations to combine are the same as set forth above with respect to claims 14 and 5. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZAIN J AHMED whose telephone number is (571)270-0251. The examiner can normally be reached 8am - 4pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jeffrey L Nickerson can be reached at (469) 295-9235. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Jeffrey Nickerson/Supervisory Patent Examiner, Art Unit 2432
Read full office action

Prosecution Timeline

Nov 27, 2024
Application Filed
Sep 04, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month