Prosecution Insights
Last updated: October 01, 2026
Application No. 18/216,772

ROBUST GRAPH REPRESENTATION OF CAUSAL RELATIONSHIPS EXPRESSED IN A NATURAL LANGUAGE DOCUMENT

Final Rejection §101§103
Filed
Jun 30, 2023
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
2 (Final)
51%
Grant Probability
Moderate
3-4
OA Rounds
1y 2m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
78 granted / 152 resolved
-3.7% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
40 currently pending
Career history
216
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
4.6%
-35.4% vs TC avg
§112
24.0%
-16.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 152 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This Office Action is responsive to Applicants' Amendment filed on July 15, 2026, in which claims 1-4, 8, 11-13, and 17-20 are currently amended. Claims 1-20 are currently pending. Response to Arguments The rejections to claims 1-20 under 35 U.S.C. § 112(b) are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-20 under 35 U.S.C. 101 based on amendment have been considered, however, are not persuasive. With respect to Applicant’s arguments on p. 9 of the Remarks submitted 7/15/2026 that "constructing a first graph representing the cause-effect pairs extracted using the first LLM" and "constructing a second graph representing the cause-effect pairs extracted", "intentionally introduce[ing] variation for evaluating graph structure stability", "measuring a graph similarity between the first graph and the second graph", and "incorporating the first graph into a knowledge base when the measured graph similarity exceeds an acceptance threshold” “are not practically performable in the human mind”, Examiner respectfully disagrees. The human mind is readily capable and routinely performs graph construction and evaluation with or without the assistance of tools such as pen and paper. “Using the first LLM” and “using the second LLM” amount to mere instructions to apply the judicial exception using generic computer components and do not integrate the judicial exception into a practical application. The generic computer components are recited at a high level of generality and recited to apply it rather than providing objective technical improvement (MPEP 2106.07(a)(II) "employing well-known computer functions to execute an abstract idea, even when limiting the use of the idea to one particular environment, does not integrate the exception into a practical application" MPEP 2106.05(a) also recites "An important consideration in determining whether a claim improves technology is the extent to which the claim covers a particular solution to a problem or a particular way to achieve a desired outcome, as opposed to merely claiming the idea of a solution or outcome.”). With respect to Applicant’s arguments on p. 9 of the Remarks submitted 7/15/2026 that “extracting cause-effect pairs from a natural language document using a first large language model(LLM)” and “extracting the cause-effect pairs from the natural language document using a second LLM” are “not practically performed in the mind”: while Examiner does not necessarily agree, these limitations have been interpreted as insignificant extra-solution activity of gathering and outputting data (See MPEP 2106.05(g)) which is well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i)) and does not integrate the judicial exception into a practical application. With respect to Applicant’s arguments that “the extraction step is not insignificant extra-solution activity because the extracted cause-effect pairs are transformed into graph representations that are subsequently subjected to graph similarity analysis and threshold-based validation”, Examiner respectfully disagrees. The cause-effect graph transformation process is a mental process subsequent to data gathering which can be readily performed entirely in the mind with the assistance of tools such as pen and paper. For at least these reasons and those further detailed below Examiner asserts that the rejection under 35 U.S.C. 101 is reasonable and should be maintained. Applicant’s arguments with respect to rejection of claims 1-20 under 35 U.S.C. 103 based on amendment have been considered and are persuasive. The argument is moot in view of a new ground of rejection set forth below. Claim Rejections - 35 USC § 101 101 Rejection 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 USC § 101 because the claimed invention is directed to non-statutory subject matter. Regarding Claim 1: Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories. Step 2A Prong One Analysis: Claim 1 under its broadest reasonable interpretation is a series of mental processes. For example, but for the generic computer components language, the above limitations in the context of this claim encompass machine learning processing, including the following: constructing a first graph representing the set of cause-effect pairs extracted using the first LLM, each node in the first graph representing a phrase in the pair of phrases, each edge in the first graph representing a cause-effect relationship between nodes connected by an edge (observation, evaluation, and judgement), constructing a second graph representing the cause-effect pairs extracted using the second LLM, wherein the second graph is constructed using the second LLM to intentionally introduce variation for evaluating graph structure stability; (observation, evaluation, and judgement) measuring a graph similarity between the first graph and a second graph (observation, evaluation, and judgement) responsive to determining that the graph similarity is above an acceptance threshold, incorporating the first graph into a knowledge base (observation, evaluation, and judgement) Therefore, claim 1 recites an abstract idea which is a judicial exception. Step 2A Prong Two Analysis: Claim 1 recites additional elements “A computer-implemented method”. However, these additional features are computer components recited at a high-level of generality, such that they amount to no more than mere instructions to apply the judicial exception using a generic computer component. An additional element that merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, does not integrate the judicial exception into a practical application (See MPEP 2106.05(f)). Claim 1 also recites additional elements “extracting, from a natural language document, using a first large language model (LLM), each cause-effect pair comprising a pair of phrases, each phrase comprising a portion of the natural language document” and “extracting the cause-effect pairs from the natural language document using a second LLM” which amounts to gathering and outputting data which is insignificant extra-solution activity (See MPEP 2106.05(g)). Therefore, claim 1 is directed to a judicial exception. Step 2B Analysis: Claim 1 does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the lack of integration of the abstract idea into a practical application, the additional elements recited in claim 1 amount to no more than mere instructions to apply the judicial exception using a generic computer component and insignificant extra-solution activity. The gathering and outputting of data is considered well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i)). For the reasons above, claim 1 is rejected as being directed to non-patentable subject matter under §101. This rejection applies equally to independent claims 8 and 17, which recite a computer program product and a system, respectively, as well as to dependent claims 2-7, 9-16, and 18-20. Independent claim 8 recites additional instructions to apply the judicial exception using generic computer components “A computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by a processor to cause the processor to perform operations comprising”. Independent claim 17 recites additional instructions to apply the judicial exception using generic computer components The additional limitations of the dependent claims are addressed briefly below: Dependent claims 2, 11, and 18 recite additional observation, evaluation, and judgement “wherein the first graph comprises a combined node, the combined node representing two phrases in the pair of phrases with a similarity score above a first clustering threshold value” Dependent claims 3, 12, and 19 recite additional observation, evaluation, and judgement “the first clustering threshold value is set according to similarity scores of phrases in the pair of phrases” Dependent claims 4, 13, and 20 recite additional observation, evaluation, and judgement “the similarity score is computed by computing a cosine similarity between sentence embeddings, each sentence embedding comprising a numerical representation of a phrase in the pair of phrases, each sentence embedding computed using a first embedding model” Dependent claims 5 and 14 recite additional observation, evaluation, and judgement “the second graph is constructed using a different embedding model from the first embedding model” Dependent claims 6 and 15 recite additional observation, evaluation, and judgement “the second graph is constructed using a different clustering threshold value from the first clustering threshold value” Dependent claims 7 and 16 recites additional observation, evaluation, and judgement “the graph similarity comprises a combination of a node similarity measurement and an edge similarity measurement between the first graph and the second graph” Dependent claim 9 recites additional instructions to apply the judicial exception using a generic computer component “wherein the stored program instructions are stored in a computer readable storage device in a data processing system” as well as additional insignificant extra-solution activity of gathering and outputting data (See MPEP 2106.05(g)) “wherein the stored program instructions are transferred over a network from a remote data processing system” which is well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i)) Dependent claim 10 recites additional instructions to apply the judicial exception using generic computer components “wherein the stored program instructions are stored in a computer readable storage device in a server data processing system” as well as additional insignificant extra-solution activity of gathering and outputting data (See MPEP 2106.05(g)) “and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system, further comprising” which is well-understood, routine, and conventional in the art (See MPEP 2106.05(d)(II)(i)). Claim 10 also recites additional observation, evaluation, and judgement “to meter use of the program instructions associated with the request” and “to generate an invoice based on the metered use” Therefore, when considering the elements separately and in combination, they do not add significantly more to the inventive concept. Accordingly, claims 1-20 are rejected under 35 U.S.C. § 101. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-3, 6-8, 11, 12, and 15-19 are rejected under U.S.C. §103 as being unpatentable over the combination of Maisonnave (“Causal graph extraction from news: a comparative study of time-series causality learning techniques”, 2022) and Li (“gBuilder: A Scalable Knowledge Graph Construction System for Unstructured Corpus”, 2022) as evidenced by Chen (“Expected Returns and Large Language Models”, 2022). PNG media_image1.png 744 1148 media_image1.png Greyscale FIG. 1 of Maisonnave Regarding claim 1, Maisonnave teaches A computer-implemented method comprising:([p. 4] "The data and full code of the methods used by the framework and to carry out the experiments are made available to allow reproducibility") extracting cause-effect pairs, from a natural language document using a first [large] language model [(LLM)] ([p. 8] "The framework takes as input a topic description and a corpus of news articles. It then applies eight steps aimed at building a causal graph associated with the topic of interest." [p. 2] "Causal modeling aims to determine the cause–effect relations among a set of variables" [p. 3] "The proposed approach combines methods coming from information retrieval, natural language processing, machine learning, and Econometrics into a framework that extracts variables from large volumes of text to build highly interpretable causal models" [p. 19] "Note that for each of the 45 evaluated pairs {v1,v2} it was possible to derive two Boolean assessments: (1) v1 causes v2 or v1 does not cause v2 and (2) v2 causes v1 or v2 does not cause v1. As a result, we obtained a total of 90 Boolean labels from each annotator") each cause-effect pair comprising a pair of phrases, each phrase comprising a portion of the natural language document;([p. 8] "Select terms from the topic-relevant sentences. The selection of relevant terms (unigrams, bigrams, and trigrams) from the given sentences relies on FDDb,a supervised term-weighting scheme proposed and evaluated" [p. 10] "A causal graph is learned where the nodes are the variables (terms and event clusters) identified in steps 2 and 5. The edges of the graph are the causal relations learned by applying a causal structure learning technique to the time series generated" [p. 12] "‘‘chemical biological weapons’’ ‘‘military action Iraq’’") constructing a first graph representing the cause-effect pairs extracted using the first LLM, ([p. 8] "Select terms from the topic-relevant sentences. The selection of relevant terms (unigrams, bigrams, and trigrams) from the given sentences relies on FDDb,a supervised term-weighting scheme proposed and evaluated" [p. 10] "A causal graph is learned where the nodes are the variables (terms and event clusters) identified in steps 2 and 5. The edges of the graph are the causal relations learned by applying a causal structure learning technique to the time series generated") each node in the first graph representing a phrase in the pair of phrases, ([p. 10] "A causal graph is learned where the nodes are the variables (terms and event clusters)" [p. 12] "a different number of variables could be used to build the time series" Maisonnave explicitly says Step 2 selects relevant terms (unigrams, bigrams, and trigrams) from the given sentences and that the nodes include the terms while edges are the causal relations. More importantly, the reference treats terms and event clusters as independently selectable variable classes, such that it is already anticipated by Maisonnave that each of the graph nodes represent a phrase in the pair of phrases. A skilled person wishing to obtain a causal graph whose nodes are textual phrases therefore needs only to select the already-disclosed term variables and omit the event cluster variable. In fact, in the exemplary embodiment in Maisonnave the graph comprises a subgraph comprising only phrase nodes.) each edge in the first graph representing a cause-effect relationship between nodes connected by an edge;([p. 8] "Select terms from the topic-relevant sentences. The selection of relevant terms (unigrams, bigrams, and trigrams) from the given sentences relies on FDDb,a supervised term-weighting scheme proposed and evaluated" [p. 10] "A causal graph is learned where the nodes are the variables (terms and event clusters) identified in steps 2 and 5. The edges of the graph are the causal relations learned by applying a causal structure learning technique to the time series generated") extracting the cause-effect pairs from the natural language document using a second [L]LM;([p. 22] "An initial evaluation using synthetic data of nine state-of-the-art causal structure learning techniques allowed us to address RQ1, offering insight into which are the most promising methods for time-series causality learning." [p. 7] "These techniques take a time series of variables and generate a causal graph [...] and ensemble models. The ensemble models combine Direct-LiNGAM, PCMCI, VAR, and PC (which proved to be the four most effective techniques according to evaluations carried out on synthetic data)" See also FIG. 1) constructing a second graph representing the cause-effect pairs extracted using the second [L]LM, ([p. 7] "These techniques take a time series of variables and generate a causal graph [...] and ensemble models. The ensemble models combine Direct-LiNGAM, PCMCI, VAR, and PC (which proved to be the four most effective techniques according to evaluations carried out on synthetic data)" See also FIG. 1. Because multiple different causal-learning techniques each generate a causal graph from the same variable time series, the reference necessarily discloses multiple model-dependent graph outputs) wherein the second graph is constructed using the second [L]LM to intentionally introduce variation for evaluating graph structure stability;([p. 7] "These techniques take a time series of variables and generate a causal graph [...] and ensemble models. The ensemble models combine Direct-LiNGAM, PCMCI, VAR, and PC (which proved to be the four most effective techniques [...] The first ensemble technique, referred to as ensemble∩, adds a causal relation only when the four best techniques agree on including it. On the other hand, the second ensemble technique, called ensemble∪, adds a causal relation when any of the four techniques includes it" [p. 16] "to identify the most promising time-series causality learning techniques" [p. 12] "The ensemble∩ technique was used in this example because high precision is desired to only include causal relations with high confidence" Maisonnave deliberately changes (varies) the causal-learning method while holding the derived input data fixed) measuring a graph similarity between the first graph and a second graph ([p. 7] "The first ensemble technique, referred to as ensemble∩, adds a causal relation only when the four best techniques agree on including it" Maisonnave performs cross-model graph output agreement testing at the causal edge level, comparing similarity among graph outputs) and responsive to determining that the graph similarity is above an acceptance threshold, ([p. 12] "The ensemble∩ technique was used in this example because high precision is desired to only include causal relations with high confidence"). However, Maisonnave does not explicitly teach incorporating the first graph into a knowledge base. using a LLM Li, in the same field of endeavor, teaches incorporating the first graph into a knowledge base.([p. 4 §4.2] "A desirable approach is to unite build-in models to conform to the target KG ontology and use them directly without training custom models. However, well-tuned build-in models are often inconsistent with the target ontology, and the distribution of training data usually deviates from the target domain. To mitigate this phenomenon, we introduce and integrate multiple built-in IE models in gBuilder" [p. 5] "uniting different IE models can be regarded as merging multiple ontologies into a target ontology [...] The highest score label would be treated as the final results, and a threshold is introduced for filtering the unreliable predication […] re 𝑤𝑗 denotes the weight of the 𝑗-th classifier (this weight is specified by the user, by default it is 1/𝑚). 𝑝 𝑗 𝑖,𝑘 denotes the score that the 𝑗-th classifier predicting the 𝑖-th sample into the 𝑘-th class, 𝛿 is the threshold to accept the predicted label (by default it is 0.5), if all prediction score below this threshold, the corresponding sample will be ignored, e.g., in relation classification task it will be marked as no relation between entities.") using a LLM([p. 10] "The conducted pipeline models are as follows [...] Among them, BERT series models can be used to evaluate the application of PLMs in gBuilder (RoBERTa is the largest PLM among BERT series)"). Maisonnave as well as Li are directed towards ensembling of natural language models for relation extraction. Therefore, Maisonnave as well as Li are analogous art in the same field of endeavor. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of Maisonnave with the teachings of Li by using Li's LLM's and . Li provides as additional motivation for combination ([p. 4 §4.2] "A desirable approach is to unite build-in models to conform to the target KG ontology and use them directly without training custom models. However, well-tuned build-in models are often inconsistent with the target ontology, and the distribution of training data usually deviates from the target domain. To mitigate this phenomenon, we introduce and integrate multiple built-in IE models in gBuilder"). While Li already discloses using BERT, and while one of ordinary skill in the art would recognize that BERT could be considered an LLM, the combination of Maisonnave and Li does not explicitly teach that BERT is an LLM. Chen, in the same field of endeavor, is introduced to reinforce that BERT is an LLM ([p. 5] "We examine six large-scale pre-trained LLMs: BERT (developed by Google), RoBERTa (by Meta)"). The combination of Maisonnave and Li as well as Chen are directed towards natural language processing. Therefore, the combination of Maisonnave and Li as well as Chen are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Maisonnave and Li with the teachings of Chen by recognizing BERT as an LLM. Chen provides as additional motivation for combination ([p. 5] "We examine six large-scale pre-trained LLMs: BERT (developed by Google), RoBERTa (by Meta)"). This motivation for combination also applies to the remaining claims which depend on this combination. Regarding claim 2, the combination of Maisonnave, Li, and Chen teaches The computer-implemented method of claim 1, wherein the first graph comprises a combined node, the combined node representing two phrases in the pair of phrases (Maisonnave [p. 10] "An event-phrase embedding representation based on GloVe vectors 1 (with dimension 300) is built for each event trigger ek in each sentences or phrase P") with a similarity score above a first clustering threshold value.(Li [p. 4 §4.2] "A desirable approach is to unite build-in models to conform to the target KG ontology and use them directly without training custom models. However, well-tuned build-in models are often inconsistent with the target ontology, and the distribution of training data usually deviates from the target domain. To mitigate this phenomenon, we introduce and integrate multiple built-in IE models in gBuilder" [p. 5] "uniting different IE models can be regarded as merging multiple ontologies into a target ontology [...] The highest score label would be treated as the final results, and a threshold is introduced for filtering the unreliable predication […] re 𝑤𝑗 denotes the weight of the 𝑗-th classifier (this weight is specified by the user, by default it is 1/𝑚). 𝑝 𝑗 𝑖,𝑘 denotes the score that the 𝑗-th classifier predicting the 𝑖-th sample into the 𝑘-th class, 𝛿 is the threshold to accept the predicted label (by default it is 0.5), if all prediction score below this threshold, the corresponding sample will be ignored, e.g., in relation classification task it will be marked as no relation between entities."). Regarding claim 3, the combination of Maisonnave, Li, and Chen teaches The computer-implemented method of claim 2, wherein the first clustering threshold value is set according to similarity scores of phrases in the pair of phrases.(Li [p. 4 §4.2] ") Chuck Extraction-based model 𝐶𝐸, which outputs a set of fact phrase chunks [...] A desirable approach is to unite build-in models to conform to the target KG ontology and use them directly without training custom models. However, well-tuned build-in models are often inconsistent with the target ontology, and the distribution of training data usually deviates from the target domain. To mitigate this phenomenon, we introduce and integrate multiple built-in IE models in gBuilder" [p. 5] "uniting different IE models can be regarded as merging multiple ontologies into a target ontology [...] The highest score label would be treated as the final results, and a threshold is introduced for filtering the unreliable predication […] re 𝑤𝑗 denotes the weight of the 𝑗-th classifier (this weight is specified by the user, by default it is 1/𝑚). 𝑝 𝑗 𝑖,𝑘 denotes the score that the 𝑗-th classifier predicting the 𝑖-th sample into the 𝑘-th class, 𝛿 is the threshold to accept the predicted label (by default it is 0.5), if all prediction score below this threshold, the corresponding sample will be ignored, e.g., in relation classification task it will be marked as no relation between entities."). Regarding claim 6, the combination of Maisonnave, Li, and Chen teaches The computer-implemented method of claim 2, wherein the second graph is constructed using a different clustering threshold value from the first clustering threshold value.(Maisonnave [p. 9] "The tunable parameter β is a positive real factor that offers a means to favor descriptive relevance over discriminative relevance (by using a β value higher than 1) or the other way around (by using a β value smaller than 1). Human-subject studies reported by Maisonnave et al. (2021a) indicate that a β =0.477 offers a good balance between descriptive and discriminative power" [p. 10] "In Step 2, the user can configure the β value according to the specific needs. In Step 5, the user should analyze different K values for the KMeans algorithm to choose the one that better suits the use case under analysis" [p. 11] "Step5. MiniBatch KMeans is applied to group the 498,560 event-phrase embedding representations built in Step 4 into 1,000 clusters. The value K =1,000 is selected by applying the Elbow method (Thorndike, 1953). Only six highly cohesive clusters with a clear semantic and containing a large number of event mentions are selected to define event cluster variables" Maisonnave K and β are interpreted as second threshold clustering values different than the first clustering threshold value). Regarding claim 7, the combination of Maisonnave, Li, and Chen teaches The computer-implemented method of claim 1, wherein the graph similarity comprises a combination of a node similarity measurement (Maisonnave [p. 10] "The EPER representation allows to create a phrase embedding that accounts for the GloVe representation of each word wi in P with a quadratic penalization based on the distance of wi to ek." distance of wi to ek interpreted as node similarity measurement) and an edge similarity measurement between the first graph and the second graph.(Maisonnave [p. 7] "The first ensemble technique, referred to as ensemble∩, adds a causal relation only when the four best techniques agree on including it"). Regarding claim 8, claim 8 is directed towards a computer program product for performing the method of claim 1. Therefore, the rejection applied to claim 1 also applies to claim 8. Claim 8 also recites additional elements A computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by a processor to cause the processor to perform operations comprising (Maisonnave [p. 4] "The data and full code of the methods used by the framework and to carry out the experiments are made available to allow reproducibility"). Similarly, regarding claims 11, 12, 15, and 16, claims 11, 12, 15, and 16 are directed towards a computer program product for performing the method of claims 2, 3, 6, and 7, respectively. Therefore the rejections applies to claims 2, 3, 6, and 7 also apply to claims 11, 12, 15, and 16. Regarding claims 17-19, claims 17-19 are directed towards a system for performing the method of claims 1-3, respectively. Therefore, the rejections applied to claims 1-3 also apply to claims 17-19. Claims 4, 5, 9, 13, 14, and 20 are rejected under U.S.C. §103 as being unpatentable over the combination of Maisonnave, Li, Chen, and in further view of Bjorkqvist (US20220207240A1). Regarding claim 4, the combination of Maisonnave, Li, and Chen teaches each sentence embedding comprising a numerical representation of a phrase in the pair of phrases, (Maisonnave [p. 10] "A causal graph is learned where the nodes are the variables (terms and event clusters)" [p. 12] "a different number of variables could be used to build the time series" Maisonnave explicitly says Step 2 selects relevant terms (unigrams, bigrams, and trigrams) from the given sentences and that the nodes include the terms while edges are the causal relations. More importantly, the reference treats terms and event clusters as independently selectable variable classes, such that it is already anticipated by Maisonnave that each of the graph nodes represent a phrase in the pair of phrases. A skilled person wishing to obtain a causal graph whose nodes are textual phrases therefore needs only to select the already-disclosed term variables and omit the event cluster variable. In fact, in the exemplary embodiment in Maisonnave the graph comprises a subgraph comprising only phrase nodes.) each sentence embedding computed using a first embedding model.(Maisonnave [p. 8] "The framework takes as input a topic description and a corpus of news articles. It then applies eight steps aimed at building a causal graph associated with the topic of interest." [p. 2] "Causal modeling aims to determine the cause–effect relations among a set of variables" [p. 3] "The proposed approach combines methods coming from information retrieval, natural language processing, machine learning, and Econometrics into a framework that extracts variables from large volumes of text to build highly interpretable causal models" [p. 19] "Note that for each of the 45 evaluated pairs {v1,v2} it was possible to derive two Boolean assessments: (1) v1 causes v2 or v1 does not cause v2 and (2) v2 causes v1 or v2 does not cause v1. As a result, we obtained a total of 90 Boolean labels from each annotator"). However, the combination of Maisonnave, Li, and Chen doesn't explicitly teach The computer-implemented method of claim 2, wherein the similarity score is computed by computing a cosine similarity between sentence embeddings. Bjorkqvist, in the same field of endeavor, teaches The computer-implemented method of claim 2, wherein the similarity score is computed by computing a cosine similarity between sentence embeddings, ([¶0089] "Cosine similarity is one possible criterion for similarity of graphs or vectors derived therefrom"). The combination of Maisonnave, Li, and Chen as well as Bjorkqvist are directed towards natural language models for relation extraction. Therefore, the combination of Maisonnave, Li, and Chen as well as Bjorkqvist are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Maisonnave, Li, and Chen with the teachings of Bjorkqvist by using cosine similarity between sentence embeddings for the similarity calculation. Bjorkqvist provides as additional motivation for combination ([¶0089] “Cosine similarity is one possible criterion for similarity of graphs or vectors derived therefrom”). Regarding claim 5, the combination of Maisonnave, Li, Chen, and Bjorkqvist teaches The computer-implemented method of claim 4, wherein the second graph is constructed using a different embedding model from the first embedding model.(Maisonnave [p. 7] "These techniques take a time series of variables and generate a causal graph [...] and ensemble models. The ensemble models combine Direct-LiNGAM, PCMCI, VAR, and PC (which proved to be the four most effective techniques [...] The first ensemble technique, referred to as ensemble∩, adds a causal relation only when the four best techniques agree on including it. On the other hand, the second ensemble technique, called ensemble∪, adds a causal relation when any of the four techniques includes it" [p. 16] "to identify the most promising time-series causality learning techniques" [p. 12] "The ensemble∩ technique was used in this example because high precision is desired to only include causal relations with high confidence" Maisonnave deliberately changes (varies) the causal-learning method while holding the derived input data fixed). Regarding claim 9, the combination of Maisonnave, Li, and Chen teaches The computer program product of claim 8, wherein the stored program instructions are stored in a computer readable storage device in a data processing system, (Maisonnave [p. 4] "The data and full code of the methods used by the framework and to carry out the experiments are made available to allow reproducibility"). However, the combination of Maisonnave, Li, and Chen doesn't explicitly teach and wherein the stored program instructions are transferred over a network from a remote data processing system. Bjorkqvist, in the same field of endeavor, teaches and wherein the stored program instructions are transferred over a network from a remote data processing system.([¶0065] "The terms “data storage unit/means”, “processing unit/means” and “user interface unit/means” refer primarily to software means, i.e. computer-executable code, that are adapted to carry out the specified functions, that is, storing of digital data, allowing user to interact with the data, and processing the data, respectively. All of these components of the system can be carried in a software run by either a local computer or a web server, through a locally installed web browser, for example, supported by suitable hardware for running the software components"). The combination of Maisonnave, Li, and Chen as well as Bjorkqvist are directed towards natural language models for relation extraction. Therefore, the combination of Maisonnave, Li, and Chen as well as Bjorkqvist are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Maisonnave, Li, and Chen with the teachings of Bjorkqvist by using cosine similarity between sentence embeddings for the similarity calculation. Bjorkqvist provides as additional motivation for combination ([¶0089] “Cosine similarity is one possible criterion for similarity of graphs or vectors derived therefrom”). Regarding claims 13 and 14, claims 13 and 14 are directed towards a computer program product for performing the method of claims 4 and 5, respectively. Therefore the rejections applies to claims 4 and 5 also apply to claims 13 and 14. Regarding claim 20, claim 20 is directed towards a system for performing the method of claim 4. Therefore, the rejection applied to claim 4 also applies to claim 20. Claim 10 is rejected under U.S.C. §103 as being unpatentable over the combination of Maisonnave, Li, Chen, Bjorkqvist and in further view of Harrison (US20080046378A1). Regarding claim 10, the combination of Maisonnave, Li, and Chen teaches The computer program product of claim 8. However, the combination of Maisonnave, Li, and Chen doesn't explicitly teach, wherein the stored program instructions are stored in a computer readable storage device in a server data processing system, and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system, further comprising: program instructions to meter use of the program instructions associated with the request; and program instructions to generate an invoice based on the metered use. Bjorkqvist, in the same field of endeavor, teaches The computer program product of claim 8, wherein the stored program instructions are stored in a computer readable storage device in a server data processing system ([¶0065] "The terms “data storage unit/means”, “processing unit/means” and “user interface unit/means” refer primarily to software means, i.e. computer-executable code, that are adapted to carry out the specified functions, that is, storing of digital data, allowing user to interact with the data, and processing the data, respectively. All of these components of the system can be carried in a software run by either a local computer or a web server, through a locally installed web browser, for example, supported by suitable hardware for running the software components"). The combination of Maisonnave, Li, and Chen as well as Bjorkqvist are directed towards natural language models for relation extraction. Therefore, the combination of Maisonnave, Li, and Chen as well as Bjorkqvist are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Maisonnave, Li, and Chen with the teachings of Bjorkqvist by using cosine similarity between sentence embeddings for the similarity calculation. Bjorkqvist provides as additional motivation for combination ([¶0089] “Cosine similarity is one possible criterion for similarity of graphs or vectors derived therefrom”). However, the combination of Maisonnave, Li, and Chen as well as Bjorkqvist doesn’t explicitly teach wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system. Harrison, in the same field of endeavor, teaches and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system, further comprising:([¶0033] "download the software to the entity based on a specific request" [¶0042] "access the data base remotely through a wide area network, such as the Internet, to learn about the modules available for equipment and systems being purchased or already owned") program instructions to meter use of the program instructions associated with the request; and ([¶0027] "the installed software may track and report the use of each of the software modules either directly or indirectly to the manufacturer for the purpose of billing.") program instructions to generate an invoice based on the metered use.([¶0051] "A summary of the usage of the module may be used to generate an invoice"). The combination of Maisonnave, Li, Chen, and Bjorkqvist as well as Harrison are directed towards software systems. Therefore, the combination of Maisonnave, Li, Chen, and Bjorkqvist as well as Harrison are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to combine the teachings of the combination of Maisonnave, Li, Chen, and Bjorkqvist with the teachings of Harrison by implementing the system in Maisonnave, Li, Chen, and Bjorkqvist in the pay-per-use software system in Harrison. While the financial benefits of doing this would be obvious to one of ordinary skill in the art, Harrison provides as additional motivation for combination ([¶0005] “Providing a means for the interested users of the customer to try the software may increase license sales.”). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Friedman (“From Unstructured Text to Causal Knowledge Graphs: A Transformer-Based Approach”, 2022) is directed towards causal relation extraction using LLM’s for knowledge graph generation. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Jun 30, 2023
Application Filed
Apr 29, 2026
Non-Final Rejection mailed — §101, §103
Jun 09, 2026
Interview Requested
Jun 29, 2026
Applicant Interview (Telephonic)
Jun 29, 2026
Examiner Interview Summary
Jul 15, 2026
Response Filed
Sep 15, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
51%
Grant Probability
86%
With Interview (+35.1%)
4y 5m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 152 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month