Prosecution Insights
Last updated: October 02, 2026
Application No. 17/961,798

METHOD FOR TEMPORAL KNOWLEDGE GRAPH REASONING BASED ON DISTRIBUTED ATTENTION

Non-Final OA §101§103§112
Filed
Oct 07, 2022
Priority
Jun 10, 2022 — CN 202210658296.0
Examiner
KAWSAR, ABDULLAH AL
Art Unit
2100
Tech Center
2100 — Computer Architecture & Software
Assignee
Huazhong University of Science and Technology
OA Round
2 (Non-Final)
79%
Grant Probability
Favorable
2-3
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
321 granted / 407 resolved
+23.9% vs TC avg
Strong +56% interview lift
Without
With
+56.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
4 currently pending
Career history
416
Total Applications
across all art units

Statute-Specific Performance

§101
16.3%
-23.7% vs TC avg
§103
44.0%
+4.0% vs TC avg
§102
12.2%
-27.8% vs TC avg
§112
23.5%
-16.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 407 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement(s) submitted on 04/30/2025 and 06/13/2025 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification The abstract of the disclosure is objected to because the length of the abstract is over 150 words. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b). The disclosure is objected to because of the following informalities: Multiple patent documents are referenced in the specification, however they were not provided in an IDS. Appropriate correction is required. The listing of references in the specification is not a proper information disclosure statement. 37 CFR 1.98(b) requires a list of all patents, publications, or other information submitted for consideration by the Office, and MPEP § 609.04(a) states, "the list may not be incorporated into the specification but must be submitted in a separate paper." Therefore, unless the references have been cited by the examiner on form PTO-892, they have not been considered. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function. Such claim limitation(s) is/are: a scheduling unit [...]; a processing unit […];an adjusting unit […]; and a training unit as well as their accompanying tasks that they are configured for in claim 26. Because this/these claim limitation(s) is/are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are not being interpreted to cover only the corresponding structure, material, or acts described in the specification as performing the claimed function, and equivalents thereof. If applicant intends to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to remove the structure, materials, or acts that performs the claimed function; or (2) present a sufficient showing that the claim limitation(s) does/do not recite sufficient structure, materials, or acts to perform the claimed function. Claim Objections Claim 24 objected to because of the following informalities: Extra period at the end of the claim. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 24 recites the limitation "these" in and exclude these sequences from an attention operation. There is insufficient antecedent basis for this limitation in the claim. There are multiple types of sequences mentioned in this claim, such as filled sequences and the longest sequence. Please clarify what sequences are being excluded in the claim language. Claim 25 is also rejected as it is dependent on Claim 24 and inherent the same issues. Claim 26 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as failing to set forth the subject matter which the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the applicant regards as the invention. Claim limitation “a scheduling unit [...]; a processing unit […];an adjusting unit […]; and a training unit” invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. 0028 and 0053 introduce the units, but the units themselves do not contain any structural limitation such that they can perform as an "unit". The written description nor the claims themselves fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 28 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. As per claim 28, the claim limitation recite “storage medium”. However, the usage of the phrase “storage medium” is broad enough to include both “non-transitory” and “transitory” media. The specification further explicitly does not limit the utilization of a non-transitory computer-usable storage medium (see Specification paragraph [0077], “computer readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems“ ). Also, extrinsic evidence suggests that “social medium” covers a signal per se. Therefore, when the broadest reasonable interpretation of a claim covers a signal per se, the claim must be rejected under 35 U.S.C 101 as covering non-statutory subject matter. See In re Nuijten, 500 F.3d 1346, 1356-57 (Fed Cir. 2007) (transitory embodiments are not directed to statutory subject matter). Therefore Claim 28 are not statutory. A suggestion is made that the Applicant amends the claim to recite non-transitory storage medium. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 21, 22, 24, and 26 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu TIE: A Framework for Embedding-based Incremental Temporal Knowledge Graph Completion and further in view of Yang ST-LBAGAN: Spatio-temporal learnable bidirectional attention generative adversarial networks for missing traffic data imputation Regarding Claim 21, Wu teaches: A method for temporal knowledge graph reasoning based on distributed attention, the method comprising: (Wu teaches: “We introduce a new task, incremental TKGC, and propose TIE, a training and evaluation framework that integrates incremental learning with TKGC. TIE combines TKG representation learning, experience replay, temporal regularization to improve model performance and alleviate catastrophic forgetting.” [Page 2] which teaches Temporal Knowledge Graph Completion (TKGC) through fine tuning a model with weights (weights read on attention).) recombining a temporal knowledge graph in a temporal serialization manner according to an order of timestamps in the temporal knowledge graph, and storing distribution of historical timestamp subgraphs into a sparse matrix; (Wu teaches: “We define a time window ranging from 𝑡 −𝜏𝑑 to 𝑡−1 to limit the scope of evaluation. For every quadruple (𝑠, 𝑟, 𝑜, 𝑡), we aim to find and then evaluate the related deleted facts from this time window. “[Page 4] which teaches recombining a first set of knowledge graph data (s, r, o, t) into a time window (reads on temporal serialization manner as explained in the spec, see the quotation below, as the time window is a subgraph of the original graph) and defined on the order of timestamps. Wu further teaches storing the data into a collection set O,s,r,t which is a representation of data points of a set (reads on matrix) which only includes negative numbers: “where 𝑂 ′ 𝑠,𝑟,𝑡 is the collection of negative objects and 𝑍𝑡 is the normalizing constant“ [Page 4], therefore it can be considered a sparse matrix as the matrix only includes nonzero elements. Temporal serialization as defined by the applicant’s specification: “A temporal knowledge graph, which contains an entity set ϵ having a size of N, a relation set having a size of P, and a timestamp set having a size of T, is partitioned into a sequence of temporal subgraphs.” [0053] as the data of Wu is converted into subsets (which are representations of subgraphs) it is understood that they are portioned in a temporal serialization manner.) constructing facts of predicted timestamps using an attention mechanism and assigning initial first-layer attention to the facts that are historically repeated; (Wu teaches: “The parameter set is 𝜽 𝑡−1 = {𝑬 𝑡−1 , 𝑹 𝑡−1 , 𝑾𝑡−1 , 𝑩 𝑡−1 }, where 𝑾 and 𝑩 are the matrices of weight and bias parameters in Equation (1)” [Page 6] which teaches the first layer of the neural network containing weights W for a predicted timestamp (t-1) for the mechanism of being updated by training the model (attention mechanism) which are initialized before training (therefore initial by “Glorot Initialization [10] if an entity appears for the first time” [Page 6]) and are assigned to the weight matrix that are tied to the weights of the facts that are historically repeated.) building second-layer attention based on statistics of historical frequency information that evolves with the timestamps, and adjusting a score of the initial first-layer attention according to updates in knowledge; (Wu teaches calculating a second layer of weights (attention) based on added facts on Page 7 Section 5.6 (Optimization). The weights are created to control the “the relative emphasis placed on each loss term” and one score, LCE as shown above is based on a score from added facts which are represented as facts (still temporal data, therefore historical frequency information as the data stores the frequency of the added facts) that change (evolve) according to each timestep. Furthermore, the first-layer attention weights are updated based on the set of input data by training the model, therefore are “adjusted” according to updated in knowledge (see Page 5).) and according to a parameter training strategy, training a model with multi-class tasks based on cross entropy loss; (Wu teaches cross-entropy being used to train the model with multi-class tasks (class tasks is understood to be multiple calculation tasks, which is understood to be the class-tasks of pattern recognition for the s, r, and o terms. See the provided figure which shows the RCE step being used for training iterations.)) PNG media_image1.png 575 729 media_image1.png Greyscale wherein the step of constructing facts of predicted timestamps using an attention mechanism and assigning initial first-layer attention to the facts that are historically repeated comprises: (Wu teaches construction facts of predicted timestamps using an attention mechanism and a first-layer based on historical repeated facts, see the early claim limitations.) computing a multi-headed attention from a query matrix Q (Wu teaches computing a multi-headed attention task using parameters (s,r,o) using the query matrix Q (which is D_test as taught by Wu): “For each quadruple (𝑠, 𝑟, 𝑜, 𝑡) ∈ 𝐷 𝑡 𝑡𝑒𝑠𝑡, we evaluate an object query (𝑠, 𝑟, ?, 𝑡) and a subject query (?, 𝑟, 𝑜, 𝑡)” [Page 3].) performing layer normalization and residual connection on outputs from both the multi- headed attention (Wu teaches: “Inspired by previous work [29], we propose temporal regularization on the parameter space to alleviate catastrophic forgetting. We impose an 𝐿2 regularization constraint (reads on layer normalization) in the context of TKGC to smooth drastic change in the current representations compared to the previous task’s parameters.” [Page 6] which teaches layer normalization on the output matrix of the neural network. Residual connection is understood under BRI as a connection of the output of one earlier layer to the input of another future layer. As described by Figure 1, the Replay Facts layer output is connected to both the current model and the previous model, therefore it is understood to be a residual connection and uses the outputs from the multi-headed attention (replay facts are defined by the s, r, o elements).) PNG media_image2.png 579 989 media_image2.png Greyscale wherein the step of building second-layer attention based on statistics of historical frequency information that evolves with the timestamps, and adjusting a score of the initial first- layer attention according to updates in knowledge comprises: (Wu teaches Figure 1 which shows that the parameters as described on page 8 are updated and passed on from model to model, therefore the attention is adjusted according to updates in knowledge (this is also supported in section 5.2, when the model parameters are updated based on the current past data see Page 4.) Furthermore, the second-layer attention weights are based on the new facts (historical information) that are trained in part by data not provided before: “We propose a novel training strategy that uses only the added facts at each time step” [Page 7 section 5.5].) superimposing frequency information statistics contained in new historical information; (Wu teaches Figure 2, which shows a representation of the superimposed added new historical information statistics added in each timestep of the model. The information is also considered frequency information, as the information as a whole describes the frequency of new facts.) PNG media_image3.png 349 646 media_image3.png Greyscale representing the updates in knowledge according to updated statistics of the historical frequency information, so as to adjust the initial first-layer attention; (Wu teaches Figure 2 (provided above), which shows how the facts at time step t update which then updates the training data of the attention weights as described earlier.) based on the updated statistics of the historical frequency information, assigning an attention punishment to any fact that has never appeared historically; (Wu teaches section 5.5 which teach the updated statistics of the historical frequency information being used based on added (information that has not appeared historically) to generate a loss function (attention punishment) which is used to update the model’s attention parameters. (See Algorithm 2 where the training of the model is done with the loss functions.)) and based on the updated statistics of the historical frequency information, assigning an attention reward to each of facts that have appeared historically. (Wu teaches section 5.4 which teach the updated statistics of the historical frequency information being used based on deleted (information that appeared historically and that are now removed) to generate a loss function (attention reward) which update the model’s attention parameters. (See Algorithm 2 where the training of the model and update is done with loss functions.)) Wu does not distinctly disclose: performing mask filling on sequences of historically repeated facts corresponding to each of queries in a same batch; to a key matrix K after said mask filling supplementing deep semantic information using a fully connected feedforward neural network that comprises plural hidden units; However, Yang teaches: performing mask filling on sequences of historically repeated facts corresponding to each of queries in a same batch; “In order to simulate irregular traffic flow distribution in reality, we generate a random mask for the traffic flow at each moment. In the mask, 1 indicates the data that exists, and 0 indicates the missing data. Next, we mark Mask as M and use it as the input of the forward attention map. M is gradually updated and applied to the corresponding encoder layer to simulate the traffic data changing over location and time significantly and non-linearly.“ [Page 3, Section 3.1.2] which teach performing mask-filling on sequences of Spatio-temporal data from the same traffic data (reads on same batch). (“Historically repeated facts” are understood to just be represented to the data of the system, in which temporal data are understood to be historical facts that have a frequency.) The matrix generated from the mask-filling would therefore also transform an original matrix with the new updated set of data. (Yang teaches: “spatiotemporal characteristics of traffic data and external influence factors to fill the missing parts, and because of the introduction of the spatio-temporal learnable bidirectional attention map, ST-LBAGAN only focuses on mask filling” [Page 3] where the original query matrix is converted to a filled (key) matrix after mask filling.) Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to modify the replay facts layer to include a mask filling step as taught by Yang for the improvement of filling missing parts of data. This improvement is taught by Yang: “The ST-LBAGAN architecture makes reasonable use of the spatiotemporal characteristics of traffic data and external influence factors to fill the missing parts, and because of the introduction of the spatio-temporal learnable bidirectional attention map, ST-LBAGAN only focuses on mask filling.” [Page 3] Wu as modified by Yang does not distinctly disclose the full connected feedforward neural network. However, Yang further teaches: supplementing deep semantic information using a fully connected feedforward neural network that comprises plural hidden units; Yang teaches Figure 4 which teach the top order of data from input to output, which formulate a straight line from Input -> Fin -> Fout -> Output and therefore represents a fully connected feedforward neural network with multiple hidden (F) layers. PNG media_image4.png 507 996 media_image4.png Greyscale Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to modify the architecture of the neural network as taught by Wu with a feed-forward network as taught by Yang for the improvement of being able to use the feedforward mapping as an encoder. Yang teaches: “As shown in Fig. 4, we divide the spatio-temporal learnable bidirectional attention map into the learnable forward attention and the learnable reverse attention and map them to the encoder features and the decoder features, respectively” [Page 4] Regarding Claim 22, Wu as modified by Yang teaches all the limitations of Claim 21, and Wu further teaches: The method of claim 21, wherein the step of recombining a temporal knowledge graph in a temporal serialization manner according to an order of timestamps in the temporal knowledge graph, and storing distribution of historical timestamp subgraphs into a sparse matrix comprises: partitioning the temporal knowledge graph into a series of knowledge subgraph sequences in a chronological order, so as to dimensionally reduce representation of the temporal knowledge graph from quadruples to triples; (Wu teaches: “We extend the notion of pattern frequency introduced in [34] to gauge the sampling probability of each quadruple in the time window and develop two approaches dubbed frequency-based sampling and inverse frequency-based sampling. A pattern of the triple (𝑠, 𝑟, 𝑜) refers to a regular expression with some of the elements replaced with the wildcard symbol ‘∗’.” [Page 4, Section 5.2.1] teaches generating triples from the quadruples in the time window which are understood to be the knowledge subgraph sequences as shown in Claim 21 above. The data is understood to be based on the time-windows, therefore still maintain the chronological order of the original data, now just based on the sub-graph time-window. Finally, it is understood that this reduce the representation of the data from 4D to 3D (goes from a quadruple set of values to a triple).) and according to records of the sparse matrix, predicting historical patterns of a to-be-predicted event in similar scenarios over time, and converting time consumption for historical queries into quantified space consumption. (Wu teaches: “The DF is analogous to the false positive rate in the classification setting, measuring the rank of the deleted triples’ current time step as their time attributes. A lower DF value suggests that a model has a better capability to exclude deleted facts from the top 10 results. The RRD is defined as the pairwise difference of reciprocal ranks between each positive quadruple in the test set and each deleted fact in the previous data.” [Page 4] which teaches predicting the capacity of the model to exclude facts from the top 10 results (not including the data in a query.). This then reads on prediction historical patterns of the event across queries (reads on similar scenarios of searching) such that it won’t be suggested often. Time consumption for the historical queries does not have much support in the spec other being recited [0021], therefore under BRI it is understood to be the related to the times that it shows in queries. The RRD scores for which the sparse matrix converts is used to analyze the deleted queries, therefore it is understood that it is being used to demonstrate time consumption for the historical queries and how it relates to lower space consumption of useless queries.) Regarding Claim 24, Wu as modified by Yang teaches all the limitations of Claim 21, and Yang further teaches: The method of claim 21, wherein the step of performing mask filling on sequences of historically repeated facts corresponding to each of the queries in the same batch comprises: using a relation-entity pair that has never appeared globally to fill any of the sequences that are shorter than a longest sequence in the batch, (Yang teaches: “In order to simulate irregular traffic flow distribution in reality, we generate a random mask (reads on has never appeared globally) for the traffic flow at each moment. In the mask, 1 indicates the data that exists, and 0 indicates the missing data” [Page 3] which teach generating a relation-entity pair of a value of 0 (a pair of data (entity) and the mask with value of 0) for the relation and the generated entity. Yang also teaches it being used to fill the missing parts of the data as well as: “ST-LBAGAN’s end-to-end training mechanism also makes it possible to improve the filling quality of traffic data.” [Page 3] which teaches it being used to improve to fill the data (reads on filling sequences that are shorter) as well as support on Page 2: “The generator of GAN focuses on filling the missing parts of the data, and discriminator feeds back the generator’s output. We construct the missing part of a matrix as a mask, and input the incomplete matrix and mask into the generator currently. [..] we superimpose the updated mask with the incomplete matrix to get the complete matrix” which teaches it filling up the matrix sequences with the masked data.) so as to generate a mask matrix using identification marks, and exclude these sequences from an attention operation. (Yang teaches the identification mark of 0 for the missing data and mark of 1 for data that exists on page 3 and Yang further teaches: “In the training phase, the model only focuses on the completion of the mask part, which allows us to greatly improve the accuracy of the task of filling missing data.” [Page 3] which teaches that the model focusses on training based on the masked entities, so sequences that are not masked are excluded from the attention operation.) The motivation for this combination is the same as the one found in Claim 21. Regarding Claim 26, Wu teaches: A system for temporal knowledge graph reasoning based on distributed attention, the method comprising: (Wu teaches: “We introduce a new task, incremental TKGC, and propose TIE, a training and evaluation framework that integrates incremental learning with TKGC. TIE combines TKG representation learning, experience replay, temporal regularization to improve model performance and alleviate catastrophic forgetting.” [Page 2] which teaches Temporal Knowledge Graph Completion (TKGC) through fine tuning a model with weights (weights read on attention).) a scheduling unit, configured to recombine a temporal knowledge graph in a temporal serialization manner according to an order of timestamps in the temporal knowledge graph, and store distribution of historical timestamp subgraphs into a sparse matrix; (Wu teaches: “We define a time window ranging from 𝑡 −𝜏𝑑 to 𝑡−1 to limit the scope of evaluation. For every quadruple (𝑠, 𝑟, 𝑜, 𝑡), we aim to find and then evaluate the related deleted facts from this time window. “[Page 4] which teaches recombining a first set of knowledge graph data (s, r, o, t) into a time window (reads on temporal serialization manner as explained in the spec, see the quotation below, as the time window is a subgraph of the original graph) and defined on the order of timestamps. Wu further teaches storing the data into a collection set O,s,r,t which is a representation of data points of a set (reads on matrix) which only includes negative numbers: “where 𝑂 ′ 𝑠,𝑟,𝑡 is the collection of negative objects and 𝑍𝑡 is the normalizing constant“ [Page 4], therefore it can be considered a sparse matrix as the matrix only includes nonzero elements. Temporal serialization as defined by the applicant’s specification: “A temporal knowledge graph, which contains an entity set ϵ having a size of N, a relation set having a size of P, and a timestamp set having a size of T, is partitioned into a sequence of temporal subgraphs.” [0053] as the data of Wu is converted into subsets (which are representations of subgraphs) it is understood that they are portioned in a temporal serialization manner. Finally, as regarded about in the Claim Interpretation, “unit” as referred to in the claim has been understood under BRI to just be representation of a step that performs the operation as there is no clear distinction in the limitation or the specification of the structure of these units. Therefore, the recitation of the unit is understood to be found in the prior art if the step performs the operations of the unit.) a processing unit, configured to construct facts of predicted timestamps using an attention mechanism and assign initial first-layer attention to the facts that are historically repeated; (Wu teaches: “The parameter set is 𝜽 𝑡−1 = {𝑬 𝑡−1 , 𝑹 𝑡−1 , 𝑾𝑡−1 , 𝑩 𝑡−1 }, where 𝑾 and 𝑩 are the matrices of weight and bias parameters in Equation (1)” [Page 6] which teaches the first layer of the neural network containing weights W for a predicted timestamp (t-1) for the mechanism of being updated by training the model (attention mechanism) which are initialized before training (therefore initial by “Glorot Initialization [10] if an entity appears for the first time” [Page 6]) and are assigned to the weight matrix that are tied to the weights of the facts that are historically repeated.) an adjusting unit, configured to build second-layer attention based on statistics of historical frequency information that evolves with the timestamps, and adjust a score of the initial first-layer attention according to updates in knowledge; (Wu teaches calculating a second layer of weights (attention) based on added facts on Page 7 Section 5.6 (Optimization). The weights are created to control the “the relative emphasis placed on each loss term” and one score, LCE as shown above is based on a score from added facts which are represented as facts (still temporal data, therefore historical frequency information as the data stores the frequency of the added facts) that change (evolve) according to each timestep. Furthermore, the first-layer attention weights are updated based on the set of input data by training the model, therefore are “adjusted” according to updated in knowledge (see Page 5).) and a training unit, configured to train a model with multi-class tasks based on cross entropy loss, according to a parameter training strategy; (Wu teaches cross-entropy being used to train the model with multi-class tasks (class tasks is understood to be multiple calculation tasks, which is understood to be the class-tasks of pattern recognition for the s, r, and o terms. See the provided figure which shows the RCE step being used for training iterations.)) PNG media_image1.png 575 729 media_image1.png Greyscale wherein the step of constructing facts of predicted timestamps using an attention mechanism and assigning initial first-layer attention to the facts that are historically repeated comprises: (Wu teaches construction facts of predicted timestamps using an attention mechanism and a first-layer based on historical repeated facts, see the early claim limitations.) computing a multi-headed attention from a query matrix Q (Wu teaches computing a multi-headed attention task using parameters (s,r,o) using the query matrix Q (which is D_test as taught by Wu): “For each quadruple (𝑠, 𝑟, 𝑜, 𝑡) ∈ 𝐷 𝑡 𝑡𝑒𝑠𝑡, we evaluate an object query (𝑠, 𝑟, ?, 𝑡) and a subject query (?, 𝑟, 𝑜, 𝑡)” [Page 3].) performing layer normalization and residual connection on outputs from both the multi- headed attention (Wu teaches: “Inspired by previous work [29], we propose temporal regularization on the parameter space to alleviate catastrophic forgetting. We impose an 𝐿2 regularization constraint (reads on layer normalization) in the context of TKGC to smooth drastic change in the current representations compared to the previous task’s parameters.” [Page 6] which teaches layer normalization on the output matrix of the neural network. Residual connection is understood under BRI as a connection of the output of one earlier layer to the input of another future layer. As described by Figure 1, the Replay Facts layer output is connected to both the current model and the previous model, therefore it is understood to be a residual connection and uses the outputs from the multi-headed attention (replay facts are defined by the s, r, o elements).) PNG media_image2.png 579 989 media_image2.png Greyscale wherein the step of building second-layer attention based on statistics of historical frequency information that evolves with the timestamps, and adjusting a score of the initial first- layer attention according to updates in knowledge comprises: (Wu teaches Figure 1 which shows that the parameters as described on page 8 are updated and passed on from model to model, therefore the attention is adjusted according to updates in knowledge (this is also supported in section 5.2, when the model parameters are updated based on the current past data see Page 4.) Furthermore, the second-layer attention weights are based on the new facts (historical information) that are trained in part by data not provided before: “We propose a novel training strategy that uses only the added facts at each time step” [Page 7 section 5.5].) superimposing frequency information statistics contained in new historical information; (Wu teaches Figure 2, which shows a representation of the superimposed added new historical information statistics added in each timestep of the model. The information is also considered frequency information, as the information as a whole describes the frequency of new facts.) PNG media_image3.png 349 646 media_image3.png Greyscale representing the updates in knowledge according to updated statistics in the historical frequency information, so as to adjust the initial first-layer attention; (Wu teaches Figure 2 (provided above), which shows how the facts at time step t update which then updates the training data of the attention weights as described earlier.) based on the updated statistics in the historical frequency information, assigning an attention punishment to any fact that has never appeared historically; (Wu teaches section 5.5 which teach the updated statistics of the historical frequency information being used based on added (information that has not appeared historically) to generate a loss function (attention punishment) which is used to update the model’s attention parameters. (See Algorithm 2 where the training of the model is done with the loss functions.)) and based on the updated statistics in the historical frequency information, assigning an attention reward to each of facts that have appeared historically. (Wu teaches section 5.4 which teach the updated statistics of the historical frequency information being used based on deleted (information that appeared historically and that are now removed) to generate a loss function (attention reward) which update the model’s attention parameters. (See Algorithm 2 where the training of the model and update is done with loss functions.)) Wu does not distinctly disclose: performing mask filling on sequences of historically repeated facts corresponding to each of queries in a same batch; to a key matrix K after said mask filling supplementing deep semantic information using a fully connected feedforward neural network that comprises plural hidden units; However, Yang teaches: performing mask filling on sequences of historically repeated facts corresponding to each of queries in a same batch; “In order to simulate irregular traffic flow distribution in reality, we generate a random mask for the traffic flow at each moment. In the mask, 1 indicates the data that exists, and 0 indicates the missing data. Next, we mark Mask as M and use it as the input of the forward attention map. M is gradually updated and applied to the corresponding encoder layer to simulate the traffic data changing over location and time significantly and non-linearly.“ [Page 3, Section 3.1.2] which teach performing mask-filling on sequences of Spatio-temporal data from the same traffic data (reads on same batch). (“Historically repeated facts” are understood to just be represented to the data of the system, in which temporal data are understood to be historical facts that have a frequency.) The matrix generated from the mask-filling would therefore also transform an original matrix with the new updated set of data. (Yang teaches: “spatiotemporal characteristics of traffic data and external influence factors to fill the missing parts, and because of the introduction of the spatio-temporal learnable bidirectional attention map, ST-LBAGAN only focuses on mask filling” [Page 3] where the original query matrix is converted to a filled (key) matrix after mask filling.) Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to modify the replay facts layer to include a mask filling step as taught by Yang for the improvement of filling missing parts of data. This improvement is taught by Yang: “The ST-LBAGAN architecture makes reasonable use of the spatiotemporal characteristics of traffic data and external influence factors to fill the missing parts, and because of the introduction of the spatio-temporal learnable bidirectional attention map, ST-LBAGAN only focuses on mask filling.” [Page 3] Wu as modified by Yang does not distinctly disclose the full connected feedforward neural network. However, Yang further teaches: supplementing deep semantic information using a fully connected feedforward neural network that comprises plural hidden units; Yang teaches Figure 4 which teach the top order of data from input to output, which formulate a straight line from Input -> Fin -> Fout -> Output and therefore represents a fully connected feedforward neural network with multiple hidden (F) layers. PNG media_image4.png 507 996 media_image4.png Greyscale Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to modify the architecture of the neural network as taught by Wu with a feed-forward network as taught by Yang for the improvement of being able to use the feedforward mapping as an encoder. Yang teaches: “As shown in Fig. 4, we divide the spatio-temporal learnable bidirectional attention map into the learnable forward attention and the learnable reverse attention and map them to the encoder features and the decoder features, respectively” [Page 4] Claim(s) 23, 27– 28 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu TIE: A Framework for Embedding-based Incremental Temporal Knowledge Graph Completion and further in view of Yang ST-LBAGAN: Spatio-temporal learnable bidirectional attention generative adversarial networks for missing traffic data imputation and finally in view of Ji (U.S. Publication US20200293874A1) Regarding Claim 23, Wu as modified by Yang teaches all the limitations of Claim 22, and Wu further teaches: The method of claim 22, wherein the step of according to a parameter training strategy, training a model with multi-class tasks based on cross entropy loss comprises: initializing at least one or more learnable parameters of a query transformation matrix, a key transformation matrix, and a linear transformation coefficient offset; (Wu teaches an entity embedding matrix (E^t) on Page 6, which represents a key transformation matrix (as it is a matrix that transforms the entities/keys through embedding and matrix transformations). Wu also teaches matrix B which is the bias parameter matrix (linear transformation coefficient offset) and matrix W which is a query transformation matrix (contains the weights for the model that are multiplied to generate the query).) treating reasoning-based completion of the temporal knowledge graph as the multi-class tasks each having a number of classes equal to a size of an entity set of the multi-class task; (Wu teaches Frequency-based Sampling which teach reasoning based completion (see the equation of P, where there are some values that are replaced with the wildcard symbol on Page 4) and then scores are calculated and reasoned for the system with the pattern p. s,r,o are all values that are taken from the first quadruple set that formed the knowledge graph as shown in Claim 21. The multi-class task, as explained in claim 21, is the pattern task using the multi-classes of the s, r, and o terms which each have a cardinality (number of terms) equal to a size of the entity set defined in the time-window (therefore reads on: each having a number of classes equal to a size of an entity set of the multi-class task). Algorithm 2 on page 6, teaches Step 3 which forms the replay buffer B^t which is the set of quadruples in the specified time-set, and as taught by Wu: “We extract 𝑃𝑡, a set of replay samples with positive labels, from 𝐵𝑡” [Page 4] has then the same size of the entity set that generates the multi-class task.) using a cross entropy loss function (Wu teaches using a cross-entropy loss function: “The replay knowledge distillation (RKD) loss and replay cross-entropy (RCE)at each iteration are defined as follows:” [Page 6] which teaches the Loss replay cross-entropy function, LRCE which forms the final loss function. This loss function is used to learn parameters of the model as taught by Algorithm 2 [Page 6], for training the model parameters which include the multi-class parameters (s, r, o) and therefore updates the multi-class task that uses these parameters. Furthermore, Wu teaches ranking them which determines the highest score: “Regarding the object query, we calculate the scores for all known entities, i.e., 𝜙 (𝑠, 𝑟, 𝑜′ , 𝑡), ∀𝑜 ′ ∈ 𝐸 𝑡 . The ranks are obtained by sorting the scores in descending order.” [Page 3] which would be ranked with the updated s, r, o scores from the cross-entropy loss function. “regard said fact as a result of future event prediction” does not have clear meaning, therefore it is considered under BRI as referring to regarding the fact as important in future event prediction. As Wu teaches: “For a quadruple (𝑠, 𝑟, 𝑜, 𝑡) ∈ 𝐷 𝑡 𝑡𝑒𝑠𝑡 and its related object query (𝑠, 𝑟, ?, 𝑡), the goal of TKGC is ranking 𝑜 as high as possible.” [Page 3] therefore Wu teaches trying to predict said highest score for each query regarding that object in future object (event) prediction.) However, Wu as modified by Yang does not distinctly disclose: an AMSGrad optimizer Wu teaches using an optimizer, in this case: A-Gem [Page 12], however does not teach using an AMSGrad optimizer. However, Ji teaches using an AMSGrad Optimizer: “The AMSGrad variant of Adam initialized with a learning rate of 0.0001 can be employed to optimize the model weights.”[0078] which teaches a model that can be optimized with the AMSGrad optimizer. Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to replace the A-Gem optimizer with the AMSGrad optimizer as it is a known variant optimizer which has the known result of optimizing parameters. Regarding Claim 27, Wu as modified by Yang and as modified by Ji teaches all the limitations of Claim 21, but Wu does not disclose: An electronic device, characterized in that it comprises: one or more processors; a memory, for storing one or more computer programs; wherein the one or more computer programs, when executed by the one or more processors, cause the one or more processors to perform the method for temporal knowledge graph reasoning based on distributed attention of claim 21. However, Ji discloses: “The example of the machine 1000 includes at least one processor 1002 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), advanced processing unit (APU), or combinations thereof), one or more memories such as a main memory 1004, a static memory 1006, or other types of memory, which communicate with each other via link 1008. “ [0123]. Which teach using a processor and a memory as the machine architecture. Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to run the system of Wu as modified by Yang and further modified by Ji with the processor and memory as taught by Ji for the known result of being a method to perform the system on a computer as the processor and memory allows for the system to perform the operations described above remotely on a device that will be able to store and process data. Regarding Claim 28, Wu as modified by Yang and as modified by Ji teaches all the limitations of Claim 21, but Wu does not disclose: A storage medium comprising computer-executable instructions, characterized in that the computer-executable instructions are used, when executed by a computer processor, to perform the method for temporal knowledge graph reasoning based on distributed attention of claim 21. However, Ji discloses: “As used herein, the terms “machine-storage medium,” “device-storage medium,” “computer-storage medium” mean the same thing and may be used interchangeably in this disclosure. The terms refer to a single or multiple storage devices and/or media (e.g., a centralized or distributed database, and/or associated caches and servers) that store executable instructions and/or data. “[0126] Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to run the system of Wu as modified by Yang and further modified by Ji storage medium as taught by Ji for the known result of being a method to perform the system on a computer as this allows for the system to perform the operations described above remotely on a device that will be able to store data. Claim(s) 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu TIE: A Framework for Embedding-based Incremental Temporal Knowledge Graph Completion and further in view of Yang ST-LBAGAN: Spatio-temporal learnable bidirectional attention generative adversarial networks for missing traffic data imputation and finally in view of Galassi Attention in Natural Language Processing Regarding Claim 25, Wu as modified by Yang teaches all the limitations of Claim 24, and Wu as modified by Yang further teaches: The method of claim 24, wherein the step of computing a multi-headed attention from (Wu teaches: “The parameter set is 𝜽 𝑡−1 = {𝑬 𝑡−1 , 𝑹 𝑡−1 , 𝑾𝑡−1 , 𝑩 𝑡−1 }, where 𝑾 and 𝑩 are the matrices of weight and bias parameters in Equation (1)” [Page 7] which teach the value matrix which stores the initial distrusted attention assigned to each historically repeated facts, which is the W^t-1 as taught by Wu which contains the initial weights.) However, Galassi teaches: performing a scaled dot-product attention operation, calculating a dot product using the query matrix Q and the key matrix K, and dividing the dot product by a scaling factor to obtain a weight matrix; (See Table 4 of Galassi, which teach the “Scaled multiplicative” operator which calculates a dot product and then divides it by a scaling factor which outputs a matrix. This is taught to be a known mathematical process for attention calculation: “Finally, a weighted sum of M1 according to the relevance of query elements is computed through the dot product between M1 and aQ, obtaining the document’s attention distribution over the keys, aK” [Page 4301] which teaches using the dot product calculation for an attention matrix (weight and attention are analogous as both refer to values that scale vectors in a neural network). Galassi teaches directly using dot product, however the scaled version is also taught to be a variation and outputs a weight matrix: “A variation of this model is scaled multiplicative attention, where a scaling factor is introduced to improve performance with large keys [36].” [Page 4297]) PNG media_image5.png 836 570 media_image5.png Greyscale Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to modify the step of mask filling as taught by Wu as modified by Yang with the added step of calculating dot products between matrices as taught by Galassi for the improved performance with large keys. This improvement is taught by Galassi: “A variation of this model is scaled multiplicative attention, where a scaling factor is introduced to improve performance with large keys [36].” [Page 7] Wu as modified by Yang and further modified by Galassi in the modification as mentioned above does not distinctly disclose the dot product using the weight matrix and a value matrix. However, Galassi further teaches in another step: and calculating a dot product using the weight matrix and a value matrix V so as to obtain a value matrix associated with a representation attention (Galassi teaches: “and the contextual information. V and a are thus combined to obtain a new set Z of weighted representations of V (reads on representation attention of the value matrix) [Page 4294] as well as Table 2 provided below which shows that a is a “attention weights vector” (which reads on a matrix with a dimensionality n x 1) and finally equation 8 on page 4295 which shows zi being equal to the multiplication (reads on dot product as dot product is just the multiplication of the entire matrix rather than the individual elements).) PNG media_image6.png 445 1105 media_image6.png Greyscale Before the effective filing date of the invention it would have been obvious to one of ordinary skill in the art to modify the step of mask filling as taught by Wu as modified by Yang with the added step of calculating dot products between matrices as taught by Galassi for the improvement of creating merged matrices that can be used for further calculations and give more meaning to a matrix. This improvement is taught by Galassi: “V and a are thus combined to obtain a new set Z of weighted representations of V [see (8)], which are then merged together so as to produce a compact representation of Z” [Page 4294] which teaches combining the matrices and then using them for a further calculation, in which the new metric has a meaning of being “weighted representations of V”. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jonathan M Bakhit whose telephone number is (571)272-0454. The examiner can normally be reached Monday - Thursday 8:00AM - 6:00PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123 J.M.B. Examiner Art Unit 2123
Read full office action

Prosecution Timeline

Oct 07, 2022
Application Filed
Aug 27, 2025
Non-Final Rejection mailed — §101, §103, §112
Nov 27, 2025
Response Filed
Sep 30, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743481
TWO-STAGE FREQUENCY SELECTION METHOD AND DEVICE FOR MICROWAVE FREQUENCY SWEEP DATA
3y 11m to grant Granted Sep 22, 2026
Patent 12717626
ELECTRONIC DEVICE AND ACCESS EVENT AUDIO-VISUALIZATION METHOD
3y 4m to grant Granted Aug 25, 2026
Patent 12632550
SMART INCENTIVIZATION FOR ACHIEVING COLLABORATIVE MACHINE LEARNING
3y 7m to grant Granted May 19, 2026
Patent 12572799
METHODS FOR RELIABLE OVER-THE-AIR COMPUTATION AND FEDERATED EDGE LEARNING
3y 10m to grant Granted Mar 10, 2026
Patent 12541568
Method, System, and Computer Program Product for Recurrent Neural Networks for Asynchronous Sequences
4y 3m to grant Granted Feb 03, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+56.5%)
4y 6m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 407 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month