Prosecution Insights
Last updated: October 02, 2026
Application No. 18/283,131

TRAINING GRAPH NEURAL NETWORKS USING A DE-NOISING OBJECTIVE

Final Rejection §103
Filed
Sep 20, 2023
Priority
May 28, 2021 — provisional 63/194,851 +1 more
Examiner
TRAN, TAN H
Art Unit
2141
Tech Center
2100 — Computer Architecture & Software
Assignee
DeepMind Technologies Limited
OA Round
2 (Final)
61%
Grant Probability
Moderate
3-4
OA Rounds
5m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
195 granted / 320 resolved
+5.9% vs TC avg
Strong +33% interview lift
Without
With
+32.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 6m
Avg Prosecution
46 currently pending
Career history
374
Total Applications
across all art units

Statute-Specific Performance

§101
13.4%
-26.6% vs TC avg
§103
59.8%
+19.8% vs TC avg
§102
16.5%
-23.5% vs TC avg
§112
6.3%
-33.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 320 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 2. This Office Action is sent in response to Applicant’s Communication received on 07/07/2026 for application number 18/283,131. Response to Amendments 3. The Amendment filed 07/07/2026 has been entered. Claims 1, 7, 8, 21, and 22 have been amended. Claim 23 has been added. Claims 1-5, 7-18, and 21-23 remain pending in the application. Response to Arguments Applicant argues that Fatemi does not disclose the amended features in claim 1. However, the argument is moot since this is a newly presented limitation, thus changing the scope of the claim. However, a newly found reference, You, is applied. Claim Rejections – 35 USC § 103 4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 5. Claims 1, 3, 4, 9-12, 18, 21, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Fatemi et al. (U.S. Patent Application Pub. No. US 20220101103 A1) in view of You et al. (When Does Self-Supervision Help Graph Convolutional Networks, arXiv, published 2020, pages 1-10). Claim 1: Fatemi teaches a method for training a neural network that includes one or more graph neural network layers (i.e. the first neural network comprises a graph neural network (GNN); para. [0021, 0023]), the method comprising: generating data defining a graph that comprises: (i) a set of nodes, (ii) a node embedding for each node (i.e. node embeddings; para. [0066]), and (iii) a set of edges that each connect a respective pair of nodes (i.e. A graph can be represented as a data structure consisting of two components: nodes (vertices) and edges. A graph is often represented by an adjacency matrix. Graph neural networks (GNNs) are a class of machine learning models designed to perform inference on data described by graphs. GNNs can often be directly applied to graphs, and provide an easy way to do node-level, edge-level, and graph-level prediction tasks; para. [0046]), comprising: obtaining a respective initial feature representation for each node (i.e. The plurality of node features 120 may include a feature for each node, and each feature may include one or more data elements (e.g., a vector) associated with the respective node; para. [0004, 0072]), original node features for each node; generating a respective final feature representation for each node, wherein, for each of one or more of the nodes, the respective final feature representation is a modified feature representation that is generated from the respective feature representation for the node using respective noise (i.e. noisy node features 145, which are generated by adding noise 143 to the node features 120. noise 143 can be added by setting the 1s in the selected mask to 0s, and L is the binary cross-entropy loss. For datasets where the input features are continuous numbers, idx consists of r percent of the indices of X selected uniformly at random in each epoch, noise 143 can be added by either replacing the values at idx with 0 or by adding independent Gaussian noises to each of the node features 120; para. [0055, 0071, 0101]), noisy node features are generated from node features by masking or adding noise, including Gaussian noise; and generating the data defining the graph using the respective final feature representations of the nodes (i.e. generating an adjacency matrix based on a plurality of node features; generating a plurality of noisy node features based on the plurality of node features; generating a plurality of denoised node features using the neural network based on the plurality of noisy node features and the adjacency matrix; and updating the adjacency matrix based on the plurality of denoised node features; para. [0071, 0141]); processing the data defining the graph using one or more of the graph neural network layers of the neural network (i.e. a denoising auto-encoder 148 can be implemented using a GNN; para. [0065, 0099]) to generate a respective updated node embedding of each node (i.e. updated node embeddings; para. [0066]), GNN/GCN layers producing updated node embeddings and the denoising autoencoder itself can be implemented using GNN; processing, for each of one or more of the nodes having modified feature representations, the updated node embedding of the node to generate a respective de-noising prediction for the node that characterizes a de-noised feature representation for the node that does not include the noise used to generate the modified feature representation of the node (i.e. a self-supervision approach disclosed herein masks some input features (or adds noise to them) and trains a separate GNN aiming at updating the adjacency matrix in such a way that it can recover the masked (or noisy) features. Denoising auto-encoder (GNNDAE) 148 can be trained such that it receives a noisy version {tilde over (X)} 145 of the node features X 120 as input and produces the denoised features X 149 as output; para. [0055, 0071, 0100]), selected node features are noised, a GNN denoising autoencoder processes them with adjacency information, and outputs denoised node features that recover the masked/noisy features; processing to generate a task prediction (i.e. Classifer 146, illustrated as “GNNC”, receives the node features 120 and normalized adjacency matrix A to output node classes 147, for example, a prediction of one or more classes for each node; para. [0071, 0085, 0086]) that is different from the respective de-noising predictions (i.e. Denoising auto-encoder 148, illustrated as “GNNDAE” receives noisy node features 145, which are generated by adding noise 143 to the node features 120, as well as the normalized adjacency matrix A, to output denoised (or de-noised) node features 149; para. [0071]), wherein the task prediction characterizes one or more elements represented by the graph (i.e. Classifier 146 thus takes as input node features 120 and the adjacency matrix A normalized by adjacency processor 144 and computes one or more predictions regarding which class 147 each node belongs to; para. [0086]); and determining an update to current values of neural network parameters of the neural network to optimize an objective function that measures errors in the de-noising predictions for the nodes (i.e. system to update one or more parameters of the first neural network GNNDAE by minimizing the loss function. The loss function is determined based on a binary cross-entropy loss or a mean-squared error loss. The denoising autoencoder 148 implemented using a GNNDAE model can be trained by minimizing loss; para. [0012, 0013, 0103, 0105]), a denoising loss, the mathematical loss expression comparing original and denoised features at noised indices, and parameter updates by minimizing the loss, wherein the objective function comprises (i) a de-noising loss term that measures the errors in the de-noising predictions for the nodes (i.e. PNG media_image1.png 42 29 media_image1.png Greyscale DAE is the denoising autoencoder loss (see Equation (3)); para. [0100-0103]) and (ii) a task loss term that measures an error in the task prediction (i.e. The training loss PNG media_image1.png 42 29 media_image1.png Greyscale C for the classification task can be computed by taking the softmax of the logits to produce a probability distribution for each node and then computing the cross-entropy loss; para. [0085]), and wherein optimizing the objective function comprises jointly training (i.e. concurrent training of both classifier 146 and denoising autoencoder 148; para. [0104]) both the de-noising loss term and the task loss term (i.e. structure learning model 140 is trained to minimize PNG media_image1.png 42 29 media_image1.png Greyscale = PNG media_image1.png 42 29 media_image1.png Greyscale C+λ PNG media_image1.png 42 29 media_image1.png Greyscale DAE where PNG media_image1.png 42 29 media_image1.png Greyscale C is the classification loss, PNG media_image1.png 42 29 media_image1.png Greyscale DAE is the denoising autoencoder loss (see Equation (3)), and is a hyperparameter controlling the relative importance of the two losses; para. [0103]). Fatemi does not explicitly teach processing the updated node embeddings of the nodes generated by the one or more graph neural network layers to generate a task prediction; and wherein optimizing the objective function comprises jointly training the one or more graph neural network layers. However, You teaches processing the updated node embeddings of the nodes generated by the one or more graph neural network layers to generate a task prediction (i.e. Thus GCN is decomposed into feature extraction and linear transformation as Z = fθ(X, ˆ A)Θ where parameters θ and Θ =W1 are learned from data; Section 3.1, pages 2-3) that is different from the respective de- noising predictions (i.e. the network is trained with the self-supervised task … where α1,α2 ∈ R>0 are the weights for the overall supervised loss Lsup(θ,Θ) as defined in (2) and those for the self-supervised loss Lss(θ,Θss) as defined in (3), respectively. To optimize the weighted sum of their losses, the target supervised and self-supervised tasks share the same feature extractor fθ(·,·) but have their individual linear transformation parameters Θ∗ and Θ∗ ss as in Figure 1; Section 3.2, page 3), wherein the task prediction characterizes one or more elements represented by the graph (i.e. the true label for labeled nodes; Section 3.1, page 3); and wherein the objective function comprises (i) a de-noising loss term that measures the errors in the de-noising predictions for the nodes and (ii) a task loss term that measures an error in the task prediction, and wherein optimizing the objective function comprises jointly training the one or more graph neural network layers using both the de-noising loss term and the task loss term (i.e. equation 4, where α1,α2 ∈ R>0 are the weights for the overall supervised loss Lsup(θ,Θ) as defined in (2) and those for the self-supervised loss Lss(θ,Θss) as defined in (3), respectively. To optimize the weighted sum of their losses, the target supervised and self-supervised tasks share the same feature extractor fθ(·,·) but have their individual linear transformation parameters Θ∗ and Θ∗ ss as in Figure 1; Section 3.2, page 3). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Fatemi to include the feature of You. One would have been motivated to make this modification because using the shared GCN feature extraction architecture of You so that the same GNN representations are trained by both the task and denoising objectives, thereby improving the representations used for the downstream task. Claim 3: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein for each of one or more of the nodes having modified feature representations, the respective de-noising prediction for the node predicts the respective initial feature representation of the node (i.e. “GNNDAE” receives noisy node features 145, which are generated by adding noise 143 to the node features 120, as well as the normalized adjacency matrix A, to output denoised (or de-noised) node features 149; para. [0071, 0100, 0131]). Claim 4: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein for each of one or more of the nodes having modified feature representations, the respective de-noising prediction for the node characterizes a target feature representation of the node (i.e. generating an adjacency matrix based on a plurality of node features; generating a plurality of noisy node features based on the plurality of node features; generating a plurality of denoised node features using a neural network based on the plurality of noisy node features and the adjacency matrix; and updating the adjacency matrix based on the plurality of denoised node features; para. [0018, 0098-0100]). Claim 9: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein the objective function measures, for each of a plurality of graph neural network layers of the neural network, respective errors in de-noising predictions for the nodes that are based on updated node embeddings generated by the graph neural network layer (i.e. Graph convolutional networks (GCNs) are a powerful variant of GNNs [24]. For a graph PNG media_image2.png 38 25 media_image2.png Greyscale ={ PNG media_image3.png 42 29 media_image3.png Greyscale , A, X} with degree matrix D, one layer (e.g., layer l) of the GCN architecture can be defined as follows: H (l)=σ(ÂH (l-1) W (l)) where  represents a normalized adjacency matrix, H(l-1) ∈ PNG media_image4.png 42 29 media_image4.png Greyscale nxd l1 represents the node representations in layer l−1)(H(0)=X), W(l) ∈ PNG media_image4.png 42 29 media_image4.png Greyscale d l-1 ×d l is a weight matrix, σ is an activation function such as ReLU [30], and H(l) ∈ PNG media_image4.png 42 29 media_image4.png Greyscale n×d l is the updated node embeddings; para. [0065, 0066, 0108]). Claim 10: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein for each of one or more of the nodes having modified feature representations, processing the updated node embedding of the node to generate the respective de-noising prediction for the node comprises: processing the updated node embedding of the node (i.e. the updated node embeddings; para. [0066, 0099, 0100]) using one or more neural network layers (i.e. Two-layer GCNs for both GNNC and GNNDAE are used as well as for baselines and two-layer MLPs; para. [0108]) to generate the respective de-noising prediction for the node (i.e. Denoising auto-encoder 148, illustrated as “GNNDAE” receives noisy node features 145, which are generated by adding noise 143 to the node features 120, as well as the normalized adjacency matrix A, to output denoised (or de-noised) node features 149; para. [0071, 0099]). Claim 11: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein determining the update to the current values of the neural network parameters of the neural network to optimize the objective function (i.e. the method may include updating one or more parameters of the first and second neural networks by minimizing a combined loss determined based on PNG media_image1.png 42 29 media_image1.png Greyscale C and PNG media_image1.png 42 29 media_image1.png Greyscale DAE, wherein PNG media_image1.png 42 29 media_image1.png Greyscale C represents a loss function of the second neural network, the combined loss is determined based on a combined loss function PNG media_image1.png 42 29 media_image1.png Greyscale = PNG media_image1.png 42 29 media_image1.png Greyscale C+λ PNG media_image1.png 42 29 media_image1.png Greyscale DAE; para. [0138, 0139]) comprises: backpropagating gradients of the objective function (i.e. the strucut learning model 140 can be implemented in PyTorch [13], and deep graph library (DGL) [16] can be used for the sparse operations, and Adam [8] can be used as the optimizer; para. [0107]) through neural network parameters of the graph neural network layers (i.e. the first neural network comprises a graph neural network (GNN); para. [0021, 0023, 0065]). Claim 12: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein for each of one or more of the nodes, the respective final feature representation for the node is generated by adding the respective noise to the respective feature representation for the node (i.e. Denoising auto-encoder 148, illustrated as “GNNDAE” receives noisy node features 145, which are generated by adding noise 143 to the node features 120; para. [0053, 0071]). Claim 18: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein each graph neural network layer of the graph neural network is configured to: receive a current graph (i.e. In the implementation of GNNs, an input is a set of node features and a graph structure (for example, modeled as adjacency matrix); para. [0048]); and update the current graph in accordance with current neural network parameter values of the graph neural network layer, comprising: updating a current node embedding of each of one or more nodes in the graph based on: (i) the current node embedding of the node, and (ii) a respective current node embedding of each of one or more neighbors of the node in the graph (i.e. Graph neural networks (GNNs) take as input a set of node features and an adjacency matrix corresponding to the graph structure and output an embedding for each node that captures not only the initial features of the node but also the features of its neighbours. The need for both node features and graph structure limits the applicability of GNNs in several domains. For example, one may have access to a set of node (or object) features and hypothesize that there exists some relation between the nodes, but not have access to the graph structure specifying which pairs of nodes are connected. The updated node embeddings; para. [0004, 0065, 0066, 0085]). Claims 21-22 are similar in scope to Claim 1 and are rejected under a similar rationale. 6. Claims 2, 5, and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Fatemi in view of You, and further in view of Lee et al. (U.S. Patent Application Pub. No. US 20240095499 A1). Claim 2: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein for each of one or more of the nodes having modified feature representations, the respective de-noising prediction for the node the noise used to generate the modified feature representation of the node (i.e. the loss function PNG media_image1.png 42 29 media_image1.png Greyscale DAE is represented by the function PNG media_image1.png 42 29 media_image1.png Greyscale DAE =L(X idx ,GNN DAE({tilde over (X)},A;θ GNN DAE )idx), where A represents the generated adjacency matrix, θGNN DAE represents parameters of the first neural network GNNDAE, X represents the plurality of node features, {tilde over (X)} represents the plurality of noisy node features, idx represent indices corresponding to the elements of X to which noise has been added, and Xidx represent corresponding values of elements at idx; para. [0124-0132]). Fatemi does not explicitly teach predicts the noise used to generate the modified feature representation. However, Lee teaches the respective de-noising prediction for the node predicts the noise used to generate the modified feature representation of the node (i.e. The method may include using a neural network to learn noise N in noisy input data Y; para. [0011-0013]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Lee. One would have been motivated to make this modification because it improves denoising performance relative to conventional denoising autoencoder that directly learns the original data. Claim 5: Fatemi and You teach the method of claim 4. Fatemi does not explicitly teach wherein for each of one or more of the nodes having modified feature representations, the respective de-noising prediction for the node predicts an incremental feature representation for the node that, if added to the modified feature representation for the node, results in the target feature representation of the node. However, Lee teaches wherein for each of one or more of the nodes having modified feature representations, the respective de-noising prediction for the node predicts an incremental feature representation for the node that, if added to the modified feature representation for the node, results in the target feature representation of the node (i.e. the nlDAE method may regenerate the original data from a noisy input by learning the noise through a neural network and then subtracting the regenerated noise from the input. Thus, the nlDAE method may differ from the conventional DAE, which attempts to learn the original data directly; para. [0005, 0011, 0015]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Lee. One would have been motivated to make this modification because it improves denoising performance relative to conventional DAE structures that directly reconstruct the original data. Claim 23 is similar in scope to Claim 2 and is rejected under a similar rationale. 7. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Fatemi in view of You, and further in view of Ren et al. (U.S. Patent Application Pub. No. US 20210158127 A1). Claim 7: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein: (i) the updated node embeddings of the nodes (i.e. updated node embeddings; para. [0066]), and (ii) original node embeddings of the nodes prior to being updated using the graph neural network layers (i.e. X … is a matrix whose rows correspond to node features/attributes; para. [0064, 0065]), are processed to generate the task prediction (i.e. Classifier 146 thus takes as input node features 120 and the adjacency matrix A normalized by adjacency processor 144 and computes one or more predictions regarding which class 147 each node belongs to; para. [0086]). Fatemi does not explicitly teach wherein both: (i), and (ii) are processed to generate the task prediction. However, Ren teaches wherein both: (i) the updated node embeddings of the nodes, and (ii) original node embeddings of the nodes prior to being updated using the graph neural network layers (i.e. an embedding is generated by each network layer computation, and that embedding is concatenated with the aggregated neighbor embeddings; para. [0050]), are processed to generate the task prediction (i.e. The graph neural network embedding algorithm 500 computes the node embedding for each node of the graph. To predict a specific parasitic y on a node type t, the node embedding Zt of node type t is communicated to several fully connected (FC) layers. The FC layers (except the last one) have the dimensions of the embeddings, and the last layer has one dimension. A mean square error (MSE) loss function may then be utilized to regress the predicted value and the ground truth; para. [0053]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Ren. One would have been motivated to make this modification because it improves task prediction performance. 8. Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Fatemi in view of You, and further in view of Schutt et al. (SchNet: A continuous-filter convolutional neural network for modeling quantum interactions, NeurIPS, published 2017, pages 1-11). Claim 8: Fatemi teaches the method of claim 1. Fatemi does not explicitly teach wherein the graph represents a molecule and the task prediction is a prediction of an equilibrium energy of the molecule. However, Schutt teaches wherein the graph represents a molecule (i.e. At each layer, the molecule is represented atom wise analogous to pixels in an image. A molecule in a certain conformation can be described uniquely by a set of n atoms with nuclear charges Z = (Z1,...,Zn) and atomic positions R = (r1,...rn); Section 4.1, page 4) and the task prediction is a prediction of an equilibrium energy of the molecule (i.e. SchNet improves the state-of-the-art in predicting energies for molecules in equilibrium of the QM9 benchmark; Section 6, page 8). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Schutt. One would have been motivated to make this modification because it improves robustness of molecular prediction from corrupted features. Claim 16: Fatemi and You teach the method of claim 1. Fatemi does not explicitly teach wherein the graph represents a molecule, each node in the graph represents a respective atom in the molecule, and generating the data defining the graph comprises: generating a node embedding for each node based on a type of atom represented by the node. However, Schutt teaches wherein the graph represents a molecule (i.e. a novel deep learning architecture modeling quantum interactions in molecules; page 1), each node in the graph represents a respective atom in the molecule (i.e. we use the proposed cfconv layers in R3 to model interactions of atoms at arbitrary positions in the molecule; Section 1, page 2), and generating the data defining the graph comprises: generating a node embedding for each node (i.e. The atom type embeddings aZ are initialized randomly and optimized during training; page 4) based on a type of atom represented by the node (i.e. The representation of atom i is initialized using an embedding dependent on the atom type Zi; page 4). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Schutt. One would have been motivated to make this modification because it improves robustness of molecular prediction from corrupted features. 9. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Fatemi in view of You, and further in view of Oisel et al. (U.S. Patent Application Pub. No. US 20060106816 A1). Claim 13: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein generating the data defining the graph using the respective final feature representations of the nodes comprises: determining, for each pair of nodes comprising a first node and a second node, a respective distance between the final feature representation for the first node and the final feature representation for the second node (i.e. In some embodiments where a kNN is implemented to sparsify the generated graph, blocking the gradient flow can be avoided. Let M ∈ PNG media_image4.png 42 29 media_image4.png Greyscale n×n with Mij=1 if PNG media_image3.png 42 29 media_image3.png Greyscale j is among the top k similar nodes to PNG media_image3.png 42 29 media_image3.png Greyscale i and 0 otherwise, and let S ∈ PNG media_image4.png 42 29 media_image4.png Greyscale n×n with Sij=Sim (Xi′, Xj′) for some differentiable similarity function Sim (e.g., a cosine function). Then Ã=kNN(X′)=M⊙S where ⊙ represents the Hadamard (element-wise) product. With this formulation, in the forward phase of the network, one can first compute the matrix M using a k-nearest neighbors algorithm and then compute the similarities in S only for pairs of nodes where Mij=1. In some embodiments, exact k-nearest neighbors are computed; one can approximate it using locality-sensitive hashing approaches for larger graphs (see, e.g., [13, 58]); para. [0075, 0077, 0109]); and determining that each pair of nodes corresponding to a distance that is a predefined threshold are connected by an edge in the graph (i.e. to connect pairs of nodes whose similarity surpasses some predefined threshold (see, e.g., [34]); para. [0058]). Fatemi does not explicitly teach determining that each pair of nodes corresponding to a distance that is less than a predefined threshold. However, Oisel teaches determining that each pair of nodes corresponding to a distance that is less than a predefined threshold are connected by an edge in the graph (i.e. the nodes are first connected by edges according to the sequential unfurling of the video sequence. This progression is subsequently taken into account, when adding the complementary edges, for the application of the temporal window. Then, the integer value T chosen for this step of initializing the graph signifies that only the nodes separated at most by T nodes from the node considered are taken into account for the calculation of the initial edges starting from this node. An initial edge a.sub.ij is created between node n.sub.i and node n.sub.j only if the temporal distance d.sub.T separating these nodes with which the shots P.sub.i and P.sub.j are associated is less than a threshold T; para. [0026]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Oisel. One would have been motivated to make this modification because it provides a way to impose a constraint on graph construction, so that only sufficiently close nodes are connected. 10. Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Fatemi in view of You, and further in view of Wang et al. (U.S. Patent Application Pub. No. US 20210081677 A1). Claim 14: Fatemi and You teach the method of claim 1. Fatemi does not explicitly teach a respective edge embedding for each edge. However, Wang teaches wherein the graph further comprises a respective edge embedding for each edge (i.e. For each pair of nodes included in the graph, an attention component can be utilized to generate a corresponding edge embedding (or edge representation) that captures relationship information between the nodes, and the edge embedding can be associated with an edge in the graph that connects the node pair; para. [0020]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Wang. One would have been motivated to make this modification because it provides a way to improve the graph representation used by the prediction model by encoding relationship information for edges. 11. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Fatemi in view of You, Wang, and further in view of Wu et al. (U.S. Patent Application Pub. No. US 20230334332 A1). Claim 15: Fatemi, You, and Wang teach the method of claim 14. Fatemi does not explicitly teach generating an edge embedding for each edge in the graph based at least in part on a difference between the respective final feature representations of the nodes connected by the edge. However, Wang further teaches wherein generating the data defining the graph comprises: generating an edge embedding for each edge in the graph based at least in part between the respective final feature representations of the nodes connected by the edge (i.e. For each pair of nodes included in the graph, an attention component can be utilized to generate a corresponding edge embedding (or edge representation) that captures relationship information between the nodes, and the edge embedding can be associated with an edge in the graph that connects the node pair; para. [0020]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Wang. One would have been motivated to make this modification because it provides a way to improve the graph representation used by the prediction model by encoding relationship information for edges. However, Wu teaches wherein generating the data defining the graph comprises: generating an edge for each edge in the graph based at least in part on a difference between the respective final feature representations of the nodes connected by the edge (i.e. values (e.g., edge weights) of the adjacency matrix (e.g., which may represent relationships between node pairings of the graph) may be determined based on a function (e.g., an edge estimation function) that maps a distance (e.g., a Euclidean distance) between two nodes (e.g., between two embeddings of a node pair) of the graph from one space to another space. In some embodiments, the function may be expressed such that the edge weight increases as the distance between two nodes decreases; para. [0052]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi, You, and Wang to include the feature of Wu. One would have been motivated to make this modification because it provides a way to encode pairwise relational differences between connected nodes, thereby improving the informativeness of the graph representation for denoising and task prediction. 12. Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Fatemi in view of You, and further in view of Yu et al. (U.S. Patent Application Pub. No. US 20220129688 A1). Claim 17: Fatemi and You teach the method of claim 1. Fatemi further teaches wherein the neural network includes at least graph neural network layers (i.e. the second neural network comprises a two-layer graph convolutional network (GCN); para. [0023, 0085]). Fatemi does not explicitly teach at least 10 graph neural network layers. However, Yu teaches wherein the neural network includes at least 10 graph neural network layers (i.e. Similar to the first hidden layer 704, each of the nodes 722-728 in the second hidden layer 706 mutates the attributes in a graph node, except that the mutations are based on graph nodes that are farther away from the graph node (e.g., 2 degrees of separation) in the graph. As such, in some embodiments, the graph neural network 700 may include sufficient hidden layers (e.g., 5, 10, etc.) such that all of the other graph nodes in the graph can contribute to the mutations of each graph node in the graph through the hidden layers of the graph neural network 700; para. [0088]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Fatemi and You to include the feature of Yu. One would have been motivated to make this modification because it provides a way to improve the quality of denoising and downstream task prediction by incorporating broader graph context. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Kar et al. (Pub. No. US 20200160178 A1), a generative model is used to synthesize datasets for use in training a downstream machine learning model to perform an associated task. The synthesized datasets may be generated by sampling a scene graph from a scene grammar—such as a probabilistic grammar—and applying the scene graph to the generative model to compute updated scene graphs more representative of object attribute distributions of real-world datasets. The downstream machine learning model may be validated against a real-world validation dataset, and the performance of the model on the real-world validation dataset may be used as an additional factor in further training or fine-tuning the generative model for generating the synthesized datasets specific to the task of the downstream machine learning model. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)). Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TAN H TRAN/Primary Examiner, Art Unit 2141
Read full office action

Prosecution Timeline

Sep 20, 2023
Application Filed
Apr 09, 2026
Non-Final Rejection mailed — §103
May 21, 2026
Interview Requested
Jun 30, 2026
Applicant Interview (Telephonic)
Jul 01, 2026
Examiner Interview Summary
Jul 07, 2026
Response Filed
Aug 20, 2026
Final Rejection mailed — §103
Sep 22, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748960
Analog Hardware Realization of Neural Networks
5y 7m to grant Granted Sep 29, 2026
Patent 12718079
Systems and Methods for Generating Libraries for Hardware Realization of Neural Networks
5y 5m to grant Granted Aug 25, 2026
Patent 12718088
DESIGNING LADDER AND LAGUERRE ORTHOGONAL RECURRENT NEURAL NETWORK ARCHITECTURES INSPIRED BY DISCRETE-TIME DYNAMICAL SYSTEMS
4y 8m to grant Granted Aug 25, 2026
Patent 12688413
METHODS FOR RELIABLE OVER-THE-AIR COMPUTATION WITH PULSES FOR DISTRIBUTED LEARNING AND WITH FEDERATED EDGE LEARNING WITHOUT CHANNEL STATE INFORMATION
4y 1m to grant Granted Jul 21, 2026
Patent 12682274
MODEL INTEGRATION APPARATUS, MODEL INTEGRATION METHOD, COMPUTER-READABLE STORAGE MEDIUM STORING A MODEL INTEGRATION PROGRAM, INFERENCE SYSTEM, INSPECTION SYSTEM, AND CONTROL SYSTEM
5y 0m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
61%
Grant Probability
94%
With Interview (+32.6%)
3y 6m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 320 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month