Prosecution Insights
Last updated: August 18, 2026
Application No. 18/103,559

DEEP LEARNING ENTITY MATCHING SYSTEM USING WEAK SUPERVISION

Final Rejection §101§103
Filed
Jan 31, 2023
Examiner
CHUANG, SU-TING
Art Unit
2146
Tech Center
2100 — Computer Architecture & Software
Assignee
Walmart Apollo LLC
OA Round
2 (Final)
51%
Grant Probability
Moderate
3-4
OA Rounds
1y 0m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
55 granted / 108 resolved
-4.1% vs TC avg
Strong +40% interview lift
Without
With
+39.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
22 currently pending
Career history
135
Total Applications
across all art units

Statute-Specific Performance

§101
26.3%
-13.7% vs TC avg
§103
47.6%
+7.6% vs TC avg
§102
11.4%
-28.6% vs TC avg
§112
12.3%
-27.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 108 resolved cases

Office Action

§101 §103
DETAILED ACTION This action is in response the communications filed on 04/23/2026 in which claims 1, 3, 5, 6, 8, 11, 13, 15, 16, 18 are amended, claims 4, 9, 10 14, 19, 20 are canceled, claims 21-26 are added, and claims 1-3, 5-8, 11-13, 15-18 and 21-26 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. - Claims 1-3, 5-6, 11-13, 15-16 and 21-26 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more Step 1: Claims 1-3, 5-6 and 21-23 recite a system comprising processors and non-transitory computer-readable media. Claims 11-13, 15-16 and 24-26 recite a method. Therefore, claims 1-3, 5-6 and 21-23 are directed to a process, and claims 11-13, 15-16 and 24-26 are directed to a machine. With respect to claims 1 and 11: 2A Prong 1: The claim recites a judicial exception. generating pairs of identities from a plurality of sources, including: a first pair of identities that includes a first identity and a second identity different from the first identity, wherein the first pair of identities is automatically paired together based on a first shared characteristic; and a second pair of identities that includes the first identity and a third identity that is different from the first identity and the second identity, wherein the second pair of identities is automatically paired together based on a second shared characteristic different from the first shared characteristic; (mental process – evaluation or judgement; generating pairs of identities) generating a graph that includes linkages between at least two nodes, including: (mental process – evaluation or judgement; generating a graph) determining a first match probability for the first pair of identities…; determining a second match probability for the second pair of identities…; (mental process – evaluation or judgement; determining a match probability) in accordance with a determination that the first match probability meets a predetermined threshold, linking a first node in the graph that is representative of the first identity to a second node in the graph that is representative of the second identity; and (mental process – evaluation or judgement; linking nodes) in accordance with a determination that the second match probability meets the predetermined threshold, linking the first node in the graph that is representative of the first identity to a third node in the graph that is representative of the third identity; (mental process – evaluation or judgement; linking nodes) generating… one or more clusters of identities based on the graph, wherein each cluster of the one or more clusters contains one or more identities representing a respective user, wherein the one or more clusters includes a first cluster that is representative of a first user, and includes the first identity, the second identity, and the third identity; and (mental process – evaluation or judgement; generating clusters of identities) generating a respective user profile representative of the respective user for each cluster, including generating a first user profile representative of the first user (mental process – evaluation or judgement; generating a user profile) 2A Prong 2: The judicial exception is not integrated into a practical application. execution of computing instruction configured to run on one or more processors and stored at one or more non-transitory computer-readable media (mere instructions to apply an exception – MPEP 2106.05(f), (2) invoking generic computer components) … using a deep-learning transformer-based binary classification model… using the deep-learning transformer-based binary classification model… using a connected component algorithm… (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea. 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. execution of computing instruction configured to run on one or more processors and stored at one or more non-transitory computer-readable media (mere instructions to apply an exception – MPEP 2106.05(f), (2) invoking generic computer components) … using a deep-learning transformer-based binary classification model… using the deep-learning transformer-based binary classification model… using a connected component algorithm… (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. With respect to claims 2 and 12: 2A Prong 1: The claim recites a judicial exception. generating a probabilistic set of labels for an unlabeled training dataset to output a labeled training dataset (mental process – evaluation or judgement) 2A Prong 2: The judicial exception is not integrated into a practical application. wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform functions comprising (mere instructions to apply an exception – MPEP 2106.05(f), (2) invoking generic computer components) training the deep-learning transformer-based binary classification model using the labeled training dataset (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea. 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform functions comprising (mere instructions to apply an exception – MPEP 2106.05(f), (2) invoking generic computer components) training the deep-learning transformer-based binary classification model using the labeled training dataset (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. With respect to claims 3 and 13: 2A Prong 2: The judicial exception is not integrated into a practical application. wherein the probabilistic set of labels is generated using heuristic functions, and (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception: in light of specification [0114] “heuristic functions can be written to perform labelling functions that can be processed through Snorkel's algorithm”) generating the probabilistic set of labels uses a weak supervision model (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea. 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. wherein the probabilistic set of labels is generated using heuristic functions, and (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception: in light of specification [0114] “heuristic functions can be written to perform labelling functions that can be processed through Snorkel's algorithm”) generating the probabilistic set of labels uses a weak supervision model (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. With respect to claims 5 and 15: 2A Prong 1: The claim recites a judicial exception. wherein determining the first match probability for the first pair of identities comprises: (mental process – evaluation or judgement) 2A Prong 2: The judicial exception is not integrated into a practical application. obtaining textual features for the first identity and the second identity, wherein each of the textual features comprises unique string length distributions; and (insignificant extra-solution activity – MPEP 2106.05(g), (3) data gathering and outputting) generating a first sub-model based on the textual features (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea. 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. obtaining textual features for the first identity and the second identity, wherein each of the textual features comprises unique string length distributions; and (insignificant extra-solution activity – MPEP 2106.05(g), (3) data gathering and outputting, and WURC: Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 - MPEP 2106.05(d)(II)(i)) generating a first sub-model based on the textual features (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception) Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. With respect to claims 6 and 16: 2A Prong 1: The claim recites a judicial exception. wherein determining the first match probability further comprises (mental process – evaluation or judgement) 2A Prong 2: The judicial exception is not integrated into a practical application. obtaining boolean features for the first identity and the second identity, wherein the boolean features comprise external metadata and transaction history (insignificant extra-solution activity – MPEP 2106.05(g), (3) data gathering and outputting) generating a second sub-model based on the boolean features (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception)) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea. 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. obtaining boolean features for the first identity and the second identity, wherein the boolean features comprise external metadata and transaction history (insignificant extra-solution activity – MPEP 2106.05(g), (3) data gathering and outputting, and WURC: Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 - MPEP 2106.05(d)(II)(i))) generating a second sub-model based on the boolean features (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception)) Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. With respect to claims 21 and 24: 2A Prong 1: The claim recites a judicial exception. wherein generating the pairs of identities from the plurality of sources further includes generating a third pair of identities that includes a fourth identity and a fifth identity different from the fourth identity, wherein the third pair of identities is automatically paired together based on a third shared characteristic different from the first shared characteristic and the second shared characteristic. (mental process – evaluation or judgement; generating pairs of identities) With respect to claims 22 and 25: 2A Prong 1: The claim recites a judicial exception. generating the graph that includes linkages between at least two nodes further includes: (mental process – evaluation or judgement; generating the graph including linkages) determining a third match probability for the third pair of identities… (mental process – evaluation or judgement; determining a match probability) in accordance with a determination that the third match probability meets the predetermined threshold, linking the first node in the graph that is representative of the first identity to a fifth node in the graph that is representative of the fifth identity; and (mental process – evaluation or judgement; linking nodes) 2A Prong 2: The judicial exception is not integrated into a practical application. wherein: the fourth identity is the same as the first identity; (whether additional elements meaningfully limit the judicial exception – MPEP 2106.05(e); not a meaningful limitation, no actual steps, merely additional details of the claim elements) using the deep-learning transformer-based binary classification model; and (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception; using the model) the first cluster that is representative of the first user includes the first identity, the second identity, the third identity, and the fifth identity (whether additional elements meaningfully limit the judicial exception – MPEP 2106.05(e); not a meaningful limitation, no actual steps, merely additional details of the claim elements) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea. 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. wherein: the fourth identity is the same as the first identity; (whether additional elements meaningfully limit the judicial exception – MPEP 2106.05(e); not a meaningful limitation, no actual steps, merely additional details of the claim elements) using the deep-learning transformer-based binary classification model; and (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception; using the model) the first cluster that is representative of the first user includes the first identity, the second identity, the third identity, and the fifth identity (whether additional elements meaningfully limit the judicial exception – MPEP 2106.05(e); not a meaningful limitation, no actual steps, merely additional details of the claim elements) Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. With respect to claims 23 and 26: 2A Prong 1: The claim recites a judicial exception. generating the graph that includes linkages between at least two nodes further includes: (mental process – evaluation or judgement; generating the graph including linkages) determining a third match probability for the third pair of identities… (mental process – evaluation or judgement; determining a match probability) in accordance with a determination that the third match probability meets the predetermined threshold, linking a fourth node in the graph that is representative of the fourth identity to a fifth node in the graph that is representative of the fifth identity; (mental process – evaluation or judgement; linking nodes) generating the one or more clusters of identities based on the graph further comprises generating a second cluster that is separate from the first cluster and that is representative of a second user different from the first user, and that includes the fourth identity and the fifth identity; and (mental process – evaluation or judgement; generating the clusters of identities, generating a cluster including multiple identities) generating a respective user profile representative of the respective user for each cluster further includes generating a second user profile that is different from the first user profile and is representative of the second user. (mental process – evaluation or judgement; generating a user profile) 2A Prong 2: The judicial exception is not integrated into a practical application. wherein: the fourth identity is different from the first identity, the second identity, and the third identity; (whether additional elements meaningfully limit the judicial exception – MPEP 2106.05(e); not a meaningful limitation, no actual steps, merely additional details of the claim elements) using the deep-learning transformer-based binary classification model; and (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception; using the model) Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is directed to an abstract idea. 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. wherein: the fourth identity is different from the first identity, the second identity, and the third identity; (whether additional elements meaningfully limit the judicial exception – MPEP 2106.05(e); not a meaningful limitation, no actual steps, merely additional details of the claim elements) using the deep-learning transformer-based binary classification model; and (mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception; using the model) Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 11 and 21-26 rejected under 35 U.S.C. 103 as being unpatentable over Yao ("Entity Resolution with Hierarchical Graph Attention Networks" 20220612) in view of Chen ("Iterative Deep Graph Learning for Graph Neural Networks: Better and Robust Node Embeddings" 2020) in further view of Koduri ("Cross-Device Identity Resolution using Machine Learning: A Scalable Device Graph Approach" 2021) In regard to claims 1 and 11, Yao teach: A system comprising: one or more processors; and one or more non-transitory computer-readable media storing computing instructions that, when executed on the one or more processors, cause the one or more processors to perform functions comprising: (Yao, p. 436, 6 Experiments "These models are implemented in PyTorch and evaluated on a Linux server with a V100 GPU."; a server inherently teaches all the computer components) generating pairs of identities from a plurality of sources, including: a first pair of identities that includes a first identity and a second identity different from the first identity, wherein the first pair of identities is automatically paired together based on a first shared characteristic; and a second pair of identities that includes the first identity and a third identity that is different from the first identity and the second identity, wherein the second pair of identities is automatically paired together based on a second shared characteristic different from the first shared characteristic; (Yao, p. 430, 2.1 Entity Resolution "In the ER problem, an entity often represents a real-world object, such as product, person, company, etc. Each entity e is described by pairs of <key, val> where key and val denote the name and value of an entity attribute, respectively... ER Problem: Given two collections of data entities D and D' [a plurality of sources], the goal of ER problem is to output an entity matching matrix L ⊆ DxD', where element l_ij = {(ei, ej) [pairs of identities]|ei∈D, ej∈D'} indicates whether ei and ej match."; p.3, HHG Construction "We illustrate the construction of HHG in Figure 4... Note that each distinct word becomes only one token node in the HHG even if it appears in multiple attributes or multiple entities. [a first/second shared characteristic] For example, there is only one 'framework' token node in Figure 4."; see Fig 2, a pair (eq vs. e1) and a pair (eq vs. e2) [a first/second pair of identities] include eq and e1 [a first identity and a second identity] and eq and e2 [the first identity and a third identity] respectively; see Fig. 4, a pair may have a common word, e.g. 'spark' in the title [a first shared characteristic] or 'framework' in the desc [a second shared characteristic]) PNG media_image1.png 264 583 media_image1.png Greyscale PNG media_image2.png 342 708 media_image2.png Greyscale generating a graph that includes linkages between at least two nodes, including: determining a first match probability for the first pair of identities using a deep-learning transformer-based binary classification model; determining a second match probability for the second pair of identities using the deep-learning transformer-based binary classification model; (Yao, p. 429 "we propose HierGAT, a new method for ER based on a Hierarchical Graph Attention Transformer Network, [GAT: a deep-learning transformer-based model] which can model and exploit the interdependence between different ER decisions."; p. 435, 5.2.2 Entity Comparison Layer "Give two entities to be compared, we concatenate these two entity embeddings ve_lr = (ve_l||ve_r) and use concatenated entity embedding ve_lr as the contextual information for attribute similarity. We use the graph attention mechanism to get two entities' similarity representation se_lr: hk = softmax (LeakyReLU (cT (ve_lr||Sa_k)))...(4) where... hk is k-th attribute's attention value of two entity's... [determining a first/second match probability for the first/second pair of identities] Those entity similarity embeddings can be used for a classifier to tell if two entities are a match or not. [a binary classification model]"; p. 436 5.3 Training Process "The classifier used for ER is a binary classifier. We use the cross entropy function to calculate the classification loss."; the attention uses a softmax function, which converts values into probabilities) Yao does not teach, but Chen teaches: in accordance with a determination that the first match probability meets a predetermined threshold, linking a first node in the graph that is representative of the first identity to a second node in the graph that is representative of the second identity; and in accordance with a determination that the second match probability meets the predetermined threshold, linking the first node in the graph that is representative of the first identity to a third node in the graph that is representative of the third identity; (Chen, p. 3 Graph similarity metric learning "Common options for metric learning include cosine similarity [44, 54], radial basis function (RBF) kernel [59, 34] and attention mechanisms [51, 23]."; p. 3, Graph sparsification via ε-neighborhood "Typically an adjacency matrix (computed from a metric) is supposed to be non-negative... We hence proceed to extract a symmetric sparse non-negative adjacency matrix A from S by considering only the ε-neighborhood for each node. Specifically, we mask off (i.e., set to zero) those elements in S which are smaller than a non-negative threshold ε [removing linkages if it is smaller than a threshold, i.e. keeping linkages if it is greater than a threshold]"; Yao teaches the attention value generated with a softmax function is a match probability) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao to incorporate the teachings of Chen by including ε-neighborhood mechanism. Doing so would yield a sparse, computationally lightweight graph that eliminates noise by removing unimportant edges. (Chen, p. 3, Graph sparsification via ε-neighborhood "many underlying graph structures are much more sparse than a fully connected graph which is not only computationally expensive but also might introduce noise (i.e., unimportant edges). We hence proceed to extract a symmetric sparse non-negative adjacency matrix A from S by considering only the ε-neighborhood for each node.") Yao and Chen do not teach, but Koduri teaches: generating, using a connected component algorithm, one or more clusters of identities based on the graph, wherein each cluster of the one or more clusters contains one or more identities representing a respective user, wherein the one or more clusters includes a first cluster that is representative of a first user, and includes the first identity, the second identity, and the third identity; and (Koduri, p. 2, 3 Methodology "we clustered the IDs into appropriate groupings [generating clusters] using community clustering algorithms. [a connected component algorithm]"; p. 3, 3.4 Community Clustering "After establishing the probabilistic relationships between pairs of devices, we progressed with different community clustering approaches (affinity propagation, label propagation, and connected components) to build the final device graph."; see Fig. 1, e.g. cluster c1 or user c1 [a first cluster, representative of a first user] in the device graph PNG media_image3.png 330 738 media_image3.png Greyscale contains d1, d2, d3 and d4 (including multiple identities)) generating a respective user profile representative of the respective user for each cluster, including generating a first user profile representative of the first user. (Koduri, p. 1, 1 Introduction "The deterministic approach relies on logins or other personally identifiable information available to link multiple identifiers to a single profile [a user profile representing the user] confidently... With an established device graph, we are no longer targeting devices, but actual users (Figure 1)."; see Fig. 1 c1, c2, c3 representing a user for each cluster in the device graph) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao and Chen to incorporate the teachings of Koduri by including community clustering algorithms for the same user. Doing so would allow reduce advertising wastage and make more informed decisions. (Koduri, p. 1, 1 Introduction "With this new user-centric perspective, an advertisement can now be served more judiciously to the same user across devices by reinforcing the impression with the right frequency and context. Advertising wastage is reduced and more informed media planning decisions can be made.") Claim 11 recites substantially the same limitation as claim 1, therefore the rejection applied to claim 1 also apply to claim 11. In addition, Yao teaches: A method being implemented via execution of computing instruction configured to run on one or more processors and stored at one or more non-transitory computer-readable media, the method comprising: (Yao, p. 436, 6 Experiments "These models are implemented in PyTorch and evaluated on a Linux server with a V100 GPU.") In regard to claims 21 and 24, Yao teach: wherein generating the pairs of identities from the plurality of sources further includes generating a third pair of identities that includes a fourth identity and a fifth identity different from the fourth identity, wherein the third pair of identities is automatically paired together based on a third shared characteristic different from the first shared characteristic and the second shared characteristic. (Yao, p. 430, 2.1 Entity Resolution "In the ER problem, an entity often represents a real-world object, such as product, person, company, etc. Each entity e is described by pairs of <key, val> where key and val denote the name and value of an entity attribute, respectively... ER Problem: Given two collections of data entities D and D' [the plurality of sources], the goal of ER problem is to output an entity matching matrix L ⊆ DxD', where element l_ij = {(ei, ej) [pairs of identities]|ei∈D, ej∈D'} indicates whether ei and ej match."; p.3, HHG Construction "We illustrate the construction of HHG in Figure 4... Note that each distinct word becomes only one token node in the HHG even if it appears in multiple attributes or multiple entities. [a third shared characteristic] For example, there is only one 'framework' token node in Figure 4."; see Fig 2, a pair (eq vs. e3) [a third pair of identities] include eq and e3 [a fourth identity and a fifth identity] and a pair may have a common word [a third shared characteristic]) In regard to claims 22 and 25, Yao teach: wherein: the fourth identity is the same as the first identity; (Yao, p.430, ER Problem "To solve this problem, a straightforward way is to take a query entity eq from D and then compare it with all the candidate entities in D'."; see Fig 2, solid lines, a pair (eq vs. e3) [a third pair of identities] include eq and e3 [a fourth identity and a fifth identity], where eq maps to the first and fourth identity) generating the graph that includes linkages between at least two nodes further includes: determining a third match probability for the third pair of identities using the deep-learning transformer-based binary classification model; and (Yao, p. 429 "we propose HierGAT, a new method for ER based on a Hierarchical Graph Attention Transformer Network, [GAT: a deep-learning transformer-based model] which can model and exploit the interdependence between different ER decisions."; p. 435, 5.2.2 Entity Comparison Layer "Give two entities to be compared, we concatenate these two entity embeddings ve_lr = (ve_l||ve_r) and use concatenated entity embedding ve_lr as the contextual information for attribute similarity. We use the graph attention mechanism to get two entities' similarity representation se_lr: hk = softmax (LeakyReLU (cT (ve_lr||Sa_k)))...(4) where... hk is k-th attribute's attention value of two entity's... [determining a third match probability for the third pair of identities] Those entity similarity embeddings can be used for a classifier to tell if two entities are a match or not. [a binary classification model]"; p. 436 5.3 Training Process "The classifier used for ER is a binary classifier. We use the cross entropy function to calculate the classification loss."; the attention uses a softmax function, which converts values into probabilities) Yao does not teach, but Chen teaches: in accordance with a determination that the third match probability meets the predetermined threshold, linking the first node in the graph that is representative of the first identity to a fifth node in the graph that is representative of the fifth identity; and (Chen, p. 3 Graph similarity metric learning "Common options for metric learning include cosine similarity [44, 54], radial basis function (RBF) kernel [59, 34] and attention mechanisms [51, 23]."; p. 3, Graph sparsification via ε-neighborhood "Typically an adjacency matrix (computed from a metric) is supposed to be non-negative... We hence proceed to extract a symmetric sparse non-negative adjacency matrix A from S by considering only the ε-neighborhood for each node. Specifically, we mask off (i.e., set to zero) those elements in S which are smaller than a non-negative threshold ε [removing linkages if it is smaller than a threshold, i.e. keeping linkages if it is greater than a threshold]"; Yao teaches the attention value generated with a softmax function is a match probability) The rationale for combining the teachings of Yao and Chen is the same as set forth in the rejection of claim 1. Yao and Chen do not teach, but Koduri teaches: the first cluster that is representative of the first user includes the first identity, the second identity, the third identity, and the fifth identity. (Koduri, p. 2, 3 Methodology "we clustered the IDs into appropriate groupings [generating clusters] using community clustering algorithms."; see Fig. 1, e.g. cluster c1 or user c1 [a respective user, representative of a first user] in the device graph contains d1, d2, d3 and d4 (including multiple identities)) The rationale for combining the teachings of Yao, Chen and Koduri is the same as set forth in the rejection of claim 1. In regard to claims 23 and 26, Yao teach: wherein: the fourth identity is different from the first identity, the second identity, and the third identity; (Yao, p.430, ER Problem "In this paper, we also consider collective ER, in which a query entity eq has N candidate entities and we determine their matching relationships together. As shown in Figure 2, we create a relation network to describe the matching relations between the query entity eq and N candidates."; see Fig 2, dashed lines, query can be any of the entities in the collection, which is different from the entities connected in solid lines) generating the graph that includes linkages between at least two nodes further includes:determining a third match probability for the third pair of identities using the deep-learning transformer-based binary classification model; and (Yao, p. 429 "we propose HierGAT, a new method for ER based on a Hierarchical Graph Attention Transformer Network, [GAT: a deep-learning transformer-based model] which can model and exploit the interdependence between different ER decisions."; p. 435, 5.2.2 Entity Comparison Layer "Give two entities to be compared, we concatenate these two entity embeddings ve_lr = (ve_l||ve_r) and use concatenated entity embedding ve_lr as the contextual information for attribute similarity. We use the graph attention mechanism to get two entities' similarity representation se_lr: hk = softmax (LeakyReLU (cT (ve_lr||Sa_k)))...(4) where... hk is k-th attribute's attention value of two entity's... [determining a third match probability for the third pair of identities] Those entity similarity embeddings can be used for a classifier to tell if two entities are a match or not. [a binary classification model]"; p. 436 5.3 Training Process "The classifier used for ER is a binary classifier. We use the cross entropy function to calculate the classification loss."; the attention uses a softmax function, which converts values into probabilities) Yao does not teach, but Chen teaches: in accordance with a determination that the third match probability meets the predetermined threshold, linking a fourth node in the graph that is representative of the fourth identity to a fifth node in the graph that is representative of the fifth identity; (Chen, p. 3 Graph similarity metric learning "Common options for metric learning include cosine similarity [44, 54], radial basis function (RBF) kernel [59, 34] and attention mechanisms [51, 23]."; p. 3, Graph sparsification via ε-neighborhood "Typically an adjacency matrix (computed from a metric) is supposed to be non-negative... We hence proceed to extract a symmetric sparse non-negative adjacency matrix A from S by considering only the ε-neighborhood for each node. Specifically, we mask off (i.e., set to zero) those elements in S which are smaller than a non-negative threshold ε [removing linkages if it is smaller than a threshold, i.e. keeping linkages if it is greater than a threshold]"; Yao teaches the attention value generated with a softmax function is a match probability) The rationale for combining the teachings of Yao and Chen is the same as set forth in the rejection of claim 1. Yao and Chen do not teach, but Koduri teaches: generating the one or more clusters of identities based on the graph further comprises generating a second cluster that is separate from the first cluster and that is representative of a second user different from the first user, and that includes the fourth identity and the fifth identity; and (Koduri, p. 2, 3 Methodology "we clustered the IDs into appropriate groupings [generating clusters] using community clustering algorithms."; see Fig. 1, e.g. cluster c2 or user c2 [a second cluster, representative of a second user] in the device graph contains d5 and d6 (including multiple identities)) generating a respective user profile representative of the respective user for each cluster further includes generating a second user profile that is different from the first user profile and is representative of the second user. (Koduri, p. 1, 1 Introduction "The deterministic approach relies on logins or other personally identifiable information available to link multiple identifiers to a single profile [a user profile representing the user] confidently... With an established device graph, we are no longer targeting devices, but actual users (Figure 1)."; see Fig. 1 c1, c2, c3 representing a user for each cluster in the device graph) The rationale for combining the teachings of Yao, Chen and Koduri is the same as set forth in the rejection of claim 1. Claims 2 and 12 rejected under 35 U.S.C. 103 as being unpatentable over Yao, Chen and Koduri as applied to claims 1 and 11, and in further view of Ahn ("Practical Binary Code Similarity Detection with BERT-based Transferable Similarity Learning" 20221205) In regard to claims 2 and 12, Yao teaches: wherein the computing instructions, when executed on the one or more processors, further cause the one or more processors to perform functions comprising: (Yao, p. 436, 6 Experiments "These models are implemented in PyTorch and evaluated on a Linux server with a V100 GPU.") … training the deep-learning transformer-based binary classification model using the labeled training dataset. (Yao, p. 435, 5.3 Training Process "The training procedure of HierGAT is as follows using a labeled training set.") Yao, Chen and Koduri do not tech, but Ahn teaches: generating a probabilistic set of labels for an unlabeled training dataset to output a labeled training dataset; and (Ahn, p. 362, 2 BACKGROUND "MLM randomly masks a certain portion of tokens in a given sentence (e.g., 15% in the original BERT scheme), exploiting unlabeled data [for an unlabeled training dataset] (i.e., masked positions) to yield labels (i.e., original tokens)."; p. 364, 4.3 Pre-trainer: Model for Assembly "MLM task. We take the identical strategy with the original BERT, replacing 15% of input tokens (instructions) with a mask symbol (i.e., [MASK] token). The parameters of MLM... t∈T... y^ = softmax [probabilistic] (Gm(X) (4) where t, T, y and y^ denote a token, a set of tokens, an original token before masking, and a predicted token [a probabilistic set of labels] for MLM, respectively."; p. 364, Figure 3 "Siamese neural network for building a BCSD model. Our model learns a weighted distance vector from a labeled dataset (i.e., a set of two functions and a label). [a labeled training dataset]"; the output of the pre-trained BERT model PNG media_image4.png 260 538 media_image4.png Greyscale is a labeled dataset, which is an input to a downstream task) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao, Chen and Koduri to incorporate the teachings of Ahn by including a pre-trainer. Doing so would allow to build a generic model that can be repurposed for various downstream tasks. (Ahn, p. 363, Pre-trainer "we build a generic BERT model with pre-training... Akin to NLP’s pre-trained language models, this model for an assembly language allows for repurposing it to varying downstream tasks.") Claims 3 and 13 rejected under 35 U.S.C. 103 as being unpatentable over Yao, Chen, Koduri and Ahn as applied to claims 2 and 12, and in further view of Wu ("Demonstration of Panda: A Weakly Supervised Entity Matching System" 20210923) In regard to claims 3 and 13, Yao, Chen, Koduri and Ahn do not teach, but Wu teaches: wherein the probabilistic set of labels is generated using heuristic functions, and (Wu, p. 1, Abstract "where labeling functions (LF) are user-provided programs that can generate large amounts of (somewhat noisy) labels quickly and cheaply"; p. 2 "To write LFs for EM, users need to examine tuple pairs from a specific EM task, in order to develop intuitions/heuristics that can be turned into code (LFs) [using heuristic functions (LFs)] to quickly label matches/non-matches."; p. 3, 3. Combining LFs "All possible triples ti, tj, tk form a feasible set Q for the probabilistic labels of the tuple pairs. [the probabilistic set of labels] We then enforce the transitivity constraint by projecting the estimated probabilistic labels to the feasible set... at each E-step.") generating the probabilistic set of labels uses a weak supervision model. (Wu, p. 1, Abstract "In this paper, we introduce Panda, a weakly supervised system specifically designed for EM.") It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao, Chen, Koduri and Ahn to incorporate the teachings of Wu by including heuristic labeling functions. Doing so would generate large amounts of (somewhat noisy) labels quickly and cheaply. (Wu, p. 1, Abstract "where labeling functions (LF) are user-provided programs that can generate large amounts of (somewhat noisy) labels quickly and cheaply") Claims 5-6 and 15-16 rejected under 35 U.S.C. 103 as being unpatentable over Yao, Chen and Koduri as applied to claims 1 and 11, and in further view of Wilcke ("End-to-End Entity Classification on Multimodal Knowledge Graphs" 20200325) In regard to claims 5 and 15, Yao teaches: wherein determining the first match probability for the first pair of identities comprises: (Yao, p. 435, 5.2.2 Entity Comparison Layer "We use the graph attention mechanism to get two entities' similarity representation se_lr: hk = softmax (LeakyReLU (cT (ve_lr||Sa_k)))...(4) where... hk is k-th attribute's attention value of two entity's... [determining the first match probability for the first pair of identities]"; the attention uses a softmax function, which converts values into probabilities) Yao, Chen and Koduri do not teach, but Wilcke teaches: obtaining textual features for the first identity and the second identity, wherein each of the textual features comprises unique string length distributions; and (Wilcke, p. 9, Table 4 "Distribution of datatypes in the datasets…. Textual information includes strings and its subsets, as well as raw URIs (e.g. links). [unique string length distributions]"; URIs (Uniform Resource Identifier) are unique identifiers) PNG media_image5.png 462 657 media_image5.png Greyscale generating a first sub-model based on the textual features. (Wilcke, p. 4, Figure 2 "Solid circles represent entities, whereas open shapes represent literals of different modalities. The nodes’ feature embeddings are learned using dedicated (neural) encoders (here f, g, and h) [a first sub-model]"; p. 5, 4.1.3 Textual Information "Vector representations for textual attributes with the datatype XSD: string or any subtype thereof...") It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao, Chen and Koduri to incorporate the teachings of Wilcke by including embeddings for node features belonging to five different types of modalities. Doing so would help our models obtain a better overall performance. (Wickle, p. Abstract "Our model uses dedicated (neural) encoders to naturally learn embeddings for node features belonging to five different types of modalities, including images and geometries, which are projected into a joint representation space together with their relational information... Our result supports our hypothesis that including information from multiple modalities can help our models obtain a better overall performance.") In regard to claims 6 and 16, Yao teaches: wherein determining the first match probability further comprises: (Yao, p. 435, 5.2.2 Entity Comparison Layer "We use the graph attention mechanism to get two entities' similarity representation se_lr: hk = softmax (LeakyReLU (cT (ve_lr||Sa_k)))...(4) where... hk is k-th attribute's attention value of two entity's... [determining the first match probability for the first pair of identities]"; the attention uses a softmax function, which converts values into probabilities) … wherein the boolean features comprise external metadata and transaction history; and (Yao, p. 436, 6 EXPERIMENTS "we use 9 publicly available evaluation sets... We also use four additional dirty datasets that are publicly available [external metadata] from DeepMatcher... the title attribute may contain the price information... We use the WDC product matching data, which are extracted from several e-commerce websites and categorized into four different domains: computer, camera, watch, and shoe... For the DI2KG dataset, the samples are collected from multiple e-commerce websites and categorized into: camera and monitor"; data from e-commerce websites include transaction history; in light of specification [0101] "other qualities of an identity, such as external metadata and transaction history, are stored as Boolean features"; price is numerical information) Yao, Chen and Koduri do not teach, but Wilcke teaches: obtaining boolean features for the first identity and the second identity, (Wilcke, p. 9, Table 4 "Distribution of datatypes in the datasets. Numerical information includes all subsets of real numbers, as well as booleans...") generating a second sub-model based on the boolean features. (Wilcke, p. 4, Figure 2 "Solid circles represent entities, whereas open shapes represent literals of different modalities. The nodes’ feature embeddings are learned using dedicated (neural) encoders (here f, g, and h) [a second sub-model]"; p. 5, 4.1.1 Numerical Information "We also include values of the type XSD: boolean into this category...") The rationale for combining the teachings of Yao, Chen and Koduri and Wilcke is the same as set forth in the rejection of claim 5. Claims 7 and 17 rejected under 35 U.S.C. 103 as being unpatentable over Yao, Chen, Koduri and Wilcke as applied to claims 6 and 16, and in view of Hou ("Token Dropping for Efficient BERT Pretraining" 20220324) in further view of Reimers ("Sentence-bert: Sentence embeddings using siamese bert-networks" 20190827) In regard to claims 7 and 17, Yao, Chen and Koduri do not teach, but Wilcke teaches: wherein generating the first sub-model comprises: (Wilcke, p. 4, Figure 2 "Solid circles represent entities, whereas open shapes represent literals of different modalities. The nodes’ feature embeddings are learned using dedicated (neural) encoders (here f, g, and h) [the first sub-model]"; p. 5, 4.1.3 Textual Information "Vector representations for textual attributes with the datatype XSD: string or any subtype thereof...") generating character-level encodings to convert the textual features into numeric representations; (Wilcke, p. 5, 4.1.3 Textual Information "Vector representations for textual attributes with the datatype XSD: string or any subtype thereof, are created using a character-level encoding, [character-level encodings] proposed in [16] Hereto, we let Es be a |Ω|×|s| matrix representing string s using vocabulary Ω, such that Es ij = 1.0 if sj = Ωi, and 0.0 otherwise. [numeric representations] A character-level representation enables our models to be language agnostic and independent of controlled vocabularies...") The rationale for combining the teachings of Yao, Chen and Koduri and Wilcke is the same as set forth in the rejection of claim 5. PNG media_image6.png 412 646 media_image6.png Greyscale Yao, Chen, Koduri and Wilcke do not teach, but Hou teaches: sending the character-level encodings into a first embedding layer to generate a first embedding, wherein the first embedding layer is trained to remove sparsity from the character-level encodings; (Hou, p. 1, Abstract "We develop a simple but effective 'token dropping' method to accelerate the pretraining of transformer models, such as BERT, without degrading its performance on downstream tasks. In short, we drop unimportant tokens starting from an intermediate layer in the model to make the model focus on important tokens; the dropped tokens are later picked up by the last layer of the model so that the model still produces full length sequences."; p. 3, 3 Token-Dropping "Using sparse tensors can address the issue of having a different number of important tokens, but sparse tensor related operations in practice are slow."; dropping tokens will keep the first intermediate layer dense, i.e. [removing sparsity from the encodings]) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao, Chen, Koduri and Wilcke to incorporate the teachings of Hou by including a token dropping method. Doing so would reduce the pretraining cost of BERT while achieving similar overall fine-tuning performance. (Hou, p. 1, Abstract "In our experiments, this simple approach reduces the pretraining cost of BERT by 25% while achieving similar overall fine-tuning performance on standard downstream tasks.") PNG media_image7.png 448 406 media_image7.png Greyscale Yao, Chen, Koduri, Wilcke and Hou do not teach, but Reimers teaches: sending the first embedding to an encoder block to generate final encodings, wherein the encoder block comprises a transformer using multi-head attention and a first fully connected layer using a Siamese architecture in which the encoder block is shared between two textual features; (Reimers, p. 3 "Figure 1: SBERT architecture with classification objective function, e.g., for fine-tuning on SNLI dataset. The two BERT networks have tied weights (siamese network structure)."; p. 2, 2 Related Work "BERT (Devlin et al., 2018) is a pre-trained transformer network... Multi-head attention over 12 (base-model) or 24 layers (large-model) is applied and the output is passed to a simple regression function to derive the final label."; 'BERT' block in Fig. 1 is [the encoder block] comprising shared/same Siamese architecture, i.e. [a first fully connected] and u and v are [embeddings]; Hou also teaches BERT Siamese architecture, 'FFW' in Fig. 2) calculating an absolute difference between the final encodings; and (Reimers, p. 3, 3 Model "Classification Objective Function. We concatenate the sentence embeddings u and v with the element-wise difference |u−v|[an absolute difference]") passing each difference of each textual feature encoding into a second fully connected layer. (Reimers, p. 3, 3 Model "softmax(Wt(u, v, |u − v|))"; softmax layer is a second fully connected layer) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao, Chen, Koduri, Wilcke and Hou to incorporate the teachings of Reimers by including Sentence-BERT to derive semantically meaningful sentence embeddings. Doing so would efficiently find the most similar pair while maintaining the accuracy. (Reimers, p. 1 "In this publication, we present Sentence-BERT (SBERT), a modification of the pretrained BERT network that use siamese and triplet network structures to derive semantically meaningful sentence embeddings... This reduces the effort for finding the most similar pair from 65 hours with BERT / RoBERTa to about 5 seconds with SBERT, while maintaining the accuracy from BERT.") Claims 8 and 18 rejected under 35 U.S.C. 103 as being unpatentable over Yao, Chen, Koduri, Wilcke , Hou and Reimers as applied to claims 7 and 17, and in view of Nie ("Deep Sequence-to-Sequence Entity Matching for Heterogeneous Entity Resolution" 20191103) In regard to claims 8 and 18, Yao teaches: weights of deep-learning transformer-based binary classification model are tuned using a binary cross-entropy loss function. (Yao, p. 436 5.3 Training Process "The classifier used for ER is a binary classifier. We use the cross entropy function to calculate the classification loss."; Nie also teache a binary cross-entropy loss function in eq(8) on p. 634) Yao, Chen and Koduri do not teach, but Wilcke teaches: wherein generating the second sub-model comprises: (Wilcke, p. 4, Figure 2 "Solid circles represent entities, whereas open shapes represent literals of different modalities. The nodes’ feature embeddings are learned using dedicated (neural) encoders (here f, g, and h) [the second sub-model]"; p. 5, 4.1.1 Numerical Information "We also include values of the type XSD: boolean into this category...") processing the boolean features using multiple fully connected layers; (Wilcke, p. 4, 3.2 Message Passing Neural Networks "A message passing neural network [3] is a graph neural network model that uses trainable functions to propagate node embeddings over the edges of the neural network."; p. 3, 2 Related Work "our approach includes a message passing layer, allowing multimodal information to be propagated through the graph, several hops, [multiple fully connected layers] before being used for classification.") The rationale for combining the teachings of Yao, Chen, Koduri and Wilcke is the same as set forth in the rejection of claim 5. Yao, Chen, Koduri, Wilcke, Hou and Reimers do not teach, but Nie teaches: determining the match probability further comprises: (Nie, p. 632, 4.1 Seq2Seq Entity Matching Network "Prediction Layer... a two layer fully-connected layer followed by a softmax classifier to get the final similarity score of the entity pair (S, T). [ the match probability]"; the output of a softmax layer is a probability) concatenating each output of the first sub-model and the second sub-model to generate a combined output; (Nie, p. 632, 4.1 Seq2Seq Entity Matching Network "Prediction Layer. The prediction layer performs similarity assessment based on the two feature vectors generated in the previous step. Specifically, taking two feature vectors as input, we first concatenate them and then pass the resultant vector...") passing the combined output, as concatenated, into a final fully connected layer; and (Nie, p. 632, 4.1 Seq2Seq Entity Matching Network "Prediction Layer... Specifically, taking two feature vectors as input, we first concatenate them and then pass the resultant vector to a two layer fully-connected layer followed by a softmax classifier...") PNG media_image8.png 334 479 media_image8.png Greyscale outputting the match probability; and (Nie, p. 632, 4.1 Seq2Seq Entity Matching Network "Prediction Layer... a two layer fully-connected layer followed by a softmax classifier to get the final similarity score of the entity pair (S, T). [ the match probability]"; the output of a softmax layer is a probability) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Yao, Chen, Koduri, Wilcke, Hou and Reimers to incorporate the teachings of Nie by including heterogeneous data and a softmax output layer. Doing so would effectively solve the heterogeneous problem and achieve remarkable performance improvements on entity resolution tasks. (Nie, p. 629, Abstract "… effectively solve the heterogeneous and dirty problems by modeling ER as a token-level sequence-to-sequence matching task... our Seq2Seq entity matching model can achieve remarkable performance improvements on 9 standard entity resolution benchmarks.") Response to Arguments Applicant's arguments with respect to the rejection of the claims under 35 U.S.C. 101 have been fully considered but they are not persuasive: Argument: (p. 10) Step 2A, Prong One… The instant claims recite "determining a first match probability for the first pair of identities using a deep-learning transformer-based binary classification model"; "determining a second match probability for the second pair of identities using the deep-learning transformer-based binary classification model"; "generating, using a connected component algorithm, one or more clusters of identities based on the graph"; and "generating a respective user profile representative of the respective user for each cluster, including generating a first user profile representative of the first user." Such steps cannot be reasonably performed in the human mind… Response: A human with aids of pen and paper can “determine a match probability for the pair, generate clusters of identities, and generating a user profile.” If a claim recites a limitation that can practically be performed in the human mind, with-- or without the use of a physical aid such as pen and paper, the limitation falls within the mental processes grouping, and the claim recites an abstract idea – MPEP 2106.04(a)(2)(III)(B). Further, the limitation “using the deep-learning transformer-based binary classification model” is an additional element evaluated at step 2A prong Two and step 2B, not evaluated at step 2A prong One, and it is mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception. Argument: (p. 13-14) Step 2A, Prong Two… The claims, as currently amended, address these existing weaknesses in automated entity matching. For example, Claim 1 recites pairing identities… allows for a more complete picture and more accurate representation of a user, rather than having disparate, separate representations of the same user based on different characteristics that may represent the same user. Furthermore, Claim 1 recites creating linkages… Independent claims 1 and 11 provide an improvement to existing methods for automated entity matching. Claims that represent an improvement to a technology or technical field… Response: Each of the "generating pairs, generating a graph that includes linkages, linking a first node… to a second node, and linking the first node… to a third node" steps is evaluated at step 2A prong One as a mental process. Therefore, those steps are directly to mental processes, which are not sufficient to provide an improvement. If the limitation is directed to an exception, it cannot provide an improvement. (MPEP 2106.05(a): “It is important to note, the judicial exception alone cannot provide the improvement…”) Further, the limitation “using the deep-learning transformer-based binary classification model” is mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception. Argument: (p. 14-15) Step 2B… Applicant has amended the independent claims to recite limitations in which a first identity is paired with a second identity based on a first shared characteristic… Accordingly, the claims "provide an inventive concept," and recite limitations other than what is well-understood, routine, or conventional in the art. Response: Each of the "generating pairs, generating a graph that includes linkages, linking a first node… to a second node, and linking the first node… to a third node" steps is evaluated at step 2A prong One as a mental process. Therefore, those steps are directly to mental processes, which are not sufficient to provide an improvement, and therefore are not indicative of an inventive concept. Further, the limitation “using the deep-learning transformer-based binary classification model” is mere instructions to apply an exception – MPEP 2106.05(f), (3) The particularity or generality of the application of the judicial exception. Applicant's arguments with respect to the rejection of the claims under 35 U.S.C. 103 have been fully considered but they are moot: Argument: (p. 17-18) … Applicant has amended the independent claims to recite limitations… Applicant respectfully submits that such features are not disclosed anywhere in the currently cited references. None of the cited references appear to disclose a system in which identity pairings are generated based on multiple different types of shared characteristics, and then the identity pairings are represented in a single, combined graph that links nodes together based on linkages that represent those multiple different types of shared characteristics… Response: the arguments do not apply to the references (Yao) being used in the current rejection. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SU-TING CHUANG whose telephone number is (408)918-7519. The examiner can normally be reached Monday - Thursday 8-5 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.C./Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Jan 31, 2023
Application Filed
Dec 23, 2025
Non-Final Rejection mailed — §101, §103
Apr 07, 2026
Applicant Interview (Telephonic)
Apr 08, 2026
Examiner Interview Summary
Apr 23, 2026
Response Filed
Jul 17, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12645997
INDIVIDUALIZED CLASSIFICATION THRESHOLDS FOR MACHINE LEARNING MODELS
3y 3m to grant Granted Jun 02, 2026
Patent 12626164
SYSTEM AND METHOD FOR REDUCTION OF DATA TRANSMISSION BY DATA RECONSTRUCTION
4y 0m to grant Granted May 12, 2026
Patent 12626106
MACHINE LEARNING MODELS FOR BEHAVIOR UNDERSTANDING
3y 11m to grant Granted May 12, 2026
Patent 12626140
SYSTEMS AND METHODS FOR ONLINE TIME SERIES FORCASTING
3y 9m to grant Granted May 12, 2026
Patent 12619890
LEARNING PATTERN DICTIONARY FROM NOISY NUMERICAL DATA IN DISTRIBUTED NETWORKS
6y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
51%
Grant Probability
90%
With Interview (+39.5%)
4y 6m (~1y 0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 108 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month