Prosecution Insights
Last updated: October 02, 2026
Application No. 18/998,957

COMPUTER-IMPLEMENTED PERCEPTION OF 2D OR 3D SCENES

Non-Final OA §103
Filed
Jan 27, 2025
Priority
Jul 27, 2022 — GB 2210981.3 +1 more
Examiner
YANG, WEI WEN
Art Unit
Tech Center
Assignee
Five AI Limited
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
560 granted / 684 resolved
+21.9% vs TC avg
Moderate +12% lift
Without
With
+11.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
32 currently pending
Career history
705
Total Applications
across all art units

Statute-Specific Performance

§101
7.8%
-32.2% vs TC avg
§103
75.0%
+35.0% vs TC avg
§102
9.3%
-30.7% vs TC avg
§112
7.8%
-32.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 684 resolved cases

Office Action

§103
DETAILED ACTION Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 4, 12-15, 17, 19 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over XU (US 20230103305 A1), and in view of Rahman (US Per-frame mAP Prediction for Continuous Performance Monitoring of Object Detection During Deployment, 2021 IEEE Winter Conference on Applications of Computer Vision Workshops (WACVW); pages 152-160; as provided in IDS). Re Claim 1, XU discloses A computer-implemented method of assessing performance of perception component, the perception component for interpreting structure in a scene (see XU: e.g., --the disclosed systems can reduce complexity and improve performance by utilizing an efficient first-order graph matching model optimized via contrastive learning to determine a correspondence between labels and objects portrayed in a digital image. Moreover, the disclosed systems can utilize this correspondence to learn parameters of a scene graph model. In this manner, the disclosed systems can efficiently and flexibly generate accurate semantic scene graphs for digital images.,--, in [0002], [0035]; and, --[0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process.--, in [0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]); the method comprising: receiving a set of multiple computed outputs obtained by applying the perception component to the scene, wherein each computed output comprises a confidence score (see XU: e.g., --the graph matching system 102 utilizes a graph neural network that explicitly encodes edge type features into the node representation. For instance, an initial node state is an input node embedding h.sub.i.sup.0=e.sub.i. At the kth iteration, the graph neural network generates a confidence score to measure the confidence of whether an edge exists pointing from node i to node j…the graph matching system 102 augments the predicate attention with edge confidence such that it attends to the background class, and obtains the attended predicate representation from the augmented attention….the graph matching system 102 formulates the above message passing information with soft attention to the predicate type, suitable for the visual graph (which excludes the predicate category). Because the label graph has the determined relation type, the graph matching system 102 adopts hard attention instead of soft attention. This results in β.sub.ij in the measured confidence score--, in [0073]-[0076]; and, --generating label graph embeddings from an ungrounded label graph. For example, act 702 involves generating label graph embeddings from connected entity labels in an ungrounded label graph corresponding to a digital image. Act 702 can involve encoding features from an entity label of the ungrounded label graph into a label graph embedding utilizing a label embedding model. For example, the label embedding model comprises a multilayer perceptron network. Alternatively, the label embedding model comprises a graph neural network. Accordingly, act 702 can involve encoding, utilizing the graph neural network, information from an entity label and relationship information associated with one or more entity relationships involving the entity label into a label graph embedding according to one or more confidence scores for the one or more entity relationships.--, in [0104]); generating, from the set of multiple computed outputs, multiple pseudo-ground truth sets, wherein each pseudo-ground truth set comprises, for each computed output, a pseudo- ground truth output sampled from a set of possible ground truth outputs based on a probability distribution defined by the confidence score of the computed output (see XU: e.g., --the graph matching system generates a semantic scene graph for training a scene graph generation neural network. For example, the graph matching system aligns an ungrounded label graph and a visual graph to generate a ground-truth semantic scene graph for a digital image. The graph matching system then utilizes the ground-truth semantic scene graph to determine a scene graph generation loss by comparing the ground-truth semantic scene graph to a semantic scene graph generated by the scene graph generation neural network. Additionally, the graph matching system modifies parameters of the scene graph generation neural network based on the scene graph generation loss.--, in [0019]; and, --[0053] In one or more additional embodiments, the graph matching system 102 generates the semantic scene graph 304 based on the correspondences between the visual graph and the ungrounded label graph. …. the semantic scene graph 304 includes a scene graph that serves as a pseudo ground-truth semantic scene graph for training the scene graph generation neural network 306. More specifically, the scene graph generation neural network 306 generates semantic scene graphs from digital images for performing visual reasoning tasks based on, for example, scene construction, object detection, or object relationships. Accordingly, the graph matching system 102 utilizes the semantic scene graph 304 to train the scene graph generation neural network 306 to more accurately generate a semantic scene graph for the digital image 300… [0055] FIG. 4 illustrates a detailed diagram of the graph matching system 102 utilizing graph matching to generate semantic scene graphs for digital images. In particular, the graph matching system 102 utilizes a plurality of operations in the graph matching process for generating a semantic scene graph based on a digital image 400 and an image description 402 associated with the digital image 400. Additionally, FIG. 4 illustrates that the graph matching system 102 utilizes contrastive learning to learn embedding models for more accurate graph matching. [0056] As illustrated in FIG. 4, the graph matching system 102 generates a visual graph 404 of the digital image 400.--, in 0053]-[0056]); computing a performance for the perception component applied to the scene with respect to each pseudo-ground truth set, by comparing the set of multiple outputs with that pseudo-ground truth set (see XU: e.g., --the disclosed systems can reduce complexity and improve performance by utilizing an efficient first-order graph matching model optimized via contrastive learning to determine a correspondence between labels and objects portrayed in a digital image. Moreover, the disclosed systems can utilize this correspondence to learn parameters of a scene graph model. In this manner, the disclosed systems can efficiently and flexibly generate accurate semantic scene graphs for digital images.,--, in [0002], [0035]; and, --[0078] FIG. 5 illustrates an embodiment in which the graph matching system 102 utilizes a graph matching process to generate semantic scene graphs for training a scene graph generation neural network 500. In particular, the graph matching system 102 utilizes a digital image 502 and an image description 504 associated with the digital image 502 to generate a ground-truth semantic scene graph 506. For instance, the graph matching system 102 utilizes a graph matching process to determine a one-to-one correspondence between nodes in a visual graph generated from the digital image 502 and nodes in a label graph generated from the image description 504, as described above. The graph matching system 102 then generates the ground-truth semantic scene graph based on the one-to-one correspondences between the nodes, resulting in entity classes for entities in the digital image 502 and predicate classes for relationships between the entities based on the information in the image description 504. [0079] After generating the ground-truth semantic scene graph 506, the graph matching system 102 utilizes the scene graph generation neural network 500 to generate a predicted semantic scene graph 508 for the digital image 502. Specifically, the scene graph generation neural network 500 processes the digital image to predict object classes and predicate classes for entities in the digital image 502. The graph matching system 102 then determines a scene graph generation loss 510 for the predicted semantic scene graph 508. In particular, the graph matching system 102 determines the scene graph generation loss 510 by comparing the predicted semantic scene graph 508 to the ground-truth semantic scene graph 506. In some embodiments, the graph matching system 102 utilizes a cross-entropy loss L.sub.sgg as the scene graph generation loss 510 by comparing the predicted object classes and predicate classes in the predicted semantic scene graph 508 to the object classes and predicate classes included in the ground-truth semantic scene graph 506. [0080] In one or more embodiments, the graph matching system 102 then utilizes the scene graph generation loss 510 to modify the scene graph generation neural network 500. For instance, the graph matching system 102 updates parameters of the scene graph generation neural network 500 to reduce the scene graph generation loss 510 (e.g., by reducing distances between the object/predicate classes in the ground-truth semantic scene graph 506 and the predicted object/predicate classes in the predicted semantic scene graph 508 ). In at least some embodiments, the total loss for the graph matching system 102 and the scene graph generation neural network 500 is L=L.sub.gm+L.sub.sgg, which includes the contrastive loss for the weakly-supervised graph matching process and the scene graph generation loss.[ 0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process. For example, instance-level recall (R.sub.inst) indicates that a bounding region is correctly matched if the bounding region is matched with the correct node in the label graph and is correctly located (i.e., has more than 50% intersection-over-union (“IoU”) with the ground-truth bounding region). The evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]); and Xu although discloses computing a performance for the perception component applied to the scene (see XU: e.g., as above cited “Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process.”, and, evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081]), Xu however does not explicitly disclose determine a performance score Rahman discloses computing a performance score (see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detec tors.”, in page 154); XU and Rahman are combinable as they are in the same field of endeavor: evaluate the performance of object detection/matching neural network model. Therefore it would have been obvious to one of ordinary skill in the art at the time the invention was made to modify XU’s method using Rahman’s teachings by including computing a performance score to XU’s performance computing {e.g., Xu’s “the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image} in order to achieve assessment capability to detect instances of performance for continuous performance monitoring of object detectors (see Rahman: e.g., in Fig. 1, pages 152-154, and abstract), XU as modified by Rahman further disclose computing an overall performance score for the perception component applied to the scene, by aggregating the performance scores computed with respect to the multiple pseudo- ground truth sets (see XU: e.g., --the disclosed systems can reduce complexity and improve performance by utilizing an efficient first-order graph matching model optimized via contrastive learning to determine a correspondence between labels and objects portrayed in a digital image. Moreover, the disclosed systems can utilize this correspondence to learn parameters of a scene graph model. In this manner, the disclosed systems can efficiently and flexibly generate accurate semantic scene graphs for digital images.,--, in [0002], [0035]; and, --[0078] FIG. 5 illustrates an embodiment in which the graph matching system 102 utilizes a graph matching process to generate semantic scene graphs for training a scene graph generation neural network 500. In particular, the graph matching system 102 utilizes a digital image 502 and an image description 504 associated with the digital image 502 to generate a ground-truth semantic scene graph 506. For instance, the graph matching system 102 utilizes a graph matching process to determine a one-to-one correspondence between nodes in a visual graph generated from the digital image 502 and nodes in a label graph generated from the image description 504, as described above. The graph matching system 102 then generates the ground-truth semantic scene graph based on the one-to-one correspondences between the nodes, resulting in entity classes for entities in the digital image 502 and predicate classes for relationships between the entities based on the information in the image description 504. [0079] After generating the ground-truth semantic scene graph 506, the graph matching system 102 utilizes the scene graph generation neural network 500 to generate a predicted semantic scene graph 508 for the digital image 502. Specifically, the scene graph generation neural network 500 processes the digital image to predict object classes and predicate classes for entities in the digital image 502. The graph matching system 102 then determines a scene graph generation loss 510 for the predicted semantic scene graph 508. In particular, the graph matching system 102 determines the scene graph generation loss 510 by comparing the predicted semantic scene graph 508 to the ground-truth semantic scene graph 506. In some embodiments, the graph matching system 102 utilizes a cross-entropy loss L.sub.sgg as the scene graph generation loss 510 by comparing the predicted object classes and predicate classes in the predicted semantic scene graph 508 to the object classes and predicate classes included in the ground-truth semantic scene graph 506. [0080] In one or more embodiments, the graph matching system 102 then utilizes the scene graph generation loss 510 to modify the scene graph generation neural network 500. For instance, the graph matching system 102 updates parameters of the scene graph generation neural network 500 to reduce the scene graph generation loss 510 (e.g., by reducing distances between the object/predicate classes in the ground-truth semantic scene graph 506 and the predicted object/predicate classes in the predicted semantic scene graph 508 ). In at least some embodiments, the total loss for the graph matching system 102 and the scene graph generation neural network 500 is L=L.sub.gm+L.sub.sgg, which includes the contrastive loss for the weakly-supervised graph matching process and the scene graph generation loss. [0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process. For example, instance-level recall (R.sub.inst) indicates that a bounding region is correctly matched if the bounding region is matched with the correct node in the label graph and is correctly located (i.e., has more than 50% intersection-over-union (“IoU”) with the ground-truth bounding region). The evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]) . Re Claim 2, XU as modified by Rahman further disclose wherein the perception component is an object detector and wherein the set of multiple computed outputs is a set of object detections (see XU: e.g., --the graph matching system 102 utilizes a graph neural network that explicitly encodes edge type features into the node representation. For instance, an initial node state is an input node embedding h.sub.i.sup.0=e.sub.i. At the kth iteration, the graph neural network generates a confidence score to measure the confidence of whether an edge exists pointing from node i to node j…the graph matching system 102 augments the predicate attention with edge confidence such that it attends to the background class, and obtains the attended predicate representation from the augmented attention….the graph matching system 102 formulates the above message passing information with soft attention to the predicate type, suitable for the visual graph (which excludes the predicate category). Because the label graph has the determined relation type, the graph matching system 102 adopts hard attention instead of soft attention. This results in β.sub.ij in the measured confidence score--, in [0073]-[0076]; and, --[0080] In one or more embodiments, the graph matching system 102 then utilizes the scene graph generation loss 510 to modify the scene graph generation neural network 500. For instance, the graph matching system 102 updates parameters of the scene graph generation neural network 500 to reduce the scene graph generation loss 510 (e.g., by reducing distances between the object/predicate classes in the ground-truth semantic scene graph 506 and the predicted object/predicate classes in the predicted semantic scene graph 508 ). In at least some embodiments, the total loss for the graph matching system 102 and the scene graph generation neural network 500 is L=L.sub.gm+L.sub.sgg, which includes the contrastive loss for the weakly-supervised graph matching process and the scene graph generation loss. [0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process. For example, instance-level recall (R.sub.inst) indicates that a bounding region is correctly matched if the bounding region is matched with the correct node in the label graph and is correctly located (i.e., has more than 50% intersection-over-union (“IoU”) with the ground-truth bounding region). The evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081]; and, --generating label graph embeddings from an ungrounded label graph. For example, act 702 involves generating label graph embeddings from connected entity labels in an ungrounded label graph corresponding to a digital image. Act 702 can involve encoding features from an entity label of the ungrounded label graph into a label graph embedding utilizing a label embedding model. For example, the label embedding model comprises a multilayer perceptron network. Alternatively, the label embedding model comprises a graph neural network. Accordingly, act 702 can involve encoding, utilizing the graph neural network, information from an entity label and relationship information associated with one or more entity relationships involving the entity label into a label graph embedding according to one or more confidence scores for the one or more entity relationships.--, in [0104]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in page 154). Re Claim 4, XU as modified by Rahman further disclose wherein each object detection defines an object location and an object extent (see XU: e.g., --[0023] The disclosed graph matching system provides a number of advantages over conventional systems. For example, the graph matching system improves the efficiency of computing systems that train and/or implement scene graph generation neural networks for digital image processing. Specifically, in contrast to conventional systems that rely on expensive annotations of object locations and relations in digital images, the graph matching system utilizes a lightweight, weakly-supervised process for generating semantic scene graphs. More specifically, by utilizing a weakly-supervised approach with relaxed annotation requirements, the graph matching system provides more efficient scene graph generation while utilizing fewer computing resources and data verification time. In particular, the graph matching system is able to obtain entity/relation information for digital images from image descriptions (e.g., captions) using efficient natural language parsing models.--, in [0023]). Re Claim 12, XU as modified by Rahman further disclose applied to multiple scenes to obtain respective overall performance scores for the multiple scenes, the method further comprising using the overall performance scores to identify and mitigate a performance issue in the perception component (see XU: e.g., --the disclosed systems can reduce complexity and improve performance by utilizing an efficient first-order graph matching model optimized via contrastive learning to determine a correspondence between labels and objects portrayed in a digital image. Moreover, the disclosed systems can utilize this correspondence to learn parameters of a scene graph model. In this manner, the disclosed systems can efficiently and flexibly generate accurate semantic scene graphs for digital images.,--, in [0002], [0035]; and, --[0078] FIG. 5 illustrates an embodiment in which the graph matching system 102 utilizes a graph matching process to generate semantic scene graphs for training a scene graph generation neural network 500. In particular, the graph matching system 102 utilizes a digital image 502 and an image description 504 associated with the digital image 502 to generate a ground-truth semantic scene graph 506. For instance, the graph matching system 102 utilizes a graph matching process to determine a one-to-one correspondence between nodes in a visual graph generated from the digital image 502 and nodes in a label graph generated from the image description 504, as described above. The graph matching system 102 then generates the ground-truth semantic scene graph based on the one-to-one correspondences between the nodes, resulting in entity classes for entities in the digital image 502 and predicate classes for relationships between the entities based on the information in the image description 504. [0079] After generating the ground-truth semantic scene graph 506, the graph matching system 102 utilizes the scene graph generation neural network 500 to generate a predicted semantic scene graph 508 for the digital image 502. Specifically, the scene graph generation neural network 500 processes the digital image to predict object classes and predicate classes for entities in the digital image 502. The graph matching system 102 then determines a scene graph generation loss 510 for the predicted semantic scene graph 508. In particular, the graph matching system 102 determines the scene graph generation loss 510 by comparing the predicted semantic scene graph 508 to the ground-truth semantic scene graph 506. In some embodiments, the graph matching system 102 utilizes a cross-entropy loss L.sub.sgg as the scene graph generation loss 510 by comparing the predicted object classes and predicate classes in the predicted semantic scene graph 508 to the object classes and predicate classes included in the ground-truth semantic scene graph 506. [0080] In one or more embodiments, the graph matching system 102 then utilizes the scene graph generation loss 510 to modify the scene graph generation neural network 500. For instance, the graph matching system 102 updates parameters of the scene graph generation neural network 500 to reduce the scene graph generation loss 510 (e.g., by reducing distances between the object/predicate classes in the ground-truth semantic scene graph 506 and the predicted object/predicate classes in the predicted semantic scene graph 508 ). In at least some embodiments, the total loss for the graph matching system 102 and the scene graph generation neural network 500 is L=L.sub.gm+L.sub.sgg, which includes the contrastive loss for the weakly-supervised graph matching process and the scene graph generation loss.[ 0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process. For example, instance-level recall (R.sub.inst) indicates that a bounding region is correctly matched if the bounding region is matched with the correct node in the label graph and is correctly located (i.e., has more than 50% intersection-over-union (“IoU”) with the ground-truth bounding region). The evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]). Re Claim 13, XU as modified by Rahman further disclose herein the perception component is a trained machine learning component, mitigating the performance issue comprises re-training the perception component based on a subset of the multiple scenes selected based on their overall performance scores (see XU: e.g., --the disclosed systems can reduce complexity and improve performance by utilizing an efficient first-order graph matching model optimized via contrastive learning to determine a correspondence between labels and objects portrayed in a digital image. Moreover, the disclosed systems can utilize this correspondence to learn parameters of a scene graph model. In this manner, the disclosed systems can efficiently and flexibly generate accurate semantic scene graphs for digital images.,--, in [0002], [0035]; and, --[0078] FIG. 5 illustrates an embodiment in which the graph matching system 102 utilizes a graph matching process to generate semantic scene graphs for training a scene graph generation neural network 500. In particular, the graph matching system 102 utilizes a digital image 502 and an image description 504 associated with the digital image 502 to generate a ground-truth semantic scene graph 506. For instance, the graph matching system 102 utilizes a graph matching process to determine a one-to-one correspondence between nodes in a visual graph generated from the digital image 502 and nodes in a label graph generated from the image description 504, as described above. The graph matching system 102 then generates the ground-truth semantic scene graph based on the one-to-one correspondences between the nodes, resulting in entity classes for entities in the digital image 502 and predicate classes for relationships between the entities based on the information in the image description 504. [0079] After generating the ground-truth semantic scene graph 506, the graph matching system 102 utilizes the scene graph generation neural network 500 to generate a predicted semantic scene graph 508 for the digital image 502. Specifically, the scene graph generation neural network 500 processes the digital image to predict object classes and predicate classes for entities in the digital image 502. The graph matching system 102 then determines a scene graph generation loss 510 for the predicted semantic scene graph 508. In particular, the graph matching system 102 determines the scene graph generation loss 510 by comparing the predicted semantic scene graph 508 to the ground-truth semantic scene graph 506. In some embodiments, the graph matching system 102 utilizes a cross-entropy loss L.sub.sgg as the scene graph generation loss 510 by comparing the predicted object classes and predicate classes in the predicted semantic scene graph 508 to the object classes and predicate classes included in the ground-truth semantic scene graph 506. [0080] In one or more embodiments, the graph matching system 102 then utilizes the scene graph generation loss 510 to modify the scene graph generation neural network 500. For instance, the graph matching system 102 updates parameters of the scene graph generation neural network 500 to reduce the scene graph generation loss 510 (e.g., by reducing distances between the object/predicate classes in the ground-truth semantic scene graph 506 and the predicted object/predicate classes in the predicted semantic scene graph 508 ). In at least some embodiments, the total loss for the graph matching system 102 and the scene graph generation neural network 500 is L=L.sub.gm+L.sub.sgg, which includes the contrastive loss for the weakly-supervised graph matching process and the scene graph generation loss.[ 0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process. For example, instance-level recall (R.sub.inst) indicates that a bounding region is correctly matched if the bounding region is matched with the correct node in the label graph and is correctly located (i.e., has more than 50% intersection-over-union (“IoU”) with the ground-truth bounding region). The evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in page 154). Re Claim 14, XU as modified by Rahman further disclose applied to a time- sequence of multiple scenes to obtain respective overall performance scores for the multiple scenes, the method further comprising generating a graphical user interface that comprises a timeline of the overall performance scores and a visualization of the multiple scenes (see XU: e.g., --the disclosed systems can reduce complexity and improve performance by utilizing an efficient first-order graph matching model optimized via contrastive learning to determine a correspondence between labels and objects portrayed in a digital image. Moreover, the disclosed systems can utilize this correspondence to learn parameters of a scene graph model. In this manner, the disclosed systems can efficiently and flexibly generate accurate semantic scene graphs for digital images.,--, in [0002], [0035]; and, --[0078] FIG. 5 illustrates an embodiment in which the graph matching system 102 utilizes a graph matching process to generate semantic scene graphs for training a scene graph generation neural network 500. In particular, the graph matching system 102 utilizes a digital image 502 and an image description 504 associated with the digital image 502 to generate a ground-truth semantic scene graph 506. For instance, the graph matching system 102 utilizes a graph matching process to determine a one-to-one correspondence between nodes in a visual graph generated from the digital image 502 and nodes in a label graph generated from the image description 504, as described above. The graph matching system 102 then generates the ground-truth semantic scene graph based on the one-to-one correspondences between the nodes, resulting in entity classes for entities in the digital image 502 and predicate classes for relationships between the entities based on the information in the image description 504. [0079] After generating the ground-truth semantic scene graph 506, the graph matching system 102 utilizes the scene graph generation neural network 500 to generate a predicted semantic scene graph 508 for the digital image 502. Specifically, the scene graph generation neural network 500 processes the digital image to predict object classes and predicate classes for entities in the digital image 502. The graph matching system 102 then determines a scene graph generation loss 510 for the predicted semantic scene graph 508. In particular, the graph matching system 102 determines the scene graph generation loss 510 by comparing the predicted semantic scene graph 508 to the ground-truth semantic scene graph 506. In some embodiments, the graph matching system 102 utilizes a cross-entropy loss L.sub.sgg as the scene graph generation loss 510 by comparing the predicted object classes and predicate classes in the predicted semantic scene graph 508 to the object classes and predicate classes included in the ground-truth semantic scene graph 506. [0080] In one or more embodiments, the graph matching system 102 then utilizes the scene graph generation loss 510 to modify the scene graph generation neural network 500. For instance, the graph matching system 102 updates parameters of the scene graph generation neural network 500 to reduce the scene graph generation loss 510 (e.g., by reducing distances between the object/predicate classes in the ground-truth semantic scene graph 506 and the predicted object/predicate classes in the predicted semantic scene graph 508 ). In at least some embodiments, the total loss for the graph matching system 102 and the scene graph generation neural network 500 is L=L.sub.gm+L.sub.sgg, which includes the contrastive loss for the weakly-supervised graph matching process and the scene graph generation loss.[ 0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process. For example, instance-level recall (R.sub.inst) indicates that a bounding region is correctly matched if the bounding region is matched with the correct node in the label graph and is correctly located (i.e., has more than 50% intersection-over-union (“IoU”) with the ground-truth bounding region). The evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in page 154). Re Claim 15, XU as modified by Rahman further disclose wherein each object detection comprises a bounding box or other bounding object defining an object location and an object extent (see XU: e.g., Fig. 1, and, -- [0031] Additionally, in one or more embodiments, a visual graph includes a set of nodes corresponding to positions of entities in a digital image. In particular, a visual graph includes a set of entity bounding regions that define positions of entities for specific portions (e.g., groups of pixels) of a digital image. For example, a visual graph includes a bounding box (or other shape) that encloses a portion of a digital image in which a particular entity is located. Furthermore, in one or more embodiments, a visual graph excludes entity classes and predicate classes corresponding to entities in a digital image. [0032] According to one or more embodiments, an embedding model includes a computer representation that encodes data into one or more digital embeddings (e.g., from a dimensional space to a lower dimensional space). For example, an embedding model converts data into a vector representation (e.g., feature vectors). Thus, in one or more embodiments, a label embedding model encodes information from an entity label of a label graph into a label graph embedding, which represents the information from the entity label in a feature representation (e.g., a feature vector in a different dimensional space). Additionally, in one or more embodiments, a visual embedding model encodes information from an entity bounding region of a visual graph into a visual graph embedding, which represents the information from the entity bounding region in a different dimensional space. In some embodiments, an embedding model includes a neural network with learnable parameters for encoding features of visual or textual data.--, in [0031]-[0032], and [0047]-[0049]; and, Fig. 4, in [0056], and [0062]). Re Claims 17,19 and 21, claims 17, 19 and 21 are the corresponding system claims to claims 1-2, and 4 respectively. Claims 17, 19 and 21 thus are rejected for similar reasons for claims 1-2, and 4. See above discussions about claims 1-2, and 4 respectively. Furthermore, XU as modified by Rahman further disclose computer system for assessing performance of an object detector on a scene, the computer system comprising: at least one memory storing computer-readable instructions; and at least one processor coupled to the at least one memory and configured to execute the computer-readable instructions, which upon execution cause the at least one processor to implement operations (see Xu: e.g., -- 0094] In some embodiments, the components of the graph matching system 102 include software, hardware, or both. For example, the components of the graph matching system 102 include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the computing device(s) 600 ). When executed by the one or more processors, the computer-executable instructions of the graph matching system 102 can cause the computing device(s) 600 to perform the operations described herein. Alternatively, the components of the graph matching system 102 can include hardware, such as a special purpose processing device to perform a certain function or group of functions. Additionally, or alternatively, the components of the graph matching system 102 can include a combination of computer-executable instructions and hardware. [0095] Furthermore, the components of the graph matching system 102 performing the functions described herein with respect to the graph matching system 102 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and/or as a cloud-computing model. Thus, the components of the graph matching system 102 may be implemented as part of a stand-alone application on a personal computing device or a mobile device.--, in [0094]-[0095]). Claims 3, 5-11, 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over XU as modified by Rahman, and further in view of TAN (US 20230005173 A). Re Claim 3, XU as modified by Rahman further disclose wherein each pseudo-ground truth output comprises either a positive existence indicator or a negative existence indicator, wherein the performance score for each pseudo-ground truth set is a perception hardness score, evaluated based on one or both of false positive detections and false negative detections with respect to that pseudo-ground truth set (see Xu: e.g., --[0018] According to one or more embodiments, after generating a semantic scene graph for a digital image, the graph matching system further trains one or more of the embedding models based on the semantic scene graph. Specifically, the graph matching system utilizes contrastive learning to compare label graph embeddings to positive and negative visual graph embeddings. The graph matching system then modifies parameters of one or more of the embedding models such that the label graph embeddings are closer to positive visual graph embeddings (e.g., reduced distance metrics) and further from negative visual graph embeddings (e.g., increased distance metrics). In additional embodiments, the graph matching system utilizes contrastive learning that compares visual graph embeddings to positive and negative samples and modifies parameters of the embedding models accordingly. [0019] In additional embodiments, the graph matching system generates a semantic scene graph for training a scene graph generation neural network. For example, the graph matching system aligns an ungrounded label graph and a visual graph to generate a ground-truth semantic scene graph for a digital image. The graph matching system then utilizes the ground-truth semantic scene graph to determine a scene graph generation loss by comparing the ground-truth semantic scene graph to a semantic scene graph generated by the scene graph generation neural network. Additionally, the graph matching system modifies parameters of the scene graph generation neural network based on the scene graph generation loss.--, in [0018]-[0019]; and, -- the graph matching system 102 determines the entity bounding regions of the visual graph 220 for entities by utilizing an object detection neural network or other image processing models. Specifically, the graph matching system 102 determines pixel regions corresponding to detected entities. Additionally, in some embodiments, the graph matching system 102 determines an entity bounding region by determining a set of pixels of the digital image 200 that encompasses a detected entity. For example, the entity bounding region includes a bounding box with a minimum size to include the detected entity. Alternatively, the entity bounding region includes a different shape or size for encompassing the detected entity. --, in [0048]; also see Rahman: Fig. 1, and, “mAP is computed for the ten frames at a time. The dashed line represents a predefined critical threshold. We can see that mAP drops below this threshold from time to time. The second row shows some samples from the low mAP regions. Green and Cyan boxes represent false negative and false positive errors made by the object detector.” in caption); XU as modified by Rahman however still do not explicitly disclose false positive detections are object detections whose confidence scores satisfy a minimum confidence threshold but which have a negative existence indicator in that pseudo-ground truth set, wherein false negative detections are object detections whose confidence scores do not satisfy the minimum confidence threshold but which have a positive existence indicator of that pseudo-ground truth set; TAN discloses false positive detections are object detections whose confidence scores satisfy a minimum confidence threshold but which have a negative existence indicator in that pseudo-ground truth set, wherein false negative detections are object detections whose confidence scores do not satisfy the minimum confidence threshold but which have a positive existence indicator of that pseudo-ground truth set (see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]); XU (as modified by Rahman) and TAN are combinable as they are in the same field of endeavor: evaluation of the performance of object detection/matching neural network model. Therefore it would have been obvious to one of ordinary skill in the art at the time the invention was made to modify XU’s method using Tan’s teachings by including false positive detections are object detections whose confidence scores satisfy a minimum confidence threshold but which have a negative existence indicator in that pseudo-ground truth set, wherein false negative detections are object detections whose confidence scores do not satisfy the minimum confidence threshold but which have a positive existence indicator of that pseudo-ground truth set to XU (as modified by Rahman)’s performance computing in order to evaluate the machine learning statistical measures (see Tan: e.g., in [0128]-[0129], [0138]-[0139], and [0150]). Re Claim 5, XU as modified by Rahman further disclose wherein each object detection defines an object location and an object extent, and wherein each pseudo-ground truth output comprises either a positive existence indicator or a negative existence indicator (see XU: e.g., -- [0018] According to one or more embodiments, after generating a semantic scene graph for a digital image, the graph matching system further trains one or more of the embedding models based on the semantic scene graph. Specifically, the graph matching system utilizes contrastive learning to compare label graph embeddings to positive and negative visual graph embeddings. The graph matching system then modifies parameters of one or more of the embedding models such that the label graph embeddings are closer to positive visual graph embeddings (e.g., reduced distance metrics) and further from negative visual graph embeddings (e.g., increased distance metrics). In additional embodiments, the graph matching system utilizes contrastive learning that compares visual graph embeddings to positive and negative samples and modifies parameters of the embedding models accordingly.--, in [0018]; and, --[0023] The disclosed graph matching system provides a number of advantages over conventional systems. For example, the graph matching system improves the efficiency of computing systems that train and/or implement scene graph generation neural networks for digital image processing. Specifically, in contrast to conventional systems that rely on expensive annotations of object locations and relations in digital images, the graph matching system utilizes a lightweight, weakly-supervised process for generating semantic scene graphs. More specifically, by utilizing a weakly-supervised approach with relaxed annotation requirements, the graph matching system provides more efficient scene graph generation while utilizing fewer computing resources and data verification time. In particular, the graph matching system is able to obtain entity/relation information for digital images from image descriptions (e.g., captions) using efficient natural language parsing models.--, in [0023]), the method comprising: for each pseudo-ground truth set: generating for each positive existence indicator, a pseudo-ground truth object that defines an object location and object extent (see XU: e.g., --the graph matching system generates a semantic scene graph for training a scene graph generation neural network. For example, the graph matching system aligns an ungrounded label graph and a visual graph to generate a ground-truth semantic scene graph for a digital image. The graph matching system then utilizes the ground-truth semantic scene graph to determine a scene graph generation loss by comparing the ground-truth semantic scene graph to a semantic scene graph generated by the scene graph generation neural network. Additionally, the graph matching system modifies parameters of the scene graph generation neural network based on the scene graph generation loss.--, in [0019]; and, --[0053] In one or more additional embodiments, the graph matching system 102 generates the semantic scene graph 304 based on the correspondences between the visual graph and the ungrounded label graph. …. the semantic scene graph 304 includes a scene graph that serves as a pseudo ground-truth semantic scene graph for training the scene graph generation neural network 306. More specifically, the scene graph generation neural network 306 generates semantic scene graphs from digital images for performing visual reasoning tasks based on, for example, scene construction, object detection, or object relationships. Accordingly, the graph matching system 102 utilizes the semantic scene graph 304 to train the scene graph generation neural network 306 to more accurately generate a semantic scene graph for the digital image 300… [0055] FIG. 4 illustrates a detailed diagram of the graph matching system 102 utilizing graph matching to generate semantic scene graphs for digital images. In particular, the graph matching system 102 utilizes a plurality of operations in the graph matching process for generating a semantic scene graph based on a digital image 400 and an image description 402 associated with the digital image 400. Additionally, FIG. 4 illustrates that the graph matching system 102 utilizes contrastive learning to learn embedding models for more accurate graph matching. [0056] As illustrated in FIG. 4, the graph matching system 102 generates a visual graph 404 of the digital image 400.--, in 0053]-[0056]), and attempting to associate each object detection with a pseudo-ground truth object based on relative intersection therebetween(see XU: e.g., --the graph matching system 102 utilizes a graph neural network that explicitly encodes edge type features into the node representation. For instance, an initial node state is an input node embedding h.sub.i.sup.0=e.sub.i. At the kth iteration, the graph neural network generates a confidence score to measure the confidence of whether an edge exists pointing from node i to node j…the graph matching system 102 augments the predicate attention with edge confidence such that it attends to the background class, and obtains the attended predicate representation from the augmented attention….the graph matching system 102 formulates the above message passing information with soft attention to the predicate type, suitable for the visual graph (which excludes the predicate category). Because the label graph has the determined relation type, the graph matching system 102 adopts hard attention instead of soft attention. This results in β.sub.ij in the measured confidence score--, in [0073]-[0076]; and, --[0080] In one or more embodiments, the graph matching system 102 then utilizes the scene graph generation loss 510 to modify the scene graph generation neural network 500. For instance, the graph matching system 102 updates parameters of the scene graph generation neural network 500 to reduce the scene graph generation loss 510 (e.g., by reducing distances between the object/predicate classes in the ground-truth semantic scene graph 506 and the predicted object/predicate classes in the predicted semantic scene graph 508 ). In at least some embodiments, the total loss for the graph matching system 102 and the scene graph generation neural network 500 is L=L.sub.gm+L.sub.sgg, which includes the contrastive loss for the weakly-supervised graph matching process and the scene graph generation loss. [0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process. For example, instance-level recall (R.sub.inst) indicates that a bounding region is correctly matched if the bounding region is matched with the correct node in the label graph and is correctly located (i.e., has more than 50% intersection-over-union (“IoU”) with the ground-truth bounding region). The evaluation determined the instance-level recall as the ratio of the correctly matched bounding regions to all of the ground-truth bounding boxes for each image, and the overall recall is averaged across all images.--, in [0078]-[0081]; and, --generating label graph embeddings from an ungrounded label graph. For example, act 702 involves generating label graph embeddings from connected entity labels in an ungrounded label graph corresponding to a digital image. Act 702 can involve encoding features from an entity label of the ungrounded label graph into a label graph embedding utilizing a label embedding model. For example, the label embedding model comprises a multilayer perceptron network. Alternatively, the label embedding model comprises a graph neural network. Accordingly, act 702 can involve encoding, utilizing the graph neural network, information from an entity label and relationship information associated with one or more entity relationships involving the entity label into a label graph embedding according to one or more confidence scores for the one or more entity relationships.--, in [0104]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in page 154); wherein the performance score for each pseudo-ground truth set is a perception hardness score, evaluated based on one or both of false positive detections and false negative detections with respect to that pseudo-ground truth set (see Xu: e.g., --[0018] According to one or more embodiments, after generating a semantic scene graph for a digital image, the graph matching system further trains one or more of the embedding models based on the semantic scene graph. Specifically, the graph matching system utilizes contrastive learning to compare label graph embeddings to positive and negative visual graph embeddings. The graph matching system then modifies parameters of one or more of the embedding models such that the label graph embeddings are closer to positive visual graph embeddings (e.g., reduced distance metrics) and further from negative visual graph embeddings (e.g., increased distance metrics). In additional embodiments, the graph matching system utilizes contrastive learning that compares visual graph embeddings to positive and negative samples and modifies parameters of the embedding models accordingly. [0019] In additional embodiments, the graph matching system generates a semantic scene graph for training a scene graph generation neural network. For example, the graph matching system aligns an ungrounded label graph and a visual graph to generate a ground-truth semantic scene graph for a digital image. The graph matching system then utilizes the ground-truth semantic scene graph to determine a scene graph generation loss by comparing the ground-truth semantic scene graph to a semantic scene graph generated by the scene graph generation neural network. Additionally, the graph matching system modifies parameters of the scene graph generation neural network based on the scene graph generation loss.--, in [0018]-[0019]; and, -- the graph matching system 102 determines the entity bounding regions of the visual graph 220 for entities by utilizing an object detection neural network or other image processing models. Specifically, the graph matching system 102 determines pixel regions corresponding to detected entities. Additionally, in some embodiments, the graph matching system 102 determines an entity bounding region by determining a set of pixels of the digital image 200 that encompasses a detected entity. For example, the entity bounding region includes a bounding box with a minimum size to include the detected entity. Alternatively, the entity bounding region includes a different shape or size for encompassing the detected entity. --, in [0048]; also see Rahman: Fig. 1, and, “mAP is computed for the ten frames at a time. The dashed line represents a predefined critical threshold. We can see that mAP drops below this threshold from time to time. The second row shows some samples from the low mAP regions. Green and Cyan boxes represent false negative and false positive errors made by the object detector.” in caption); XU as modified by Rahman however still do not explicitly disclose wherein false positive detections are object detections whose confidence scores satisfy a minimum confidence threshold but which are not successfully associated with any pseudo-ground truth object of that pseudo-ground truth set, wherein false negative detections are object detections whose confidence scores do not satisfy the minimum confidence threshold but which have been successfully associated with a pseudo-ground truth object of that pseudo-ground truth set; TAN discloses wherein false positive detections are object detections whose confidence scores satisfy a minimum confidence threshold but which are not successfully associated with any pseudo-ground truth object of that pseudo-ground truth set, wherein false negative detections are object detections whose confidence scores do not satisfy the minimum confidence threshold but which have been successfully associated with a pseudo-ground truth object of that pseudo-ground truth set (see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]); XU (as modified by Rahman) and TAN are combinable as they are in the same field of endeavor: evaluation of the performance of object detection/matching neural network model. Therefore it would have been obvious to one of ordinary skill in the art at the time the invention was made to modify XU’s method using Tan’s teachings by including wherein false positive detections are object detections whose confidence scores satisfy a minimum confidence threshold but which are not successfully associated with any pseudo-ground truth object of that pseudo-ground truth set, wherein false negative detections are object detections whose confidence scores do not satisfy the minimum confidence threshold but which have been successfully associated with a pseudo-ground truth object of that pseudo-ground truth set to XU (as modified by Rahman)’s performance computing in order to evaluate the machine learning statistical measures (see Tan: e.g., in [0128]-[0129], [0138]-[0139], and [0150]). Re Claim 6, XU as modified by Rahman and TAN further disclose wherein the performance score for each pseudo-ground truth set is: a count of false positive detections for that pseudo-ground truth set (see XU: e.g., --[0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process.--, in [0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detec tors.”, in page 154; and see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]), a count of false negative detections for that pseudo-ground truth set (see XU: e.g., --[0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process.--, in [0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detec tors.”, in page 154; and see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]), or a count of both false positive and false negative detections for that pseudo-ground truth set (see XU: e.g., --[0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process.--, in [0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detec tors.”, in page 154; and see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]). Re Claim 7, XU as modified by Rahman and TAN further disclose the performance score for each pseudo-ground truth set comprises computing, for each object detection of an error set, a weighted error, which is an object size as a fraction of a size of the scene, wherein the performance score is computed by summing the weighted errors, and wherein the error set consists of all false positive detections for that pseudo-ground truth set, all false negative detections for that pseudo-ground truth set, or all false positive detections and all false negative detections for that pseudo-ground truth set (see XU: e.g., --[0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process.--, in [0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detec tors.”, in page 154; and see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]). Re Claim 8, XU as modified by Rahman further disclose wherein the scene is a 2D image, wherein each object detection comprises a 2D bounding object defining the object location and the object extent, and wherein the performance score is computed as: PixelAdj,(e(ý,y))= Σ bEe(y.y) where y denotes the pseudo-ground truth set, ý denotes the set of object detections, e(ý,y) denotes the error set, X denotes the scene, and b denotes a 2D bounding object (see XU: e.g., Fig. 1, and, -- [0031] Additionally, in one or more embodiments, a visual graph includes a set of nodes corresponding to positions of entities in a digital image. In particular, a visual graph includes a set of entity bounding regions that define positions of entities for specific portions (e.g., groups of pixels) of a digital image. For example, a visual graph includes a bounding box (or other shape) that encloses a portion of a digital image in which a particular entity is located. Furthermore, in one or more embodiments, a visual graph excludes entity classes and predicate classes corresponding to entities in a digital image. [0032] According to one or more embodiments, an embedding model includes a computer representation that encodes data into one or more digital embeddings (e.g., from a dimensional space to a lower dimensional space). For example, an embedding model converts data into a vector representation (e.g., feature vectors). Thus, in one or more embodiments, a label embedding model encodes information from an entity label of a label graph into a label graph embedding, which represents the information from the entity label in a feature representation (e.g., a feature vector in a different dimensional space). Additionally, in one or more embodiments, a visual embedding model encodes information from an entity bounding region of a visual graph into a visual graph embedding, which represents the information from the entity bounding region in a different dimensional space. In some embodiments, an embedding model includes a neural network with learnable parameters for encoding features of visual or textual data.--, in [0031]-[0032], and [0047]-[0049]; and, Fig. 4, in [0056], and [0062]; and, also see: see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]). Re Claim 9, XU as modified by Rahman and TAN further disclose computing the performance score for each pseudo-ground truth object set comprises computing, for each object detection of an error set, an occlusion value, which is a measure of intersection between the object detection and any true positive detection as a fraction of object size, wherein the performance score is computed by summing the occlusion values, and wherein the error set consists of all false positive detections for that pseudo- ground truth set, all false negative detections for that pseudo-ground truth set, or all false positive detections and all false negative detections for that pseudo-ground truth set, true positives being detections whose confidence score satisfies the minimum confidence threshold and which have a positive existence indicator in that pseudo-ground truth set or which have been successfully associated with a pseudo-ground truth object of that pseudo- ground truth set (see XU: e.g., --[0081] Experimenters have conducted one or more evaluations (hereinafter, “the evaluation”) of embodiments of the graph matching system 102 relative to existing systems for generating semantic scene graphs for a dataset of images with scene graph annotations. Specifically, experimenters evaluated different label preprocessing strategies for a set of frequent object categories and predicate types. Experimenters identified differences in an instance-level recall, an object-level recall, and a predicate-level recall as metrics to measure the performance of the graph matching process.--, in [0081], Table 2, and Table. 3 show the performance of the example embodiment of the scene graph generation neural network with and without scene graph constraint, respectively. And, --Additionally, Table 7 illustrates comparisons with the original scene graph cut into sub-graphs (at 10% and 50% of nodes remaining). Fewer nodes resulted in decreased performance, indicating that the one-to-one mapping constraint is more important to improved performance than errors due to mismatch propagation.--, in [0089]; also see Rahman: e.g., Fig. 1, “mAP is computed for the ten frames at a time.” In caption, and, -- Performance monitoring of object detection… the performance fluctuates as a function of the deployment conditions.--, in abstract, and, -- The standard practice to prepare an object detection model for deployment is to train and evaluate the model using training and evaluation split of some dataset to measure the accuracy and generalization capacity. Here, the assumption is the training and evaluation data are representative of the real operating environment. However, this assumption does not hold in the context of autonomous vehicles where the operating environment is continuously evolving and might change unexpectedly. Consequently, object detection performance fluctuates without any prior notification. Moreover, the performance might drop below any critical threshold, which can cause a fatal incident. See Figure 1 for an overview. One possible solution is to develop an exceptionally accurate and domain adaptive object detection system for autonomous vehicles. However, it is impossible in most practical circumstances to account for all imaginable future deployment conditions during training. Another approach is to identify when the performance of the deployed object detector drops below a critical threshold. So without the need to increase the detection accuracy directly, a performance drop identifier can protect the autonomous vehicle by providing crucial alerts during periods of silent failure. However, measuring the performance drop directly during deployment is impractical due to the absence of ground-truth data in this phase. Therefore, we advocate equipping object detectors with self-assessment capability to detect instances of performance drop during deployment.--, in right col. of page 152; also see: “the use of per-frame mAP prediction for continuous performance monitoring of object detectors.”, in left col. page 153, and, “the use of per-frame mAP prediction for continuous performance monitoring of object detec tors.”, in page 154; and see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]). . Re Claim 10, XU as modified by Rahman and TAN further disclose wherein the scene is a 2D image, wherein each object detection comprises a 2D bounding object defining the object location and the object extent, and wherein the performance score is computed as: OccAwarex(e(y,y)) = = Σ inter(b,b) 1 bEe(yiy) b'Etp(x) where y denotes the pseudo-ground truth set, y denotes the set of object detections, e(ý,y) denotes the error set for the pseudo-ground truth object set y, X denotes the scene, b denotes a 2D bounding object, and tp(x) denotes the set of all true positives (see XU: e.g., Fig. 1, and, -- [0031] Additionally, in one or more embodiments, a visual graph includes a set of nodes corresponding to positions of entities in a digital image. In particular, a visual graph includes a set of entity bounding regions that define positions of entities for specific portions (e.g., groups of pixels) of a digital image. For example, a visual graph includes a bounding box (or other shape) that encloses a portion of a digital image in which a particular entity is located. Furthermore, in one or more embodiments, a visual graph excludes entity classes and predicate classes corresponding to entities in a digital image. [0032] According to one or more embodiments, an embedding model includes a computer representation that encodes data into one or more digital embeddings (e.g., from a dimensional space to a lower dimensional space). For example, an embedding model converts data into a vector representation (e.g., feature vectors). Thus, in one or more embodiments, a label embedding model encodes information from an entity label of a label graph into a label graph embedding, which represents the information from the entity label in a feature representation (e.g., a feature vector in a different dimensional space). Additionally, in one or more embodiments, a visual embedding model encodes information from an entity bounding region of a visual graph into a visual graph embedding, which represents the information from the entity bounding region in a different dimensional space. In some embodiments, an embedding model includes a neural network with learnable parameters for encoding features of visual or textual data.--, in [0031]-[0032], and [0047]-[0049]; and, Fig. 4, in [0056], and [0062]; and, also see: see TAN: e.g., --[0138] Machine learning statistical measures are implemented to determine error based uncertainties. Machine learning statistical measures include, but are not limited to, false positive false negative (FP+FN), precision, recall, and F1 score, or any combinations thereof. Generally, the machine learning statistical measures are based on a ground truth compared with a prediction. In evaluating the machine learning statistical measures, a false positive (FP) is an error that indicates a condition exists when it actually does not exist. A false negative (FN) is an error that incorrectly indicates that a condition does not exist. A true positive is a correctly indicated positive condition, and a true negative is a correctly indicated negative condition. Accordingly, for a FP+FN statistical measure, a false positive is a prediction that does not have a sufficiently high IOU with any ground truth box. When a confidence score of a detection that is to detect a ground-truth is lower than a predetermined threshold, a false negative occurs. The number of FP+FN errors are counted between the pseudo-ground truth modality predictions and other modality predictions. [0139] Generally, the precision is the number of true positives divided by the sum of true positives and false positives. Subtracting the precision from one results in an active learning score where the higher the precision, the lower the resulting inconsistency. A lower inconsistency computation indicates that the associated projections are consistent. Similarly, the recall is the number of true positives divided by the sum of true positives and false positives. Subtracting the recall from one results in an active learning score where the higher the recall, the lower the inconsistency computation. An F1 score is a balanced F-score and is the harmonic mean of precision and recall. In an embodiment, the F1 score is a measure of accuracy. Accuracy is the probability that a randomly chosen instance (positive or negative, relevant or irrelevant) will be correct. --, in [0138]-[0139], and [1050];also see: -- [0128] For example, the modification determines an IoU for convex polygons to account for rotations between the projected bounding boxes. In the modified IoU determination, the proposal with the highest confidence (e.g., box A) is iteratively selected from a list of projected bounding boxes (e.g., list of B) added to the final list of projections (e.g., list F). The IoU of box A with all the proposals in the list of B is found, and again the boxes which have an IoU higher than the IoU threshold are removed. This process is repeated until there are no more proposals left in in the list of projected bounding boxes (B). When determining the IoU between the box A and the bounding boxes in the list of B, all corners of box A that are contained in box B are found. All corners of box B that are contained in box A are found. Intersection points between box A and box B are found, and all points are sorted in a clockwise manner using arctan2. [0129] In an embodiment, post processing 1318 and post processing 1320 enable post-processing for a heatmap representation. If more than one box is assigned to the same cell of the heatmap, the box with the highest confidence score is selected as the final bounding box associated with that cell. In this manner, the projected bounding boxes that do not satisfy a threshold for the highest confidence score are removed.--, in [1028]-[0129]). Re Claim 11, XU as modified by Rahman and TAN further disclose wherein each object detection comprises an object class, and the object detections are classified as false positive or false negatives with respect to a particular object class (see Xu: e.g., --[0031] Additionally, in one or more embodiments, a visual graph includes a set of nodes corresponding to positions of entities in a digital image. In particular, a visual graph includes a set of entity bounding regions that define positions of entities for specific portions (e.g., groups of pixels) of a digital image. For example, a visual graph includes a bounding box (or other shape) that encloses a portion of a digital image in which a particular entity is located. Furthermore, in one or more embodiments, a visual graph excludes entity classes and predicate classes corresponding to entities in a digital image.--, in [0030]-[0031], [0048]-[0049], [0068], [0077]-[0078], and, --[0079] After generating the ground-truth semantic scene graph 506, the graph matching system 102 utilizes the scene graph generation neural network 500 to generate a predicted semantic scene graph 508 for the digital image 502. Specifically, the scene graph generation neural network 500 processes the digital image to predict object classes and predicate classes for entities in the digital image 502. The graph matching system 102 then determines a scene graph generation loss 510 for the predicted semantic scene graph 508. In particular, the graph matching system 102 determines the scene graph generation loss 510 by comparing the predicted semantic scene graph 508 to the ground-truth semantic scene graph 506. In some embodiments, the graph matching system 102 utilizes a cross-entropy loss L.sub.sgg as the scene graph generation loss 510 by comparing the predicted object classes and predicate classes in the predicted semantic scene graph 508 to the object classes and predicate classes included in the ground-truth semantic scene graph 506.-- [0082] Additionally, the evaluation utilized (1) predicate classification (“PredCls”), (2) scene graph classification (“SGCls”), (3) scene graph detection (“SGGEN”), and (4) phrase detection (“PhrDet”). For predicate classification, given ground-truth object bounding regions and object labels, experimenters utilized the evaluated systems to predict relationship types of object pairs. For scene graph classification, given ground-truth bounding regions, the evaluated systems predicted object categories and relationship types. For scene graph detection, given an image, the evaluated systems predicted the bounding regions, categories of region proposals, and relation types of object pairs. A correctly detected entity was determined if the labels of the subject-relation-object triplet were correctly classified, and the bounding regions of subject and object had more than 50% IoU with the ground-truth. Also, for phrase detection, given an image, experimenters utilized the evaluated systems to predict the relationship triplet with a union bounding region enclosing both the object and subject. A correctly detected phrase was indicated if the labels of the triplet were correct and the union region match with the ground-truth union region with IoU greater than 0.5. Additionally, the evaluation computed the recall of the above metrics for each image and then averages over the dataset, leading to recall at K metrics (K−[20, 50, 100]). Moreover, in the triplet ranking process, the evaluation applied a constraint that the same object pair cannot predict multiple predicates in the default setting--, in [0079]-[0082], and [0086]). Re Claim 16, claim 16 is the corresponding medium claim to claims 1, and 3, respectively. Claim 16 thus is rejected for similar reasons for claims 1, and 3. See above discussions about claims 1, and 3 respectively. Furthermore, XU as modified by Rahman and TAN further disclose a non-transitory computer readable medium embodying computer program instructions, the computer program instructions configured so as when executed on one or more hardware processors, to implement operations (see Xu: e.g., -- 0094] In some embodiments, the components of the graph matching system 102 include software, hardware, or both. For example, the components of the graph matching system 102 include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the computing device(s) 600). When executed by the one or more processors, the computer-executable instructions of the graph matching system 102 can cause the computing device(s) 600 to perform the operations described herein. Alternatively, the components of the graph matching system 102 can include hardware, such as a special purpose processing device to perform a certain function or group of functions. Additionally, or alternatively, the components of the graph matching system 102 can include a combination of computer-executable instructions and hardware.--, in [0094]-[0095]). Re Claim 20, claim 20 is the corresponding system claim to claim 3, respectively. Claim 20 thus is rejected for similar reasons for claim 3. See above discussions about claim 3 respectively. Furthermore, XU as modified by Rahman and TAN further disclose computer system for assessing performance of an object detector on a scene, the computer system comprising: at least one memory storing computer-readable instructions; and at least one processor coupled to the at least one memory and configured to execute the computer-readable instructions, which upon execution cause the at least one processor to implement operations (see Xu: e.g., -- 0094] In some embodiments, the components of the graph matching system 102 include software, hardware, or both. For example, the components of the graph matching system 102 include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the computing device(s) 600 ). When executed by the one or more processors, the computer-executable instructions of the graph matching system 102 can cause the computing device(s) 600 to perform the operations described herein. Alternatively, the components of the graph matching system 102 can include hardware, such as a special purpose processing device to perform a certain function or group of functions. Additionally, or alternatively, the components of the graph matching system 102 can include a combination of computer-executable instructions and hardware. [0095] Furthermore, the components of the graph matching system 102 performing the functions described herein with respect to the graph matching system 102 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and/or as a cloud-computing model. Thus, the components of the graph matching system 102 may be implemented as part of a stand-alone application on a personal computing device or a mobile device.--, in [0094]-[0095]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIWEN YANG whose telephone number is (571)270-5670. The examiner can normally be reached on Monday-Friday 8:30am-4:30pm east. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WEI WEN YANG/Primary Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Jan 27, 2025
Application Filed
Aug 27, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749347
METHOD AND SYSTEM FOR IDENTIFYING A USER AND NON-TRANSITORY COMPUTER READABLE STORAGE MEDIUM
2y 8m to grant Granted Sep 29, 2026
Patent 12743775
BIOMARKERS OF COLLAGEN FIBER ARCHITECTURE IN EPITHELIAL OVERIAN CANCER (EOC) PATIENTS
2y 11m to grant Granted Sep 22, 2026
Patent 12737879
Machine Learning for Detection of Diseases from External Anterior Eye Images
3y 9m to grant Granted Sep 15, 2026
Patent 12737880
SYSTEMS AND METHODS OF ANALYZING MICROBIOMES USING ARTIFICIAL INTELLIGENCE
3y 4m to grant Granted Sep 15, 2026
Patent 12738057
CUT-PASTE TRAINING AUGMENTATION FOR MACHINE LEARNING MODELS
2y 7m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
93%
With Interview (+11.5%)
2y 5m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 684 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month