Prosecution Insights
Last updated: August 15, 2026
Application No. 18/438,449

MULTIPLE CAMERA AND MULTIPLE THREE-DIMENSIONAL OBJECT TRACKING ON THE MOVE FOR AUTONOMOUS VEHICLES

Non-Final OA §101§103§112
Filed
Feb 10, 2024
Priority
Feb 12, 2023 — provisional 63/444,971
Examiner
JONES, ANDREW B
Art Unit
2667
Tech Center
2600 — Communications
Assignee
The Board of Trustees of the University of Arkansas
OA Round
2 (Non-Final)
71%
Grant Probability
Favorable
2-3
OA Rounds
5m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 71% — above average
71%
Career Allowance Rate
59 granted / 83 resolved
+9.1% vs TC avg
Strong +17% interview lift
Without
With
+17.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
31 currently pending
Career history
108
Total Applications
across all art units

Statute-Specific Performance

§101
9.3%
-30.7% vs TC avg
§103
52.8%
+12.8% vs TC avg
§102
13.8%
-26.2% vs TC avg
§112
19.5%
-20.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 83 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendment filed 18 June, 2026 has been entered. The amendment of claims 1, 5, 6, 11, 15, 16, and 20 has been acknowledged. Response to Arguments Applicant’s arguments, see page 7, section “Rejections under 35 U.S.C. § 112(b)”, filed 18 June, 2026 with respect to the rejection of claims 1 - 20 have been fully considered and are persuasive. The rejection of claims 1 - 20 under 35 U.S.C. § 112(b) has been withdrawn. However, upon further examination, a new rejection is made for claim 20 under 35 U.S.C. § 112(b). Applicant’s arguments, see page 7, section “Rejections under 35 U.S.C. § 103”, filed 18 June, 2026 with respect to the rejection of claims 1 - 19 have been fully considered and are persuasive. The rejection of claims 1 - 19 under 35 U.S.C. § 103 has been withdrawn. However, upon further examination, a new rejection is made for claims 1 – 20 under 35 U.S.C. § 103. Applicant’s arguments, see page 7, section “Rejections under 35 U.S.C. § 112(b)”, filed 18 June, 2026 with respect to the rejection of claims 1 - 19 have been fully considered and are persuasive. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 20 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 20 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being incomplete for omitting essential steps, such omission amounting to a gap between the steps. See MPEP § 2172.01. Claim 20 states on line 5 “receiving detection outcomes”, and later states on line 7 and 9 “responsive to the receiving, maintaining a graph comprising nodes (Line 7)… wherein the nodes represent received tracked objects (line 9, emphasis added)…” however there has been no step of receiving “tracked objects”. The limitation of line 5 receives “detection outcomes” which is distinctly different from “tracked objects” as there is nothing in this claim which ties them together. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1 – 20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. When reviewing independent claim 20, and based upon consideration of all of the relevant factors with respect to the claim as a whole, claim 20 is held to claim an abstract idea without reciting elements that amount to significantly more than the abstract idea and is/are therefore rejected as ineligible subject matter under 35 U.S.C. 101. The Examiner will analyze claim 20, and similar rationale applies to independent claims 1 and 11. The rationale, under MPEP § 2106, for this finding is explained below: The claimed invention (1) must be directed to one of the four statutory categories, and (2) must not be wholly directed to subject matter encompassing a judicially recognized exception, as defined below. The following two step analysis is used to evaluate these criteria. Step 1: Is the claim directed to one of the four patent-eligible subject matter categories: process, machine, manufacture, or composition of matter? When examining the claim under 35 U.S.C. 101, the Examiner interprets that the claims is related to a process since the claim is directed to a method of three-dimensional object tracking across cameras. Step 2a, Prong 1: Does the claim wholly embrace a judicially recognized exception, which includes laws of nature, physical phenomena, and abstract ideas, or is it a particular practical application of a judicial exception? The Examiner interprets that the judicial exception applies since claim 1 limitations of “receiving detection outcomes generated by a three-dimensional object detector from a plurality of synchronized camera inputs”, “responsive to the receiving, maintaining a graph comprising nodes and weighted edges between at least a portion of the nodes, wherein the nodes represent received tracked objects and comprise appearance features and motion features, and wherein the weighted edges are computed based on node similarity”, “executing appearance modeling of the nodes via a self-attention layer of a graph transformer network to yield resultant appearance-modeling data”, and “executing motion modeling of the nodes using the resultant appearance- modeling data via a cross-attention layer of the graph transformer network, wherein the motion modeling yields resultant motion-modeling data for tracking the tracked objects across the plurality of synchronized camera inputs.” are directed to an abstract idea. The claim is related to mathematical concept by each step comprising actions which amount to simple data transformations or decisions made based on mathematic computations. If the claim recites a judicial exception (i.e., an abstract idea enumerated in MPEP § 2106.04(a), a law of nature, or a natural phenomenon), the claim requires further analysis in Prong Two. Step 2a, Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? The Examiner interprets that claim 1 limitations do not provide additional elements or combination of additional elements to a practical application since the claim is generally linking the use of the judicial exception to a particular technological environment or field of use – see MPEP 2106.05(h). See, MPEP §2106.04(a), Because a judicial exception is not eligible subject matter, Bilski, 561 U.S. at 601, 95 USPQ2d at 1005-06 (quoting Chakrabarty, 447 U.S. at 309, 206 USPQ at 197 (1980)), if there are no additional claim elements besides the judicial exception, or if the additional claim elements merely recite another judicial exception, that is insufficient to integrate the judicial exception into a practical application. See, e.g., RecogniCorp, LLC v. Nintendo Co., 855 F.3d 1322, 1327, 122 USPQ2d 1377 (Fed. Cir. 2017) ("Adding one abstract idea (math) to another abstract idea (encoding and decoding) does not render the claim non-abstract"). OR Genetic Techs. v. Merial LLC, 818 F.3d 1369, 1376, 118 USPQ2d 1541, 1546 (Fed. Cir. 2016) (eligibility "cannot be furnished by the unpatentable law of nature (or natural phenomenon or abstract idea) itself."). For a claim reciting a judicial exception to be eligible, the additional elements (if any) in the claim must "transform the nature of the claim" into a patent-eligible application of the judicial exception, Alice Corp., 573 U.S. at 217, 110 USPQ2d at 1981, either at Prong Two or in Step 2B. If there are no additional elements in the claim, then it cannot be eligible. In such a case, after making the appropriate rejection (see MPEP § 2106.07 for more information on formulating a rejection for lack of eligibility), it is a best practice for the examiner to recommend an amendment, if possible, that would resolve eligibility of the claim. Step 2b: If a judicial exception into a practical application is not recited in the claim, the Examiner must interpret if the claim recites additional elements that amount to significantly more than the judicial exception. The Examiner interprets that the claims do not amount to significantly more since the claim only comprises steps which fall within the judicial exception Furthermore, the generic computer components memory and a processor recited as performing generic computer functions that are well-understood, routine and conventional activities amount to no more than implementing the abstract idea with a computerized system. Claims 2 – 10 and 12 - 19 depending on the independent claims 1 and 11 include all the limitation of the independent claims. The Examiner finds that claim 2 – 10 and 12 - 19 does not state significantly more since the claim only recites “wherein the nodes represent tracked objects comprising at least one of appearance features or motion features.” in claims 2 and 12, “wherein the weighted edges are computed based at least in part on node similarity.” in claims 3 and 13, “wherein the node similarity is computed based on at least one of appearance similarity or location similarity between the tracked objects.” in claims 4 and 14, “wherein the self-attention layer provides output embeddings.” in claims 5 and 15, “wherein the tracked objects are tracklets.” in claims 6 and 16, “wherein the motion modeling yields resultant motion-modeling data.” in claims 7 and 17, “comprising post-processing the resultant motion-modeling data via motion propagation and node merging.” in claims 8 and 18, “wherein the post-processing comprises adding a node to the graph via link prediction.” in claims 9, “wherein the post-processing comprises removing a node from the graph via link prediction.” in claim 10, and “herein the post-processing comprises at least one of adding a node to the graph or removing a node from the graph via link prediction.” in claim 19. Thus, claims 2 – 10 and 12 - 19 recite the same abstract idea and therefore are not drawn to the eligible subject matter as they are directed to the abstract idea without significantly more. Therefore, the Examiner interprets that the claims are rejected under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3, 5, 7 – 11, 13, 15, and 17 – 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ramezani et al (U.S. Patent No. 11361449 B2, hereinafter “Ramezani”) in view of Chu et al (P. Chu "TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking," 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2023, pp. 4859-4869, doi: 10.1109/WACV56688.2023.00485., hereinafter “Chu”). Regarding claim 1, Ramezani teaches a method of three-dimensional object tracking across cameras, comprising, by a computer system: receiving detection outcomes generated by a three-dimensional object detector from a plurality of synchronized camera inputs (Col. 4, Lines 45 – 48: Referring to FIG. 1A, a sensor system 10 can capture scenes S1, S2, S3 of the environment 12 at times t1, t2, and t3 respectively, and generate a sequence of images I1, I2, I3, etc.; Col 4, Lines 54 – 57: In some implementations… each point has certain coordinates in a 3D space… ; Col. 4, Lines 61 – 64: After the sensor system 10 generates the images I1, I2, I3, etc., an object detector 12 can detect features F1, F2, F3, and F4 in image I1, features F’1, F’2, F’3, and F’4 in image I2, and features F’’1, F’’2, F’’3, and F’’4 in image I3.); responsive to the receiving, maintaining a graph comprising nodes representing the received detection outcomes and weighted edges between at least a portion of the nodes (Figure 2A; Col. 5, Line 11 – 25: As illustrated in FIG. 1A, the message passing graph 50 generates layer 52A for the image I1 generated at time t1, layer 52B for the image I2 generated at time t2 and layer 52C for the image I3 generated at time t3. In general, the message passing graph 50 can include any suitable number of layers, each with any suitable number of features (however, as discussed below, the multi-object tracker 14 can apply a rolling window to the message passing graph 50 so as to limit the number of layers used at any one time). Each of the layers 52A-C includes feature nodes corresponding to the features the object detector 12 identified in the corresponding image (represented by oval shapes). The features nodes are interconnected via edges which, in this example implementation, also include edge nodes (represented by rectangular shapes); Col 6, Line 15 – 18: In particular, feature nodes and d10, d20 and d30 can be considered predicted true positive detection nodes connected to other predicted true positive detection nodes (e.g., d30 to d11)in later-in-time layers. The predicted true positive detection nodes are connected across layers in a pairwise manner via finalized edges 80, generated after data association.); executing appearance modeling of the nodes via (Figure 1A, Including Ref. No 12, F1, F2, F3, and F4; Col. 3, Line 55- 58 : The system then can implement message passing in accordance with graph neural networks message propagation techniques, so that the graph operates as a message passing graph (with two classes of nodes, feature nodes and edge nodes).; Col. 7, Line 6 – 16: As indicated above, each layer of the graph 50 or 70 represents a timestep, i.e., corresponds to a different instance in time. The data association can be understood as a graph problem in which a connection between observations in time is an edge. The multi-object tracker 14 implements a neural network that learns from examples how to make these connections, and further learns from examples whether an observation is true (e.g., the connection between nodes d30 observation is true (e.g., the connection between nodes d30 and d11) or false (e.g., resulting in the node d21 not being connected to any other nodes at other times, and thus becoming a false positive feature node).; Examiner’s note: As the claim does not explicitly define what “appearance-modeling” is, the examiner is interpreting it in view of figure 3 and ¶ 0046 and 0062 – 0064 of the applicant’s disclosure filed 10 February, 2026. As such, the examiner concludes under broadest reasonable interpretation that “appearance-modeling data” is merely any node data output from a self-attention layer of a graph transformer network.); and executing motion modeling of the nodes using the resultant appearance-modeling data via (Col. 6, Line 15 – 23: In particular, feature nodes and d10, d20 and d30 can be considered predicted true positive detection nodes connected to other predicted true positive detection nodes ( e.g., d30 to d11) in later-in-time layers. The predicted true positive detection nodes are connected across layers in a pairwise manner via finalized edges 80, generated after data association. Thus, the multi-object tracker 14 determines where the same object is in the imagery collected at time t and in the imagery collected at time t+1.). Ramezani does not explicitly teach executing appearance modeling of the nodes via a self-attention layer of a graph transformer network to yield resultant appearance-modelling data; and executing motion modeling of the nodes using the resultant appearance-modeling data via a cross-attention layer of the graph transformer network to yield resultant motion-modeling data for tracking three-dimensional objects across the plurality of synchronized camera inputs. However, Chu does teach executing appearance modeling of the nodes via a self-attention layer of a graph transformer network (Page 4861, Col. 1, Section 3: “Our method, named Trans-MOT (Spatial-temporal graph Transformer for MOT)…) to yield resultant appearance-modelling data (Figure 2; Page 4861, Col. 2, Section 4: We propose graph multi-head attention to model spatial relationship of the tracklets and candidates using the self-attention mechanism.; page 4862, Col. 1, ¶ 2: Inside the layer, a multi-head graph attention module is utilized to generate self-attention for the input graph series. This module takes feature tensor Fs and the graph weights w-Xt-1 to generate self-attention weights for the i-th head: PNG media_image1.png 37 330 media_image1.png Greyscale ; Examiner’s note: As the claim does not explicitly define what “appearance-modeling” is, the examiner is interpreting it in view of figure 3 and ¶ 0046 and 0062 – 0064 of the applicant’s disclosure filed 10 February, 2026. As such, the examiner concludes under broadest reasonable interpretation that “appearance-modeling data” is merely any node data output from a self-attention layer of a graph transformer network.); and executing motion modeling of the nodes using the resultant appearance-modeling data via a cross-attention layer of the graph transformer network (Figure 3) to yield resultant motion-modeling data for tracking three-dimensional objects (Figure 4; Page 4863, Col. 1, ¶ 2: Multi-head cross attention is calculated for Fde’att and Fen’out to generate unnormalized attention weights. The output is passed through a feed-forward layer and a normalization layer to generate the output tensor R M t + 1 x N t - 1 + 1 x D that corresponds to the matching between the tracklets and the candidates. The output of the spatial graph decoder can be passed through a linear layer and a Softmax layer to generate the assignment matrix A - t ϵ R M t + 1 x N t - 1 + 1 x D ). Chu is considered to be analogous art it pertains to multi-object tracking using a graph transformation network. Therefore, it would have been obvious to one of ordinary skill in the art to combine the neural network for object detection (as taught by Ramezani) and the spatial-temporal graph transformer for multiple object tracking (as taught by Chu) before the effective filing date of the claimed invention. The motivation for this combination of references would be the system of Chu utilizes a re-match stage which associates un-matched high confidence detections with remaining tracklets which improves the recall of the tracking association. (See Page 4863, Col. 2, Section 4.4). This motivation for the combination of Ramezani and Chu is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Regarding claim 3, the Ramezani and Chu combination teaches the method of claim 1. Additionally, Ramezani teaches wherein the weighted edges are computed based at least in part on node similarity (Col. 6, Line 15 – 23: In particular, feature nodes and d10, d20 and d30 can be considered predicted true positive detection nodes connected to other predicted true positive detection nodes ( e.g., d30 to d11) in later-in-time layers. The predicted true positive detection nodes are connected across layers in a pairwise manner via finalized edges 80, generated after data association.). Regarding claim 5, the Ramezani and Chu combination teaches the method of claim 1. Additionally, Chu teaches wherein the self-attention layer provides output embeddings (Figure 2; Page 4861, Col. 2, Section 4: We propose graph multi-head attention to model spatial relationship of the tracklets and candidates using the self-attention mechanism.; page 4862, Col. 1, ¶ 2: Inside the layer, a multi-head graph attention module is utilized to generate self-attention for the input graph series. This module takes feature tensor Fs and the graph weights w-Xt-1 to generate self-attention weights for the i-th head: PNG media_image1.png 37 330 media_image1.png Greyscale Regarding claim 7, the Ramezani and Chu combination teaches the method of claim 1. Additionally, Chu teaches wherein the tracked objects are tracklets (Page 4860, Col. 2, ¶ 1: The proposed TransMOT also constructs a spatial graph for the objects within the same frame, but it exploits the Transformer networking architecture to jointly learn the spatial and temporal relationship of the tracklets and candidates for efficient association.). Regarding claim 8, the Ramezani and Chu combination teaches the method of claim 1. Additionally, Ramezani teaches comprising post-processing the resultant motion-modeling data via motion propagation and node merging (Col. 6, Line 29 – 35: Features d1t, d2t . . . d-Nt+T within the rolling window 72 30 can be considered active detection nodes, with at least some of the active detection nodes interconnected by active edges 82, generated with learned feature representation. Unlike the nodes in the earlier-in-time layers 74, the values (e.g., GRU outputs) of the nodes and edges continue to change while 35 these nodes are within the rolling window 72.; Col. 7, Line 11 – 16: The multi-object tracker 14 implements a neural network that learns from examples how to make these connections, and further learns from examples whether an observation is true (e.g., the connection between nodes d30 and d10) or false (e.g., resulting in the node d21 not being connected to any other nodes at other times, and thus becoming a false positive feature node).). Regarding claim 9, the Ramezani and Chu combination teaches the method of claim 8. Additionally, Ramezani teaches wherein the post-processing comprises adding a node to the graph via link prediction (Col. 8, Line 6 – 16: update graph( ): This function is called after every timestep, to add new nodes (detections) and corresponding edges to the end of the currently active part of the graph (e.g., the part of the graph within a sliding time window, as discussed below), and fix parameters of and exclude further changes to the oldest set of nodes and edges from the currently active part of the graph, as the sliding time window no longer includes their layer. This essentially moves the sliding time window one step forward.). Regarding claim 9, the Ramezani and Chu combination teaches the method of claim 8. Additionally, Ramezani teaches wherein the post-processing comprises removing a node from the graph via link prediction (Col. 8, Line 17 – 22: prune graph( ): This function removes low probability edges and nodes from the currently active part of the graph using a user specified threshold. This function can be called whenever memory/compute requirements exceed what is permissible ( e.g., exceed a predetermined value(s) for either memory or processing resources).). Regarding claim 11, claim 11 has been analyzed with regard to claim 1 and is rejected for the same reasons of obviousness as used above as well as in accordance with Ramezani’s further teaching on: Memory (Col. 2, Lines 14 – 17: In another embodiment, a non-transitory computer-readable medium stores thereon instructions executable by one or more processors to implement a multi-object tracking architecture.); At least one processor coupled to the memory and configured to implement a method (Col. 2, Lines 14 – 17: In another embodiment, a non-transitory computer-readable medium stores thereon instructions executable by one or more processors to implement a multi-object tracking architecture.)… Regarding claim 13, claim 13 has been analyzed with regard to respective claim 3 and is rejected for the same reasons of obviousness as used above. Regarding claim 15, claim 15 has been analyzed with regard to respective claim 5 and is rejected for the same reasons of obviousness as used above. Regarding claim 17, claim 17 has been analyzed with regard to respective claim 7 and is rejected for the same reasons of obviousness as used above. Regarding claim 18, claim 18 has been analyzed with regard to respective claim 8 and is rejected for the same reasons of obviousness as used above. Regarding claim 19, the Ramezani and Liao combination teaches the system of claim 18. Additionally, Ramezani teaches wherein the post-processing comprises at least one of adding a node to the graph or removing a node from the graph via link prediction (Col. 8, Line 6 – 16: update graph( ): This function is called after every timestep, to add new nodes (detections) and corresponding edges to the end of the currently active part of the graph (e.g., the part of the graph within a sliding time window, as discussed below), and fix parameters of and exclude further changes to the oldest set of nodes and edges from the currently active part of the graph, as the sliding time window no longer includes their layer. This essentially moves the sliding time window one step forward.; Col. 8, Line 17 – 22: prune graph( ): This function removes low probability edges and nodes from the currently active part of the graph using a user specified threshold. This function can be called whenever memory/compute requirements exceed what is permissible ( e.g., exceed a predetermined value(s) for either memory or processing resources).). Regarding claim 20, Ramezani teaches a computer program product comprising a non-transitory computer-usable medium having computer-readable program code embodied therein, the computer-readable program code adapted to be executed to implement a method for three-dimensional object tracking across cameras (Col. 22, Line 6 – 10: As an example, the steps of a method or algorithm disclosed herein may be implemented in a processor-executable software module which may reside on a computer readable non-transitory storage medium), comprising: receiving detection outcomes generated by a three-dimensional object detector from a plurality of synchronized camera inputs (Col. 4, Lines 45 – 48: Referring to FIG. 1A, a sensor system 10 can capture scenes S1, S2, S3 of the environment 12 at times t1, t2, and t3 respectively, and generate a sequence of images I1, I2, I3, etc.; Col 4, Lines 54 – 57: In some implementations… each point has certain coordinates in a 3D space… ; Col. 4, Lines 61 – 64: After the sensor system 10 generates the images I1, I2, I3, etc., an object detector 12 can detect features F1, F2, F3, and F4 in image I1, features F’1, F’2, F’3, and F’4 in image I2, and features F’’1, F’’2, F’’3, and F’’4 in image I3.); responsive to the receiving, maintaining a graph comprising nodes and weighted edges between at least a portion of the nodes (Figure 2A), wherein the nodes represent the received tracked objects and comprise appearance features and motion features (Col. 2, Line 5 – 9: a plurality of feature nodes to represent features detected in the corresponding image, and generating edges that interconnect at least some of the feature nodes across adjacent layers of the graph neural network to represent associations between the features. Col 6, Line 15 – 18: In particular, feature nodes and d10, d20 and d30 can be considered predicted true positive detection nodes connected to other predicted true positive detection nodes (e.g., d30 to d11)in later-in-time layers. The predicted true positive detection nodes are connected across layers in a pairwise manner via finalized edges 80, generated after data association. Thus, the multi-object tracker 14 determines where the same object is in the imagery collected at time t and in the imagery collected at time t+1.; Examiner’s note: The act of determining a location of an object at a time t and a time t + 1 determines a movement of an object across the time step and is understood to be a motion feature under broadest reasonable interpretation.), and wherein the weighted edges are computed based on node similarity (Col. 2, Line 27 – 29: generating edges that interconnect at least some of the feature nodes across adjacent layers of the graph neural network to represent associations between the features.); executing appearance modeling of the nodes via (Figure 1A, Including Ref. No 12, F1, F2, F3, and F4; Col. 3, Line 55- 58 : The system then can implement message passing in accordance with graph neural networks message propagation techniques, so that the graph operates as a message passing graph (with two classes of nodes, feature nodes and edge nodes).; Col. 7, Line 6 – 16: As indicated above, each layer of the graph 50 or 70 represents a timestep, i.e., corresponds to a different instance in time. The data association can be understood as a graph problem in which a connection between observations in time is an edge. The multi-object tracker 14 implements a neural network that learns from examples how to make these connections, and further learns from examples whether an observation is true (e.g., the connection between nodes d30 observation is true (e.g., the connection between nodes d30 and d11) or false (e.g., resulting in the node d21 not being connected to any other nodes at other times, and thus becoming a false positive feature node).; Examiner’s note: As the claim does not explicitly define what “appearance-modeling” is, the examiner is interpreting it in view of figure 3 and ¶ 0046 and 0062 – 0064 of the applicant’s disclosure filed 10 February, 2026. As such, the examiner concludes under broadest reasonable interpretation that “appearance-modeling data” is merely any node data output from a self-attention layer of a graph transformer network.); and executing motion modeling of the nodes using the resultant appearance-modeling data via (Col. 6, Line 15 – 23: In particular, feature nodes and d10, d20 and d30 can be considered predicted true positive detection nodes connected to other predicted true positive detection nodes ( e.g., d30 to d11) in later-in-time layers. The predicted true positive detection nodes are connected across layers in a pairwise manner via finalized edges 80, generated after data association. Thus, the multi-object tracker 14 determines where the same object is in the imagery collected at time t and in the imagery collected at time t+1.). Ramezani does not explicitly teach executing appearance modeling of the nodes via a self-attention layer of a graph transformer network to yield resultant appearance-modelling data; and executing motion modeling of the nodes using the resultant appearance-modeling data via a cross-attention layer of the graph transformer network to yield resultant motion-modeling data for tracking three-dimensional objects across the plurality of synchronized camera inputs. However, Chu does teach executing appearance modeling of the nodes via a self-attention layer of a graph transformer network (Page 4861, Col. 1, Section 3: “Our method, named Trans-MOT (Spatial-temporal graph Transformer for MOT)…) to yield resultant appearance-modelling data (Figure 2; Page 4861, Col. 2, Section 4: We propose graph multi-head attention to model spatial relationship of the tracklets and candidates using the self-attention mechanism.; page 4862, Col. 1, ¶ 2: Inside the layer, a multi-head graph attention module is utilized to generate self-attention for the input graph series. This module takes feature tensor Fs and the graph weights w-Xt-1 to generate self-attention weights for the i-th head: PNG media_image1.png 37 330 media_image1.png Greyscale ; Examiner’s note: As the claim does not explicitly define what “appearance-modeling” is, the examiner is interpreting it in view of figure 3 and ¶ 0046 and 0062 – 0064 of the applicant’s disclosure filed 10 February, 2026. As such, the examiner concludes under broadest reasonable interpretation that “appearance-modeling data” is merely any node data output from a self-attention layer of a graph transformer network.); and executing motion modeling of the nodes using the resultant appearance-modeling data via a cross-attention layer of the graph transformer network (Figure 3) to yield resultant motion-modeling data for tracking three-dimensional objects (Figure 4; Page 4863, Col. 1, ¶ 2: Multi-head cross attention is calculated for Fde’att and Fen’out to generate unnormalized attention weights. The output is passed through a feed-forward layer and a normalization layer to generate the output tensor R M t + 1 x N t - 1 + 1 x D that corresponds to the matching between the tracklets and the candidates. The output of the spatial graph decoder can be passed through a linear layer and a Softmax layer to generate the assignment matrix A - t ϵ R M t + 1 x N t - 1 + 1 x D ). Chu is considered to be analogous art it pertains to multi-object tracking using a graph transformation network. Therefore, it would have been obvious to one of ordinary skill in the art to combine the neural network for object detection (as taught by Ramezani) and the spatial-temporal graph transformer for multiple object tracking (as taught by Chu) before the effective filing date of the claimed invention. The motivation for this combination of references would be the system of Chu utilizes a re-match stage which associates un-matched high confidence detections with remaining tracklets which improves the recall of the tracking association. (See Page 4863, Col. 2, Section 4.4). Claims 2, 4, 12 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Ramezani et al (U.S. Patent No. 11361449 B2, hereinafter “Ramezani”) in view of Chu et al (P. Chu "TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking," 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 2023, pp. 4859-4869, doi: 10.1109/WACV56688.2023.00485., hereinafter “Chu”) and further in view of Liao et al (U.S. Patent Publication No. 2024/0242462 A1, hereinafter “Liao”). Regarding claim 2, the Ramezani and Liao combination teaches the method of claim 1. Additionally, Liao teaches wherein the nodes represent tracked objects comprising at least one of appearance features or motion features (¶ 0027: Such raw data is transformed to graph node classification model input data, which may include a node graph and a set of features for each node of the node graph.; ¶ 0028: For each node, node features ( e.g., features corresponding to American football and designed to provide accurate and robust classification by the graph node classification model such as a GCN or GNN) are generated, and the node graph and feature sets are provided to or fed to a pretrained graph node classification model such as DeepGCN to perform a node classification.). Liao is considered to be analogous art as it pertains multi-camera multi-object tracking. Therefore, it would have been obvious to one of ordinary skill in the art to combine the neural network for object detection (as taught by Ramezani) and game focus estimation in team sports for immersive video (as taught by Liao) before the effective filing date of the claimed invention. The motivation for this combination of references would be the system of Liao attains and refines single camera information by associating all single camera information so as to improve accuracy. (See ¶ 0043). This motivation for the combination of Ramezani, Chu, and Liao is supported by KSR exemplary rationale (G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention. MPEP 2141 (III). Regarding claim 4, the Ramezani and Chu combination teaches the method of claim 3. Additionally, Ramezani and Liao teach wherein the node similarity is computed based on at least one of appearance similarity or location similarity between the tracked objects (Ramezani Col. 6, Line 15 – 23: In particular, feature nodes and d10, d20 and d30 can be considered predicted true positive detection nodes connected to other predicted true positive detection nodes ( e.g., d30 to d11) in later-in-time layers. The predicted true positive detection nodes are connected across layers in a pairwise manner via finalized edges 80, generated after data association.; Liao ¶ 0045: Notably, herein a node is defined based on data within a selected region and an edge connects two nodes when the regions have a shared boundary therebetween.). Liao is considered to be analogous art as it pertains multi-camera multi-object tracking. Therefore, it would have been obvious to one of ordinary skill in the art to combine the neural network for object detection (as taught by Ramezani) and game focus estimation in team sports for immersive video (as taught by Liao) before the effective filing date of the claimed invention. The motivation for this combination of references would be the system of Liao attains and refines single camera information by associating all single camera information so as to improve accuracy. (See ¶ 0043). Regarding claim 12, claim 12 has been analyzed with regard to respective claim 2 and is rejected for the same reasons of obviousness as used above. Regarding claim 14, claim 14 has been analyzed with regard to respective claim 4 and is rejected for the same reasons of obviousness as used above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Liao et al (U.S. Patent Publication No. 2023/0377335 A1) teaches a system for key person recognition in multi-camera immersive video attained for a scene including detecting predefined person formations in the scene based on an arrangement of the persons in the scene, generating a feature vector for each person in the detected formation, and applying a classifier to the feature vectors to indicate one or more key persons in the scene. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW JONES whose telephone number is (703)756-4573. The examiner can normally be reached Monday - Friday 8:00-5:00 EST, off Every Other Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella can be reached at (571) 272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ANDREW B. JONES/Examiner, Art Unit 2667 /MATTHEW C BELLA/Supervisory Patent Examiner, Art Unit 2667
Read full office action

Prosecution Timeline

Feb 10, 2024
Application Filed
Mar 18, 2026
Non-Final Rejection mailed — §101, §103, §112
Jun 18, 2026
Response Filed
Jul 27, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700159
METHOD FOR USE IN X-RAY CT IMAGE RECONSTRUCTION
3y 2m to grant Granted Aug 04, 2026
Patent 12694514
SYSTEMS AND METHODS FOR IDENTIFYING IMAGES CONTAINING INDICATORS OF A CELIAC-LIKE DISEASE
3y 2m to grant Granted Jul 28, 2026
Patent 12682593
IMAGE SCORING APPARATUS, IMAGE SCORING METHOD AND METHOD OF MACHINE LEARNING
3y 3m to grant Granted Jul 14, 2026
Patent 12664668
METHODS AND SYSTEMS FOR REGISTERING PREOPERATIVE IMAGE DATA TO INTRAOPERATIVE IMAGE DATA OF A SCENE, SUCH AS A SURGICAL SCENE
2y 11m to grant Granted Jun 23, 2026
Patent 12651441
METHOD AND APPARATUS FOR GENERATING AN ADAPTIVE TRAINING IMAGE DATASET
3y 11m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
71%
Grant Probability
88%
With Interview (+17.0%)
2y 11m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 83 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month