Prosecution Insights
Last updated: October 02, 2026
Application No. 18/597,451

BIDIRECTIONAL OBJECT TRACKING IN COMPUTER VISION APPLICATIONS

Final Rejection §101§102§103
Filed
Mar 06, 2024
Examiner
DRYDEN, EMMA ELIZABETH
Art Unit
2677
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
68%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
19 granted / 28 resolved
+5.9% vs TC avg
Moderate +12% lift
Without
With
+12.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
18 currently pending
Career history
51
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
59.3%
+19.3% vs TC avg
§102
12.9%
-27.1% vs TC avg
§112
13.3%
-26.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 28 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendment filed 06/15/2026 has been entered. Applicant’s amendments to the specification have overcome each and every objection previously set forth in the Non-Final Office Action mailed 04/02/2026. Claims 1-20 remain pending in the application. Response to Arguments Applicant's arguments, pg. 11-13 of the Remarks filed 06/15/2026, have been fully considered but they are not persuasive. In Section I, Applicant argues that the claimed bidirectional tracking architecture cannot practically be performed in the human mind: PNG media_image1.png 579 690 media_image1.png Greyscale Examiner respectfully disagrees. MPEP 2106.04(a)(2)(III) sets forth that the courts “do not distinguish between mental processes that are performed entirely in the human mind and mental processes that require a human to use a physical aid (e.g., pen and paper or a slide rule) to perform the claim limitation” and “Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer. As the Federal Circuit has explained, "[c]ourts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind." Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015).” The recitation that object tracking is performed in each of two directions does not bring the claimed steps beyond what can practically be performed in the human mind. At most, it may require a human to use a physical aid, such as pen and paper, to perform the claim limitation. Additionally, the aggregation of accumulated information (set forth in the most recent amendment) can be tracked with pen/paper, and therefore also does not distinguish the independent claims from a mental process. Applicant further points to the USPTO’s Subject Matter Eligibility Example 39 to show that a computer vision method involving the processing of digital images does not recite a mental process: PNG media_image2.png 319 683 media_image2.png Greyscale Examiner respectfully disagrees that the provided example shows that the claims of the instant application are not operations that a human could practically perform in the human mind. USPTO’s Subject Matter Eligibility Example 39 recites claim limitations, that are not analogous to the limitations currently being examined, that contribute to the example not being considered a judicial exception. For example, Subject Matter Eligibility Example 39 recites applying transformations to digital images to modify them and steps for training a neural network. The current claims do not recite the modification of digital images, but simply the visual tracking of objects across frames. The example being directed to an application of detecting faces in image frames is not necessarily the reason it cannot practically be performed in the human mind. In Section II, Applicant argues that the claims recite a specific technical architecture that improves object tracking, confirmed by the USPTO Director’s decision in Ex parte Desjardins: PNG media_image3.png 605 552 media_image3.png Greyscale Examiner respectfully disagrees. In the USPTO Director’s decision in light of Ex Parte Desjardins, dated 12/05/2025, MPEP 2106.04(d)(1) is revised to read: PNG media_image4.png 362 683 media_image4.png Greyscale As demonstrated by the USPTO’s Subject Matter Eligibility Example 40, the differing components/steps recited in claims 1 and 2 (of the example) result in different determinations under Step 2A Prong 2. In the current claim set, the claimed steps in claims 1-6, 10-14, and 18-20 are recited at a high level of generality, and merely recite steps of obtaining, aggregating, and comparing representations of objects in image frames. Accordingly, claims 1-6, 10-14, and 18-20 do not include the components or steps of the invention that provide the improvement described in the specification (i.e., the prevention and correction of false ID switches). However, claims 7-9 and 15-17 do set forth the steps of the invention that provide the improvement described in the specification – identifying and correcting false ID switches through specific thresholding and the assigning and mapping of new object ID’s. In view of the foregoing, the rejection of claims 1-6, 10-14, and 18-20 under 35 U.S.C. 101 is maintained and the rejection of claims 7-9 and 15-17 under 35 U.S.C. 101 is withdrawn. Applicant's arguments, pg. 13-18 of the Remarks filed 06/15/2026, have been considered but are moot because the new ground of rejection does not rely on any combination of references applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. However, Section III describes Wang’s processing as inherently unidirectional, which is applicable to the new ground of rejection set forth in the rejection of claim 5 below. Applicant argues: PNG media_image5.png 469 606 media_image5.png Greyscale Examiner respectfully disagrees. Using the broadest reasonable interpretation of the recited claims, there is no directional requirement defining “bidirectional tracking” besides the requirement of certain video frames to be “upstream” or “downstream”. Though Wang does not explicitly state that their method is “bidirectional”, a three frame tracklet where the middle frame is considered the frame that is 1) downstream from a previous frame (the forward tracking being defined as the previous [Wingdings font/0xE0] middle frame) and 2) downstream from a future frame (the reverse tracking being defined as the future [Wingdings font/0xE0] middle frame) can be considered bidirectional tracking based on the current claim limitations. This interpretation can be similarly applied and demonstrated with the image frames of Figure 4 in the instant application (annotations overlaid in the example attached below). PNG media_image6.png 643 808 media_image6.png Greyscale Furthermore, the recited limitations associated with the tracking (i.e., updated state) lack further detail requiring directional characteristics between two frames that are compared. As set forth in claim 5, similarity between feature vectors associated with objects are computed, with no further requirements regarding characteristics of directional tracking. In view of the foregoing, it is maintained that Wang teaches bidirectional object tracking across image frames. Regardless, the newly introduced prior art accounts for Applicant’s narrower interpretation set forth in the Remarks and accounts for the newly entered amendments. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-6, 10-14, and 18-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Under Step 1, claims 1-6 and 10 are process/method claims and claims 11-14 and 18-20 are machine claims. Under Step 2A Prong One, all claims recite abstract ideas, specifically mental processes – concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III). These mental processes are more particularly recited in claim 1 as: obtaining digital representations of an object depicted in a plurality of video frames (i.e., human obtaining numerical values characterizing object location or visual description or obtaining images with objects visually located); and performing a bidirectional tracking of the object across the plurality of video frames (i.e., following the location of objects across frames either with pen/paper or mentally), wherein the bidirectional tracking comprises: for each of a forward direction (FD) of the bidirectional tracking and a reverse direction (RD) of the bidirectional tracking, obtaining, using (i) a current state of the object associated with an upstream video frame and (ii) the digital representation of the object for a downstream video frame, an updated state of the object associated with the downstream video frame (i.e., following the numerical values/locations of objects in both previous and subsequent frames), wherein the updated state aggregates the current state of the object with the digital representation of the object for the downstream video frame (i.e., aggregating the two sets of object descriptors with pen/paper); obtaining, using at least one of the updated state of the object for the FD or the updated state of the object for the RD, a bidirectional state of the object (i.e., determining a location of objects in one frame with respect to its location in a different frame); and determining, using the bidirectional state of the object, a trajectory of the object across the plurality of video frames (i.e., characterizing with pen/paper or mentally how an object moved across frames). Dependent claims 2-6 and 10 provide additional limitations that are further part of the abstract idea of bidirectional tracking of objects across image frames. Claims 5-6 and 10 also provide additional limitations that are considered mathematical concepts, and are thus further part of the abstract idea. It is noted that the above analysis is according to the 2019 Revised Patent Subject Matter Eligibility Guidance published in the Federal Register (84 FR 50) on January 7, 2019 and MPEP 2106.04(a)(2)(III). Consider also that “If a claim recites a limitation that can practically be performed in the human mind, with or without the use of a physical aid such as pen and paper, the limitation falls within the mental processes grouping, and the claim recites an abstract idea” as per MPEP 2106.04(a)(2)(III)(B). See also footnotes 14 and 15 of the Federal Register Notice. As detailed above, the steps for bidirectional tracking of objects across image frames may be practically performed in the human mind with or without the use of a physical aid such as a pen and paper. Under Step 2A Prong Two, this judicial exception is not integrated into a practical application because each of claims 1-6 and 10 do not recite additional elements that integrate the exception into a practical application. The additional element of “obtaining digital representations of an object” in claim 1 adds insignificant extra-solution activity, which is not indicative of integration into a practical application as per MPEP 2106.05(g). The additional element of the machine learning model of claim 2 is recited at a high level of generality and merely equate to “apply it” or otherwise merely uses a generic computer as a tool to perform an abstract idea which is not indicative of integration into a practical application, as per MPEP 2106.05(f). Additionally, the digital representations are obtained before bidirectional tracking is performed, thus a human performing the claim solely receives the data as generated. See also MPEP 2106.04(a)(2)(III) with respect to Mental Processes: “Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer”. See also MPEP 2106.04(a)(2)(III)(C)(3) “Using a computer as tool to perform a mental process” and MPEP 2106.04(a)(2)(III)(D), as well as the case law cited therein. The additional elements in claims 2-6 and 10 reciting mathematical concepts and/or abstract ideas do not integrate the judicial exception into a practical application. See MPEP 2106.04(II)(A)(2). Under Step 2B, each of claims 1-6 and 10 do not recite additional elements that are indicative of an inventive concept. The additional elements are simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception as per MPEP 2106.05(d) and 2106.07(a)III. In other words, the additional elements do not amount to significantly more than the judicial exception. Regarding claim 1, the obtaining digital representations of an object is well-known extra-solution activity (a common way for a human to identify an object in an image), and thus does not amount to significantly more (see MPEP 2106.05(g)). Regarding claim 2, use of a machine learning model to obtain data before performing the abstract idea is considered insignificant extra-solution activity (see MPEP 2106.05(g)) and amounts to merely an instruction to apply an aspect of the abstract idea using generic computer elements (see MPEP 2106.05(f) MPEP 2106.05(I)(A)). Thus, it does not integrate the judicial exception into a practical application. Regarding claims 2-4, further defining the type of data used to characterize objects in images is further part of the abstract idea of claim 1. Additionally, the tracking of these values across a plurality of image frames may be practically performed in the human mind with the use of a physical aid such as a pen and paper. Regarding claims 5-9, additional limitations are directed to the abstract ideas of mathematical calculations (comparing vectors in claim 5) and comparing numerical values to each other with the use of a threshold value (claim 6), which are considered mathematical concepts/calculations (see MPEP 2106.04(a)(2)) and mental processes (see MPEP 2106.04(a)(2)). Regarding claim 10, the generating of a human-perceivable report is further part of the abstract idea of claim 1 and may be practically performed in the human mind with the use of a physical aid such as a pen and paper. The addition of further judicial exceptions does not amount to significantly more (see MPEP 2106.05(I)). Regarding independent claims 11 and 20, the rationale provided in the rejection of claim 1, and corresponding dependent claims, is incorporated herein. The system of claim 1 corresponds to the system of claim 11 and the processor of claim 20, and performs the same steps disclosed in claim 1. In addition, the processing devices of claims 11 and 20 and the systems of claim 19 amount to merely an instruction to apply the abstract idea using generic computer elements, and does not integrate the judicial exception into a practical application (see MPEP 2106.05(d)). This does not amount to significantly more than the judicial exception. For all of the above reasons, taken alone or in combination, claims 1-6, 10-14, and 18-20 recite a non-statutory mental process. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 11, and 20 are rejected under 35 U.S.C. 102(a)(1) and (a)(2) as being anticipated by Xie (U.S. Patent No. 2019/0034700 A1). Regarding claim 1, Xie teaches a method comprising: obtaining digital representations of an object depicted in a plurality of video frames (Xie, edge feature points in video frames, para 32: “determining an initial position of a window to be tracked on a current frame image in the video stream according to the position of the window to be tracked of the reference frame, and extracting an edge feature point in the window to be tracked of the current frame”; feature points shown in FIG. 4 attached below); and performing a bidirectional tracking of the object across the plurality of video frames (Xie, para 71: “FIG. 3 is a schematic diagram of the principle of the inverse tracking and verifying according to an embodiment of the present disclosure. In the present embodiment, each of edge feature points in the window to be tracked that has been adjusted of the current frame is verified”; see FIG. 3 attached below), PNG media_image7.png 786 482 media_image7.png Greyscale PNG media_image8.png 830 676 media_image8.png Greyscale wherein the bidirectional tracking comprises: for each of a forward direction (FD) of the bidirectional tracking (Xie, clockwise direction) and a reverse direction (RD) of the bidirectional tracking (Xie, anticlockwise direction), obtaining, using (i) a current state of the object associated with an upstream video frame (Xie, feature point in frame It (FD) and It+k (RD), see para 72) and (ii) the digital representation of the object for a downstream video frame (Xie, feature point in frame It+k (FD) and It (RD), see para 72), an updated state of the object associated with the downstream video frame, wherein the updated state aggregates the current state of the object with the digital representation of the object for the downstream video frame (Xie, the feature points at each frame are aggregated for both directions – Yt, Yt+1, and Xt+k (RD) and Xt, Xt+1, and Xt+k (FD)); obtaining, using at least one of the updated state of the object for the FD or the updated state of the object for the RD, a bidirectional state of the object (Xie, Yt and Xt, the bidirectional state of the object, is determined using both directions; see FIG. 3); and determining, using the bidirectional state of the object, a trajectory of the object across the plurality of video frames (Xie, face tracking is the trajectory of the edge feature points of a face object, para 74: “in the process of face tracking, regarding the problem in the face movement that edge features may be blocked to result in incorrect tracking, the present disclosure, by verifying the edge feature points that is tracked and excluding in time the edge feature points that have incorrect tracking, ensures to conduct the subsequent tracking in the clockwise direction using the edge feature points that are correctly tracked, overcomes the problem of failed tracking resulted by misjudgment of the edge feature points, and further improves the accuracy of face tracking”; Yt and Xt, the results of the bidirectional tracking, are used to correct the tracking process via the verification process described in para 72-74; thus, determining the trajectory uses the bidirectional state of the object). Regarding claim 11, Xie teaches a system comprising: a processing device (Xie, device for face tracking that processes the video frames, para 99: “In the present embodiment, the device 50 further comprises a tracking and verifying module and a face movement judging module”). All claim limitations carried out by the processing device are met by Xie because the method steps of claim 1 are the same as the steps in claim 11. Regarding claim 20, Xie teaches a processor comprising one or more processing devices (Xie, device for face tracking that processes the video frames, para 99: “In the present embodiment, the device 50 further comprises a tracking and verifying module and a face movement judging module”). All claim limitations carried out by the one or more processing devices are met by Xie because the method steps of claim 1 are the same as the steps in claim 20. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2-3, 12, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Xie in view of Yasutomi et al. (U.S. Patent No. 2024/0119739 A1), hereinafter Yasutomi. Regarding claim 2 (dependent on claim 1), Xie teaches feature points extracted from the video frames, but fails to explicitly teach wherein the digital representations of the object comprise one or more feature vectors generated by a machine learning model for respective video frames of the plurality of video frames. However, Yasutomi similarly teaches an object tracking method (Yasutomi, abstract). Yasutomi teaches wherein the digital representations of an object comprise one or more feature vectors generated by a machine learning model for respective video frames of the plurality of video frames (Yasutomi, para 50: “Image data 226 and 227 of a plurality of objects 223a and 223b recognized as the same object by an object tracking model (to be described later) in frame images 221 (221a and 221b) at different times of unlabeled moving image data 220 (moving image) are used for contrastive learning”; see FIG. 4, attached below, wherein images are input to an object detection model before the objects are input to a tracking model; see examples of trained detection models in para 58 and “feature vectors” in para 59). Xie discloses a base method for extracting features from a plurality of video frames, but does not specify specific methods for feature extraction. Yasutomi teaches a method for object feature detection using a known technique of generating feature vectors by a machine learning model. A person having ordinary skill in the art, before the effective filing date of the claimed invention, could have applied the known technique, as taught by Yasutomi, in the same way to the method of Xie and achieved predictable results of accurately identifying objects in image frames using a trained and validated model. PNG media_image9.png 472 568 media_image9.png Greyscale Regarding claim 3 (dependent on claim 1), Xie teaches feature points extracted from the video frames, but fails to explicitly teach wherein the current state of the object comprises: a feature vector representative of a depiction of the object in at least the upstream video frame, and a location of the object associated with the upstream video frame. However, Yasutomi similarly teaches an object tracking method (Yasutomi, abstract). Yasutomi teaches wherein extracting the current state of an object in image frames comprises: a feature vector representative of a depiction of the object in a video frame, and a location of the object associated with the video frame (Yasutomi, para 50: “Image data 226 and 227 of a plurality of objects 223a and 223b recognized as the same object by an object tracking model (to be described later) in frame images 221 (221a and 221b) at different times of unlabeled moving image data 220 (moving image) are used for contrastive learning”; see FIG. 4, attached above and examples of trained detection models in para 58 and “feature vectors” in para 59; bounding boxes depict the location of the object). Xie discloses a base method for extracting features from a plurality of video frames, but does not specify specific methods for feature extraction and data format of the output. Yasutomi teaches a method for object feature detection using a known technique of generating feature vectors and object location using a machine learning model. A person having ordinary skill in the art, before the effective filing date of the claimed invention, could have applied the known technique, as taught by Yasutomi, in the same way to the method of Xie and achieved predictable results of accurately identifying objects, represented by feature vectors, in image frames using a trained and validated model. Regarding claim 12 (dependent on claim 11), Xie teaches feature points extracted from the video frames, but fails to explicitly teach wherein the current state of the object comprises: a feature vector representative of a depiction of the object in at least the upstream video frame, a size of the object, and a location of the object associated with the upstream video frame. However, Yasutomi similarly teaches an object tracking method (Yasutomi, abstract). Yasutomi teaches wherein extracting the current state of an object in image frames comprises: a feature vector representative of a depiction of the object in a video frame (Yasutomi, para 50: “Image data 226 and 227 of a plurality of objects 223a and 223b recognized as the same object by an object tracking model (to be described later) in frame images 221 (221a and 221b) at different times of unlabeled moving image data 220 (moving image) are used for contrastive learning”; see FIG. 4, attached above and examples of trained detection models in para 58 and “feature vectors” in para 59), a size of the object (Yasutomi, bounding boxes associated with object features, para 55: “The boundary position information may include one plane coordinate of a height, a width, and a vertex of each of the bounding boxes 202”), and a location of the object associated with the video frame (Yasutomi, bounding boxes depict the location of the object). Xie discloses a base method for extracting features from a plurality of video frames, but does not specify specific methods for feature extraction and data format of the output. Yasutomi teaches a method for object feature detection using a known technique of generating feature vectors and object size/location using a machine learning model. A person having ordinary skill in the art, before the effective filing date of the claimed invention, could have applied the known technique, as taught by Yasutomi, in the same way to the method of Xie and achieved predictable results of accurately identifying characteristics of objects, represented by feature vectors, in image frames using a trained and validated model. Regarding claim 19 (dependent on claim 11), Xie fails to explicitly teach wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system for generating or presenting at least one of augmented reality content, virtual reality content, or mixed reality content; a system implemented using a robot; a system for performing conversational AI operations; a system for generating synthetic data using AI operations; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. However, Yasutomi similarly teaches an object tracking method (Yasutomi, abstract). Yasutomi teaches wherein the system performing the object tracking method is comprised in a system for performing deep learning operations (Yasutomi, the system includes a deep neural network, see para 81). Xie discloses a base method for extracting features from a plurality of video frames in a target tracking system, but does not specify a specific type of system for carrying out the operations. Yasutomi teaches a method for multi-target tracking using a known technique of implementing the method in a system performing deep learning operations to carry out the method. A person having ordinary skill in the art, before the effective filing date of the claimed invention, could have applied the known technique of utilizing deep learning operations to perform computer vision tasks, as taught by Yasutomi, in the same way to the system of Xie and achieved predictable results of accurately and efficiently identifying and tracking objects across image frames using a trained and validated model. Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Xie in view of Yasutomi, in further view of Kirsch et al. (U.S. Patent No. 2019/0353775 A1), hereinafter Kirsch. Regarding claim 4 (dependent on claim 3), Xie in view of Yasutomi teaches wherein the current state of the object further comprises: a size of the object (Yasutomi, bounding boxes associated with object features, para 55: “The boundary position information may include one plane coordinate of a height, a width, and a vertex of each of the bounding boxes 202”), but fails to explicitly teach further comprising a velocity of the object associated with the upstream video frame, and wherein the location of the object, the size of the object and the velocity of the object are determined using a statistical filter. However, Kirsch teaches a method for tracking objects in images, including determining the location of the object, the size of the object and the velocity of the object using a statistical filter (Kirsch, para 100: “a process 500 is shown for detecting and tracking an object. The process 500 can be performed by the camera manager 318 with…a Kalman filter”; para 121: “the camera manager 318 can predict an object bounding box using a Kalman filter”; para 125: “The camera manager 318 can retrieve a predicted speed of the object from the Kalman filter for each frame of a sequence of frames analyzed by the camera manager 318”). It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the determining of the velocity of the object and use of a Kalman filter, as taught by Kirsch, with the method of Xie in view of Yasutomi in order to improve the accuracy of tracking objects by identifying the speed of objects over times and predicting movement/locations over time using a Kalman filter (Kirsch, para 121: “The prediction by the Kalman filter can be made based on one or multiple past known locations of the object (e.g., past bounding boxes). The Kalman filter can track one or multiple different objects, generating a predicted location for each.”; para 112: “the camera manager 318 can perform filtering of the tracks of the objects generated in the step 804 (or over multiple iterations of the steps 802 and 804) based on speed and size of the objects”). Claims 5, 7-9, 13, and 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Xie in view of Wang et al. (U.S. Patent No. 2024/0037757 A1), hereinafter Wang. Regarding claim 5 (dependent on claim 1), Xie teaches computing a similarity between feature points Yt and Xt, to determine where object occlusions occur (Xie, para 73: “comparing the position of the edge feature point in the reference frame which is obtained by verifying and the position of the edge feature point in the reference frame image which is acquired in advance, and if the comparison result is consistent, the verifying succeeds; and if the comparison result is not consistent, the verifying fails”; para 74: “It can be known from the above that, in the process of face tracking, regarding the problem in the face movement that edge features may be blocked to result in incorrect tracking…”), but fails to explicitly teach wherein obtaining the bidirectional state of the object comprises: computing an FD similarity between an FD feature vector associated with the current state of the object for the FD of the bidirectional tracking and a feature vector associated with the digital representation of the object for the downstream frame; computing an RD similarity between an RD feature vector associated with the current state of the object for the RD of the bidirectional tracking and the feature vector associated with the digital representation of the object for the downstream frame; and obtaining the bidirectional state of the object using the FD similarity and the RD similarity. However, Wang teaches a method for tracking an object across a plurality of video frames based on computing a similarity between object features (Wang, para 37: “determine a candidate identification switch image patch based on similarities of adjoining image patch pairs”; FIG. 2e and 3, attached below – frame t’ is the downstream frame for both a previous, t’’, and future, t, frame; para 32 describes detecting incorrect target identifications between time frames t, t’, and t’’, referred to throughout the disclosure as an image patch sequence in para 85, for example). PNG media_image10.png 514 753 media_image10.png Greyscale PNG media_image11.png 694 712 media_image11.png Greyscale Wang further teaches computing an FD similarity between an FD feature vector associated with the current state of the object for the FD of the bidirectional tracking and a feature vector associated with the digital representation of the object for the downstream frame (Wang, process of determining whether adjoining frames contain a “special feature similarity”, para 39, determined using cosine similarity, para 38: “adjoining image patch pairs can be obtained, and a j-th feature similarity Sim[j] is a similarity Sim(F[j+1], F[j]) between a re-identification feature F[j+1] of an image patch Patch[j+1] and a re-identification feature F[j] of an image patch Patch[j]. The similarity can be a cosine similarity between re-identification feature”); computing an RD similarity between an RD feature vector associated with the current state of the object for the RD of the bidirectional tracking and the feature vector associated with the digital representation of the object for the downstream frame (Wang, same processing above is computed for the number of patch sets in a sequence – para 38: “feature similarities of re-identification feature pairs of a plurality of adjoining image patch pairs in the image patch sequence SqPatch[i] are determined.”; para 40: “it is possible to find at a time all special feature similarities in the plurality of feature similarities, and to designate image patches associated therewith as candidate identification switch image patches”); and obtaining the bidirectional state of the object using the FD similarity and the RD similarity (Wang, para 39: “it is determined whether a candidate identification switch image patch is present in the tracklet Trk[j] according to whether a special feature similarity Simp less than a predetermined similarity threshold sTh is present in the plurality of feature similarities”). Similar to the bidirectional method taught by Xie, Wang teaches object tracking across three video frames. While Xie teaches computing a similarity between features of the same frame, Wang teaches computing a similarity for two directions (similarity between a current frame and an upstream frame and a similarity between the current frame and a downstream frame). Thus, Xie and Wang each disclose an object tracking method for identifying incorrect object tracking across frames based on object occlusions or other computer vision errors. A person of ordinary skill in the art, before the effective filing date of the claimed invention, would have recognized that the similarity method taught by Xie could have been substituted by the similarity method taught by Wang because both serve the purpose of identifying incorrect object tracking across frames. Furthermore, a person of ordinary skill in the art would have been able to carry out the substitution. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to substitute the similarity method of Xie for the similarity method of Wang according to known methods to yield the predictable result of identifying incorrect object tracking across frames and correcting object trajectories accordingly. Regarding claim 7 (dependent on claim 5), Xie in view of Wang teaches wherein obtaining the bidirectional state of the object using the FD similarity and the RD similarity comprises: determining that the FD similarity is above a threshold similarity and that the RD similarity is below the threshold similarity (Wang, a case when only the RD comparison contains a special feature similarity; para 39: “whether a special feature similarity Simp less than a predetermined similarity threshold sTh is present in the plurality of feature similarities”); resetting the state of the object for the RD (Wang, candidate for ID switch, para 39: “when a special feature similarity Simp is present, it is determined that a candidate identification switch image patch is present in the tracklet Trk[j], and, an image patch associated with the special feature similarity Simp is designated as the candidate identification switch image patch. For example, when the special feature similarity Simp is Sim[j] (i.e., Sim[j]<sTh), the image patch Patch[j] is designated as the candidate identification switch image patch.”); assigning a new object identification (ID) as an object ID for the RD of the bidirectional tracking (Wang, identification switch for the object in the patch by correcting the tracklet; tracklets are associated with target object ID, see FIG. 2, so splitting the tracklets corrects the IDs for all patches in each tracklet, see FIG. 1); and mapping the object ID for the RD of the bidirectional tracking to an object ID for the FD of the bidirectional tracking (Wang, correcting tracklets results in correct IDs in both directions, see para 53-54). Regarding claim 8 (dependent on claim 5), Xie in view of Wang teaches wherein obtaining the bidirectional state of the object using the FD similarity and the RD similarity comprises: determining that the RD similarity is above a threshold similarity and that the FD similarity is below the threshold similarity (Wang, a case when only the FD comparison contains a special feature similarity; para 39: “whether a special feature similarity Simp less than a predetermined similarity threshold sTh is present in the plurality of feature similarities”); resetting the state of the object for the FD (Wang, candidate for ID switch, para 39: “when a special feature similarity Simp is present, it is determined that a candidate identification switch image patch is present in the tracklet Trk[j], and, an image patch associated with the special feature similarity Simp is designated as the candidate identification switch image patch. For example, when the special feature similarity Simp is Sim[j] (i.e., Sim[j]<sTh), the image patch Patch[j] is designated as the candidate identification switch image patch.”); assigning a new object ID as the object for the FD of the bidirectional tracking (Wang, identification switch for the object in the patch by correcting the tracklet; tracklets are associated with target object ID, see FIG. 2, so splitting the tracklets corrects the IDs for all patches in each tracklet, see FIG. 1); and mapping the object ID for the FD of the bidirectional tracking to an object ID for the RD of the bidirectional tracking (Wang, correcting tracklets results in correct IDs in both directions, see para 53-54). Regarding claim 9 (dependent on claim 5), Xie in view of Wang teaches wherein obtaining the bidirectional state of the object using the FD similarity and the RD similarity comprises: determining that each of the FD similarity and the RD similarity is below a threshold similarity (Wang, a case when both the FD and RD are special feature similarities due to both being below the similarity threshold, sTh, see para 39); resetting the updated state of the object for the FD; resetting the updated state of the object for the RD (Wang, candidate for ID switch for both sets of patches outlined in claim 1, see para 39); and assigning new object IDs to each of the FD of the bidirectional tracking and the RD of the bidirectional tracking (Wang, identification switch for the object in each patch set by correcting the tracklet; tracklets are associated with target object ID, see FIG. 2, so splitting the tracklets corrects the IDs for all patches in each tracklet, see FIG. 1). Regarding claim 13 (dependent on claim 11), all claim limitations are met and rendered obvious by Xie in view of Wang because the method steps of claim 5 are the same as the steps in claim 13. Regarding claim 15 (dependent on claim 13), all claim limitations are met and rendered obvious by Xie in view of Wang because the method steps of claim 7 are the same as the steps in claim 15. Regarding claim 16 (dependent on claim 13), all claim limitations are met and rendered obvious by Xie in view of Wang because the method steps of claim 8 are the same as the steps in claim 16. Regarding claim 17 (dependent on claim 13), all claim limitations are met and rendered obvious by Xie in view of Wang because the method steps of claim 9 are the same as the steps in claim 17. Claims 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Xie in view of Wang, in further view of Zhang (CN Patent No. 110400329 A). Regarding claim 6 (dependent on claim 5), Xie in view of Wang teaches wherein obtaining the bidirectional state of the object using the FD similarity and the RD similarity comprises: determining that each of the FD similarity and the RD similarity is above a threshold similarity (Wang, a case when both the FD and RD are not special feature similarities due to both being above the similarity threshold, sTh, see para 39), but fails to explicitly teach modifying the updated state for the FD using the updated state for the RD; and modifying the updated state for the RD using the updated state for the FD. However, Zhang teaches a similar FD/RD tracking method wherein when both directions meet a threshold condition, modifying the updated state for the FD using the updated state for the RD; and modifying the updated state for the RD using the updated state for the FD (Zhang, pg. 35, para 71: “For the objects in the forward detection process and the inverse detection process of the reference area and the tracking process have very similar tracking values, they will be updated to be the same object. On the other hand, the forward tracking value and the inverse tracking value in the reference area may not be in a very close range, but both meet the preset tracking threshold condition.”; modifying the state by updating to the same object). It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the method for modifying the updated states based on a threshold value, as taught by Zhang, with the method of Xie in view of Wang in order to correctly identify the objects as being the same state in both directions when they are determined to be the same object in the downstream frame (See Zhang citation above). Regarding claim 14 (dependent on claim 13), all claim limitations are met and rendered obvious by Xie in view of Wang and Zhang because the method steps of claim 6 are the same as the steps in claim 14. Claims 10 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Xie in view of Roshtkhari et al. (U.S. Patent No. 2016/0335502 A1), hereinafter Roshtkhari. Regarding claim 10 (dependent on claim 1), Xie fails to explicitly teach further comprising: generating at least one of: a human-perceivable report that is based, at least in part, on the trajectory of the object, a statistical report that is based, at least in part, on the trajectory of the object, or a video that depicts, at least a portion of the trajectory of the object. However, Roshtkhari teaches a similar method for tracking objects (Roshtkhari, abstract), further comprising: generating a human-perceivable report that is based, at least in part, on the trajectory of the object (Roshtkhari, final tracking result in FIG. 3d, attached below, and para 38; para 34: “The tracking system also includes a trajectory creation module 20 for using the tracklets to generate trajectories for the objects being tracked”). It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to have combined the generated human-perceivable result, taught by Roshtkhari, with the method of Xie in order to visualize the complete trajectories of objects (Roshtkhari, see para 34, 38, and FIG 2-3d). PNG media_image12.png 337 592 media_image12.png Greyscale Regarding claim 18 (dependent on claim 11), all claim limitations are met and rendered obvious by Xie in view of Roshtkhari because the method steps of claim 10 are the same as the steps in claim 18. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Teaches bi-directional object tracking: U.S. Patent No. 2002/0114394 A1 U.S. Patent No. 6,724,915 B1 U.S. Patent No. 7,817,822 B2 Firouzi, H., & Najjaran, H. (2012, October). Multiple object tracking via a two-way confidence-based correspondence algorithm. In 2012 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (pp. 745-749). IEEE. Luo, H., & Zeng, Z. (2023). Real-time multi-object tracking based on bi-directional matching. arXiv preprint arXiv:2303.08444. Javed, S., Zhang, X., Seneviratne, L., Dias, J., & Werghi, N. (2020, July). Deep bidirectional correlation filters for visual object tracking. In 2020 IEEE 23rd International Conference on Information Fusion (FUSION) (pp. 1-8). IEEE. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to EMMA E DRYDEN whose telephone number is (571)272-1179. The examiner can normally be reached M-F 8-4 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ANDREW BEE can be reached at (571) 270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /EMMA E DRYDEN/Examiner, Art Unit 2677 /ANDREW W BEE/Supervisory Patent Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Mar 06, 2024
Application Filed
Apr 02, 2026
Non-Final Rejection mailed — §101, §102, §103
Jun 09, 2026
Applicant Interview (Telephonic)
Jun 09, 2026
Examiner Interview Summary
Jun 15, 2026
Response Filed
Sep 15, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743830
GENERATING DIGITAL MATERIALS FROM DIGITAL IMAGES USING A CONTROLLED DIFFUSION NEURAL NETWORK
3y 1m to grant Granted Sep 22, 2026
Patent 12731271
ACTIVE LEARNING SYSTEM AND METHOD
2y 7m to grant Granted Sep 08, 2026
Patent 12705722
Real Time Inconsistency Detection During Composite Material Manufacturing
2y 6m to grant Granted Aug 11, 2026
Patent 12664680
LOCALIZATION AND MAPPING BY A GROUP OF MOBILE COMMUNICATIONS DEVICES
3y 8m to grant Granted Jun 23, 2026
Patent 12632966
METHOD, ELECTRONIC DEVICE, AND COMPUTER PROGRAM PRODUCT FOR RECOGNIZING OBJECT REGIONS IN IMAGE
2y 11m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
68%
Grant Probability
80%
With Interview (+12.5%)
3y 1m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 28 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month