Prosecution Insights
Last updated: October 02, 2026
Application No. 18/895,133

OBJECT REPRESENTATION VIA STATE DIAGRAMS FOR OBJECT DETECTION AND TRACKING

Non-Final OA §102§103
Filed
Sep 24, 2024
Examiner
MENDEZ MUNIZ, DYLAN JOHN
Art Unit
2675
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
19 granted / 24 resolved
+17.2% vs TC avg
Strong +28% interview lift
Without
With
+27.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
23 currently pending
Career history
44
Total Applications
across all art units

Statute-Specific Performance

§101
9.4%
-30.6% vs TC avg
§103
54.9%
+14.9% vs TC avg
§102
18.3%
-21.7% vs TC avg
§112
17.4%
-22.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 24 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) was filed on 09/24/2024. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-6, 9, 14 and 15-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Stearns et. al., hereafter Stearns (US Pub. No. 2024/0013409 A1). As per claim 1, Stearns teaches “An apparatus comprising: one or more memories; and one or more processors, coupled to the one or more memories, configured to cause the apparatus to:” (See paragraph 5. Stearns) “obtain a first frame, associated with a first time point, the first frame comprising a plurality of points corresponding to one or more first objects in a scene at the first time point;” (See paragraphs 32-37, it obtains a history of frames that are time stamped of one or more objects “[0032] Referring to FIG. 2, a block diagram of a tracking pipeline 200 of the Spatiotemporal multiple-object tracking system 100 is depicted. The block diagram depicts the interconnection of input data, intermediate data, and models. The tracking pipeline 200 may perform tracking of one or more objects on the current frame t based on the current object data, such as current time-stamped points Pt 137 and the current bounding boxes bt 147, and historical data in the historical frames (t− K≤i≤t− 1), such as historical tracklets Tt−1 including historical object point cloud segments Qt−1 127 and historical tracklet states St−1 117. The tracking pipeline 200 may then generate current tracklets Tt 229 comprising the current object point cloud segments Qt 221 and current tracklet states St 227. The current object point cloud segments Qt 221 and current tracklet states St 227 may then be used to update the historical data for the object tracking in the next frame.” Stearns) “obtain a final state diagram comprising a respective final time series sequence of predicted object states, for each second object of one or more second objects, associated with a first plurality of time points prior to the first time point or after the first time point; and” (See paragraph 30, The sequence refinement module generates a directed acyclic graph (DAG), which is interpreted as a state diagram and are also interconnected as seen on fig. 2. “[0030] The sequence refinement module 142 may be trained and provided machine learning capabilities via a neural network as described herein. By way of example, and not as a limitation, the neural network may utilize one or more artificial neural networks (ANNs). In ANNs, connections between nodes may form a directed acyclic graph (DAG). ANNs may include node inputs, one or more hidden activation layers, and node outputs, and may be utilized with activation functions…ANNs are trained by applying such activation functions to training data sets to determine an optimized solution from adjustable weights and biases applied to nodes within the hidden activation layers to generate one or more outputs… The one or more ANN models may utilize one to one, one to many, many to one, and/or many to many (e.g., sequence to sequence) sequence modeling.” Since the sequence refinement module (ANN) produces a DAG of all nodes, it therefore produces a DAG for the tracklets, the tracklets contain a time series of predicted object states “[0037] In embodiments, the historical object point cloud segments Qt−1 127, historical tracklet states St−1 117, and current object point cloud segments Qt 221 are input into the prediction module 122 to perform prediction. The current tracklets {circumflex over (T)}t are predicted based on the historical tracklets Tt−1. The current tracklets {circumflex over (T)}t comprise predicted tracklet states Ŝt 223 and the historical object point cloud segments Qt−1 127. The predicted tracklet states Ŝt 223 may then be input into the association module 132 along with the current time-stamped points Pt 137 and the current bounding boxes bt 147. The association module may associate the current detected data and the predicted data by comparing these data to generate associated tracklets Tt comprising tracklet states St 225 and the current object point cloud segments Qt 221. The associated tracklets Tt may then be input into the sequence refinement module 142 to conduct a posterior tracklet update via a neural network module 242. The posterior tracklet update generates a refined current tracklet states St 227. The Spatiotemporal multiple-object tracking system 100 may then output a current tracklet Tt 229 comprising Qt 221 and St 227.” See also paragraph 43-45 “he SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St… ” Examiner interprets “final state diagram” as any created state diagram such as the current (the current final state is an aggregation of all states). See also paragraphs 37-40. See also paragraph 21. See also fig. 2 along with unit 242 shows the interconnected nodes of a state diagram. See all of paragraph 39. Stearns) “process the first frame and the final state diagram to detect the one or more first objects in the scene at the first time point.” (The presented tracklets already contain the detected objects. See also paragraphs 55, “0055] Still referring to FIG. 4, the SSR module 142 includes a decoder 420 that outputs a refined, ordered sequence of object states St that is amenable to association in subsequent frames. To output object state trajectories, a decoder without an explicit prior, such as a decoder that directly predicts the ordered sequence of bounding boxes in one forward pass is used.” See also paragraph 61. See also paragraph 23 “[0023] As described in more detail herein, embodiments of the present disclosure provide systems and methods of spatiotemporal object tracking by actively maintaining the history of both object-level point clouds and bounding boxes for each tracked object. The method disclosed herein provides embodiments that efficiently maintain an active history of object-level point clouds and bounding boxes for each tracklet. At each frame, new object detections are associated with these maintained past sequences of object points and tracklet status. The sequences are then updated using a 4D backbone to refine the sequence of bounding boxes and to predict the current tracklets, both of which are used to further forecast the tracklet of the object into the next frame.” See also paragraphs 28, and 32-46, along with paragraphs 58-62 and 64. “[0059] In embodiments, with the predictions, two post-processed representations can be computed. First, using the per-frame center estimates for an extended sequence, a quadratic regression is implemented to obtain a second-order motion approximation of the object. Second, using the refined object centers, yaws, and segmentation from the generated predictions for the current tracklet states St 227, the original object points can be transformed into the rigid-canonical reference frame. This yields a canonical, aggregated representation of the object which can be referenced for shape analysis, such as to determine shape size, configuration, dimensions, and the like.” Unit 142 contains unit 242. See all of paragraph 39. Stearns) Claim 15 is rejected under the same analysis as claim 1. Claim 20 is rejected under the same analysis as claim 1. (Paragraph 6 shows non transitory computer readable media. Stearns) As per claim 2, Stearns already teaches “the apparatus of claim 1, wherein each respective final time series sequence of predicted object states is represented as a respective plurality of interconnected final nodes in the final state diagram.” (The sequence refinement module which contains the ANN, already has interconnected nodes. See paragraph 30, The sequence refinement module generates a directed acyclic graph (DAG), which is interpreted as a state diagram. “[0030] The sequence refinement module 142 may be trained and provided machine learning capabilities via a neural network as described herein. By way of example, and not as a limitation, the neural network may utilize one or more artificial neural networks (ANNs). In ANNs, connections between nodes may form a directed acyclic graph (DAG). ANNs may include node inputs, one or more hidden activation layers, and node outputs, and may be utilized with activation functions…ANNs are trained by applying such activation functions to training data sets to determine an optimized solution from adjustable weights and biases applied to nodes within the hidden activation layers to generate one or more outputs… The one or more ANN models may utilize one to one, one to many, many to one, and/or many to many (e.g., sequence to sequence) sequence modeling.” Since the sequence refinement module (ANN) produces a DAG of all nodes, it therefore produces a DAG for the tracklets, the tracklets contain a time series of predicted object states (all are nodes in the ANN). See also paragraphs 44, 29 and 47. See also fig. 2, unit 242 shows the interconnected nodes of a state diagram. Unit 142 contains unit 242. See all of paragraph 39. Stearns ) Claim 16 is rejected under the same analysis as claim 2. As per claim 3, Stearns already teaches “the apparatus of claim 1, wherein the one or more second objects comprise at least the one or more first objects.” (The presented system already performs object detection for all of the one or more objects. This is within a BRI (broadest reasonable interpretation) of a first and second object. See all of paragraph 39 “…The spatiotemporal multiple-object tracking system 100 generates object tracklet states for each detected objects, for example St(1) for object 1 and St(2) for object 2 in 301. The spatiotemporal multiple-object tracking system 100 accesses a historical tracklet states St−1 303 with M historical detected objects. For example, as illustrated, three historical detected objects are found in the historical frames… Through the comparison, the spatiotemporal multiple-object tracking system 100 may detect that one or more historical detected objects are not observed in the current detection… For example, in FIG. 3, the sequence refinement refines the tracklets for N objects, where N≤M. ” See also paragraph “0032] Referring to FIG. 2, a block diagram of a tracking pipeline 200 of the Spatiotemporal multiple-object tracking system 100 is depicted. The block diagram depicts the interconnection of input data, intermediate data, and models. The tracking pipeline 200 may perform tracking of one or more objects on the current frame t based on the current object data, such as current time-stamped points Pt 137 and the current bounding boxes bt 147, and historical data in the historical frames (t− K≤i≤t− 1), such as historical tracklets Tt−1 including historical object point cloud segments Qt−1 127 and historical tracklet states St−1 117.” Stearns) Claim 17 is rejected under the same analysis as claim 3. As per claim 4, Stearns already teaches “the apparatus of claim 1, wherein the one or more processors are configured to cause the apparatus to: process the first frame and the final state diagram to detect at least one of the one or more second objects in the scene at the first time point.” (The processes shown prior for the rejection of claim 1 already teach the claim limitations. All objects at all time points are detected. See all of paragraph 39 “… The current environment data 301 may include the current time-stamped points Pt 137, the current bounding boxes bt 147. The spatiotemporal multiple-object tracking system 100 generates object tracklet states for each detected objects, for example St(1) for object 1 and St(2) for object 2 in 301. The spatiotemporal multiple-object tracking system 100 accesses a historical tracklet states St−1 303 with M historical detected objects. For example, as illustrated, three historical detected objects are found in the historical frames. Each St−1 117 comprises (K−1) sequence states Si from St−1 to St−k in the historical frames… The spatiotemporal multiple-object tracking system 100 compares the predicted tracklet states Ŝt 223 with the current time-stamped points Pt 137 and the current bounding boxes bt 147 to associate the prediction data and current detected data and arrive at an associated detection 311. Through the comparison, the spatiotemporal multiple-object tracking system 100 may detect that one or more historical detected objects are not observed in the current detection.” . The presented tracklets already contain the detected objects. See also paragraphs 55, “0055] Still referring to FIG. 4, the SSR module 142 includes a decoder 420 that outputs a refined, ordered sequence of object states St that is amenable to association in subsequent frames. To output object state trajectories, a decoder without an explicit prior, such as a decoder that directly predicts the ordered sequence of bounding boxes in one forward pass is used.” See also paragraph 61. See also paragraph 23 “[0023] As described in more detail herein, embodiments of the present disclosure provide systems and methods of spatiotemporal object tracking by actively maintaining the history of both object-level point clouds and bounding boxes for each tracked object. The method disclosed herein provides embodiments that efficiently maintain an active history of object-level point clouds and bounding boxes for each tracklet. At each frame, new object detections are associated with these maintained past sequences of object points and tracklet status. The sequences are then updated using a 4D backbone to refine the sequence of bounding boxes and to predict the current tracklets, both of which are used to further forecast the tracklet of the object into the next frame.” See also paragraphs 28, and 32-46, along with paragraphs 58-62 and 64. “[0059] In embodiments, with the predictions, two post-processed representations can be computed. First, using the per-frame center estimates for an extended sequence, a quadratic regression is implemented to obtain a second-order motion approximation of the object. Second, using the refined object centers, yaws, and segmentation from the generated predictions for the current tracklet states St 227, the original object points can be transformed into the rigid-canonical reference frame. This yields a canonical, aggregated representation of the object which can be referenced for shape analysis, such as to determine shape size, configuration, dimensions, and the like.” Unit 142 contains unit 242. See all of paragraph 39. Stearns) Claim 18 is rejected under the same analysis as claim 4. As per claim 5, Stearns teaches “the apparatus of claim 1, wherein the final state diagram comprises a graph neural network.” (See paragraph 30 “[0030] The sequence refinement module 142 may be trained and provided machine learning capabilities via a neural network as described herein. By way of example, and not as a limitation, the neural network may utilize one or more artificial neural networks (ANNs). In ANNs, connections between nodes may form a directed acyclic graph (DAG).” See also fig. 2. The neural network produces graphs. Stearns) Claim 19 is rejected under the same analysis as claim 5. As per claim 6, Stearns teaches “the apparatus of claim 1, wherein each respective final time series sequence of predicted object states is associated with the first plurality of time points prior to the first time point.” (The presented paragraphs in claim 1 already cover the limitations of this claim which already associates object states with all time points, prior and after. See paragraph 43 “[0043] Referring to FIG. 4, an illustrative block diagram of a spatiotemporal sequence-to-sequence refinement (SSR) module 142 is depicted. The SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St. The SSR module 142 first processes the sequential information with a 4D backbone to extract per-point context features. In a decoding stage, it predicts a global object size across all frames, as well as per-frame time-relevant object attributes including center, pose, velocity, and confidence…”. See paragraph 37 “” See also paragraph 34 “0034] In embodiments, the current object point cloud segments Qt 221 of an object at frame t represent the spatiotemporal information in the form of time-stamped points in the history frames and the current frame… The current object point cloud segments Qt 221 encode the spatiotemporal information from raw sensor observation in the form of time-stamped points in the frames (t−K≤i≤t).” Therefore the time stamped points of all frames are utilized. Stearns) As per claim 9, Stearns already teaches “the apparatus of claim 1, wherein each respective final time series sequence of predicted object states is associated with the first plurality of time points after the first time point. (The presented paragraphs in claim 1 already cover the limitations of this claim, which already associates object states with all time points, prior and after. See paragraph 43 “[0043] Referring to FIG. 4, an illustrative block diagram of a spatiotemporal sequence-to-sequence refinement (SSR) module 142 is depicted. The SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St. The SSR module 142 first processes the sequential information with a 4D backbone to extract per-point context features. In a decoding stage, it predicts a global object size across all frames, as well as per-frame time-relevant object attributes including center, pose, velocity, and confidence…”. See paragraph 37 “” See also paragraph 34 “0034] In embodiments, the current object point cloud segments Qt 221 of an object at frame t represent the spatiotemporal information in the form of time-stamped points in the history frames and the current frame… The current object point cloud segments Qt 221 encode the spatiotemporal information from raw sensor observation in the form of time-stamped points in the frames (t−K≤i≤t).” Therefore the time stamped points of all frames are utilized. Stearns) As per claim 12, Stearns already teaches “the apparatus of claim 1, wherein to process the first frame and the final state diagram to detect the one or more first objects, the one or more processors are configured to cause the apparatus to:… than all of a respective plurality of interconnected final nodes associated with at least one respective final time series sequence of object states.”, however Stearns also teaches “process less…” (See paragraph 62 “[0062] It should be understood that blocks of the aforementioned process may be omitted or performed in a variety of orders while still achieving the object of the present disclosure.” See also paragraph 39 “The spatiotemporal multiple-object tracking system 100 accesses a historical tracklet states St−1 303 with M historical detected objects. For example, as illustrated, three historical detected objects are found in the historical frames. Each St−1 117 comprises (K−1) sequence states Si from St−1 to St−k in the historical frames, where K indicates a pre-determined length of maximum history and is greater than or equals to 2. In embodiments, K may be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, or 40. ” Since K-1 states are used, therefore less than all are processed. Stearns) As per claim 14, Stearns already teaches “the apparatus of claim 1, wherein each respective predicted object state of each respective final time series sequence of predicted object states associated with each respective second object comprises at least one of:”, however Stearns also teaches as least one of “a size of the respective second object; a location of the respective second object in the scene; an orientation of the respective second object; a pose estimation of the respective second object; one or more shape descriptors associated with the respective second object; one or more visual features of the respective second object; a velocity of the respective second object; an acceleration of the respective second object; a heading of the respective second object; a semantic class associated with the respective second object; a semantic class confidence score; a trajectory score associated with the respective second object; one or more confidence scores; a trajectory standard deviation; time elapsed since a last detection of the respective second object; one or more dynamics of the scene; an occlusion state of the respective second object; one or more interaction features; an environmental context; an appearance change rate; a measure of a consistency of the respective second object; a tracking history of the respective second object; a predicted future position of the respective second object; a sensor modality confidence score; scene flow information; or optical flow information.” (Paragraph 35 shows object size. Paragraph 59 shows object shape and dimensions. Paragraph 43 shows velocity and confidence. Paragraphs 4-6, 60 shows tracking history. Paragraphs 66 and 68 show occlusions. Paragraphs 38 and 68 shows future predictions. See also paragraph 20. The objects are visually detected. Stearns) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Stearns in view of Saalbach et. al. (US Pub. No. 20210177296 A1) . As per claim 7, Stearns teaches “the apparatus of claim 6, wherein to obtain the final state diagram, the one or more processors are configured to cause the apparatus to: obtain a time series sequence of frames for the scene associated with a second plurality of time points prior to the first time point;” (See paragraph 34 “[0034] In embodiments, the current object point cloud segments Qt 221 of an object at frame t represent the spatiotemporal information in the form of time-stamped points in the history frames and the current frame. The current object point cloud segments Qt 221 may comprise the cropped point cloud regions {circumflex over (P)}i in the historical frames and the current frame (t−K≤i≤t), where a cropped point cloud regions {circumflex over (P)}i at frame i are the points cropped from the time-stamped points Pi according to an associated bounding boxes bi. The current object point cloud segments Qt 221 encode the spatiotemporal information from raw sensor observation in the form of time-stamped points in the frames (t−K≤i≤t).” The time stamp is used from all frames in the sequence, the history frames are the prior frames. Stearns) “divide the time series sequence of frames into a plurality of time series subsequences of frames, wherein each time series subsequence of frames is associated with a respective subset of the plurality of second time points;” (Paragraphs 55-56 show a predicted sequence of bounding boxes in a forward pass that uses subsequent frames, the object attributes are also measured (which contain the subsequences as seen in paragraph 43). See also paragraphs 34 and 43, the time series of sequence of frames contain the time series subsequences of frames as they have all the time related time stamps for object segments. They are already associated with a subset of time points as they are time stamped such as the time stamps for the object attributes. Therefore there is also a division of time series. “[0034] In embodiments, the current object point cloud segments Qt 221 of an object at frame t represent the spatiotemporal information in the form of time-stamped points in the history frames and the current frame. The current object point cloud segments Qt 221 may comprise the cropped point cloud regions {circumflex over (P)}i in the historical frames and the current frame (t−K≤i≤t), where a cropped point cloud regions {circumflex over (P)}i at frame i are the points cropped from the time-stamped points Pi according to an associated bounding boxes bi. The current object point cloud segments Qt 221 encode the spatiotemporal information from raw sensor observation in the form of time-stamped points in the frames (t−K≤i≤t).” “[0043] Referring to FIG. 4, an illustrative block diagram of a spatiotemporal sequence-to-sequence refinement (SSR) module 142 is depicted. The SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St. The SSR module 142 first processes the sequential information with a 4D backbone to extract per-point context features. In a decoding stage, it predicts a global object size across all frames, as well as per-frame time-relevant object attributes including center, pose, velocity, and confidence…”. See also all of paragraph 45 “…The encoder 410 appends the bounding-box information in the associated tracklet states St to each frame in the current object point cloud segments Qt 221 to yield a set of object-aware features. The object-aware features include time-stamped locations of the time-stamped points and parameters of bounding boxes…” Stearns) “for each time series subsequence of frames of the plurality of time series subsequences of frames: generate a respective state diagram comprising a respective time series sequence of object states for at least one second object of the one or more second objects over the respective subset of the plurality of second time points… a respective last time point, wherein each respective time series sequence of object states is represented as a respective plurality of interconnected nodes in the respective state diagram; and” (See paragraph 30, The sequence refinement module generates a directed acyclic graph (DAG), which is interpreted as a state diagram and are also interconnected as seen on fig. 2. “[0030] The sequence refinement module 142 may be trained and provided machine learning capabilities via a neural network as described herein. By way of example, and not as a limitation, the neural network may utilize one or more artificial neural networks (ANNs). In ANNs, connections between nodes may form a directed acyclic graph (DAG). ANNs may include node inputs, one or more hidden activation layers, and node outputs, and may be utilized with activation functions…ANNs are trained by applying such activation functions to training data sets to determine an optimized solution from adjustable weights and biases applied to nodes within the hidden activation layers to generate one or more outputs… The one or more ANN models may utilize one to one, one to many, many to one, and/or many to many (e.g., sequence to sequence) sequence modeling.” Since the sequence refinement module (ANN) produces a DAG of all nodes, it therefore produces a DAG for the tracklets, the tracklets contain a time series of predicted object states “[0037] In embodiments, the historical object point cloud segments Qt−1 127, historical tracklet states St−1 117, and current object point cloud segments Qt 221 are input into the prediction module 122 to perform prediction. The current tracklets {circumflex over (T)}t are predicted based on the historical tracklets Tt−1. The current tracklets {circumflex over (T)}t comprise predicted tracklet states Ŝt 223 and the historical object point cloud segments Qt−1 127. The predicted tracklet states Ŝt 223 may then be input into the association module 132 along with the current time-stamped points Pt 137 and the current bounding boxes bt 147. The association module may associate the current detected data and the predicted data by comparing these data to generate associated tracklets Tt comprising tracklet states St 225 and the current object point cloud segments Qt 221. The associated tracklets Tt may then be input into the sequence refinement module 142 to conduct a posterior tracklet update via a neural network module 242. The posterior tracklet update generates a refined current tracklet states St 227. The Spatiotemporal multiple-object tracking system 100 may then output a current tracklet Tt 229 comprising Qt 221 and St 227.” See also paragraph 43-45 “he SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St… ” Examiner interprets “final state diagram” as any created state diagram such as the current (the current final state is an aggregation of all states). See also paragraphs 37-43. See also paragraph 21. See also fig. 2 along with unit 242 shows the interconnected nodes of a state diagram. See all of paragraph 39. The points are time stamped as seen on paragraph 34. Stearns) “perform forward motion forecasting to determine predicted object states for the at least one second object at the respective last time point based on the respective state diagram; and concatenate the predicted object states determined for the plurality of time series subsequences of frames.” (Paragraphs 3 and 23 show forward motion forecasting by predicting object trajectory in a forward aspect. Paragraphs 55-56 show a predicted sequence of bounding boxes in a forward pass that uses subsequent frames, the object attributes are also measured (which contain the subsequences as seen in paragraph 43). “[0055] Still referring to FIG. 4, the SSR module 142 includes a decoder 420 that outputs a refined, ordered sequence of object states St that is amenable to association in subsequent frames. To output object state trajectories, a decoder without an explicit prior, such as a decoder that directly predicts the ordered sequence of bounding boxes in one forward pass is used.” See paragraphs 30, 37, “[0037] In embodiments, the historical object point cloud segments Qt−1 127, historical tracklet states St−1 117, and current object point cloud segments Qt 221 are input into the prediction module 122 to perform prediction. The current tracklets {circumflex over (T)}t are predicted based on the historical tracklets Tt−1. The current tracklets {circumflex over (T)}t comprise predicted tracklet states Ŝt 223 and the historical object point cloud segments Qt−1 127. The predicted tracklet states Ŝt 223 may then be input into the association module 132 along with the current time-stamped points Pt 137 and the current bounding boxes bt 147. The association module may associate the current detected data and the predicted data by comparing these data to generate associated tracklets Tt comprising tracklet states St 225 and the current object point cloud segments Qt 221. The associated tracklets Tt may then be input into the sequence refinement module 142 to conduct a posterior tracklet update via a neural network module 242. The posterior tracklet update generates a refined current tracklet states St 227. The Spatiotemporal multiple-object tracking system 100 may then output a current tracklet Tt 229 comprising Qt 221 and St 227.” See also paragraph 43-45 “The SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St… ” See also paragraph 23 “[0023] As described in more detail herein, embodiments of the present disclosure provide systems and methods of spatiotemporal object tracking by actively maintaining the history of both object-level point clouds and bounding boxes for each tracked object. The method disclosed herein provides embodiments that efficiently maintain an active history of object-level point clouds and bounding boxes for each tracklet. At each frame, new object detections are associated with these maintained past sequences of object points and tracklet status. The sequences are then updated using a 4D backbone to refine the sequence of bounding boxes and to predict the current tracklets, both of which are used to further forecast the tracklet of the object into the next frame.” See also paragraphs 28, and 32-46, along with paragraphs 58-62 and 64. “[0059] In embodiments, with the predictions, two post-processed representations can be computed. First, using the per-frame center estimates for an extended sequence, a quadratic regression is implemented to obtain a second-order motion approximation of the object. Second, using the refined object centers, yaws, and segmentation from the generated predictions for the current tracklet states St 227, the original object points can be transformed into the rigid-canonical reference frame. This yields a canonical, aggregated representation of the object which can be referenced for shape analysis, such as to determine shape size, configuration, dimensions, and the like.” Unit 142 contains unit 242. Examiner interprets “final state diagram” as any created state diagram such as the current (the current final state is an aggregation of all states). See also paragraphs 37-40. See also paragraph 21. See also fig. 2 along with unit 242 shows the interconnected nodes of a state diagram. See all of paragraph 39. Stearns), Stearns also implicitly teaches “omitting a respective last time point” (See paragraphs 36-38, the the frame i-1 and sequence states k-1 are utilized and therefore implicitly omits a respective time point. Stearns) however Stearns does not completely teach “omitting a respective… time point” Saalbach teaches “omitting a respective… time point” (See paragraph 45 “0045] The image Im.sub.i may be subjected to the motion classifier provided by the first machine learning resulting in a prediction of a motion level L.sub.i. If L.sub.i>L.sub.i−1, i.e. if the artifact level increases due to including the latest dataset of time period P.sub.i, then it is likely that the patient has moved during the respective time period P.sub.i…. According to embodiments, it may be checked, whether the inclusion of the MRI data acquired during the following time period P.sub.i+1 instead of the MRI data of time period P.sub.i also results in an increased artifact level. In other words, a reduced image rIm.sub.i+1 is reconstructed taking into account MRI data from all time periods P.sub.k with k≤i+1 except of k=1. In other words, the MRI data from time period P.sub.i is omitted.” Saalbach) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Stearns with the teachings of Saalbach to omit a specific timepoint. The modification would have been motivated by the desire to increase prediction of motion of an object and reduce images, as suggested by Saalbach (See paragraphs 45-46 “…According to embodiments, it may be checked, whether the inclusion of the MRI data acquired during the following time period P.sub.i+1 instead of the MRI data of time period P.sub.i also results in an increased artifact level. In other words, a reduced image rIm.sub.i+1 is reconstructed taking into account MRI data from all time periods P.sub.k with k≤i+1 except of k=1. In other words, the MRI data from time period P.sub.i is omitted. [0046] In case, an omission of the MRI data from time period P.sub.i results in an increase of the prediction for the motion level L.sub.i+1 relative to the prediction for L.sub.i−1, the patient may have reached a new posture…” Saalbach) As per claim 8, Stearns in view of Saalbach already teaches “the apparatus of claim 7, wherein each respective state diagram comprises a graph neural network.” (See paragraph 30 “[0030] The sequence refinement module 142 may be trained and provided machine learning capabilities via a neural network as described herein. By way of example, and not as a limitation, the neural network may utilize one or more artificial neural networks (ANNs). In ANNs, connections between nodes may form a directed acyclic graph (DAG).” See also fig. 2. The neural network produces graphs. Stearns) Claims 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Stearns in view of Li et. al. (Li, Yingwei, et al. "Modar: Using motion forecasting for 3d object detection in point cloud sequences." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.) and further in view of in view of Saalbach et. al. (US Pub. No. 20210177296 A1) . As per claim 10, Stearns already teaches “the apparatus of claim 9, wherein to obtain the final state diagram, the one or more processors are configured to cause the apparatus to: obtain a time series sequence of frames for the scene associated with a second plurality of time points after the first time point;” (See paragraph 34 “[0034] In embodiments, the current object point cloud segments Qt 221 of an object at frame t represent the spatiotemporal information in the form of time-stamped points in the history frames and the current frame. The current object point cloud segments Qt 221 may comprise the cropped point cloud regions {circumflex over (P)}i in the historical frames and the current frame (t−K≤i≤t), where a cropped point cloud regions {circumflex over (P)}i at frame i are the points cropped from the time-stamped points Pi according to an associated bounding boxes bi. The current object point cloud segments Qt 221 encode the spatiotemporal information from raw sensor observation in the form of time-stamped points in the frames (t−K≤i≤t).” The time stamp is used from all frames in the sequence, the history frames are the prior frames. Stearns) “divide the time series sequence of frames into a plurality of time series subsequences of frames, wherein each time series subsequence of frames is associated with a respective subset of the second plurality of time points;” (Paragraphs 55-56 show a predicted sequence of bounding boxes in a forward pass that uses subsequent frames, the object attributes are also measured (which contain the subsequences as seen in paragraph 43). See also paragraphs 34 and 43, the time series of sequence of frames contain the time series subsequences of frames as they have all the time related time stamps for object segments. They are already associated with a subset of time points as they are time stamped such as the time stamps for the object attributes. Therefore there is also a division of time series. “[0034] In embodiments, the current object point cloud segments Qt 221 of an object at frame t represent the spatiotemporal information in the form of time-stamped points in the history frames and the current frame. The current object point cloud segments Qt 221 may comprise the cropped point cloud regions {circumflex over (P)}i in the historical frames and the current frame (t−K≤i≤t), where a cropped point cloud regions {circumflex over (P)}i at frame i are the points cropped from the time-stamped points Pi according to an associated bounding boxes bi. The current object point cloud segments Qt 221 encode the spatiotemporal information from raw sensor observation in the form of time-stamped points in the frames (t−K≤i≤t).” “[0043] Referring to FIG. 4, an illustrative block diagram of a spatiotemporal sequence-to-sequence refinement (SSR) module 142 is depicted. The SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St. The SSR module 142 first processes the sequential information with a 4D backbone to extract per-point context features. In a decoding stage, it predicts a global object size across all frames, as well as per-frame time-relevant object attributes including center, pose, velocity, and confidence…”. See also all of paragraph 45 “…The encoder 410 appends the bounding-box information in the associated tracklet states St to each frame in the current object point cloud segments Qt 221 to yield a set of object-aware features. The object-aware features include time-stamped locations of the time-stamped points and parameters of bounding boxes…” Stearns) “for each time series subsequence of frames of the plurality of time series subsequences of frames: generate a respective state diagram comprising a respective time series sequence of object states for at least one second object of the one or more second objects over the respective subset of the second plurality of time points… a respective first time point, wherein each respective time series sequence of object states is represented as a respective plurality of interconnected nodes in the respective state diagram; and” (See paragraph 30, The sequence refinement module generates a directed acyclic graph (DAG), which is interpreted as a state diagram and are also interconnected as seen on fig. 2. “[0030] The sequence refinement module 142 may be trained and provided machine learning capabilities via a neural network as described herein. By way of example, and not as a limitation, the neural network may utilize one or more artificial neural networks (ANNs). In ANNs, connections between nodes may form a directed acyclic graph (DAG). ANNs may include node inputs, one or more hidden activation layers, and node outputs, and may be utilized with activation functions…ANNs are trained by applying such activation functions to training data sets to determine an optimized solution from adjustable weights and biases applied to nodes within the hidden activation layers to generate one or more outputs… The one or more ANN models may utilize one to one, one to many, many to one, and/or many to many (e.g., sequence to sequence) sequence modeling.” Since the sequence refinement module (ANN) produces a DAG of all nodes, it therefore produces a DAG for the tracklets, the tracklets contain a time series of predicted object states “[0037] In embodiments, the historical object point cloud segments Qt−1 127, historical tracklet states St−1 117, and current object point cloud segments Qt 221 are input into the prediction module 122 to perform prediction. The current tracklets {circumflex over (T)}t are predicted based on the historical tracklets Tt−1. The current tracklets {circumflex over (T)}t comprise predicted tracklet states Ŝt 223 and the historical object point cloud segments Qt−1 127. The predicted tracklet states Ŝt 223 may then be input into the association module 132 along with the current time-stamped points Pt 137 and the current bounding boxes bt 147. The association module may associate the current detected data and the predicted data by comparing these data to generate associated tracklets Tt comprising tracklet states St 225 and the current object point cloud segments Qt 221. The associated tracklets Tt may then be input into the sequence refinement module 142 to conduct a posterior tracklet update via a neural network module 242. The posterior tracklet update generates a refined current tracklet states St 227. The Spatiotemporal multiple-object tracking system 100 may then output a current tracklet Tt 229 comprising Qt 221 and St 227.” See also paragraph 43-45 “he SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St… ” Examiner interprets “final state diagram” as any created state diagram such as the current (the current final state is an aggregation of all states). See also paragraphs 37-43. See also paragraph 21. See also fig. 2 along with unit 242 shows the interconnected nodes of a state diagram. See all of paragraph 39. The points are time stamped as seen on paragraph 34. Stearns) “perform… motion forecasting to determine predicted object states for the at least one second object at the respective first time point based on the respective state diagram; and concatenate the predicted object states determined for the plurality of time series subsequences of frames.” (Paragraphs 3 and 23 show motion forecasting by predicting object trajectory. Paragraphs 55-56 show a predicted sequence of bounding boxes in a forward pass that uses subsequent frames, the object attributes are also measured (which contain the subsequences as seen in paragraph 43). “[0055] Still referring to FIG. 4, the SSR module 142 includes a decoder 420 that outputs a refined, ordered sequence of object states St that is amenable to association in subsequent frames. To output object state trajectories, a decoder without an explicit prior, such as a decoder that directly predicts the ordered sequence of bounding boxes in one forward pass is used.” See paragraphs 30, 37, “[0037] In embodiments, the historical object point cloud segments Qt−1 127, historical tracklet states St−1 117, and current object point cloud segments Qt 221 are input into the prediction module 122 to perform prediction. The current tracklets {circumflex over (T)}t are predicted based on the historical tracklets Tt−1. The current tracklets {circumflex over (T)}t comprise predicted tracklet states Ŝt 223 and the historical object point cloud segments Qt−1 127. The predicted tracklet states Ŝt 223 may then be input into the association module 132 along with the current time-stamped points Pt 137 and the current bounding boxes bt 147. The association module may associate the current detected data and the predicted data by comparing these data to generate associated tracklets Tt comprising tracklet states St 225 and the current object point cloud segments Qt 221. The associated tracklets Tt may then be input into the sequence refinement module 142 to conduct a posterior tracklet update via a neural network module 242. The posterior tracklet update generates a refined current tracklet states St 227. The Spatiotemporal multiple-object tracking system 100 may then output a current tracklet Tt 229 comprising Qt 221 and St 227.” See also paragraph 43-45 “The SSR module 142 takes the associated tracklet states St and the time-stamped object point cloud segments Qt as input and outputs refined final tracklet states St… ” See also paragraph 23 “[0023] As described in more detail herein, embodiments of the present disclosure provide systems and methods of spatiotemporal object tracking by actively maintaining the history of both object-level point clouds and bounding boxes for each tracked object. The method disclosed herein provides embodiments that efficiently maintain an active history of object-level point clouds and bounding boxes for each tracklet. At each frame, new object detections are associated with these maintained past sequences of object points and tracklet status. The sequences are then updated using a 4D backbone to refine the sequence of bounding boxes and to predict the current tracklets, both of which are used to further forecast the tracklet of the object into the next frame.” See also paragraphs 28, and 32-46, along with paragraphs 58-62 and 64. “[0059] In embodiments, with the predictions, two post-processed representations can be computed. First, using the per-frame center estimates for an extended sequence, a quadratic regression is implemented to obtain a second-order motion approximation of the object. Second, using the refined object centers, yaws, and segmentation from the generated predictions for the current tracklet states St 227, the original object points can be transformed into the rigid-canonical reference frame. This yields a canonical, aggregated representation of the object which can be referenced for shape analysis, such as to determine shape size, configuration, dimensions, and the like.” Unit 142 contains unit 242. Examiner interprets “final state diagram” as any created state diagram such as the current (the current final state is an aggregation of all states). See also paragraphs 37-40. See also paragraph 21. See also fig. 2 along with unit 242 shows the interconnected nodes of a state diagram. See all of paragraph 39. Stearns), Stearns also implicitly teaches “omitting a respective first time point” (See paragraphs 36-38, the the frame i-1 and sequence states k-1 are utilized and therefore implicitly omits a respective time point. Stearns). however Stearns does not teach “omitting a respective… time point” and “backwards motion forecasting” Li teaches “backwards motion forecasting” (See page 2 column 1 paragraph 2 “In an offboard/offline detection setup, we can use both forward prediction and reverse pre diction (use future frames as input to the forecasting model) to combine information from the past and the future.” See also page 7 column 1 paragraphs 1-2 Li), Li also teaches subsequences of frames that concatenates. (See page 3 fig. 3 along with the accompanying paragraph. Li) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Stearns with the teachings of Li to perform backwards motion forecasting to concatenate the subsequence of frames. The modification would have been motivated by the desire to have better detection performance, as suggested by Li (See page 7 column 1 paragraphs 1-2 “We observe that for both setting, adding MoDARs from more past or future predictions generally lead to better detection and this improvement does not saturate until using MoDARs from 80 past and 80 future predictions. It is also noteworthy that the future frames provide unique information that significantly improves the results compared to only using past frames.” Li) Saalbach teaches “omitting a respective… time point” (See paragraph 45 “0045] The image Im.sub.i may be subjected to the motion classifier provided by the first machine learning resulting in a prediction of a motion level L.sub.i. If L.sub.i>L.sub.i−1, i.e. if the artifact level increases due to including the latest dataset of time period P.sub.i, then it is likely that the patient has moved during the respective time period P.sub.i…. According to embodiments, it may be checked, whether the inclusion of the MRI data acquired during the following time period P.sub.i+1 instead of the MRI data of time period P.sub.i also results in an increased artifact level. In other words, a reduced image rIm.sub.i+1 is reconstructed taking into account MRI data from all time periods P.sub.k with k≤i+1 except of k=1. In other words, the MRI data from time period P.sub.i is omitted.” Saalbach) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Stearns with the teachings of Saalbach and Li to omit a specific timepoint. The modification would have been motivated by the desire to increase prediction of motion of an object and reduce images, as suggested by Saalbach (See paragraphs 45-46 “…According to embodiments, it may be checked, whether the inclusion of the MRI data acquired during the following time period P.sub.i+1 instead of the MRI data of time period P.sub.i also results in an increased artifact level. In other words, a reduced image rIm.sub.i+1 is reconstructed taking into account MRI data from all time periods P.sub.k with k≤i+1 except of k=1. In other words, the MRI data from time period P.sub.i is omitted. [0046] In case, an omission of the MRI data from time period P.sub.i results in an increase of the prediction for the motion level L.sub.i+1 relative to the prediction for L.sub.i−1, the patient may have reached a new posture…” Saalbach) As per claim 11, Stearns in view of Li and Saalbach already teaches “the apparatus of claim 10, wherein each respective state diagram comprises a graph neural network.” (See paragraph 30 “[0030] The sequence refinement module 142 may be trained and provided machine learning capabilities via a neural network as described herein. By way of example, and not as a limitation, the neural network may utilize one or more artificial neural networks (ANNs). In ANNs, connections between nodes may form a directed acyclic graph (DAG).” See also fig. 2. The neural network produces graphs. Stearns) Claims 13 is rejected under 35 U.S.C. 103 as being unpatentable over Stearns in view of Liu et. al. (US Pub. No. 20250232510 A1) . As per claim 13, Stearns already teaches “the apparatus of claim 1, wherein the first frame comprises a… point cloud.”, however Stearns does not teach “a sparse point cloud” Liu teaches “a sparse point cloud” (See paragraph 43 “[0043] In block 420, a sparse point cloud for the target object is determined based on the plurality of key image frames. A point cloud is a data structure to represent the morphology of an object in three-dimensional space…Compared with a dense point cloud, a sparse point cloud uses fewer points to show various visual information of the three-dimensional model. Generating sparse point clouds requires less computation, has lower requirements on hardware configuration, and has faster generation speed.” Liu) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Stearns with the teachings of Liu to utilize a sparse dense point cloud for frames. The modification would have been motivated by the desire to require less computation power and have faster generation speed, as suggested by Liu (See paragraph 43 “[0043] In block 420, a sparse point cloud for the target object is determined based on the plurality of key image frames. A point cloud is a data structure to represent the morphology of an object in three-dimensional space…Compared with a dense point cloud, a sparse point cloud uses fewer points to show various visual information of the three-dimensional model. Generating sparse point clouds requires less computation, has lower requirements on hardware configuration, and has faster generation speed.” Liu) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DYLAN J MENDEZ MUNIZ whose telephone number is (703)756-5672. The examiner can normally be reached M-F, 8AM - 5PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Moyer can be reached at (571) 272-9523. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DYLAN JOHN MENDEZ MUNIZ/Examiner, Art Unit 2675 /ANDREW M MOYER/Supervisory Patent Examiner, Art Unit 2675
Read full office action

Prosecution Timeline

Sep 24, 2024
Application Filed
Jun 29, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743796
A METHOD FOR CALCULATING INFORMATION RELATIVE TO A RELATIVE SPEED BETWEEN AN OBJECT AND A CAMERA, A CONTROL METHOD FOR A VEHICLE, A COMPUTER PROGRAM, A COMPUTER-READABLE RECORDING MEDIUM, AN OBJECT MOTION ANALYSIS SYSTEM AND A CONTROL SYSTEM
3y 7m to grant Granted Sep 22, 2026
Patent 12688600
Establishing Interactions Between Dynamic Objects and Quasi-Static Objects
3y 3m to grant Granted Jul 21, 2026
Patent 12688710
REARWARD WHITE LINE INFERENCE DEVICE, TARGET RECOGNITION DEVICE, AND METHOD
2y 7m to grant Granted Jul 21, 2026
Patent 12670692
TRANSFER LEARNING BY DOWNSCALING AND UPSCALING
2y 7m to grant Granted Jun 30, 2026
Patent 12664637
METHOD AND APPARATUS FOR ANALYZING AN IMAGE OF A MICROLITHOGRAPHIC MICROSTRUCTURED COMPONENT
4y 1m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+27.8%)
2y 11m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 24 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month