Prosecution Insights
Last updated: October 01, 2026
Application No. 18/605,121

MULTI-SENSOR SUBJECT TRACKING FOR MONITORED ENVIRONMENTS FOR REAL-TIME AND NEAR-REAL-TIME SYSTEMS AND APPLICATIONS

Final Rejection §103
Filed
Mar 14, 2024
Examiner
BEZUAYEHU, SOLOMON G
Art Unit
2674
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
480 granted / 634 resolved
+13.7% vs TC avg
Strong +30% interview lift
Without
With
+29.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
37 currently pending
Career history
667
Total Applications
across all art units

Statute-Specific Performance

§101
17.2%
-22.8% vs TC avg
§103
52.7%
+12.7% vs TC avg
§102
12.7%
-27.3% vs TC avg
§112
10.0%
-30.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 634 resolved cases

Office Action

§103
DETAILED ACTION Response to Arguments Applicant's arguments filed with respect to claims 1-20 have been fully considered but are moot in view of the new ground(s) of rejection. The rejections are necessitated due to claim amendments. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 8, 12, 13, 18-20 rejected under 35 U.S.C. 103 as being unpatentable over Othman et al. (Pub. No. US 2021/0304421) in view of Ding et al. (Pub US 2008/0111730). Regarding claim 1, Othman teaches One or more processors (processing units 14) comprising processing circuitry (processor 15A) to: Compute (identify) a plurality of representations (detections) corresponding to a behavior (visual features and motion data) of one or more subjects (person) within an environment (monitored location) based on a first time frame (current video frame) of streaming sensor data (sequence of video frames) [Para. 19, 31, and 33 “ The visual processing unit 14 is adapted to receive the video frames 50 as input, and has a detection module 56 which is adapted to identify an image of a person 30 within each video frame 50.”], wherein the streaming sensor data (sequence of video frames) includes behavior data (visual features and motion data) of one or more subjects (person) within an environment (monitored location) captured from (video captured by) a plurality of optical sensors (plurality of cameras 12) [Para. 9 and 30]; Associate (match), based at least on trajectory tracking data (motion data) for the one or more subjects (person), one or more previously initialized behavior states (incumbent tracks) of a first plurality of previously initialized behavior states with one or more first representations (detection) of the plurality of representations (detections) [Para. 38 and 39]. Othman teaches assigning the plurality of representations (detections) based at least on at least one of the plurality of representations not having (if no corresponding) an associated (match) previously initialized behavior state (incumbent track) from the first plurality of previously initialized behavior states (incumbent tracks) [Para. 38 and 41]. However, Othman doesn’t explicitly teach assigning one or more second previously initialized behavior states of a second plurality of previously initialized behavior states to one or more second representations. Ding teaches assign (associated) one or more second previously initialized behavior states (track state) of a second plurality of previously initialized behavior states (unobservable track) to one or more second representations (remaining measurements) of the plurality of representations (measurements) based at least on at least one of the plurality of representations not having an associated previously initialized behavior state from the first plurality of previously initialized behavior states (confirmed tracks) [Para. 94 “A track state is updated only if it is associated with a measurement. The track state is a combination of all track information including target position, target speed, target strength, track quality, history, number of misses, number of updates, track layer, measurement association statistics and the like” and Para. 68]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman’s detection-to-track matching by retraining unmatched or deactivated incumbent tracks as Ding’s unobservable tracks and after associating detections with the active incumbent tracks, associating the remaining detections with those remained unobservable tracks. This modification improves Othman by maintaining a subject’s track identity across temporary detection gaps and reducing unnecessary creation of a new track for a previously tracked subject. Othman also teaches updating (is updated accordingly) at least one of the first plurality of previously initialized behavior (incumbent track) states or the second plurality of previously initialized behavior states based on the plurality of representations (detections) to generate updated behavior states (maintained as an incumbent track) [Para. 46 “Any incumbent track 59 or new track 59N which is matched to one of the detections 58 will be maintained as an incumbent track 59 when the next video frame is processed by the tracking module 64, and the motion data 59P for each incumbent track 59 is updated accordingly”]. Regarding claims 8 and 18, Othman teaches generate the plurality of representations using a machine learning model trained to detect characteristics representing the one or more subjects based at least on the streaming sensor data [para. 5, 11, 36, and 37]. Regarding claims 12 and 19, Othman teaches wherein the one or more processors are comprised in at least one of: a control system for an autonomous or semi-autonomous machine [Para. 3, 4, and 50]. Regarding claim 13, Othman teaches a system (system 10) comprising one or more processors (processor 15A) to: generate (are identified) a plurality of representations (detections 58) computed based on optical image streaming data (video frames 50) representing one or more tracked subjects (each person 30 with a track identity) within a monitored area (monitored location 34) [Para. 30, 31, 33, and 44]; Associate (match), based at least on trajectory tracking data (motion data), a first set of individual representations (detections 58) of the plurality of representations with one or more first previously initialized behavior states of a first plurality of previously initialized behavior states (incumbent tracks 59) computed from the optical image streaming data (at least one video frame 50 prior to the current video frame 50) [Para. 38, and 39]. Othman teaches assigning, using an assignment process (combinatorial optimization algorithm) and based at least on at least one of the plurality of representations (detection 58) not having an associated (no corresponding) first previously initialized behavior state (incumbent track 59) from the first plurality of previously initialized behavior states (incumbent track 59) [Para. 38 and 42]. However, Othman doesn’t explicitly teach one or more second previously initialized behavior states of a second plurality of previously initialized behavior states to a second set of individual representations of the plurality of representations. Ding teaches assigning (associated with), using an assignment process (auction method (i.e. 2-D assignment)), one or more second previously initialized (still retained) behavior states (track state) of a second plurality of previously initialized behavior states (unobservable tracks) to a second set of individual representations of the plurality of representations (measurement from the first updated list) based at least on at least one of the plurality of representations not having an associated first previously initialized behavior state (track state) from the first plurality of previously initialized behavior states (confirmed tracks) [Para. 68, 91, and 92]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman’s detection-to-track matching by retraining unmatched or deactivated incumbent tracks as Ding’s unobservable tracks and after associating detections with the active incumbent tracks, associating the remaining detections with those remained unobservable tracks. This modification improves Othman by maintaining a subject’s track identity across temporary detection gaps and reducing unnecessary creation of a new track for a previously tracked subject. Otheman also teaches update (is updated accordingly) at least one of the first plurality of previously initialized behavior states (incumbent track 59) and the second plurality of previously initialized behavior states based on the plurality of representations (detections 58) to generate updated behavior states (maintained as an incumbent track 59) [Para. 46]. Claim 20 is rejected for the same reasons as claim 1. Furthermore, Othman teaches a method to perform the claim limitations [see abstract and brined summary]. Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Othman et al. (Pub. No. US 2021/0304421) in view of Ding et al. (Pub US 2008/0111730) in view of Taylor (Pub. No. US 2012/0249802). Regarding claim 2, Othman in view of Ding doesn’t explicitly teach the claim limitation. However, Taylor teaches wherein the processing circuitry is further to: define one or more anchors (track) to associate individual previously initiated behavior (state vector) states from at least one of the first plurality of previously initiated behavior states and the second plurality of previously initiated behavior states with a global [unique] identifier (ID) [52]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding to teach the claim limitations above, feature as taught by Taylor; because the modification enables the system to improves reliable, accurate multi-camera target tracking over a wide area by using self-localizing distributed smart cameras that share compact sighting/track data and maintain persistent globally unique track identifiers as objects move through the network. Claims 3, 4, 6, 14, 15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Othman et al. (Pub. No. US 2021/0304421) in view of Ding et al. (Pub US 2008/0111730) in view of Zhu et al. (Pub. No. US 2018/0204093). Regarding claims 3 and 14, Othman in view of Ding doesn’t explicitly teach the claim limitation. However, Zhu teaches cluster the at least one of the plurality of representations (bounding boxes) not having the associated previously initiated behavior state (existing identity) from the first plurality of previously initiated behavior states to generate one or more clusters based at least on similarity (pairwise similarity metrics) [Para. 31, claim 3 “image has not yet been associated with an existing identity of the entity database”]; and assign, using a matching algorithm (assignment matrix), the second plurality of previously initiated behavior states/identities to the one or more clusters [Para. 75]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding to teach the claim limitations above, feature as taught by Zhu; because the modification enables the system to improve a person reidentification accuracy by clustering similar detections before matching to an identity database and creating a new identity when no existing identity matches, reducing missed/incorrect identity associations. Regarding claims 4 and 15, Othman in view of Ding doesn’t explicitly teach the claim limitation. However, Zhu teaches initialize (create) a new anchor (addition identity) and corresponding behavior state (identity database) for at least one cluster (image clusters) of the one or more clusters (image clusters) based on the at least one cluster not being assigned (has not yet been associated) at least one of the second plurality of previously initiated behavior states by the matching algorithm (assignment Matrix) [Para. 70, claim 3 “determining that at least one of the one or more image clusters represents an image of a person whose image has not yet been associated with an existing identity of the entity database”]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding to teach the claim limitations above, feature as taught by Zhu; because the modification enables the system to improve a person reidentification accuracy by clustering similar detections before matching to an identity database and creating a new identity when no existing identity matches, reducing missed/incorrect identity associations. Regarding claims 6 and 17, Othman teaches a Hungarian matching algorithm [Para. 42]. Claims 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Othman et al. (Pub. No. US 2021/0304421) in view of Ding et al. (Pub US 2008/0111730) in view of in view of Zhu et al. (Pub. No. US 2018/0204093) and further in view of Campos et al. (Pub. No. US 2003/0212702). Regarding claims 5 and 16, Othman in view of Ding in view of Zhu doesn’t explicitly teach the claim limitation. However, Campos teaches wherein the processing circuitry is further to cluster the at least one of the plurality of representations (training data) not having the associated previously initiated behavior state, based on a hierarchical clustering process that performs operations to: based at least on a similarity (Euclidean distance) of representations in the plurality of representations, determine a subject prediction number (maximum number of clusters) based on a first clustering process that clusters (core k-means process) at least one of the plurality of representations not having the associated previously initiated behavior state [Para. 122, 125 and 127]; and based at least on the prediction number (maximum number of clusters), apply to the at least one of the plurality of representations (data point) not having the associated previously initiated behavior state, a second clustering process (k-means process) to cluster the at least one of the plurality of representations not having the associated previously initiated behavior state [Para. 122 and 127]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding in view of Zhu to teach the claim limitations above, feature as taught by Campos; because the modification enables the system to improve the efficiency of clustering large datasets by using a hierarchical k-means partitioning approach that splits data in stages while controlling the maximum number of clusters. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Othman et al. (Pub. No. US 2021/0304421) in view of Ding et al. (Pub US 2008/0111730) in view of in view of Yu et al. (Pub. No. US 2022/0245835). Regarding claim 7, Othman in view of Ding doesn’t explicitly teach the claim limitation. Yu teaches wherein the behavior data (update tracking data) comprises at least one of appearance data (appearance) or spatiotemporal data (geometry and motion) represented by the plurality of representations (geo-motion embedding, appearance embedding) [Para. 36, 37, and 41]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding to teach the claim limitations above, feature as taught by Yu; because the modification enables the system to improve multi-object tracking data association by using learned geo-motion and appearance embeddings that encode an object’s motion/geometry and visual appearance over time to more accurately match new detections to existing tracks. Claims 9 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Othman et al. (Pub. No. US 2021/0304421) in view of Ding et al. (Pub US 2008/0111730) in view of in view of Narayanswamy et al. (Pub. No. US 2017/0134619). Regarding claim 9, Othman in view of Ding doesn’t explicitly teach the claim limitation. However, Narayanswamy teaches wherein the streaming sensor data comprises synchronized optical image streaming data (synchronized image data) that includes individual image feeds (streaming video streams) from the plurality of optical sensors, wherein the plurality of optical sensors is synchronized to capture the individual image feeds at the same time [Para. 22, 44, 43]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding in view of Zhu to teach the claim limitations above, feature as taught by Narayanswamy; because the modification enables the system to improve multi-sensor camera capture by providing hardware-trigged synchronization and timestamped aggregation so corresponding frames from multiple cameras can be reliably aligned in time for downstream processing. Regarding claim 10, Othman in view of Ding doesn’t explicitly teach the claim limitation. However, Narayanswamy teaches wherein the processing circuitry is further to: map the plurality of representations to a global image coordinate system based at least on camera calibration parameters (rotation matrix)/(translation vector) associated with the plurality of optical sensors [Para. 11, 24, and 25]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding to teach the claim limitations above, feature as taught by Narayanswamy; because the modification enables the system to improve multi-sensor camera capture by providing hardware-trigged synchronization and timestamped aggregation so corresponding frames from multiple cameras can be reliably aligned in time for downstream processing. Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Othman et al. (Pub. No. US 2021/0304421) in view of Ding et al. (Pub US 2008/0111730) in view of in view of Shen et al. (Pub. No. US 2022/0358314). Regarding claim 11, Othman in view of Ding doesn’t explicitly teach the claim limitation. Shen teaches wherein the processing circuitry is further to cause a display of a computer vision-based view of the one or more subjects for at least part of the environment based at least on the updated behavior states [Para. 37, 38, 49, and 42]. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Othman in view of Ding to teach the claim limitations above, feature as taught by Shen; because the modification enables the system to improve multi-sensor camera capture by providing hardware-trigged synchronization and timestamped aggregation so corresponding frames from multiple cameras can be reliably aligned in time for downstream processing. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOLOMON G BEZUAYEHU whose telephone number is (571)270-7452. The examiner can normally be reached on Monday-Friday 10 AM-8 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached on 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 888-786-0101 (IN USA OR CANADA) or 571-272-4000. /SOLOMON G BEZUAYEHU/ Primary Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Mar 14, 2024
Application Filed
Apr 07, 2026
Non-Final Rejection mailed — §103
Jul 07, 2026
Response Filed
Sep 11, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749214
THREE-DIMENSIONAL POSE ESTIMATION USING TWO-DIMENSIONAL IMAGES
3y 2m to grant Granted Sep 29, 2026
Patent 12731374
SAFE CONTROL/MONITORING OF A COMPUTER-CONTROLLED SYSTEM
3y 2m to grant Granted Sep 08, 2026
Patent 12725271
IMAGE QUALITY ENHANCING
4y 2m to grant Granted Sep 01, 2026
Patent 12721537
SYSTEM AND METHOD FOR MEASURING BLOOD FLOW VELOCITY ON A MICROFLUIDIC CHIP
3y 3m to grant Granted Sep 01, 2026
Patent 12718592
INFORMATION COLLECTION SYSTEM, SERVER, AND INFORMATION COLLECTION METHOD
3y 3m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
99%
With Interview (+29.9%)
3y 2m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 634 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month