Prosecution Insights
Last updated: October 02, 2026
Application No. 18/921,667

REAL-TIME PERSISTENT OBJECT TRACKING FOR INTELLIGENT VIDEO ANALYTICS SYSTEMS

Non-Final OA §103§112§DOUBLEPATENT
Filed
Oct 21, 2024
Priority
Oct 15, 2021 — continuation of 12/125,277
Examiner
ABDI, AMARA
Art Unit
Tech Center
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
697 granted / 840 resolved
+23.0% vs TC avg
Minimal -7% lift
Without
With
+-7.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
22 currently pending
Career history
859
Total Applications
across all art units

Statute-Specific Performance

§101
11.0%
-29.0% vs TC avg
§103
64.5%
+24.5% vs TC avg
§102
9.7%
-30.3% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 840 resolved cases

Office Action

§103 §112 §DOUBLEPATENT
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12,125,277. Although the claims at issue are not identical, they are not patentably distinct from each other because: -- Claims 1, 11, and 17 of the instant Application, recite common subject matter with the patent claims 1, 10, and 15; -- Whereby claims 1, 11, and 17 of the instant application, which recite the open-ended transitional phrase “comprising”, do not preclude the additional elements recited by patent claims 1, 10, and 15, and -- Whereby the elements of claims 1, 11, and 17 of the instant Application are fully anticipated by patent claim 1, 10, and 15. Instant Application Comparison US-Patent 12,125,277 1. A method comprising: tracking a first object in an environment depicted by a first set of images; obtaining one or more predicted future states of the first object in the environment; detecting a second object in the environment depicted by a second set of images, wherein a number of images of the second set of images exceeds a threshold number of images; determining whether a current state of the second object corresponds to at least one of the one or more predicted future states of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, updating state data for the first object based on the determined current state of the second object. 1. (Currently amended) A method comprising:, based on a first set of images depicting an environment, a state of a first object included in the environment, wherein the first set of images is generated during a first time period; determining that the first object is not detected in the environment depicted in a second set of images generated during a second time period that is subsequent to the first time period; obtaining one or more predicted future states of the first object in view of the state of the first object in the environment depicted in the first set of images;cluded in the environment depicted in a third set of images generated during a third time period that is subsequent to the second time period, wherein a number of images of the third set of images exceeds a threshold number of images;n identifier associated with the second object to correspond to an identifier associated with the first object. 1. (Currently amended) A method comprising: tracking, based on a first set of images depicting an environment, a state of a first object included in the environment, wherein the first set of images is generated during a first time period; determining that the first object is not detected in the environment depicted in a second set of images generated during a second time period that is subsequent to the first time period; obtaining one or more predicted future states of the first object in view of the state of the first object in the environment depicted in the first set of images; detecting a second object included in the environment depicted in a third set of images generated during a third time period that is subsequent to the second time period, wherein a number of images of the third set of images exceeds a threshold number of images; determining whether a current state of the second object corresponds to at least one of the one or more predicted future states of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, updating an identifier associated with the second object to correspond to an identifier associated with the first object. 5. The method of claim 1, wherein obtaining the one or more predicted future states of the first object comprises: determining a current state of the first object in the environment depicted by the first set of images; calculating a path that the first object is expected to follow in the environment during a future time period based on the determined current state; and determining the one or more predicted future states of the first object based on the calculated path. 2. (Original) The method of claim 1, wherein obtaining the one or more predicted future states of the first object comprises:obtaining state data associated with the first object based on the state of the first object in the environment depicted in each of the first set of images;obtained data; and determining the one or more predicted future states of the first object based on the calculated path. 2. The method of claim 1, wherein obtaining the one or more predicted future states of the first object comprises: obtaining state data associated with the first object based on the state of the first object in the environment depicted in each of the first set of images; calculating a path that the first object is expected to follow in the environment during a future time period based on the obtained state data; and determining the one or more predicted future states of the first object based on the calculated path. 6. The method of claim 5, wherein determining the current state of the first object in the environment comprises: providing, as an input to a machine learning model, an indication of at least one of: a prior state of the first object in the environment, or the current state of the first object in the environment; extracting, from one or more outputs of the machine learning model, one or more sets of object state data and, for at least one set of object state data, a level of confidence that the at least one set of object state data corresponds to the first object; and responsive to determining that a level of confidence for the at least one set of object state data satisfies one or more confidence criteria, extracting the current state of the first object from the at least one set of object state data. 4. (Original) The method of claim 2, wherein obtaining the state data associated with the first object r a current state of the first object in the environment depicted in each of the first set of images as an input to a machine learng model; obtaining; extracting, from the one or more outputs, one or more sets of object state data and, for ch set of object state data, an indication of a level of confidence that ch set of object state data corresponds to the first object; and ntifying sociated with a level of confidence that ssfies a confidence criterion. 4. The method of claim 2, wherein obtaining the state data associated with the first object comprises: providing an indication of at least one of a prior state or a current state of the first object in the environment depicted in each of the first set of images as an input to a machine learning model; obtaining one or more outputs of the machine learning model; extracting, from the one or more outputs, one or more sets of object state data and, for each set of object state data, an indication of a level of confidence that each set of object state data corresponds to the first object; and identifying a set of object state data associated with a level of confidence that satisfies a confidence criterion. 7. The method of claim 6, wherein the machine learning model comprises a recurrent neural network. 5. (Original) The method of claim 4, wherein the machine learning model comprises a recurrent neural network. 5. The method of claim 4, wherein the machine learning model comprises a recurrent neural network. 8. The method of claim 1, wherein tracking the state of the first object comprises: obtaining the first set of images and a first set of bounding boxes associated with the first set of images, wherein the first set of bounding boxes indicate one or more regions of the first set of images that include a detected presence of the first object; and updating the state data for the first object to indicate the one or more regions of the first set of images indicated by the first set of bounding boxes. (Original) The method of claim 1, wherein tracking the state of the first object comprises:determining the state of the first object cludd in the environment depicted in the first set of images based on the first set of bounding boxes. 7. The method of claim 1, wherein tracking the state of the first object comprises: obtaining the first set of images and a first set of bounding boxes associated with the first set of images, wherein the first set of bounding boxes indicate one or more regions of the first set of images that include a detected presence of the first object; and determining the state of the first object included in the environment depicted in the first set of images based on the first set of bounding boxes. 11. A system comprising: a memory; and a set of one or more processing devices coupled to the memory, wherein the set of one or more processing devices is to: track a first object in an environment depicted by a first set of images; obtain one or more predicted future states of the first object in the environment; detect a second object in the environment depicted by a second set of images, wherein a number of images of the second set of images exceeds a threshold number of images; determine whether a current state of the second object corresponds to at least one of the one or more predicted future states of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, update state data for the first object based on the determined current state of the second object. 11. (Currently amended) A system comprising: device; and a device, wherein the processing device is to: track, based on a first set of images depicting an environment, a state of a first object included in the environment, wherein the first set of images is generated during a first time period; determine that the first object is not detected in the environment depicted in a second set of images generated during a second time period that is subsequent to the first time period; obtain one or more predicted future states of the first object in view of the state of the first object in the environmentdepicted in the first set of images; detect a second object included in the environment depicted in a third set of images generated during a third time period that is subsequent to the second time period wherein a number of images of the third set of images exceeds a threshold number of images;n identifier associated with the second object to correspond to an identifier associated with the first object. 10. A system comprising: a memory device; and a processing device coupled to the memory device, wherein the processing device is to: track, based on a first set of images depicting an environment, a state of a first object included in the environment, wherein the first set of images is generated during a first time period; determine that the first object is not detected in the environment depicted in a second set of images generated during a second time period that is subsequent to the first time period; obtain one or more predicted future states of the first object in view of the state of the first object in the environment depicted in the first set of images; detect a second object included in the environment depicted in a third set of images generated during a third time period that is subsequent to the second time period wherein a number of images of the third set of images exceeds a threshold number of images; determine whether a current state of the second object corresponds to at least one of the one or more predicted future state of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, update an identifier associated with the second object to correspond to an identifier associated with the first object. 15. The system of claim 11, wherein to obtain the one or more predicted future states of the first object, the set of one or more processing devices is to: determine a current state of the first object in the environment depicted by the first set of images; calculate a path that the first object is expected to follow in the environment during a future time period based on the determined current state; and determine the one or more predicted future states of the first object based on the calculated path. 1. (Original) The system of claim 11, wherein to obtain the one or more predicted future states of the first object comprises, the processing deviceobtain state data associated with the first object based on the state of the first object in the environment depicted in each of the first set of images;obtained data; and determine the one or more predicted future states of the first object based on the calculated path. 11. The system of claim 10, wherein to obtain the one or more predicted future states of the first object comprises, the processing device is to: obtain state data associated with the first object based on the state of the first object in the environment depicted in each of the first set of images; calculate a path that the first object is expected to follow in the environment during a future time period based on the obtained state data; and determine the one or more predicted future states of the first object based on the calculated path. 16. The system of claim 15, wherein to determine the current state of the first object in the environment, the set of one or more processing devices is to: provide, as an input to a machine learning model, an indication of at least one of a prior state of the first object in the environment or the current state of the first object in the environment; extract, from one or more outputs of the machine learning model, one or more sets of object state data and, for at least one set of object state data, a level of confidence that the at least one set of object state data corresponds to the first object; and responsive to determining that a level of confidence for the at least one set of object state data satisfies one or more confidence criteria, extract the current state of the first object from the at least one set of object state data. 1. (Original) The system of claim 12, wherein to obtainstate data associated with the first objectr a current state of the first object in the environment depicted in each of the first set of images as an input to a machine learng model; obtain one or more outputs of the machine learning model; extract, from the one or more outputs, one or more sets of object state data and, for ch set of object state data, an indication of a level of confidence that each set of object state data corresponds to the first object; and ify aassociated with a level of confidence at ssfies a confidence criterion. 13. The system of claim 11, wherein to obtain the state data associated with the first object, the processing device is to: provide an indication of at least one of a prior state or a current state of the first object in the environment depicted in each of the first set of images as an input to a machine learning model; obtain one or more outputs of the machine learning model; extract, from the one or more outputs, one or more sets of object state data and, for each set of object state data, an indication of a level of confidence that each set of object state data corresponds to the first object; and identify a set of object state data associated with a level of confidence that satisfies a confidence criterion. 17. A non-transitory computer readable storage medium comprising instructions that, when executed by a set of one or more processing devices, cause the set of one or more processing devices to: track a first object in an environment depicted by a first set of images; obtain one or more predicted future states of the first object in the environment; detect a second object in the environment depicted by a second set of images, wherein a number of images of the second set of images exceeds a threshold number of images; determine whether a current state of the second object corresponds to at least one of the one or more predicted future states of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, update state data for the first object based on the determined current state of the second object. 17. (Currently amended) A non-transitory computer readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to:track, based on a firstimages depicting an environment, a state of a first bject included in the environment, wherein the first set of images is generated during a first time period; determine that the first object is not detected in the environment depicted in a second set of images generated during a second time period that is subsequent to the first time period; obtain one or more predicted future states of the first object in view of the state of the first object in the environmentdepicted in the first set of images; detect a second object included in the environment depicted in a third set of images generated during a third time period that is subsequent to the second time periods wherein a number of images of the third set of images exceeds a threshold number of images;n identifier associated with the second object to correspond to an identifier associated with the first object. 15. A non-transitory computer readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to: track, based on a first set of images depicting an environment, a state of a first object included in the environment, wherein the first set of images is generated during a first time period; determine that the first object is not detected in the environment depicted in a second set of images generated during a second time period that is subsequent to the first time period; obtain one or more predicted future states of the first object in view of the state of the first object in the environment depicted in the first set of images; detect a second object included in the environment depicted in a third set of images generated during a third time period that is subsequent to the second time period, wherein a number of images of the third set of images exceeds a threshold number of images; determine whether a current state of the second object corresponds to at least one of the one or more predicted future state of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, update an identifier associated with the second object to correspond to an identifier associated with the first object. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-19 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites the limitation "determining that a current state" in claim 1, line 8. There is insufficient antecedent basis for this limitation in the claim. It is suggested changing the limitation "determining that a current state” into “determining that the current state” to refer to “a current state” cited on line 6. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 8-15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Cheng et al, (US-PGPUB 2019/0130582) in view of Irimoto (US-PGPUB 20140348398); and further in view of Subramanian et al, (US-PGPUB 20220292286) Regarding claim 1, Cheng discloses a method comprising: tracking a first object in an environment depicted by a first set of images, (see at least: Fig. 10A-10C, and Par. 0146-0147, and 149, the object tracking system correctly identifies an object (a group of people 1006), as a moving object, and outputs a bounding box 1004 around the group of people 1006 for tracking the group of people 1006, in the scene, [i.e., tracking a first object in an environment, “scene”, depicted by a first set of images, “video sequences of scene”]); obtaining one or more predicted future states of the first object in the environment, (see at least: Par. 0149, the predicted bounding box can predict the location of the group of people 1006 to be about in the middle of the second video frame 1010 and to the left in the third video frame 1020, based on using a predicted bounding box for the group of people 1006, determined, for example, for the first video frame 1000, [i.e., obtaining one or more predicted future states, “location”, of the first object the environment, “the blob in the scene”]); detecting a second object in the environment depicted by a second set of images, wherein a number of images of the second set of images exceeds a threshold number of images, (see at least: Fig. 10C, and Par. 0147, the object tracking system implicitly identifies the group of people 1006 as a moving object at frame 1020, generated during a third time period that is subsequent to the second time period, and outputs a bounding box 1004 around the group of people 1006 for tracking the group of people 1006, [i.e., detecting a second object in the environment depicted by a second set of images, “identifies the group of people 1006 in the scene, as a moving object at frame 1020”]). Cheng et al does not expressly disclose wherein a number of images of the second set of images exceeds a threshold number of images, determining whether a current state of the second object corresponds to at least one of the one or more predicted future states of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, updating state data for the first object based on the determined current state of the second object. However, Irimoto discloses wherein a number of images of the second set of images exceeds a threshold number of images, (see at least: Par. 0087, the clustering module 53 classifies the images into groups, based on the detected face images, …where the display controller 56 display representative images …such that the groups are distinguished between the first group in which the number of classified images is equal to or more than a threshold and the second group in which the number of classified images is less than the threshold, [i.e., wherein a number of images of the second set of images, “the number of first group images”, exceeds a threshold number of images, “equal to or more than a threshold”]). Cheng, and Irimoto are combinable because they are both concerned with object tracking. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify Cheng, to use the clustering module 53, as thought by Irimoto, in order to distinguish between the first group in which the number of classified images is equal to or more than a threshold and the second group in which the number of classified images is less than the threshold, (Irimoto, Par. 0087). The combination of Cheng and Irimoto as whole does not expressly disclose determining whether a current state of the second object corresponds to at least one of the one or more predicted future states of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, updating state data for the first object based on the determined current state of the second object. However, Subramanian discloses determining whether a current state of the second object corresponds to at least one of the one or more predicted future states of the first object; and responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, updating state data for the first object based on the determined current state of the second object, (see at least: Par. 0031, the object tracking component 122 may be configured to generate tracking information 134 indicating the trajectory of the detected objects 128(1)-(N) over the video frames 114(1)-(N). The object tracking component 122 may receive at least the bounding representations of the detected objects 128(1)-(N) from the object detection component 120 for each frame 114, and determine if the bounding representations of a current video frame 114 have corresponding bounding representations in one of the preceding video frames 114, where the object tracking component 122 may assign object identifiers to the detection objects 128(1)-(N) within the tracking information 134, “an identifier associated with the object”, [i.e., determining whether a current state of the second object, “bounding representations for the detected objects 128(1)-(N) of one of the preceding video frames 114”, corresponds to at least one of the one or more predicted future states of the first object, “bounding representations for the detected objects 128(1)-(N) of one of the current video frames 114”]. For instance, if the object tracking component 122 determines that a current bounding representation has a corresponding historic bounding representation, “responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object”, the object tracking component 122 assigns the object identifier of the corresponding historic bounding representation to the current bounding representation, “updating an identifier associated with the second object to correspond to an identifier associated with the first object”, [i.e., responsive to determining that a current state of the second object corresponds to at least one of the one or more predicted future states of the first object, “if the object tracking component 122 determines that a current bounding representation has a corresponding historic bounding representation”, updating state data for the first object based on the determined current state of the second object, “the object tracking component 122 assigns the object identifier of the corresponding historic bounding representation to the current bounding representation, based on identifier assigned to the object”]). Cheng, Irimoto, and Subramanian are combinable because they are both concerned with object tracking. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Cheng and Irimoto, to determine if the bounding representations of a current video frame have corresponding bounding representations in one of the preceding video frames, as though by Subramanian, in order to generate tracks corresponding to the trajectory of the detected objects across the video frames based on the assigned object identifiers, (Subramanian, Par. 0031). Regarding claim 2, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. Subramanian further discloses wherein updating the state data for the first object based on the determined current state of the second object comprises: updating an identifier associated with the first object to correspond to the determined current state of the second object, (Subramanian, see at least: Par. 0031, the object tracking component 122 assigns the object identifier of the corresponding historic bounding representation to the current bounding representation, based on identifier assigned to the object”). Regarding claim 3, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. Cheng further discloses wherein the first set of images is generated at a first time period and the second set of images is generated at a second time period that is subsequent to the first time period, (see at least: Par. 0146-0149, the first frame image 1000 of the video sequences is implicitly generated during a first time period before the object arrives to tree). Regarding claim 4, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. Cheng further discloses determining that the first object is not detected in the environment depicted by a third set of images generated at a third time period that is subsequent to the first time period and prior to the second time period, (see at least: Figs. 10a-10c, and Par. 0146-0148, In the second video frame 1010, illustrated in FIG. 10B, an object, (a group of people 100) walking from right to left on the path that passes behind the tree 1008, (an exclusion zone 1002), and thus is legitimately lost in the scene, [i.e., determining that the first object is not detected in the environment, “is legitimately lost in the scene in the exclusion zone 1002”, depicted by a third set of images, “second frame 1010”, generated at a third time period that is subsequent to the first time period and prior to the second time period, “generated during a second time period that is subsequent to the first frame 1000”]). Regarding claim 5, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. Cheng further discloses wherein obtaining the one or more predicted future states of the first object comprises: determining a current state of the first object in the environment depicted by the first set of images, (see at least: Par. 0054, prediction of the location of the blob tracker, “current state of the first object in the environment”, in the current frame can be based on the location of the blob in the previous frame. See also, Par. 0088, the location, “state data”, of a blob tracker in a current frame may be predicted based on information, “the state of the first object”, from a previous frame); calculating a path that the first object is expected to follow in the environment during a future time period based on the determined current state, (see at least: Par. 0088, a blob tracker can employ a Kalman filter to measure its trajectory as well as predict its future location(s), [i.e., the trajectory is implicitly based on the location, “determined current state”, of a blob tracker in a current frame]); and determining the one or more predicted future states of the first object based on the calculated path, (see at least: Par. 0088, a blob tracker can employ a Kalman filter to measure its trajectory as well as predict its future location(s), [i.e., implicitly predict its future location(s), “one or more predicted future states of the first object”, based on the measured trajectory, “based on the calculated path”]). Regarding claim 8, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. Cheng further discloses wherein tracking the state of the first object comprises: obtaining the first set of images and a first set of bounding boxes associated with the first set of images, wherein the first set of bounding boxes indicate one or more regions of the first set of images that include a detected presence of the first object, (see at least: Fig. 1, Par. 0051, video analytics system 100 receives video frames 102 from a video source 130, “obtaining the first set of images”, and bounding box can be associated with a blob, “a first set of bounding boxes associated with the first set of images”. Par. 0053, A tracker can also be represented by a tracker bounding region, “wherein the first set of bounding boxes indicate one or more regions of the first set of images that include a detected presence of the first object”); and updating the state data for the first object to indicate the one or more regions of the first set of images indicated by the first set of bounding boxes, (see at least: Par. 0054, A blob tracker can be associated with a tracker bounding box and can be assigned a tracker identifier (ID). The prediction of the location of the blob tracker in the current frame can be based on the location of the blob in the previous frame, [i.e., implicitly updating the state data for the first object to indicate the one or more regions of the first set of images indicated by the first set of bounding boxes, “the prediction of the location of the blob tracker in the current frame. See also, Par. 0133, (different embodiment), Par. 0135, the trajectory, “state of the first object”, can be determined from the history of bounding boxes included in the tracker, where the history includes bounding boxes for a blob from previous input frames, “based on the first set of bounding boxes”). Regarding claim 9, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. Cheng further discloses wherein the one or more predicted future states of the first object comprise at least one of a predicted future position of the first object in the environment, a predicted future location of the first object in the environment, a predicted future size of the first object in the environment, a predicted future scale of the first object in the environment, or a predicted future velocity of the first object in the environment, (Par. 0149, the predicted bounding box can predict the location of the group of people 1006, “i.e., the location corresponds to the predicted future states of the first object”. See also, Par. 0031, the object tracking component 122 may be configured to generate tracking information 134 indicating the trajectory of the detected objects 128(1)-(N) over the video frames 114(1)-(N), “i.e., the trajectory corresponds also to the predicted future states of the first object”). Regarding claim 10, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. wherein the current state of the second object comprises at least one of a current position of the second object in the environment, a current location of the second object in the environment, a current size of the second object in the environment, a current scale of the second object in the environment, or a current velocity of the second object in the environment, (see at least: Par. 0149, predicted bounding box can predict the location of the group of people 1006, “i.e., the location implicitly corresponds to the current states of the first object”). Regarding claim 11, claim 11 recites substantially similar limitations as set forth in claim 1. As such, claim 11 is rejected for at least similar rational. The Examiner further acknowledged the following additional limitation(s): “a system comprising: a memory device; and a processing device coupled to the memory device”. However, Cheng et al discloses a system comprising: a memory device; and a processing device coupled to the memory device, (see at least: Par. 0008, an apparatus is provided that includes a memory configured to store video data and a processor, “system comprising: a memory device; and a processing device coupled to the memory device”). Regarding claim 12, claim 12 recites substantially similar limitations as set forth in claim 2. As such, claim 12 is rejected for at least similar rational. Regarding claim 13, claim 13 recites substantially similar limitations as set forth in claim 3. As such, claim 13 is rejected for at least similar rational. Regarding claim 14, claim 14 recites substantially similar limitations as set forth in claim 4. As such, claim 14 is rejected for at least similar rational. Regarding claim 15, claim 15 recites substantially similar limitations as set forth in claim 5. As such, claim 15 is rejected for at least similar rational. Regarding claim 17, claim 17 recites substantially similar limitations as set forth in claim 1. As such, claim 17 is rejected for at least similar rational. The Examiner further acknowledged the following additional limitation(s): “a non-transitory computer readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to”. However, Cheng et al discloses, (see at least: Par. 0009, a computer readable medium is provided having stored thereon instructions that when executed by a processor perform a method). Regarding claim 18, claim 18 recites substantially similar limitations as set forth in claim 2. As such, claim 18 is rejected for at least similar rational. Regarding claim 19, claim 19 recites substantially similar limitations as set forth in claim 3. As such, claim 19 is rejected for at least similar rational. Regarding claim 20, claim 20 recites substantially similar limitations as set forth in claim 4. As such, claim 20 is rejected for at least similar rational. Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Cheng, Irimoto, and Subramanian, as applied to claim 1 above; further in view of Sharma, (US-PGPUB 20230028152); and further in view of Winarski et al, (US-PGPUB 20220282980) Regarding claim 6, the combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 1. the combination of Cheng, Irimoto, and Subramanian as whole does not expressly disclose wherein determining the current state of the first object in the environment comprises: providing, as an input to a machine learning model, an indication of at least one of a prior state of the first object in the environment, or the current state of the first object in the environment; extracting, from one or more outputs of the machine learning model, one or more sets of object state data and, for at least one set of object state data, a level of confidence that the at least one set of object state data corresponds to the first object; and responsive to determining that a level of confidence for the at least one set of object state data satisfies one or more confidence criteria, extracting the current state of the first object from the at least one set of object state data. However, Sharma discloses providing, as an input to a machine learning model, an indication of at least one of a prior state of the first object in the environment, or the current state of the first object in the environment, (see at least: Par. 0022, predictive analysis module 156 may receive as input 158, “as an input to a machine learning model”, information on a location of a requestor, “a indication of at least one of a prior state of the first object in the environment”, seeking a location of an object type, descriptive information on the object, as well as any other useful information that could contribute to improved location prediction); and extracting, from one or more outputs of the machine learning model, (see at least: Par. 0022, the predictive analysis module may be trained to predict locations of an object of an object type based on a historical corpus of previous offload locations of objects of the object type based on different combinations of input, [i.e., obtaining one or more outputs of the machine learning model]). Cheng, Irimoto, Subramanian, and Sharma are combinable because they are all concerned with object tracking. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify combination of Cheng, Irimoto, and Subramanian, to use machine learning and artificial intelligence, as though by Sharma, in order to predict location of a physical object if the location cannot be determined from the tracking information 300, (Sharma, Par. 0022). The combination of Cheng, Irimoto, Subramanian, and Sharma as whole does not expressly disclose extracting, from one or more outputs of the machine learning model, one or more sets of object state data and, for at least one set of object state data, a level of confidence that the at least one set of object state data corresponds to the first object; and responsive to determining that a level of confidence for the at least one set of object state data satisfies one or more confidence criteria, extracting the current state of the first object from the at least one set of object state data. Winarski discloses extracting, from one or more outputs of the machine learning model, one or more sets of object state data and, for at least one set of object state data, a level of confidence that the at least one set of object state data corresponds to the first object; (see at least: Par. 0047-0049, at block 204, a future location of each of the one or more other people at a future point in time (i.e., that is after the first point in time) is predicted and a confidence level is assigned to the prediction, and at block 206, a travel route is generated from the first location to a destination location based at least in part on the predicted future locations of the other people and their assigned confidence levels, [i.e., extracting, from one or more outputs of the machine learning model, one or more sets of object state data, “a travel route”, and, for at least one set of object state data, a level of confidence that the at least one set of object state data corresponds to the first object, “assigned confidence levels”]; and responsive to determining that a level of confidence for the at least one set of object state data satisfies one or more confidence criteria, extracting the current state of the first object from the at least one set of object state data, (see at least: Par. 0049, the predicted future locations having an assigned confidence level higher than a threshold value (e.g., 40%, 50%, 75%, 90%, etc.) are used when generating the travel route, [i.e., implicitly identifying a set of object state data, “predicted future locations”, responsive to determining that a level of confidence for the at least one set of object state data satisfies one or more confidence criteria, “based on the threshold value”]). Cheng, Irimoto, Subramanian, Sharma, and Winarski are combinable because they are all concerned with object tracking. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Cheng, Irimoto, Subramanian, and Sharma, to use the block 204 and 206, as though by Winarski, in order to predict future locations having an assigned confidence level higher than a threshold value, (Winarski, Par. 0049). Regarding claim 16, claim 16 recites substantially similar limitations as set forth in claim 6. As such, claim 16 is rejected for at least similar rational. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Cheng, Irimoto, Subramanian, Sharma, and Winarski, as applied to claim 6 above; further in view of Tusch et al, (US-PGPUB 20210279475) The combination of Cheng, Irimoto, and Subramanian as whole discloses the limitations of claim 6. The combination of Cheng, Irimoto, Subramanian, Sharma, and Winarski as whole does not expressly disclose wherein the machine learning model comprises a recurrent neural network. However, Tusch et al discloses wherein the machine learning model comprises a recurrent neural network, (see at least: Par. 0069-0070, engine applies a convolutional or recurrent neural network or another object detection algorithm to do to continuously monitors the motion of individuals in the scene and predicts their next location to enable reliable tracking even when the subject is temporarily lost or passes behind another object, [i.e., implicitly using recurrent neural network for predicting individuals next location]). Cheng, Irimoto, Subramanian, Sharma, Winarski, and Tusch et al are combinable because they are all concerned with object tracking. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Cheng, Irimoto, Subramanian, Sharma, and Winarski, to substitute the Sharma’s machine learning, with the recurrent neural network, as though by Tusch et al, in order to enable reliable tracking even when the subject is temporarily lost or passes behind another object, (Tusch et al, Par. 0070). Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMARA ABDI whose telephone number is (571)272-0273. The examiner can normally be reached 9:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMARA ABDI/Primary Examiner, Art Unit 2668 09/19/2026
Read full office action

Prosecution Timeline

Oct 21, 2024
Application Filed
Sep 22, 2026
Non-Final Rejection mailed — §103, §112, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749345
MONITORING AND ANALYZING BODY LANGUAGE WITH MACHINE LEARNING, USING ARTIFICIAL INTELLIGENCE SYSTEMS FOR IMPROVING INTERACTION BETWEEN HUMANS, AND HUMANS AND ROBOTS
2y 7m to grant Granted Sep 29, 2026
Patent 12749346
FINGER ENCODING BASED POSE CLASSIFICATION
2y 6m to grant Granted Sep 29, 2026
Patent 12749307
CAMERA APPARATUS AND METHOD OF ENHANCED FOLIAGE DETECTION
2y 2m to grant Granted Sep 29, 2026
Patent 12743792
SYSTEMS AND METHODS FOR IMAGE PROCESSING
2y 11m to grant Granted Sep 22, 2026
Patent 12728441
ROBOTIC REPAIR CONTROL SYSTEMS AND METHODS
3y 6m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
76%
With Interview (-7.3%)
2y 6m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 840 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month