Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1, 13 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. (US 20230112584 A1) in view of Li et al (US 20220262142 A1) further in view of Ning et al. (US 20220020158 A1).
Regarding claim 1, Shah et al. teaches a method of tracking persons in an environment, the method comprising: obtaining a plurality of camera views from a corresponding plurality of cameras positioned at different locations around the environment (see para [0003]; “tracking a person at a location across multiple cameras”, see also para [0022]; “The cameras 102 may be positioned at different positions within a location 111. As an example, the cameras 102 may operate as part of a security system at a retail location 111, wherein the cameras 102 may be positioned within and have fields of view encompassing respective parts of the retail location 111. The cameras 102 may operate to obtain image data associated with their field of view and communicate the image data to one or more other subsystems”) and configured to obtain frames of video data containing persons within the environment; for each camera view obtained from an individual camera (see para [0022]; “The image data may include a plurality of video frames captured in sequence according to a frame rate. For example, image data may capture an image of a person 101 or multiple people 101 in any given video frame. In some examples, at least two cameras 102a,b may have overlapping fields of view, where an image of a person 101 may be captured concurrently by the at least two cameras 102a,b”); (i) detecting one or more persons within the camera view by processing the video data from the individual camera via a neural network (see para [0025]; “for a single camera 102, a deep sort algorithm may be used to track people 101 and identify them within the location 111. A convolutional neural network may be used that extracts visual features from a detected bounding box”), and (ii) generating a bounding box for each person detected in the camera view (see para [0033]; “generate one or more bounding boxes 303 to include each detected person 101”); generating one or more sets of associated bounding boxes, wherein each set of associated bounding boxes includes two or more bounding boxes from two or more different camera views that correspond to a same person detected in the different camera views (see para [0041]; “match a camera-specific track 411 associated with a person 101 received from a first camera 102a to a global track 504 of the person 101….a camera-specific track 411 may be linked to a last-updated camera-specific track 411 included in a global track 504, wherein the last-updated camera-specific track 411 may be from a same camera 102 or another camera 102”). However, Shah et al does not teach and for each timeframe of a plurality of timeframes, (i) generating a bounding cuboid for each set of associated bounding boxes for the timeframe, and (ii) for each generated bounding cuboid, associating the generated bounding cuboid with one of a plurality of tracklets for the timeframe based on (a) a position of the generated bounding cuboid during the timeframe and (b) predicted future positions for each of the plurality of tracklets for the timeframe, wherein each tracklet corresponds to a set of historical positions of one of the persons detected in the plurality of camera views.
In the same field of endeavor, Li et al. teaches and for each timeframe of a plurality of timeframes, (i) generating a bounding cuboid for each set of associated bounding boxes for the timeframe (see para [0036]; “to create a 3D cube model for each player, all 2D bounding boxes for each player are associated with one another…… player 3D cube model generation is conducted at illustrated block 52 based on the results of block 50. In general, as 2D player bounding box and association data becomes available in block 50, a 3D bounding box (e.g., cube) model may be constructed for each player”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking f Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image in order to automatically generate a 3D bounding box around the 3D object (see para [0036]). However, the combination of Shah et al. and Li et al. as a whole does not teach and (ii) for each generated bounding cuboid, associating the generated bounding cuboid with one of a plurality of tracklets for the timeframe, based on (a) a position of the generated bounding cuboid during the timeframe, and (b) predicted future positions for each of the plurality of tracklets for the timeframe, wherein each tracklet corresponds to a set of historical positions of one of the persons detected in the plurality of camera views.
In the same field of endeavor, Ning et al. teaches and (ii) for each generated bounding cuboid, associating the generated bounding cuboid with one of a plurality of tracklets for the timeframe (see para [0089]; “Use the Hungarian algorithm to associate the correct track instance (detected measurements, trajectory history) to the predictions based on the cost”, see also para [0086; “assign the correct detected measurements to predicted tracks”) based on (a) a position of the generated bounding cuboid during the timeframe (see para [0087]; “the state vector is the 2D/3D object location, represented by a vector of coordinates”, see also para [0088]; “The cost is defined as the Euclidean distance of detections and predictions, which is weighted on the 2D and 3D coordinates”) and (b) predicted future positions for each of the plurality of tracklets for the timeframe (see para [0088]; “Calculate the predictions of the Kalman filters based on tracking history and then calculate the cost between the predictions and the current detections from the detection module”), wherein each tracklet corresponds to a set of historical positions of one of the persons detected in the plurality of camera views (see para [0087]; “a track is an entity consisting of several attributes, including a) the prediction vector representing both the 2D and 3D locations, b) the unique track ID, c) the Kalman filter instances for both 2D and 3D objects, d) the trajectory history of this track, represented by a list of prediction vectors” see also para [0088]; “Calculate the predictions of the Kalman filters based on tracking history”, and para [0093]; “Kalman filter only takes into account the location histories”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking f Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. in order to achieve the predictable result of more accurate multi-camera a person tracking (see para [0089]).
Regarding claim 13, the rejection of claim 1 is fully incorporated herein.
Ning et al. in the combination further teach further comprising: for each tracklet in the plurality of tracklets, using a Kalman filter to generate the predicted future positions for each of the plurality of tracklets for the timeframe (see para [0064]; “the 3D object tracking module 262 is performed using a 3D Kalman filter 263 and a Hungarian algorithm 264”, see also para [0087]; “the prediction vector representing both the 2D and 3D locations”), wherein each tracklet corresponds to a set of historical cuboids of one of the persons detected in the plurality of camera views (see para [0087]; “the trajectory history of this track, represented by a list of prediction vectors”).
Regarding claim 20, the scope of claim 20 is fully applicable inhere, the rejection analysis of claim 1 is equally applicable here.
Clam 2 is rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al in view of Ning et al. as applied in claim 1 above, and further in view of Yu et al. (US 20190244366 A1).
Regarding claim 2, the rejection of claim 1 is fully incorporated herein.
Shah et al. in the combination further teaches wherein for each camera view obtained from an individual camera, detecting one or more persons within the camera view by processing the video data from the individual camera via a neural network (see Abstract; “The system may track a person at a location across multiple cameras in real time, while maintaining identification of the person through their visit. The person may not be personally identified; however, the person's presence in video frames captured by different cameras may be linked together”). However, the combination of Shah et al., Li et al. and Ning et al. as a whole does not teach comprises: for frames of video data obtained from an individual camera, generating input tensor data based on an anchor frame selected from the frames of video data obtained from the individual camera; detecting the one or more persons via the neural network based on the input tensor data; and generating output tensor data corresponding to each person detected within the anchor frame.
In the same field of endeavor, Yu et al. teach comprises: for frames of video data obtained from an individual camera, generating input tensor data based on an anchor frame selected from the frames of video data obtained from the individual camera (see para [0054]; “a previous frame as the “reference-frame” and subtract the reference from each frame may be selected to generate a subtracted 4D tensor 504. The subtracted 4D tensor 504 may be used as an input into 3D ConvNets 506 and 508”, see also para [0060]; “the input of 3D ConvNets may be a 15×90×160×3 tensor”, and para [0029]; “The processed images may be utilized to construct a 4D tensor of the down-sampled video. The 4D tensor may be variously used as, for example, the input of a neural network, such as a 3D convolutional neural network”); detecting the one or more persons via the neural network based on the input tensor data (see para [0029]; “whether there is any relevant motion in the video and/or whether the motion is caused by person/vehicles/pets and so on”, see also para [0040]; “One or more object detectors may be used to detect objects. One or more method may comprise applying the object detectors based on deep convolutional neural networks (CNNs) to identify objects of interest”); and generating output tensor data corresponding to each person detected within the anchor frame (see para [0056]; “Each spatial-temporal location of the output tensor from pool 510 may have a binary prediction ……the output tensor from pool 510 may have a binary prediction”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of apparatuses for detecting relevant motion of objects of interest in surveillance videos of Yu et al. in order to generate a plurality of prediction results of relevant motion of the objects of interest (see para [0054]).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al. in view of Ning et al. and Yu et al. as applied in claims 1-2 above, and further in view of Pierce et al. (US 20170154425 A1).
Regarding claim 3, the rejection of claim 2 is fully incorporated herein. The combination of Shah et al., Li et al. Ning et al. and Yu et al. does not teach wherein for each camera view obtained from an individual camera, generating a bounding box for each person detected in the camera view comprises: for each person detected within the anchor frame, generating the bounding box corresponding to the person detected within the anchor frame based on the output tensor data corresponding to the person detected within the anchor frame.
In the same field of endeavor, Pierce et al. teaches wherein for each camera view obtained from an individual camera, generating a bounding box for each person detected in the camera view (see para [0045]; “neural network 300 outputs a first order tensor with five dimensions corresponding to the smallest bounding box around the object of interest”). comprises: for each person detected within the anchor frame, generating the bounding box corresponding to the person detected within the anchor frame (see para [0060]; “neural network 100 will output a bounding box 525 around the predicted object of interest at 523”) based on the output tensor data corresponding to the person detected within the anchor frame (see para [0069]; “The bounding box 625 may be sized and located based on the dimensions 621 contained in first-order tensor 124-OB”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of apparatuses for detecting relevant motion of objects of interest in surveillance videos of Yu et al. and systems for object detection by a neural network comprising a convolution-nonlinearity step and a recurrent step of Pierce et al. in order to increase accuracy by utilizing improved computational operations (see para [0045]).
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al. in view of Ning et al. and Yu et al. as applied in claims 1-2 above, and further in view of Cota et al. (US 20250131697 A1).
Regarding claim 4, the rejection of claim 2 is fully incorporated herein. The combination of Shah et al., Li et al. Ning et al. and Yu et al. does not teach wherein for each camera view obtained from an individual camera, generating a bounding box for each person detected in the camera view comprises: when the camera view includes two or more persons, and when the neural network has generated a set of bounding boxes comprising two or more bounding boxes for each detected person, applying a non-max suppression filtering process with an Intersection over Minimum Area (IoMA) metric to the set of bounding boxes.
In the same field of endeavor, Cota et al. wherein for each camera view obtained from an individual camera, generating a bounding box for each person detected in the camera view comprises: when the camera view includes two or more persons, and when the neural network has generated a set of bounding boxes comprising two or more bounding boxes for each detected person, applying a non-max suppression filtering process with an Intersection over Minimum Area (IoMA) metric to the set of bounding boxes (see para [0025]; “”The subphase of non-maximum suppression or suppression of non-maximums for some classes of findings in which, where there are two bounding boxes of elements belonging to classes that cannot share the same area, if those boxes overlap with an intersection value greater than or equal to a predefined threshold, the box with lower confidence is eliminated; the subphase of mask non-maximum suppression or suppression of non-maximum masks based on intersection over minimum area (class-independent) for some classes of findings, to eliminate findings that are also partially contained in others and cannot share the same area”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking f Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of apparatuses for detecting relevant motion of objects of interest in surveillance videos of Yu and a method for analyzing oral X-Ray pictures and its analysis system of Cota et al. in order to detect the position and shape of the main anatomical parts of the body zone (see para [0025]).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al. in view of Ning et al. and Yu et al. as applied in claims 1-2 above, and further in view of Zhou et al. (US 20190130189 A1) and Chen et al. (US 20220391621 A1).
Regarding claim 5, the rejection of claim 2 is fully incorporated herein. The combination of Shah et al., Li et al. Ning et al. and Yu et al. does not teach wherein for each camera view obtained from an individual camera, generating a bounding box for each person detected in the camera view comprises, when the camera view includes two or more persons, and when the neural network has generated a set of bounding boxes comprising two or more bounding boxes for each detected person: when an individual camera view includes an object that at least partially obstructs the individual camera’s view of one or more persons, applying a non-max suppression filtering process with an Intersection over Minimum Area (IoMA) metric to the set of bounding boxes; and when the individual camera view does not include an object that at least partially obstructs the individual camera’s view of one or more persons, applying a non-max suppression filtering process with an Intersection over Union (IoU) metric to the set of bounding boxes.
In the same field of endeavor, Zhou et al. teach wherein for each camera view obtained from an individual camera, generating a bounding box for each person detected in the camera view comprises, when the camera view includes two or more persons (see Fig. 5D; disclose camera view includes two or more persons), and when the neural network has generated a set of bounding boxes comprising two or more bounding boxes for each detected person (see para [0108]; “the complex object detector system may generate duplicated bounding boxes for a single object from the same video frame. FIG. 5A illustrates examples of duplicated bounding boxes. As shown in FIG. 5A, a complex object detector may generate, from a video frame 500A, detector bounding boxes 502 and 504 for an object 506 (a person)”, see also para [0211]; “FIG. 33 is an illustrative example of a deep learning neural network 3300 that can be used by complex object detector system 608”): applying a non-max suppression filtering process with an Intersection over Minimum Area (IoMA) metric to the set of bounding boxes; and when the individual camera view does not include an object that at least partially obstructs the individual camera’s view of one or more persons, applying a non-max suppression filtering process with an Intersection over Union (IoU) metric to the set of bounding boxes (see para [0110]; “duplicated bounding boxes can be removed based on non-maximum suppression (NMS). With NMS, the video analytics system can compute an intersection-over-union (IoU) ratio for a pair of bounding boxes. If the IoU ratio is higher than a threshold, the video analytics system may determine that the two bounding boxes are likely to be associated with a single detected object. FIG. 5B is a diagram showing an example of an intersection I and union U of two bounding boxes, including bounding box BB.sub.A 522 and bounding box BB.sub.B 524. Both bounding box BB.sub.A 522 and bounding box BB.sub.B 524 can be detector bounding boxes generated on the same video frame. Intersecting region 528 includes the overlapped region between bounding box BB.sub.A 522 and bounding box BB.sub.B 524”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking f Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of apparatuses for detecting relevant motion of objects of interest in surveillance videos of Yu and for tracking objects in one or more video frames of Zhou et al. in order to provide efficient and robust video sequence processing (see para [0108]). However, the combination of Shah et al., Li et al. Ning et al. Yu et al. and Zhou et al. does not teach when an individual camera view includes an object that at least partially obstructs the individual camera’s view of one or more persons.
In the same field of endeavor, Chen et al. teach when an individual camera view includes an object that at least partially obstructs the individual camera’s view of one or more persons (see para [0020]; “MOT system 100 is able to track objects even when the objects cannot be detected by object detection subsystem 110 in one or more of the image frames (e.g., because of being fully or partially occluded or moving fully or partially out of the frame)”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking f Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of apparatuses for detecting relevant motion of objects of interest in surveillance videos of Yu and for tracking objects in one or more video frames of Zhou et al. and a system for tracking a target object across a plurality of image frames of Chen et al. in order to automatically estimate a bounding box for the target object in the target frame based on the occlusion center (see para [0020]).
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al in view of Ning et al. as applied in claim 1 above, and further in view of Fang et al. (US 20230237801 A1).
Regarding claim 6, the rejection of claim 1 is fully incorporated herein. The combination of Shah et al., Li et al. and Ning et al. as a whole does not teach wherein generating one or more sets of associated bounding boxes comprises: for each bounding box in each camera view, generating a ground point for the bounding box that corresponds to a point on a ground plane of the environment where a ray projected from the camera that obtained the camera view intersects a midpoint along a bottom of the bounding box.
In the same field of endeavor, Fang et al. teaches wherein generating one or more sets of associated bounding boxes (see para [0041]; “the player in each bounding box may be viewable from many or all views of scene 210 and, therefore, the bounding boxes correspond as including the same player”) comprises : for each bounding box in each camera view (see para [0045]; “bounding box correspondence build module 301 generates a location on the ground plane of the 3D coordinate system for each bounding box in each video frame at each time instance”), generating a ground point for the bounding box that corresponds to a point on a ground plane of the environment where a ray projected from the camera that obtained the camera view intersects a midpoint along a bottom of the bounding box (see para [0042]; “the location of each such bounding box on a ground plane in the 3D space is determined and used for correspondence. The location of each bounding box on the ground plane may be generated using any suitable technique or techniques. In some embodiments, the distance of foot points for different bounding boxes are used as the constraint to build the correspondence of bounding boxes”, see also para [0044]; “Each foot point 406 (e.g., one for each bounding box in each video frame corresponding to each view) is then projected all foot points from the respective camera view to a ground plane of the 3D coordinate system using the projection matrices for the camera views”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of a technique to performing object or person association or correspondence in multi-view video of Fang et al. in order to provide new and immersive user experiences from multi-view video becomes more widespread (see para [0041]).
Claims 7-9 are rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al and Ning et al. in view of Fang et al as applied in claims 1, and 6 above, and further in view of López-Cifuentes et al. NPL “Semantic-driven multi-camera pedestrian detection” herein after Lopez.
Regarding claim 7, the rejection of claim 6 is fully incorporated herein. The combination of Shah et al., Li et al. Ning et al. and Fang et al. as a whole does not teach wherein generating one or more sets of associated bounding boxes further comprises: mapping ground points corresponding to bounding boxes in each of the camera views to a graph within the ground plane; performing a connected components analysis on the ground points to identify clusters of ground points, wherein an individual cluster includes ground points within a threshold distance of each other on the ground plane; and within each cluster, assigning each bounding box in the cluster to one set of associated bounding boxes corresponding to a detected person based on the ground point of the bounding box.
In the same field of endeavor, Fang et al. teaches wherein generating one or more sets of associated bounding boxes further comprises: mapping ground points corresponding to bounding boxes in each of the camera views to a graph within the ground plane (see page 1219, 3.3; “Everycamera single detection is considered a vertex of a disconnected graph located in the reference plane”); performing a connected components analysis on the ground points to identify clusters of ground points (see page 1219, 3.3; “Vertices are then joined generating connected components Cm, each representing a joint 3D global detection……. The outcome of the fusion process for K cameras is a set of M connected components {Cm, m =1,...,M},each containing Km detections: | Cm |= Km ≤ K,whereKm < K when a person is occluded or not detected in one or more cameras”), wherein an individual cluster includes ground points within a threshold distance of each other on the ground plane ( see page 1219, 3.3, 1; “That vertices in a connected component are close enough. Thel2-norm between any two vertices in Cm shall be smaller than a predefined distance R1: ∥Pj,k,Pj′,k′∥2 ≤ R1 (Fig. 5a). R1 maybefixedintheintervalbetween2.5and3.5 with no influence in the results. We experimentally set R1 = 3 meters to: 1)”); and within each cluster, assigning each bounding box in the cluster to one set of associated bounding boxes corresponding to a detected person based on the ground point of the bounding box (see page 1219, 3.3; “As each connected component is assumed to represent a single person”, see also page 1219, 3.4; “in each camera, ground plane detections need to be back-projected to each camera and 2D bounding boxes enclosing pedestrians need to be outlined based on these projections”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of a technique to performing object or person association or correspondence in multi-view video of Fang et al. and a multi-camera approach to globally combine pedestrian detections leveraging automatically extracted scene context of Lopez in order to signify the versatility and robustness of the proposed method without requiring ad hoc annotations nor human-guided configuration (see page 1219, 3.3).
Regarding claim 8, the rejection of claim 7 is fully incorporated herein.
Fang et al. in the combination further teach wherein within each cluster, assigning each bounding box in the cluster to one set of associated bounding boxes corresponding to a detected person based on the ground point of the bounding box comprises: for a first pair of camera views comprising a first camera view obtained from a first camera and a second camera view obtained from a second camera (see para [0035]; “each camera of camera array 101 has a particular view of scene 210. For example, camera 102 has a first view of scene 210 and camera 103 has a second view of a scene and so on”), wherein the first camera view and the second camera view include ground points within a cluster, associating a first bounding box from the cluster in the first camera view with a second bounding box from the cluster in the second camera view by using a Hungarian algorithm approach based on Euclidean distances between ground points of the bounding boxes in the first camera view and ground points of the bounding boxes in the second camera view (see para [0065-0066]; “groups of tracklets are generated from each frame as a starting frame. Given L.sub.i associated players in a frame i, the location correspondence in subsequent frames may be generated using any suitable technique or techniques such as application of a combinational optimization algorithm inclusive of the Hungarian algorithm… a distance threshold, ε, is used to select a best location pair. For example, given a distance d in a location pair selected using the Hungarian algorithm, only a pair satisfying a distance constraint of d < ε is selected. If no pairing meets the threshold, a location may be added to the frame. Repeating such processing, correspondences in each frame buffer are generated for all players beginning with the first frame in the buffer. Similarly, all player correspondences are generated beginning from frame 2, from frame three and so on to generate groups of tracklets beginning at each frame”).
Regarding claim 9, the rejection of claim 8 is fully incorporated herein.
Fang et al. in the combination further teach wherein within each cluster, assigning each bounding box in the cluster to one set of associated bounding boxes based on the ground point of the bounding box further comprises: for a second pair of camera views comprising a third camera view obtained from a third camera and a fourth camera view obtained from a fourth camera (see para [0037]; “the association lists 122 are provided to optional multi-camera tracking module 115 ….Such data is provided to multi-camera association module 114 to associate or provide correspondence for players from different views of the scene. Multi-camera tracking is applied to fuse all the player information (e.g., team, jersey) together in temporal and spatial domain to generate the player trajectory”), wherein the third camera view and the fourth camera view include ground points within the cluster, associating a first bounding box from the cluster in the third camera view with a second bounding box from the cluster in the fourth camera view by using the Hungarian algorithm approach based on Euclidean distances between ground points of the bounding boxes in the third camera view and ground points of the bounding boxes in the fourth camera view (see para [0042]; “the location of each such bounding box on a ground plane in the 3D space is determined and used for correspondence. The location of each bounding box on the ground plane may be generated using any suitable technique or techniques. In some embodiments, the distance of foot points for different bounding boxes are used as the constraint to build the correspondence of bounding boxes”, see also para [0065]; “groups of tracklets are generated from each frame as a starting frame. Given L.sub.i associated players in a frame i, the location correspondence in subsequent frames may be generated using any suitable technique or techniques such as application of a combinational optimization algorithm inclusive of the Hungarian algorithm. For example, the location of a player x in start frame s may be represented as [00007] L and a correspondence location in a subsequent frame t may be represented as…of player x using the Hungarian algorithm for example. Furthermore, a distance threshold, ε, is used to select a best location pair. For example, given a distance d in a location pair selected using the Hungarian algorithm, only a pair satisfying a distance constraint of d < ε is selected”); generating a first set of composite ground points, wherein each composite ground point in the first set of composite ground points is based on an average of two ground points of two bounding boxes from the cluster from the first camera view and the second camera view; generating a second set of composite ground points, wherein each composite ground point in the second set of composite ground points is based on an average of two ground points of two bounding boxes from the cluster from the second camera view and the third camera view (see para [0063]; “FIG. 12, each frame indicates locations of corresponding bounding boxes at a particular time instance. For example, each player dot may be an average (or other aggregation) of the locations of the players from each of the camera views at the time instance based on the bounding boxes deemed to correspond as discussed above. That is, each dot may be an average location on the ground plane in the 3D coordinate system of the locations of corresponding bounding boxes”); and associating composite ground points in the first set of composite ground points with ground points in the second set of composite ground points by using the Hungarian algorithm approach based on Euclidean distances between the composite ground points in the first set of composite ground points and the ground points in the second set of composite ground points (see para [0067]; “Based on application of the Hungarian algorithm and distance threshold, actual player location 1214 is selected and added to the tracklet for player location 1211. Processing continues with generation of a predicted location 1215 of the player is generated as discussed with respect to Equation (6) and as indicated by L′.sub.x,14 (location, predicted, of player x based on start frame 1 in frame 4). Based on application of the Hungarian algorithm and distance threshold, actual player location 1216 is selected and added to the tracklet for player location 1211. Such processing continues through frame n such that a tracklet 1241 is generated for player location 1211”).
Claims 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al in view of Ning et al. as applied in claim 1 above, and further in view of Gao et al. (US 20180047193 A1).
Regarding claim 11, the rejection of claim 1 is fully incorporated herein.
Ning et al. in the combination further teach wherein for each generated bounding cuboid, associating the generated bounding cuboid with one of a plurality of tracklets for the timeframe based on (a) a position of the generated bounding cuboid during the timeframe (see para [0089]; “Use the Hungarian algorithm to associate the correct track instance (detected measurements, trajectory history) to the predictions based on the cost”, see also para [0086; “assign the correct detected measurements to predicted tracks”, and para [0087]; “the state vector is the 2D/3D object location, represented by a vector of coordinates”, see also para [0088]; “The cost is defined as the Euclidean distance of detections and predictions, which is weighted on the 2D and 3D coordinates”) and (b) predicted future positions for each of the plurality of tracklets for the timeframe (see para [0088]; “Calculate the predictions of the Kalman filters based on tracking history and then calculate the cost between the predictions and the current detections from the detection module”), wherein each tracklet corresponds to a set of historical positions of one of the persons detected in the plurality of camera views (see para [0087]; “a track is an entity consisting of several attributes, including a) the prediction vector representing both the 2D and 3D locations, b) the unique track ID, c) the Kalman filter instances for both 2D and 3D objects, d) the trajectory history of this track, represented by a list of prediction vectors” see also para [0088]; “Calculate the predictions of the Kalman filters based on tracking history”, and para [0093]; “Kalman filter only takes into account the location histories”). However, the combination of Shah et al., Li et al. and Ning et al. as a whole does not teach which comprises: creating one or more cuboid groups, wherein an individual cuboid group contains two or more generated bounding cuboids that have centroids within a threshold distance of each other; and within each cuboid group, assigning each generated bounding cuboid in the cuboid group to one tracklet of the plurality of tracklets based on distances between (i) positions of centroids of the generated bounding cuboids in the cuboid group and (ii) positions of predicted centroids for each tracklet in the plurality of tracklets.
In the same field of endeavor, Gao et al. teaches which comprises: creating one or more cuboid groups, wherein an individual cuboid group contains two or more generated bounding cuboids that have centroids within a threshold distance of each other (see para [0005]; “one object may be detected as two or more blobs in a video frame …..a bounding box merge process for grouping such blobs, and producing a single bounding box that describes the group of blobs as one object”, see also para [0013]; “determining a first distance between the first bounding box and the second bounding box and comparing the first distance to a first distance threshold”, and para [0091]; “calculating the Euclidean distance between the centroid of the tracker (e.g., the bounding box for the tracker) and the centroid of the bounding box of the foreground blob”); and within each cuboid group, assigning each generated bounding cuboid in the cuboid group to one tracklet of the plurality of tracklets (see para [0092]; “The terms (t.sub.x, t.sub.y) and (b.sub.x, b.sub.y) are the center locations of the blob tracker and blob bounding boxes, respectively. As noted herein, in some examples, the bounding box of the blob tracker can be the bounding box of a blob associated with the blob tracker in a previous frame”) based on distances between (i) positions of centroids of the generated bounding cuboids in the cuboid group (see para [0091]; “the cost determination engine 412 can measure the cost between a blob tracker and a blob by calculating the Euclidean distance between the centroid of the tracker (e.g., the bounding box for the tracker) and the centroid of the bounding box of the foreground blob”), and (ii) positions of predicted centroids for each tracklet in the plurality of tracklets (see para [0069]; “prediction of a location of the blob tracker in the next frame (which is the current frame in this example). The prediction of the location of the blob tracker in the current frame can be based on the location of the blob in the previous frame”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of for content-adaptive bounding box merging of Gao et al. in order to more accurately merge (or not merge) bounding boxes and their associated blobs (see para [0005]).
Regarding claim 12, the rejection of claim 11 is fully incorporated herein.
Ning et al. in the combination further teach wherein within each cuboid group, assigning each generated bounding cuboid in the cuboid group to one tracklet of the plurality of tracklets based on distances between (i) positions of centroids of the generated bounding cuboids in the cuboid group, and (ii) positions of predicted centroids for each tracklet in the plurality of tracklets comprises: associating each generated bounding cuboid within the cuboid group with a tracklet in the plurality of tracklets by using a Hungarian algorithm approach based on Euclidean distances between centroids of the bounding cuboids in cuboid group and predicted centroids for each tracklet in the plurality of tracklets (see para [0085]; “As shown in FIG. 3, the 3D object tracking module 362 may perform the tracking using a 3D Kalman filter module 363 and a Hungarian algorithm module 364”, see also para [0086]; “the Hungarian algorithm module 364 is configured to update identities of objects using Hungarian algorithm…. the disclosure applies the Hungarian algorithm to associate identities, i.e., to assign the correct detected measurements to predicted tracks”, and para [0088]; “The cost is defined as the Euclidean distance of detections and predictions, which is weighted on the 2D and 3D coordinates”).
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al in view of Ning et al. as applied in claim 1 above, and further in view of Griesmeyer (US 9612316 B1).
Regarding claim 14, the rejection of claim 13 is fully incorporated herein. The combination of Shah et al., Li et al. and Ning et al. as a whole does not teach wherein the Kalman filter is configured to, for each tracklet in the plurality of tracklets: assume acceleration of the person within the plurality of timeframes is constant; and treat a time derivative of acceleration and a third moment of position in each timeframe of the plurality of timeframes as noise.
In the same field of endeavor, Griesmeyer teach wherein the Kalman filter is configured to, for each tracklet in the plurality of tracklets: assume acceleration of the person within the plurality of timeframes is constant (see col. 22, lines 6-11; “9 State Kalman Filter (163) A 9 state filter for a target assumes that the state of the target may be described by its ECI position, velocity and acceleration, which is assumed to be constant over a time step. Random changes in acceleration are the process error over a time step”); and treat a time derivative of acceleration and a third moment of position in each timeframe of the plurality of timeframes as noise (see col. 22 lines 33-35; “A 12 state filter for a target assumes that the state of the target can be described by its ECI position, velocity, acceleration and jerk (rate of change of acceleration)”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of a system for tracking at least one object using a plurality of pointing sensors and a tracking system of Griesmeyer in order to more accurately merge (or not merge) bounding boxes and their associated blobs (see col. 22, lines 6-11).
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al in view of Ning et al. as applied in claim 1 above, and further in view of Zhu et al. (US 20240071029 A1).
Regarding claim 16, the rejection of claim 1 is fully incorporated herein. The combination of Shah et al., Li et al. and Ning et al. as a whole does not teach wherein the neural network comprises an anchor-free, single-stage object detector.
In the same field of endeavor, Zhu et al. teaches wherein the neural network comprises an anchor-free, single-stage object detector (see [0002]; “Anchor-free object detectors are object detectors that are not reliant on anchor boxes”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of a method of soft anchor-point detection (SAPD), which implements a concise, single-stage anchor-point detector with both faster speed and higher accuracy of Zhu in order to reduce false attention by reweighting their contributions to the network loss according to their geometrical relation with the instance box (see para [0002]).
Claims 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Shah et al. and Li et al in view of Ning et al. as applied in claim 1 above, and further in view of Zhu et al. (US 20240071029 A1).
Regarding claim 17, the rejection of claim 1 is fully incorporated herein.
Li et al. in the combination further teach wherein the environment is a sports environment, wherein the persons include players (see para [0036]; “In general, player detection is performed for each camera view. In an embodiment, the output of block 46 is a 2D bounding box around each player in each camera view. A pre-defined field mask may be used to remove non-sport field areas such as audience areas, coach seating, etc., to reduce noise”). However, the combination of Shah et al., Li et al. and Ning et al. as a whole does not teach and wherein the method further comprises: for each player detected within the video data, identifying a role of the player by mapping an initial position of the player within the sports environment to a particular region within the sports environment.
In the same field of endeavor, Lucey et al. teaches and wherein the method further comprises: for each player detected within the video data (see col. 1, lines 7-9; “analyze participants engaged in an activity and, in particular, to tracking player role using non-rigid formation priors”), identifying a role of the player by mapping an initial position of the player within the sports environment to a particular region within the sports environment (see col. 8 lines 37-42; “As shown in FIG. 2, player 210(2) may move to a position the outside of player 210(5), taking the role of left wing. In response, player 210(6) may move toward the prior position of player 210(2), taking the role of left halfback. Player 210(5) then assumes the role of inside left”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of a method for assigning roles to agents in a group of agents engaging in an activity of Lucey et al. in order to tracking player role using non-rigid formation priors (see col. 1, lines 7-9).
Regarding claim 18, the rejection of claim 1 is fully incorporated herein.
Lucey et al. in the combination further teach further comprising: for at least one player, correlating the player with the data associated with the player, wherein the data associated with the player includes game statistics associated with the player (see col. 7, lines 30-34; “Identifying such emergent patterns of play may aid fans, players, coaches, and broadcasters (including commentators, camera operators, producers, and game statisticians) in understanding as the game evolves and progresses”). Accordingly, it would have been obvious to one of ordinary skills in the art before the effective filling day of the invention to modify a method for providing person re-identification and tracking of Shah et al. in view of a method for technology that obtains multi-camera video data identify an association between a first instance of a 3D object in the first 2D image and a second instance of the 3D object in the second 2D image of Li et al and a method for cooperative maneuvering and cooperative risk warning of vehicles of Ning et al. and further in view of a method for assigning roles to agents in a group of agents engaging in an activity of Lucey et al. in order to tracking player role using non-rigid formation priors (see col. 1, lines 7-9).
Allowable Subject Matter
Claims 10, 15 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding claim 10, none of the cited reference teaches the specific wherein for each timeframe of a plurality of timeframes, generating a bounding cuboid for each set of associated bounding boxes for the timeframe comprises: for every bounding box in the set of associated bounding boxes, generating a point for each corner of the bounding box; and generating a bounding cuboid corresponding to the set of associated bounding boxes that encloses substantially all of the points at each corner of all of the bounding boxes within the set of associated bounding boxes, wherein generating a bounding cuboid corresponding to the set of associated bounding boxes that encloses substantially all of the points at each corner of all of the bounding boxes within the set of associated bounding boxes comprises: estimating a centroid for the set of associated bounding boxes by triangulating mid-points from all of the bounding boxes within the set of associated bounding boxes; back-projecting all four corners of each bounding box within the set of associated bounding boxes onto a plane perpendicular to a camera view intersecting at the estimated centroid; and generating the bounding cuboid based on the back-projected corners and the estimated centroids.
Regarding claim 15, none of the cited reference teaches the specific wherein the Kalman filter comprises a third degree Kalman filter with eighteen dimensional states, wherein the eighteen dimensional states comprise (i) six coordinates that define a bounding cuboid corresponding to a person, (ii) six coordinates that define velocity of the bounding cuboid corresponding to the person, and (iii) six coordinates that define acceleration of the bounding cuboid corresponding to the person.
Regarding claim 19, none of the cited reference teaches the specific further comprising tracking one or more objects detected within the environment, wherein tracking the one or more detected objects comprises, for each detected object that has been detected by the plurality of cameras: generating a point cloud that includes a plurality of points, wherein each point of the plurality of points corresponds to a point in space where two rays projected from two of the plurality of cameras intersect the detected object; creating a selected subset of the plurality of points, wherein the selected subset of the plurality of points includes points that are within a threshold distance of a threshold number of other points of the plurality of points; generating a cuboid associated with the detected object, wherein the cuboid is based on a centroid of the points in the selected subset of points; and matching the generated cuboid associated with the detected object with one of a plurality of object tracklets, wherein each object tracklet corresponds to a tracked object, and wherein matching the generated cuboid associated with the detected object with one of the plurality of object tracklets is based on distances between (i) a position of a centroid of the generated cuboid associated with the detected object, and (ii) positions of predicted centroids for each object tracklet in the plurality of object tracklets.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WINTA GEBRESLASSIE whose telephone number is (571)272-3475. The examiner can normally be reached Monday-Friday9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at 571-270-5180. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WINTA GEBRESLASSIE/ Examiner, Art Unit 2677