DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 4 is objected to because of the following informalities:
Claim 4, line 1, should be “wherein at least one of the one or more
Appropriate correction is required.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 2, 5, 6, 7, 19, and 20 are rejected under 35 U.S.C. 102(a)(1) and (a)(2) as being anticipated by DE Patent Publication 2017 102017108254 A1, (Wang et al.) (cited from translated document attached).
Claim 1
Regarding Claim 1, Wang et al. disclose a method, comprising: receiving a respective sequence of sensor data from each of a plurality of sensors; ("system also includes a processing system to retrieve the images from the two or more cameras," pg. 2 par. 2) processing each of the respective sequences of sensor data using a neural network ("In block 340 Work is done based on the work done by the cameras 140 images obtained objects 220 to detect and track. These operations include known operations of image processing, computer vision, and machine learning, and may be performed, for example, by a deep-learning neural network," pg. 5 par. 5) to generate a respective network output that represents detection and tracking information for corresponding objects captured in the respective sequence of sensor data; ("configured to output data about the locations of the objects in the field of view of the two or more cameras," pg. 2 par. 9) for each of the respective network outputs, transforming the respective network output to a bird’s eye view (BEV) space to generate a BEV data sequence representing a set of tracked objects in the BEV space; ("The picture output 510a shows a top-down view of a vehicle 101 and five different objects 220 around the vehicle 101 shows around," pg. 6 par. 6) and generating characteristic information for at least a portion of the set of tracked objects based on the BEV data sequence ("Performing a time-based detection … For example, a determination of whether an object 220 on the vehicle 101 moved or moved away from this, affecting data taken by the controller 120 other vehicle systems," pg. 6 par. 2).
Claim 2
Regarding Claim 2, Wang et al. disclose the method of claim 1, wherein the sequences of sensor data comprise sequences of two-dimensional image frames obtained by the plurality of sensors, and wherein the plurality of sensors comprise one or more cameras ("system also includes a processing system to retrieve the images from the two or more cameras," pg. 2 par. 2).
Claim 5
Regarding Claim 5, Wang et al. disclose the method of claim 1, wherein the detection and tracking information for the respective network output comprises a network data sequence representing, for each frame of a plurality of frames in the network data sequence, a set of two-dimensional bounding boxes for one or more objects captured in the frame ("A three-dimensional bounding box (BBOX) is used to create each object 220 through the all-round vision camera system 100 is detected," pg. 7 par. 1).
Claim 6
Regarding Claim 6, Wang et al. disclose the method of claim 5, wherein the detection and tracking information further comprises a location, a dimension, a class, a heading direction, or an identifier for each of the corresponding objects captured in each frame of the plurality of frames in the network data sequence ("Performing cross-image detection is also part of the block processing 340 , One part of this cross-image detection process is that objects 220 associated with each other and made to coincide with each other by more than one camera 140 of the all-round vision camera system 100 be recorded. In essence, the position of an object 220 based on the pictures from two or more cameras 140 be triangulated," pg. 5 par. 6).
Claim 7
Regarding Claim 7, Wang et al. disclose the method of claim 5, wherein for each of the respective network outputs, transforming the respective network output to the BEV space to generate the BEV data sequence representing the set of tracked objects in the BEV space comprises: for each frame of the network data sequence,
PNG
media_image1.png
286
208
media_image1.png
Greyscale
generating a BEV bounding box for each of the corresponding objects in the respective network output ("A three-dimensional bounding box (BBOX) is used to create each object 220 through the all-round vision camera system 100 is detected, display. Color coding or pattern coding can be used to provide additional data about the objects 220 [AltContent: textbox (Figure 6 shows the clusters of bounding boxes in a BEV view created by the model.)]specify. For example, the objects can 220a through one of the side cameras 140a . 140c ( 1 ) have been detected while the objects 220b," pg. 7 par. 1).
Claim 19
Regarding Claim 19, Wang et al. disclose a system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one or more computers to perform respective operations, the operations comprising: ("The processing system 110 may include: an integrated circuit (ASIC), an electronic circuit, a shared, dedicated, or group processor with memory executing one or more software or firmware programs, a combinatorial logic circuit, and / or other suitable ones Components that provide the described functionality," pg. 4 par. 4) receiving a respective sequence of sensor data from each of a plurality of sensors; ("system also includes a processing system to retrieve the images from the two or more cameras," pg. 2 par. 2) processing each of the respective sequences of sensor data using a neural network to generate a respective network output ("In block 340 Work is done based on the work done by the cameras 140 images obtained objects 220 to detect and track. These operations include known operations of image processing, computer vision, and machine learning, and may be performed, for example, by a deep-learning neural network," pg. 5 par. 5) that represents detection and tracking information for corresponding objects captured in the respective sequence of sensor data; ("configured to output data about the locations of the objects in the field of view of the two or more cameras," pg. 2 par. 9) for each of the respective network outputs, transforming the respective network output to a bird’s eye view (BEV) space to generate a BEV data sequence representing a set of tracked objects in the BEV space; ("The picture output 510a shows a top-down view of a vehicle 101 and five different objects 220 around the vehicle 101 shows around," pg. 6 par. 6) and generating characteristic information for at least a portion of the set of tracked objects based on the BEV data sequence ("Performing a time-based detection … For example, a determination of whether an object 220 on the vehicle 101 moved or moved away from this, affecting data taken by the controller 120 other vehicle systems," pg. 6 par. 2) .
Claim 20
Regarding Claim 20, Wang et al. disclose one or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform respective operations, the respective operations comprising: ("The processing system 110 may include: an integrated circuit (ASIC), an electronic circuit, a shared, dedicated, or group processor with memory executing one or more software or firmware programs, a combinatorial logic circuit, and / or other suitable ones Components that provide the described functionality," pg. 4 par. 4) receiving a respective sequence of sensor data from each of a plurality of sensors; ("system also includes a processing system to retrieve the images from the two or more cameras," pg. 2 par. 2) processing each of the respective sequences of sensor data using a neural network ("In block 340 Work is done based on the work done by the cameras 140 images obtained objects 220 to detect and track. These operations include known operations of image processing, computer vision, and machine learning, and may be performed, for example, by a deep-learning neural network," pg. 5 par. 5) to generate a respective network output that represents detection and tracking information for corresponding objects captured in the respective sequence of sensor data; ("configured to output data about the locations of the objects in the field of view of the two or more cameras," pg. 2 par. 9) for each of the respective network outputs, transforming the respective network output to a bird’s eye view (BEV) space to generate a BEV data sequence representing a set of tracked objects in the BEV space; ("The picture output 510a shows a top-down view of a vehicle 101 and five different objects 220 around the vehicle 101 shows around," pg. 6 par. 6) and generating characteristic information for at least a portion of the set of tracked objects based on the BEV data sequence ("Performing a time-based detection … For example, a determination of whether an object 220 on the vehicle 101 moved or moved away from this, affecting data taken by the controller 120 other vehicle systems," pg. 6 par. 2).
1st Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 3, 4, 8, and 9 are rejected under 35 U.S.C. 103 as obvious over DE Patent Publication 2017 102017108254 A1, (Wang et al.) in view of US Patent Publication 2023 0316785 A1, (Habib et al.).
Claim 3
Regarding claim 3, Wang et al. teach the method of claim 1.
Wang et al. do not explicitly teach all of transforming the respective network output to the BEV space using one or more transformation matrices.
However, Habib et al. teach transforming the respective network output to the BEV space using one or more transformation matrices ("the system uses a total of 10 homography transformation matrices to synchronize between each camera pair," par. 61).
Therefore, taking the teachings of Wang et al. and Habib et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the image processing and object detection methods as taught by Wang et al. to use transformation matrices as taught by Habib et al. The suggestion/motivation for doing so would have been that, “the system utilizes a set of transformation matrices where each matrix provides a one-to-one pixel relationship between the plane representing the stage for a given camera perspective and the corresponding plane representing the stage as captured by a reference camera perspective” as noted by the Habib et al. disclosure in paragraph [0021], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that the combination would arrive at the predictable result of improved accuracy and precise pixel-mapping between different camera perspectives without unexpected complications; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 4
Regarding claim 4, Wang et al. teach and Habib et al. teach the method of claim 3 as noted above.
Wang et al. do not explicitly teach all of wherein at least one of the one or more homography transformation matrices is a homography transformation matrix, and where in the homography transformation matrix is determined by: obtaining a first set of points from a frame of the respective sequence of sensor data from a corresponding sensor of the plurality of sensors, obtaining a set of reference points in the BEV space of a BEV image, the reference points corresponding to the first set of points; projecting the first set of points to the BEV space to generate a second set of points; and determining the homography transformation matrix by minimizing a cost based on the reference points and the second set of points.
However, Habib et al. teach wherein at least one of the one or more homography transformation matrices is a homography transformation matrix, and where in the homography transformation matrix is determined by: ("the system uses a total of 10 homography transformation matrices to synchronize between each camera pair," par. 61) obtaining a first set of points from a frame of the respective sequence of sensor data from a corresponding sensor of the plurality of sensors, ("For each camera pair (e.g., camera 151 (Cam C) and camera 152 (Cam B)), there is a homography matrix which provides a one-to-one pixel relationship between the plane representing the stage for a given camera perspective (e.g., Cam B) and the corresponding plane representing the stage as captured by the reference camera perspective," par. 61) obtaining a set of reference points in the BEV space of a BEV image, the reference points corresponding to the first set of points; ("Where (x′,y′,1) represents the x,y coordinates of an item point that contacts the stage for the first camera and (x,y,1) represents the x,y coordinates of the same item point that contacts the stage for the second camera," par. 61) projecting the first set of points ("data is associated to a reference camera by performing one or more processing techniques, such as association by projection," par. 60) to the BEV space to generate a second set of points; ("the center point of the reconstructed bounding box lies somewhere close to a directed line 982 which is projected from the midpoint," par. 76) and determining the homography transformation matrix by minimizing a cost based on the reference points and the second set of points ("The cost matrix is reduced to find the best set of associated items between each camera and the reference camera," par. 78).
Wang et al. and Habib et al. are combined as per claim 3.
Claim 8
[AltContent: textbox (Figure 5 shows the characteristic information for corresponding detected objects.)]
PNG
media_image2.png
340
558
media_image2.png
Greyscale
Regarding claim 8, Wang et al. teach the method of claim 7 as noted above, further comprising: generating BEV characteristic information corresponding to the BEV bounding box ("Performing a time-based detection … For example, a determination of whether an object 220 on the vehicle 101 moved or moved away from this, affecting data taken by the controller 120 other vehicle systems," pg. 6 par. 2) wherein the BEV characteristic information comprises a location, a dimension, or a heading direction for the corresponding object in the BEV space ("The angle of each object 220 relative to the vehicle 101 is shown and indicates the direction of travel of each object 220 at," pg. 6 par. 6).
Wang et al. do not explicitly teach all of based on a homography transformation matrix.
However, Habib et al. teach based on a homography transformation matrix ("Utilizing the homography matrices, items detected in different images are associated with each other based on a proximity assessment," par. 62).
Wang et al. and Habib et al. are combined as per claim 3.
Claim 9
Regarding claim 9, Wang et al. and Habib et al. teach the method of claim 8 as noted above.
Wang et al. teach wherein the heading direction for the object in the BEV space is determined based on a BEV location of a corresponding sensor of the plurality of sensors and a heading direction of the object in the sensor data determined by the neural network ("The angle of each object 220 relative to the vehicle 101 is shown and indicates the direction of travel of each object 220 at," pg. 6 par. 6).
Wang et al. and Habib et al. are combined as per claim 3.
2nd Claim Rejections - 35 USC § 103
Claim 10 is rejected under 35 U.S.C. 103 as obvious over DE Patent Publication 2017 102017108254 A1, (Wang et al.) and US Patent Publication 2023 0316785 A1, (Habib et al.) in view of US Patent Publication 2014 0313339 A1, (Diessner).
Claim 10
Regarding claim 10, Wang et al. and Habib et al. teach the method of claim 9 as noted above.
Wang et al. teach wherein the location of the object in the BEV space is determined based on one or more BEV reference points, ("this cross-image detection process is that objects 220 associated with each other and made to coincide with each other by more than one camera 140 of the all-round vision camera system 100 be recorded. In essence, the position of an object 220 based on the pictures from two or more cameras 140 be triangulated," pg. 5 par. 6) wherein the one or more BEV reference points are determined based on ground points of the object in the corresponding sensor data ("in one transformation technique where the top-view camera is the reference camera perspective, detected items from different camera perspectives are related based on a proximity assessment of where a point representing each item's lower edge that contacts the stage from the various side-view camera perspectives lies," par. 21).
Wang et al. do not explicitly teach all of an offset value, wherein the offset value is determined based on the heading direction for the object in the BEV space and dimension data of the object.
However, Diessner teach an offset value, wherein the offset value is determined based on the heading direction for the object in the BEV space and dimension data of the object ("the vehicle position can be defined by its coordinate in the world coordinate system and vehicle angle relative to the y-axis of the world coordinate system. The camera focal point or camera coordinate is defined by a X-offset and Y-offset relative to the vehicle coordinates," par. 80).
Therefore, taking the teachings of Wang et al., Habib et al., and Diessner as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the image processing and object detection methods as taught by Wang et al. and transformation matrices as taught by Habib et al. to use an object heading offset value as taught by Diessner. The suggestion/motivation for doing so would have been that, “Upon activation of the object detection system (for example: selection of reverse gear), the vehicle coordinate system is set to the origin of a world coordinate system with an angle of 0 degrees (such as shown in FIG. 10). The vehicle main axis is lined up with the y-axis of the world coordinate system. As the vehicle is moving, the vehicle position can be defined by its coordinate in the world coordinate system and vehicle angle relative to the y-axis of the world coordinate system” as noted by the Diessner disclosure in paragraph [0079-0080], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that the combined system would achieve improved vehicle tracking and positioning accuracy by continuously aligning the vehicle coordinate system with the world coordinate system using the object heading offset; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
3rd Claim Rejections - 35 USC § 103
Claims 11, 13, 14, and 16 are rejected under 35 U.S.C. 103 as obvious over DE Patent Publication 2017 102017108254 A1, (Wang et al.) in view of US Patent Publication 2023 0316785 A1, (Habib et al.) and US Patent Publication 2024 0428566 A1, (Huang et al.).
Claim 11
Regarding claim 11, Wang et al. teach the method of claim 7 as noted above, further comprising: for each corresponding frame of the network data sequences, generating clusters of BEV bounding boxes for the corresponding objects, wherein the generating comprises ("A three-dimensional bounding box (BBOX) is used to create each object 220 through the all-round vision camera system 100 is detected, display. Color coding or pattern coding can be used to provide additional data about the objects 220 specify. For example, the objects can 220a through one of the side cameras 140a . 140c ( 1 ) have been detected while the objects 220b," pg. 7 par. 1).
Wang et al. do not explicitly teach all of obtaining the BEV bounding boxes for the corresponding frame of the network data sequences associated with the plurality of sensors; for each pair of BEV bounding boxes that are not associated with the same sensor if the plurality of sensors, computing an Intersection over Union (IoU) value; and clustering the BEV bounding boxes based on the IoU values.
However, Habib et al. teach obtaining the BEV bounding boxes for the corresponding frame of the network data sequences associated with the plurality of sensors; ("A 2D bounding box associated with each item can be created in each portion of the concatenated image (i.e., in each of the images used to create the concatenated image)," par. 41) for each pair of BEV bounding boxes that are not associated with the same sensor if the plurality of sensors, computing an Intersection over Union (IoU) value ("calculating an intersection over union value between bounding boxes of different items in different images from different cameras," par. 43).
Therefore, taking the teachings of Wang et al. and Habib et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the image processing and object detection methods as taught by Wang et al. to use the calculation of intersection over union values as taught by Habib et al. The suggestion/motivation for doing so would have been that, “Items captured in multiple images can be tracked by association by projection, for example, and inferences can be drawn and fused to provide a more accurate prediction of the identity of the items on the stage” as noted by the Habib et al. disclosure in paragraph [0047], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that calculating overlapping bounding box areas would accurately match and verify the same detected object across multiple frames or viewpoints, leading to more reliable tracking results; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Additionally, Huang et al. teach clustering the BEV bounding boxes based on the IoU values ("clustering, in the common coordinate system, the respective bounding boxes based on an Intersection over Union threshold," par. 79).
Therefore, taking the teachings of Wang et al., Habib et al., and Huang et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the image processing and object detection methods as taught by Wang et al. and the calculation of intersection over union values as taught by Habib et al. to use clustering based on intersection over union values methods as taught by Huang et al. The suggestion/motivation for doing so would have been that, “apply to each one of the plurality of pseudo-labelled 3D point clouds, a spatial consistency processing algorithm, including: transforming the respective bounding boxes of corresponding pseudo-labels to a common coordinate system; clustering, in the common coordinate system, the respective bounding boxes based on an Intersection over Union threshold” as noted by the Huang et al. disclosure in paragraph [0079], which also motivates combination because the combination would predictably have a higher efficiency as there is a reasonable expectation that clustering bounding boxes based on an intersection over union threshold in a common coordinate system would predictably reduce computational redundancy and streamline spatial consistency; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 13
Regarding claim 13, Wang et al., Habib et al., and Huang et al. teach the method of claim 11 as noted above.
Wang et al. teach initiating a set of tracked objects in the BEV space for the first frame of the BEV data sequence based on the clusters of BEV bounding boxes ("A three-dimensional bounding box (BBOX) is used to create each object 220 through the all-round vision camera system 100 is detected, display. Color coding or pattern coding can be used to provide additional data about the objects 220 specify. For example, the objects can 220a through one of the side cameras 140a . 140c ( 1 ) have been detected while the objects 220b," pg. 7 par. 1).
Wang et al., Habib et al., and Huang et al. are combined as per claim 11.
Claim 14
Regarding claim 14, Wang et al., Habib et al., and Huang et al. teach the method of claim 13 as noted above.
Wang et al. teach for each frame succeeding the first frame and for each object in the set of tracked objects, matching an identifier for the object with one of the clusters of the BEV bounding boxes in the frame, wherein the identifier is associated with the object by the neural network ("A three-dimensional bounding box (BBOX) is used to create each object 220 through the all-round vision camera system 100 is detected, display. Color coding or pattern coding can be used to provide additional data about the objects 220 specify. For example, the objects can 220a through one of the side cameras 140a . 140c ( 1 ) have been detected while the objects 220b," pg. 7 par. 1).
Wang et al. do not explicitly teach all of updating the set of tracked objects for the frame based on the matching result.
However, Huang et al. teach updating the set of tracked objects for the frame based on the matching result ("generating updated corresponding pseudo-labels," par. 43).
Wang et al., Habib et al., and Huang et al. are combined as per claim 11.
Claim 16
Regarding claim 16, Wang et al., Habib et al., and Huang et al. teach the method of claim 14 as noted above.
Wang et al. do not explicitly teach determining whether one of the set of tracked objects in the frame does not match any of the clusters of BEV bounding boxes in the frame, in response to determining that one of the set of tracked objects in the frame does not match any of the clusters of BEV bounding boxes in the frame, updating a location of the tracked object for the frame using a motion model, and calculating an IoU value using the updated location for each of the remaining clusters of the clusters of the BEV bounding boxes for the frame.
However, Habib et al. teach determining whether one of the set of tracked objects in the frame does not match any of the clusters of BEV bounding boxes in the frame, ("The process also identifies any items that are found to exist in the present scene but were absent in the previous scene at its current location," par. 89) in response to determining that one of the set of tracked objects in the frame does not match any of the clusters of BEV bounding boxes in the frame, updating a location of the tracked object for the frame using a motion model, ("When new items are detected, the system registers the added item and starts tracking the item across subsequent scenes. The process also identifies items that existed in the previous scene but are now absent in the current scene at its previous location. In other words, the user has removed an item from the stage or moved the item to a new location on the stage. For removed or moved items, the system deregisters that item at an old location and either stops tracking the item in subsequent scenes (if removed) or starts tracking the item in subsequent scenes at its new location (if moved)," par. 89) and calculating an IoU value using the updated location for each of the remaining clusters of the clusters of the BEV bounding boxes for the frame ("calculating an intersection over union value between bounding boxes of different items in different images from different cameras," par. 43).
Wang et al., Habib et al., and Huang et al. are combined as per claim 11.
4th Claim Rejections - 35 USC § 103
Claim 12 is rejected under 35 U.S.C. 103 as obvious over DE Patent Publication 2017 102017108254 A1, (Wang et al.), US Patent Publication 2023 0316785 A1, (Habib et al.), and US Patent Publication 2024 0428566 A1, (Huang et al.) in view of US Patent Publication 2026 0017925 A1, (Wei et al.).
Claim 12
Regarding claim 12, Wang et al., Habib et al., and Huang et al. teach the method of claim 11 as noted above.
Wang et al. do not explicitly teach all of summing the IoU values for each cluster of the clusters of BEV bounding boxes.
However, Wei et al. teach summing the IoU values for each cluster of the clusters of BEV bounding boxes ("a summation of the intersection of union of the first bounding box and the second bounding box ," par. 99).
Therefore, taking the teachings of Wang et al., Habib et al., Huang et al., and Wei et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the image processing and object detection methods as taught by Wang et al., the calculation of intersection over union values as taught by Habib et al., and clustering based on intersection over union values methods as taught by Huang et al. to use the summing of intersection over union values as taught by Wei et al. The suggestion/motivation for doing so would have been that, “wherein the loss function is based on a summation of the intersection of union of the first bounding box and the second bounding box and a second intersection of unions of the first bounding box” as noted by the Wei et al. disclosure in paragraph [0116], which also motivates combination because the combination would predictably have a higher reliability as there is a reasonable expectation that umming the intersection over union values provides a more accurate and stable measurement of bounding box overlap, leading to improved convergence and precision in object detection; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
5th Claim Rejections - 35 USC § 103
Claim 15 is rejected under 35 U.S.C. 103 as obvious over DE Patent Publication 2017 102017108254 A1, (Wang et al.), US Patent Publication 2023 0316785 A1, (Habib et al.), and US Patent Publication 2024 0428566 A1, (Huang et al.) in view of US Patent Publication 2023 0282000 A1, (Pang et al.).
Claim 15
Regarding claim 15, Wang et al., Habib et al., and Huang et al. teach the method of claim 14 as noted above.
Wang et al. do not explicitly teach all of determining whether one of the set of tracked objects in the frame matches with more than one of the clusters of BEV bounding boxes in the frame, in response to determining that the tracked object matches with more than one of the clusters of BEV bounding boxes in the frame, computing, for each of the matching clusters, a matching cost to represent a change of position and heading direction of the matching cluster between an immediately preceding frame and the frame, and selecting one of the matching clusters as the cluster to match with the tracked object based on the matching costs.
However, Pang et al. teach determining whether one of the set of tracked objects in the frame matches with more than one of the clusters of BEV bounding boxes in the frame, ("The first detected objects and first undetected objects can be determined by inputting the first objects and the first radar clusters into a data association algorithm," par. 18) in response to determining that the tracked object matches with more than one of the clusters of BEV bounding boxes in the frame, computing, for each of the matching clusters, a matching cost to represent a change of position and heading direction of the matching cluster between an immediately preceding frame and the frame, ("The cost function can be evaluated using a Hungarian algorithm or Murty's algorithm, which ranks all the potential assignments of object data points 810 to exiting objects according to cost and makes the assignments in increasing order of cost. A Murty algorithm minimizes a cost for k object assignments based on a cost matrix, where k is a user-determined number, and the cost matrix is based on object to object measurement distances and object probabilities," par. 55) and selecting one of the matching clusters as the cluster to match with the tracked object based on the matching costs ("A track updated with an object data point 810 confirmed in image data only has the weight increased slightly in proportion to a confidence value output by the CNN 400 when the object is detected," par. 55).
Therefore, taking the teachings of Wang et al., Habib et al., Huang et al., and Pang et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the image processing and object detection methods as taught by Wang et al., the calculation of intersection over union values as taught by Habib et al., and clustering based on intersection over union values methods as taught by Huang et al. to use the comparison of movement costs to match tracked objects with as taught by Pang et al. The suggestion/motivation for doing so would have been that, “The cost function can be evaluated using a Hungarian algorithm or Murty's algorithm, which ranks all the potential assignments of object data points 810 to exiting objects according to cost and makes the assignments in increasing order of cost. A Murty algorithm minimizes a cost for k object assignments based on a cost matrix, where k is a user-determined number, and the cost matrix is based on object to object measurement distances and object probabilities” as noted by the Pang et al. disclosure in paragraph [0055], which also motivates combination because the combination would predictably have a higher productivity as there is a reasonable expectation that such optimization techniques would successfully match tracked objects with reduced assignment error and improved computational efficiency; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
6th Claim Rejections - 35 USC § 103
Claims 17 and 18 are rejected under 35 U.S.C. 103 as obvious over DE Patent Publication 2017 102017108254 A1, (Wang et al.), US Patent Publication 2023 0316785 A1, (Habib et al.), and US Patent Publication 2024 0428566 A1, (Huang et al.) in view of US Patent Publication 2020 0134837 A1, (Varadarajan et al.).
Claim 17
Regarding claim 17, Wang et al., Habib et al., and Huang et al. teach the method of claim 16 as noted above.
Wang et al. do not explicitly teach all of determining whether the IoU value calculated for one of the remaining clusters exceeds a threshold IoU value, and in response to determining that the IoU value calculated for one of the remaining clusters exceeds the threshold IoU value, matching the track object with the cluster associated with IoU value.
However, Huang et al. teach determining whether the IoU value calculated for one of the remaining clusters exceeds a threshold IoU value ("the server 23 can be configured to cluster the SA pseudo-labels 602 (that is, bounding boxes) in the global coordinate system, based on an IoU threshold μ. To ensure the pseudo-labels are consistent across multiple training 3D point clouds, in some non-limiting embodiments of the present technology, the server 23 can further be configured to filter the clusters of the SA pseudo-labels 602 based on a number of the bounding boxes in a cluster of a predetermined size," par. 195).
Additionally, Varadarajan et al. teach in response to determining that the IoU value calculated for one of the remaining clusters exceeds the threshold IoU value, matching the track object with the cluster associated with IoU value ("The intersection over union is a ratio of (A) the intersection between all the blob(s) of the frame with all the tracked object(s) (e.g., bounding box(es) from a previous AI-based object detection) and (B) the union of all the blob(s) of the frame with all the tracked object(s). If the intersection over union value is below a threshold, the dynamic object tracker 108 determines that there is a new object in the frame," par. 18).
Therefore, taking the teachings of Wang et al., Habib et al., Huang et al., and Varadarajan et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the image processing and object detection methods as taught by Wang et al., the calculation of intersection over union values as taught by Habib et al., and clustering based on intersection over union values methods as taught by Huang et al. to use threshold intersection over union values by Varadarajan et al. The suggestion/motivation for doing so would have been that, “the object training quality metrics corresponds to intersection over union values between paired bounding boxes. In such examples, the threshold comparator 214 may increment the object detection period count when all, or most, of the object training quality metrics are greater than or equal to the threshold and decrement and/or reset the object detection period count when one or more of the object training quality metrics are less than the threshold” as noted by the Varadarajan et al. disclosure in paragraph [0030], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that using a threshold intersection over union value would successfully group and match detected objects with the correct cluster based on a clear numerical standard of overlap; and/or because doing so merely combines prior art elements according to known methods to yield predictable results.
Claim 18
Regarding claim 18, Wang et al., Habib et al., and Huang et al. teach the method of claim 16 as noted above.
Wang et al. do not explicitly teach all of determining whether the IoU value calculated for one of the remaining clusters exceeds a threshold IoU value, in response to determining that the IoU value calculated for one of the remaining clusters does not exceed the threshold IoU value, repeatedly increasing a threshold consecutive miss value and updating the IoU value based on the increased threshold consecutive miss value, and in response to determining that the threshold consecutive miss value exceeds a maximum permissible value, removing the tracked object from the set of tracked objects.
However, Huang et al. teach determining whether the IoU value calculated for one of the remaining clusters exceeds a threshold IoU value ("the server 23 can be configured to cluster the SA pseudo-labels 602 (that is, bounding boxes) in the global coordinate system, based on an IoU threshold μ. To ensure the pseudo-labels are consistent across multiple training 3D point clouds, in some non-limiting embodiments of the present technology, the server 23 can further be configured to filter the clusters of the SA pseudo-labels 602 based on a number of the bounding boxes in a cluster of a predetermined size," par. 195).
Additionally, Varadarajan et al. teach in response to determining that the IoU value calculated for one of the remaining clusters does not exceed the threshold IoU value, repeatedly increasing a threshold consecutive miss value and updating the IoU value based on the increased threshold consecutive miss value, ("the example counter 204 can increment, decrement, and/or reset the object detection period count based on the result of a comparison of the object training quality metric to the threshold," par. 30) and in response to determining that the threshold consecutive miss value exceeds a maximum permissible value, removing the tracked object from the set of tracked objects ("the dynamic object tracker 108 may generate coarse blob(s) (e.g., cluster(s) of macroblocks of the frame with area over a threshold (e.g., to filter out false positives of insignificant foreground objects) corresponding to objects of interest in the current frame," par. 18).
Wang et al., Habib et al., Huang et al., and Varadarajan et al. are combined as per claim 17.
Reference Cited
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US Patent Publication 2025 0166395 A1 to Han et al. discloses gathering feature sets from four distinct 2D views across two sensors, applying cross-attention between them, and using the results to detect 3D objects.
.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KARSTEN F LANTZ whose telephone number is (571) 272-4564. The examiner can normally be reached Monday-Friday 8:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached on 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Karsten F. Lantz/Examiner, Art Unit 2664
Date: 8/5/2026
/PING Y HSIEH/Primary Examiner, Art Unit 2664