DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/16/2025 and 07/10/2026 is being considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5-8, 10-11, 14-17 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over KARASEV (No. WO-2023158642-A1 “Karasev”) in view of WANG (No. US-20150329048-A1 “Wang”) and in further view of MACMILLAN (No. US-9325917-B2 “Macmillan”).
Regarding claim 1, Karasev teaches “A method for time of capture (ToC) compensation of bird's eye view (BEV) features, the method comprising:” (a method that includes ... using the plurality of sets of features, a fused bird’s-eye view (BEV) grid; Para 0011);
“obtaining sensor data generated by one or more sensors of a vehicle, wherein the sensor data includes images from a plurality of cameras, and wherein at least two of the images have different ToCs;” (The sensing system 110 can further include one or more cameras 118 to capture images of the driving environment 101... Some of the cameras 118 of the sensing system 110 can be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 101; Para 0029); (The input data 201 can include images and/or any other data, e.g., ... such as timestamps; Para 0043);
“extracting, from the sensor data, BEV features of the images;” (BEV grid can be generated by transforming a respective set of features extracted from the input data into a respective set of points; Para 0018); (a set of camera data features can be extracted from the camera data; Para 0019);
However, while Karasev fails to teach the limitation “comparing overlapping features of the BEV features”.
Wang teaches “comparing overlapping features of the BEV features; and” (the images from the cameras overlap at the corners of the vehicle, where the camera calibration process “stitches” the adjacent images together so that common elements in the separate images directly overlap with each other to provide the desired top-down view; Para 0009); (identifies overlap image areas for adjacent cameras. ... identifies matching feature points in the overlap areas of the images; Para 0011);
However, while Karasev and Wang fail to teach the limitation “applying ToC compensation to the BEV features based on comparing the overlapping features to generate compensated BEV features”.
Macmillan teaches “applying ToC compensation to the BEV features based on comparing the overlapping features to generate compensated BEV features.” (A time lag is determined between the image sensors based on the identified pixel shift. The pair of image sensors is calibrated based on the determined time lag or the identified pixel shift to synchronize subsequent image capture by the image sensors; Col 2, Line 60-65); (The one or more identified edges from the first image are matched to the one or more identified edges in the second image. The pixel shift is determined between the pair of images based, at least in part, on the matching between the one or more edges; Col 3, Line 5-9);
Karasev discloses a vehicle that extracts camera features and transforms them into BEV representations. Wang discloses in a multi camera vehicle, a top-down system with camera vies that overlap in areas and the matching features points within the overlap areas can be identified and used to correct alignment. Macmillan discloses that when overlapping camera images are misaligned, the corresponding features in the overlap can be matched and the result can be used to determine time lag between cameras and calibrate based on the time lag to synchronize the captures. It would be obvious to apply Macmillan features of time compensation to overlapping features used in the BEV system of Karasev and Wang in order to reduce misalignment caused by the different image capture times and improve the spatial consistency and accuracy of the BEV representation.
Karasev, Wang and Macmillan are analogous art as they are related to image sensing, BEV and vehicle.
The motivation for the above is to have accurate overlapping features to reduce misalignment and improve BEV representation.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Karasev by comparing overlapping features of the BEV features as taught by Wang and by applying ToC compensation to the BEV features based on comparing the overlapping features to generate compensated BEV features as taught by Macmillan.
Regarding claim 2, Karasev further teaches “The method of claim 1, wherein extracting, from the sensor data, the BEV features of the images comprises:
extracting, from the sensor data, perspective view features of the images; and” (A camera feature network 212 can receive the camera data 210 and extract a set of camera data features from the camera data 210... Camera data feature network 212 can use any suitable perspective backbone(s) to obtain the set of camera data features; Para 0045);
“projecting the perspective view features onto a BEV space representing an environment surrounding the vehicle.” (Each BEV grid can be generated by transforming a respective set of features extracted from the input data into a respective set of points; Para 0018); (a camera data feature can be a two-dimensional (2D) camera data feature; Para 0045); (Camera data features can be provided to a camera data feature projection component214. The camera data feature projection component 214 can utilize camera data feature projection to transform the set of camera data features into a set of pixel points; Para 0046);
Karasev discloses the feature network the extracts 2D camera features using a perspective backbone. Thus, teaching these are perspective view features extracted from the camera sensor data. The it transforms those features into the BEV grid representation thus teaching the claimed subject matter.
Regarding claim 5, while Karasev fails to teach all of claim 5, Wang teaches “The method of claim 1, wherein comparing the overlapping features of the BEV features comprises:
detecting the overlapping features of the BEV features; and” (identifies overlap image areas for adjacent cameras. .... identifies matching feature points in the overlap areas of the images; Para 0011);
“aligning the overlapping features based on a target ToC.” (A synchronization block 110 synchronizes the timing of the images 102-108 from the cameras 12-18 so that all of the images 32, 34, 36 and 38 are aligned in time before being aligned in space from the calibration process; Para 0028);
Wang discloses detecting features that correspond in overlapping areas of the vehicle’s camera views and matching feature points in the overlapping areas to align corresponding points based on time. While not explicitly mentioning target ToC, Wang teaches aligning based on the time of the images and Macmillan could supplement capture time based on alignment as well.
The motivation for the above is to have efficient detection of overlapping features for accurate alignment of BEV features.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Karasev by detecting the overlapping features of the BEV features and aligning the overlapping features based on a target ToC as taught by Wang.
Regarding claim 6, while Karasev and Wang fails to teach all of claim 6, Macmillan teaches “The method of claim 1, wherein applying the ToC compensation to the BEV features based on comparing the overlapping features comprises:
determining a ToC compensation transform for the BEV features based on comparing the overlapping features; and” (A pixel shift is identified between the captured images based on captured image data representative of the overlapping field of view. A time lag is determined between the image sensors based on the identified pixel shift; Col 2, Line 58-62);
“applying the ToC compensation transform to the BEV features.” (The pair of image sensors is calibrated based on the determined time lag or the identified pixel shift to synchronize subsequent image capture by the image sensors; Col 2, Line 62-65);
Macmillan discloses determining a shift between capture images which corresponds to comparing overlapping image information. They also disclose a time lag between the image that constitutes a ToC compensation transform. Karasev supplies applying the compensation to BEV features.
The motivation for the above is to accurate time information of image for efficient and accurate alignment of BEV features.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Karasev and Wang by determining a ToC compensation transform for the BEV features based on comparing the overlapping features and applying the ToC compensation transform to the BEV features as taught by Macmillan.
Regarding claim 7, Karasev further teaches “The method of claim 1, wherein the sensor data further includes a light detection and ranging (LiDAR) point cloud, and wherein the ToC compensation is at least partially based on the LiDAR point cloud.” (The sensing system 110 can include one or more lidars 112; Para 0027); (Lidar returns (the point cloud) have to be processed, segmented into groups associated with separate hypothesized objects, and matched with objects detected using other sensing modalities (e.g., cameras); Para 0016); (multi-scale BEV space can be four-dimensional, with three spatial dimensions (e.g., 3D voxel space) and a time dimension. Each element of multi-scale BEV space can include a voxel, a time associated with this voxel, and a combined feature vector; Para 0055);
Karasev discloses a sensing system that includes LiDAR and identifies LiDAR returns as a point cloud. Karasev establishes using LiDAR features in a BEV representation that contains time dimension. In combination with the temporal compensation of Macmillan, this discloses using LiDAR information with ToC compensation.
Regarding claim 8, Karasev further teaches “The method of claim 1, further comprising:
generating a BEV image including the compensated BEV features.” (generate, using the plurality of sets of features, a fused bird’s-eye view (BEV) grid; Para 0011); (Each BEV grid can be generated by transforming a respective set of features extracted from the input data; Para 0018);
Regarding claim 10, Karasev further teaches “An apparatus for time of capture (ToC) compensation of bird's eye view (BEV) features, the apparatus comprising:” (an apparatus for performing the methods; Para 00110); (using the plurality of sets of features, a fused bird’s-eye view (BEV) grid; Para 0011);
“a memory for storing sensor data; and
processing circuitry in communication with the memory, wherein the processing circuitry is configured to:” (a memory and a processing device, operatively coupled to the memory; Para 0011);
Claim 10 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 1. Therefore, claim 10 limitations are also rejected with the same rationale as regarding claim 1.
Claim 11 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 2. Therefore, claim 11 limitations are also rejected with the same rationale as regarding claim 2.
Claim 14 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 5. Therefore, claim 14 limitations are also rejected with the same rationale as regarding claim 5.
Claim 15 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 6. Therefore, claim 15 limitations are also rejected with the same rationale as regarding claim 6.
Claim 16 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 7. Therefore, claim 16 limitations are also rejected with the same rationale as regarding claim 7.
Claim 17 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 8. Therefore, claim 17 limitations are also rejected with the same rationale as regarding claim 8.
Regarding claim 19, Karasev further teaches “The apparatus of claim 10, further comprising a vehicle including the memory and the processing circuitry.” (autonomous vehicle (AV); Para 0022); (a memory and a processing device, .... associated with an autonomous vehicle (AV));
Regarding claim 20, Karasev teaches “A non-transitory computer-readable storage medium having instructions encoded thereon, the instructions configured to cause processing circuitry to:” (a non-transitory computer-readable storage medium having instructions stored thereon that, when executed by a processing device; Para 0012);
Claim 20 is directed to a non-transitory computer-readable storage medium and its limitations are similar in scope and functions performed by the method of claim 1. Therefore, claim 20 limitations are also rejected with the same rationale as regarding claim 1.
Claim(s) 3-4 and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over KARASEV in view of WANG and in further view of MACMILLAN and in further view of ALBERTSON (No. US-8269834-B2 “Albertson”).
Regarding claim 3, while Karasev, Wang and Macmillan fail to teach all of claim 3, Albertson teaches “The method of claim 1, further comprising:
tagging the BEV features with image source information, wherein the image source information indicates a camera that was used to capture one or more of the BEV features.” (the object detector system attaching metadata to the image frames and sensed data, and the process passes to block 1206. In one example, metadata includes data such as, but not limited to, a camera identifier, frame number, timestamp, and pixel count; Col 29-20, Line 66-67 & 1-3); (generating streams of tracked object properties with metadata from each image stream; Col 30, Line 8-9);
Albertson discloses carrying source related metadata from image streams into feature information. The “tracked object properties correspond to the derived image features rather than just raw images. In combination with Karasev’s transformation of camera features into BEV features, they suggest tagging the BEV features with the metadata corresponding to the source image information.
Karasev, Wang, Macmillan and Albertson are analogous art as they are related to image sensors and a vehicle.
The motivation for the above is to have accurate source information for improvement of capturing BEV features.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Karasev, Wang and Macmillan by tagging the BEV features with image source information, wherein the image source information indicates a camera that was used to capture one or more of the BEV features as taught by Albertson.
Regarding claim 4, while Karasev, Wang and Macmillan fail to teach all of claim 4, Albertson teaches “The method of claim 3, wherein the image source information is encoded in metadata for the BEV features.” (generating streams of tracked object properties with metadata from each image stream; Col 30, Line 8-9);
Albertson discloses metadata form each image stream which showcases source camera identification as metadata of transformed feature information.
The motivation for the above is to have accurate source information for improvement of capturing BEV features.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Karasev, Wang and Macmillan by wherein the image source information is encoded in metadata for the BEV features as taught by Albertson.
Claim 12 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 3. Therefore, claim 12 limitations are also rejected with the same rationale as regarding claim 3.
Claim 13 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 4. Therefore, claim 13 limitations are also rejected with the same rationale as regarding claim 4.
Claim(s) 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over KARASEV in view of WANG and in further view of MACMILLAN and in further view of LI (No. US-20200200906-A1 “Li”).
Regarding claim 9, while Karasev, Wang and Macmillan fail to teach all of claim 9, Li teaches “The method of claim 1, further comprising: operating an Advanced Driver Assistance System (ADAS) based on the compensated BEV features.” (an advanced driver assistance system (ADAS) for a vehicle ... receive the 3D LIDAR point cloud data, convert the 3D LIDAR point cloud data to a two-dimensional (2D) birdview projection; Para 0003); (the controller is further configured to track the detected object and control an ADAS function of the vehicle based on the tracking; Para 0007);
Li discloses a BEV representation that collects perception information from the BEV and controls the ADAS function based on the information.
Karasev, Wang, Macmillan and Li are analogous art as they are related to image sensors and a vehicle.
The motivation for the above is to have operation of ADAS to efficiently control BEV features of vehicle.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Karasev, Wang and Macmillan by operating an Advanced Driver Assistance System (ADAS) based on the compensated BEV features as taught by Li.
Claim 18 is directed to an apparatus and its limitations are similar in scope and functions performed by the method of claim 9. Therefore, claim 18 limitations are also rejected with the same rationale as regarding claim 9.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US-8295644-B2 (Shulman) – Discloses a live image and a previously acquired or generated image are superimposed or composited to represented a virtual vantage point for flying, driving or navigating a plane, vehicle or vessel.
US-11131753-B2 (Banerjee) – Discloses a vehicle obtains camera sensor data of a camera of the vehicle. The method further obtains lidar sensor data of a lidar sensor of the vehicle. The method determines information related to a motion of the vehicle. The method determines a combined image of the camera sensor data and the lidar sensor data based on the information related to the motion of the vehicle.
US-20230334699-A1 (Schulz) – Discloses a plurality of first image frames are obtained from a camera monitoring a specific geographical area. For each first image frame of the plurality of first image frames, a second image frame is generated by adjusting a viewpoint of the first image frame. A third image frame is generated by rasterizing the second image frame. Photonic content is identified in the third image frame. Invariant components are determined in a plurality of third image frames based on the photonic content identified in subsequent third image frames. Position changes of the camera are determined by identifying position changes of the invariant components in the subsequent third image frames.
US-20240320989-A1 (Zhou) – Discloses generating a plurality of DVS frames, each of which is generated by integrating DVS pixels received from a DVS mounted on a vehicle; transforming at least some of the DVS frames to bird-eye view images, so as to form a plurality of bird-eye view images; aligning the plurality of bird-eye view images according to relative positions and/or orientations of the vehicle at the capturing time of the bird-eye view images, so as to form a plurality of aligned bird-eye view images; combing the plurality of aligned bird-eye view images into one output image.
US-20240412494-A1 (Balachandran) – Discloses a method for multi-sensor fusion includes receiving first information indicative of a first set of BEV features of image data captured by an image sensor; receiving second information indicative of a second set of BEV features of non-image sensor data captured by a non-image sensor; and determining fused data that combines the image data and the non-image sensor data based on the first information, the second information, and third information indicative of differences between BEV features of training data and the first set of BEV features and the second set of BEV features.
US-20250054286-A1 (Huang) – Discloses performing, using images obtained from one or more sensors onboard a vehicle, a 2-dimensional (2D) feature extraction; performing, a 3-dimensional (3D) feature extraction on the images; detecting objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction.
US-11691650-B2 (Li; Lingyun) – Discloses input data that describes sensor data into an object detection model and receive, as an output of the object detection model, object detection data describing features of the plurality of the actors relative to the autonomous vehicle. The computing system can generate an input sequence that describes the object detection data. The computing system can analyze the input sequence using an interaction model to produce, as an output of the interaction model, an attention embedding with respect to the plurality of actors.
Hayes, S., Sharma, S., & Eising, C. (2024, August). Velocity driven vision: asynchronous sensor fusion birds eye view models for autonomous vehicles. In IET Conference Proceedings CP887 (Vol. 2024, No. 10, pp. 23-30). Stevenage, UK: The Institution of Engineering and Technology. (Year: 2024) – Discloses the challenge of radar and LiDAR sensors being asynchronous relative to the camera sensors, for various time latencies. The spatial alignment will be resolved before lifting into BEV space via the transformation of the radar/LiDAR point clouds into the new ego frame coordinate system. Only after this can we concatenate the radar/LiDAR point cloud and lifted camera features. Temporal alignment will be remedied for radar data only, we will implement a novel method of inferring the future radar point positions using the velocity information.
Xie, E., Yu, Z., Zhou, D., Philion, J., Anandkumar, A., Fidler, S., ... & Alvarez, J. M. (2022). M $^ 2$ BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation. arXiv preprint arXiv:2204.05088. (Year: 2022) – Discloses a unified framework that jointly performs3D object detection and map segmentation in the Bird’s Eye View (BEV) space with multi-camera image inputs. Unlike the majority of previous works which separately process detection and segmentation, M2BEV infers both tasks with a unified model and improves efficiency.M2BEVefficiently transforms multi-view2Dimagefeaturesintothe3DBEVfeatureinego-carcoordinates.
Lin, Z., Liu, Z., Xia, Z., Wang, X., Wang, Y., Qi, S., ... & Zhu, C. (2024, June). Rcbevdet: Radar-camera fusion in bird's eye view for 3d object detection. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 14928-14937). IEEE. – Discloses combining multi-view cameras with the economical millimeter-wave radar sensor to achieve more reliable multi-modal 3D object detection.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIGITER D PROTAZI whose telephone number is (571)272-7995. The examiner can normally be reached Monday - Friday 7:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said A Broome can be reached at 5712722931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/B.D.P./Examiner, Art Unit 2612
/Said Broome/Supervisory Patent Examiner, Art Unit 2612