CTNF 19/066,342 CTNF 89213 DETAIL ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Notice on Prior Art Rejections 07-06 AIA 15-10-15 2. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 12-151 AIA 26-51 12-51 Status of Claims 3. This Office Action is in response to the Applicant's application filed February 28, 2025. Claims 1-20 are presently pending and are presented for examination. Claim Rejections - 35 USC § 103 07-20-aia AIA 4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-21-aia AIA 5. Claim s 1, 3-11, and 13-20 are rejected under 35 U.S.C 103 as being unpatentable over Rosenblum et al, US 2024/0391494, in view of Rezvani et al. US 2018/0257666, further in view of Krehl et al. US 2024/0312218, hereinafter referred to as Rosenblum, Rezvani, and Krehl respectively . Regarding claim 1 , Rosenblum discloses a system for processing data, the system comprising: one or more memories for storing radar data from a radar system, the radar data comprising frequency domain data (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 352, 353, 354, 355, 356, 137, “processing unit 110 may combine information from a set of images with additional sensory information (e.g., information from radar, lidar, etc.) to perform the monocular image analysis”) , and image data from a plurality of camera sensors (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, “image acquisition unit 120 may include one or more image capture devices ( e.g., cameras), such as image capture device 122, image capture device 124, and image capture device 126.”) ; and one or more processors in communication with the one or more memories, the one or more processors configured to: encode the image data to generate encoded image data (See at least fig 1-33, ¶ 73, 74, 76, 77, 79, 85, 86, 87, 71, 124, 142, 125, 78, “a graphics processing unit (GPU), support circuits, digital signal processors, integrated circuits, memory, or any other types of devices for image processing and analysis. The image preprocessor may include a video processor for capturing, digitizing and processing the imagery from the image sensors”) ; encode the frequency domain data using an encoder to generate encoded radar data (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 353, 354, 355, 356, 137, 352, “the vehicle may combine information from a set of images with additional sensory information, such as RADAR or LIDAR data. For example, as described above, a vehicle may collect such information using one or more RADAR devices and one or more camera devices included in a vehicle”) ; fuse the encoded radar data and the encoded image data to generate fused data (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 353, 354, 352, 356, 137, 355, “combined analysis of RADAR and camera data described herein provides an improved system for determining road elevation profiles and/or estimating ranges to objects in the environment of a vehicle”) ; and navigate a vehicle based on the fused data (See at least fig 1-33, ¶ 2, 3, 4, 5, 6, 7, 8, 54, 55, 68, 69, 70, 138, “processing unit 110 may combine information from the first and second sets of images with additional sensory information (e.g., information from radar) to perform the stereo image analysis”) . Rosenblum fails to explicitly disclose the radar data comprising frequency domain data. However, Rezvani teaches the radar data comprising frequency domain data (See at least fig 1-26B, ¶ 17, 38, 39, 40, 42, 46, 48, 54, 59, 115, 118, 121, 112, “the data collected by every radar unit would likely need to be jointly processed, which in turn would require all units to be perfectly synchronized with each other in both the RF and digital (sampled) domains”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include the radar data comprising frequency domain data as taught by Rezvani because it would allow the system to significantly improve detection and understanding of the scenery and, in particular, to see around the comer to identify non-line-of-sight (NLOS) targets (Rezvani ¶ 44). Rosenblum fails to explicitly disclose encoded data. However, Krehl teaches encoded data (See at least fig 1-5, ¶ 15, “each vehicle can encode collected sensor data using an autoencoder, which comprises a neural network that can lower the dimensionality of the sensor data. In various applications, the neural network can comprise a bottleneck architecture that provides for the autoencoder reducing the dimensionality of the sensor data.”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include encoded data as taught by Krehl because it would allow the system to perform the dimensionality reduction and/or data compression technique (Krehl ¶ 15). Regarding claim 3 , Rosenblum discloses the system of claim 1, wherein the one or more processors are further configured to perform a query initialization on a query, the query initialization comprising lifting image features of the encoded image data into a bird's-eye-view (BEV) space to generate lifted image features (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, “monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other feature associated with an environment of a vehicle”). Regarding claim 4 , Rosenblum discloses the system of claim 3, wherein the one or more processors are further configured to perform the query, the query comprising using the encoded radar data to query the lifted image features (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, “stereo image analysis module 404 may include instructions for detecting a set of features within the first and second sets of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and the like.”). Regarding claim 5 , Rosenblum discloses the system of claim 3, wherein as part of performing the query initialization, the one or more processors are configured to refine uniformly unprojected image features of the query utilizing deformable attention (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, “Processing unit 110 may execute monocular image analysis module 402 to analyze the plurality of images at step 520, as described in further detail in connection with FIGS. SB-SD below. By performing the analysis, processing unit 110 may detect a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and the like”). Regarding claim 6 , Rosenblum discloses the system of claim 5, wherein the one or more processors are further configured to lift multiscale image features via a lifting transformer to generate lifted BEV features and wherein as part of fusing the encoded radar data and the encoded image data, the one or more processors are configured to use the lifted BEV features as input values to a fusion transformer that uses radar BEV features as the query (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, 169, “The processed information corresponding to the analysis of the first, second, and/or third plurality of images may be combined. In some embodiments, processing unit 110 may perform a combination of monocular and stereo image analyses. For example, processing unit 110 may perform monocular image analysis (e.g., via execution of monocular image analysis module 402) on the first plurality of images and stereo image analysis ( e.g., via execution of stereo image analysis module 404) on the second and third plurality of images.”). Regarding claim 7 , Rosenblum discloses the system of claim 3, wherein as part of fusing the encoded radar data and the encoded image data, the one or more processors are configured to combine the lifted image features and the encoded radar data (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, 169, 172, “processing unit 110 may combine the processed information derived from each of image capture devices 122, 124, and 126 (whether by monocular analysis, stereo analysis, or any combination of the two) and determine visual indicators ( e.g., lane markings, a detected vehicle and its location and/or path, a detected traffic light, etc.) that are consistent across the images captured from each of image capture devices 122, 124, and 126”). Regarding claim 8 , Rosenblum discloses the system of claim 1, wherein the one or more processors are configured to fuse the encoded radar data and the encoded image data based on learnable BEV queries, radar BEV queries, and learnable positional embedding (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, 169, 172, 352, “the vehicle may combine information from a set of images with additional sensory information, such as RADAR or LIDAR data. For example, as described above, a vehicle may collect such information using one or more RADAR devices and one or more camera devices included in a vehicle. This combined image and RADAR data may be used in various ways.”). Regarding claim 9 , Rosenblum discloses the system of claim 1, wherein the Doppler spectrum encoder comprises a neural network encoder (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 85, 86, 87, 71, 137, 142, 169, 172, 352, 79, 138, “stereo image analysis module 404 may implement techniques associated with a trained system (such as a neural network or a deep neural network) or an untrained system, such as a system that may be configured to use computer vision algorithms to detect and/or label objects in an environment”). Regarding claim 10 , Rosenblum discloses the system of claim 1, wherein the Doppler spectrum encoder comprises a ResNet 18 encoder, a Minkwoski Engine, or a Point2Voxel encoder (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 352, 353, 354, 355, 356, 137, “processing unit 110 may combine information from a set of images with additional sensory information (e.g., information from radar, lidar, etc.) to perform the monocular image analysis”). Regarding claim 11 , Rosenblum discloses a method for processing data, the method comprising: encoding image data from a plurality of camera sensors to generate encoded image data (See at least fig 1-33, ¶ 73, 74, 76, 77, 79, 85, 86, 87, 71, 124, 142, 125, 78, “a graphics processing unit (GPU), support circuits, digital signal processors, integrated circuits, memory, or any other types of devices for image processing and analysis. The image preprocessor may include a video processor for capturing, digitizing and processing the imagery from the image sensors”) ; encoding frequency domain data of radar data from a radar system using an encoder to generate encoded radar data (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 353, 354, 355, 356, 137, 352, “the vehicle may combine information from a set of images with additional sensory information, such as RADAR or LIDAR data. For example, as described above, a vehicle may collect such information using one or more RADAR devices and one or more camera devices included in a vehicle”) ; fusing the encoded radar data and the encoded image data to generate fused data (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 353, 354, 352, 356, 137, 355, “combined analysis of RADAR and camera data described herein provides an improved system for determining road elevation profiles and/or estimating ranges to objects in the environment of a vehicle”) ; and navigating a vehicle based on the fused data (See at least fig 1-33, ¶ 2, 3, 4, 5, 6, 7, 8, 54, 55, 68, 69, 70, 138, “processing unit 110 may combine information from the first and second sets of images with additional sensory information (e.g., information from radar) to perform the stereo image analysis”) . Rosenblum fails to explicitly disclose the radar data comprising frequency domain data. However, Rezvani teaches the radar data comprising frequency domain data (See at least fig 1-26B, ¶ 17, 38, 39, 40, 42, 46, 48, 54, 59, 115, 118, 121, 112, “the data collected by every radar unit would likely need to be jointly processed, which in turn would require all units to be perfectly synchronized with each other in both the RF and digital (sampled) domains”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include the radar data comprising frequency domain data as taught by Rezvani because it would allow the system to significantly improve detection and understanding of the scenery and, in particular, to see around the comer to identify non-line-of-sight (NLOS) targets (Rezvani ¶ 44). Rosenblum fails to explicitly disclose encoded data. However, Krehl teaches encoded data (See at least fig 1-5, ¶ 15, “each vehicle can encode collected sensor data using an autoencoder, which comprises a neural network that can lower the dimensionality of the sensor data. In various applications, the neural network can comprise a bottleneck architecture that provides for the autoencoder reducing the dimensionality of the sensor data.”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include encoded data as taught by Krehl because it would allow the system to perform the dimensionality reduction and/or data compression technique (Krehl ¶ 15). Regarding claim 13 , Rosenblum discloses the method of claim 11, further comprising performing a query initialization on a query, the query initialization comprising lifting image features of the encoded image data into a bird's-eye-view (BEV) space to generate lifted image features (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, “monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other feature associated with an environment of a vehicle”). Regarding claim 14 , Rosenblum discloses the method of claim 13, further comprising performing the query comprising using the encoded radar data to query the lifted image features (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, “stereo image analysis module 404 may include instructions for detecting a set of features within the first and second sets of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and the like.”). Regarding claim 15 , Rosenblum discloses the method of claim 13, wherein performing the query initialization, comprises refining uniformly unprojected image features of the query utilizing deformable attention (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, “Processing unit 110 may execute monocular image analysis module 402 to analyze the plurality of images at step 520, as described in further detail in connection with FIGS. SB-SD below. By performing the analysis, processing unit 110 may detect a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, and the like”). Regarding claim 16 , Rosenblum discloses the method of claim 15, further comprising lifting multiscale image features via a lifting transformer to generate lifted BEV features and wherein fusing the encoded radar data and the encoded image data comprises using the lifted BEV features as input values to a fusion transformer that uses radar BEV features as the query (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, 169, “The processed information corresponding to the analysis of the first, second, and/or third plurality of images may be combined. In some embodiments, processing unit 110 may perform a combination of monocular and stereo image analyses. For example, processing unit 110 may perform monocular image analysis (e.g., via execution of monocular image analysis module 402) on the first plurality of images and stereo image analysis ( e.g., via execution of stereo image analysis module 404) on the second and third plurality of images.”). Regarding claim 17 , Rosenblum discloses the method of claim 13, wherein fusing the encoded radar data and the encoded image data comprises combining the lifted image features and the encoded radar data (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, 169, 172, “processing unit 110 may combine the processed information derived from each of image capture devices 122, 124, and 126 (whether by monocular analysis, stereo analysis, or any combination of the two) and determine visual indicators ( e.g., lane markings, a detected vehicle and its location and/or path, a detected traffic light, etc.) that are consistent across the images captured from each of image capture devices 122, 124, and 126”). Regarding claim 18 , Rosenblum discloses the method of claim 11, wherein fusing the encoded radar data and the encoded image data is based on learnable BEV queries, radar BEV queries, and learnable positional embedding (See at least fig 1-33, ¶ 73, 74, 76, 77, 78, 79, 85, 86, 87, 71, 137, 138, 142, 169, 172, 352, “the vehicle may combine information from a set of images with additional sensory information, such as RADAR or LIDAR data. For example, as described above, a vehicle may collect such information using one or more RADAR devices and one or more camera devices included in a vehicle. This combined image and RADAR data may be used in various ways.”). Regarding claim 19 , Rosenblum discloses the method of claim 11, wherein the Doppler spectrum encoder comprises a ResNet 18 encoder, a Minkwoski Engine, or a Point2Voxel encoder (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 352, 353, 354, 355, 356, 137, “processing unit 110 may combine information from a set of images with additional sensory information (e.g., information from radar, lidar, etc.) to perform the monocular image analysis”). Regarding claim 20 , Rosenblum discloses Non-transitory computer-readable media storing instructions, which, when executed by one or more processors, cause the one or more processors to: encode image data from a plurality of camera sensors to generate encoded image data (See at least fig 1-33, ¶ 73, 74, 76, 77, 79, 85, 86, 87, 71, 124, 142, 125, 78, “a graphics processing unit (GPU), support circuits, digital signal processors, integrated circuits, memory, or any other types of devices for image processing and analysis. The image preprocessor may include a video processor for capturing, digitizing and processing the imagery from the image sensors”) ; encode frequency domain data of radar data from a radar system using an encoder to generate encoded radar data (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 353, 354, 355, 356, 137, 352, “the vehicle may combine information from a set of images with additional sensory information, such as RADAR or LIDAR data. For example, as described above, a vehicle may collect such information using one or more RADAR devices and one or more camera devices included in a vehicle”) ; fuse the encoded radar data and the encoded image data to generate fused data (See at least fig 1-33, ¶ 138, 139, 140, 186, 215, 311, 353, 354, 352, 356, 137, 355, “combined analysis of RADAR and camera data described herein provides an improved system for determining road elevation profiles and/or estimating ranges to objects in the environment of a vehicle”) ; and navigate a vehicle based on the fused data (See at least fig 1-33, ¶ 2, 3, 4, 5, 6, 7, 8, 54, 55, 68, 69, 70, 138, “processing unit 110 may combine information from the first and second sets of images with additional sensory information (e.g., information from radar) to perform the stereo image analysis”) . Rosenblum fails to explicitly disclose frequency domain data of radar data from a radar. However, Rezvani teaches frequency domain data of radar data from a radar (See at least fig 1-26B, ¶ 17, 38, 39, 40, 42, 46, 48, 54, 59, 115, 118, 121, 112, “the data collected by every radar unit would likely need to be jointly processed, which in turn would require all units to be perfectly synchronized with each other in both the RF and digital (sampled) domains”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include frequency domain data of radar data from a radar as taught by Rezvani because it would allow the system to significantly improve detection and understanding of the scenery and, in particular, to see around the comer to identify non-line-of-sight (NLOS) targets (Rezvani ¶ 44). Rosenblum fails to explicitly disclose encoded data. However, Krehl teaches encoded data (See at least fig 1-5, ¶ 15, “each vehicle can encode collected sensor data using an autoencoder, which comprises a neural network that can lower the dimensionality of the sensor data. In various applications, the neural network can comprise a bottleneck architecture that provides for the autoencoder reducing the dimensionality of the sensor data.”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include encoded data as taught by Krehl because it would allow the system to perform the dimensionality reduction and/or data compression technique (Krehl ¶ 15) . 07-21-aia AIA 6. Claim s 2 and 12 are rejected under 35 U.S.C 103 as being unpatentable over Rosenblum, in view of Rezvani, in view of Krehl, further in view of Yanik et al. US 2024/0134008, hereinafter referred to as Krehl respectively . Regarding claim 2 , Rosenblum discloses the system of claim 1. Rosenblum fails to explicitly disclose wherein the frequency domain data comprises a range- Doppler cuboid and the encoder comprises a Doppler spectrum encoder. However, Yanik teaches wherein the frequency domain data comprises a range- Doppler cuboid and the encoder comprises a Doppler spectrum encoder (See at least fig 1-14, ¶ 34, 35, 36, 77, 83, 88, 111, 7, “A signal is received. A range Fast Fourier Transform (FFT), a Doppler FFT, and an angle FFT are performed on the signal to generate a radar cube”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include wherein the frequency domain data comprises a range- Doppler cuboid and the encoder comprises a Doppler spectrum encoder as taught by Yanik because it would allow the system to received reflected signals to train an object classifier, or to classify objects corresponding to received reflected signals using a pretrained classifier (Yanik ¶ 4). Regarding claim 12 , Rosenblum discloses the method of claim 11. Rosenblum fails to explicitly disclose wherein the frequency domain data comprises a range- Doppler cuboid and the encoder comprises a Doppler spectrum encoder. However, Yanik teaches wherein the frequency domain data comprises a range- Doppler cuboid and the encoder comprises a Doppler spectrum encoder (See at least fig 1-14, ¶ 34, 35, 36, 77, 83, 88, 111, 7, “A signal is received. A range Fast Fourier Transform (FFT), a Doppler FFT, and an angle FFT are performed on the signal to generate a radar cube”) . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Rosenblum and include wherein the frequency domain data comprises a range- Doppler cuboid and the encoder comprises a Doppler spectrum encoder as taught by Yanik because it would allow the system to received reflected signals to train an object classifier, or to classify objects corresponding to received reflected signals using a pretrained classifier (Yanik ¶ 4). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to LUIS A MARTINEZ BORRERO whose email is luis.martinezborrero@uspto.gov and telephone number is (571)272-4577. The examiner can normally be reached on M-F 8:00-5:00. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, HUNTER LONSBERRY can be reached on (571)272-7298. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LUIS A MARTINEZ BORRERO/Primary Examiner, Art Unit 3665 Application/Control Number: 19/066,342 Page 2 Art Unit: 3665 Application/Control Number: 19/066,342 Page 3 Art Unit: 3665 Application/Control Number: 19/066,342 Page 4 Art Unit: 3665 Application/Control Number: 19/066,342 Page 5 Art Unit: 3665 Application/Control Number: 19/066,342 Page 6 Art Unit: 3665 Application/Control Number: 19/066,342 Page 7 Art Unit: 3665 Application/Control Number: 19/066,342 Page 8 Art Unit: 3665 Application/Control Number: 19/066,342 Page 9 Art Unit: 3665 Application/Control Number: 19/066,342 Page 10 Art Unit: 3665 Application/Control Number: 19/066,342 Page 11 Art Unit: 3665 Application/Control Number: 19/066,342 Page 12 Art Unit: 3665 Application/Control Number: 19/066,342 Page 13 Art Unit: 3665 Application/Control Number: 19/066,342 Page 14 Art Unit: 3665 Application/Control Number: 19/066,342 Page 15 Art Unit: 3665 Application/Control Number: 19/066,342 Page 16 Art Unit: 3665