DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Prior arts cited in this office action:
Yathirajam et al. (WO 2024062025 A1, hereinafter “Yathirajam”)
Beijbom et al. (CN 111160561 B, hereinafter “Beijbom”)
Oya (US 20240087100 A1, hereinafter “Oya”)
Ren et al. (US 20230319218 A1, hereinafter “Ren”)
Response to Arguments
Applicant’s Arguments/Remarks filed on 07/31/2026 have been fully considered and are moot in view of new ground of rejection set forth below.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 9, 17 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Yathirajam et al. (WO 2024062025 A1, hereinafter “Yathirajam”) in view of Beijbom et al. (CN 111160561 B, hereinafter “Beijbom”), Ren et al. (US 20230319218 A1, hereinafter “Ren”).
Regarding claims 1, 9, 17 and 25:
Yathirajam teaches a method for image processing (page 1 line 5-14, where Yathirajam teaches a computer-implemented method for automatic envi-ronmental perception based on multi-modal sensor data of a vehicle, wherein a first image of a first environmental sensor modality of the vehicle and a second image of a second environmental sensor modality of the vehicle are received), comprising:
receiving an image frame captured by a fisheye image sensor (Yathirajam page 1 line 5-14, where Yathirajam teaches a computer-implemented method for automatic envi-ronmental perception based on multi-modal sensor data of a vehicle, wherein a first image of a first environmental sensor modality of the vehicle and a second image of a second environmental sensor modality of the vehicle are received);
determining locations for one or more objects depicted within the image frame (Yathirajam page 11 lines 6-24, where Yathirajam teaches a result of the semantic segmentation task comprises a semantically segmented image. In the semantically segmented image, a respective pixel level object class, such as dynamic object, static object, road surface, lane marking, et cetera, is for example assigned to each pixel of the first image. According to several implementations, the at least one visual perception task comprises an object detection task and/or a depth estimation task);
determining bounding surfaces for the one or more objects, wherein the bounding surfaces comprise a three-dimensional representation of a region containing the one or more objects and are defined by three-dimensional spatial coordinates Yathirajam page 11 lines 19-30, Yathirajam teaches the reference coordinate system may also be a sensor coordinate system of one of the environmental sensor modalities. The extrinsic parameters may for example comprise three spatial coordinates specifying a position in the reference coordinate system and three angles specifying the orientation in the coordinate system. The angles may for example be Euler angles, which are denotes as yaw, roll and pitch angle, respectively. The result of the object detection task comprises position information for one or more bounding boxes for respective objects in the environment of the vehicle and a respective object class assigned to the object or the bounding box, respectively. This type of procedure can be applied to any other camera model or sensor model, for example for fisheye cameras. ) , the three-dimensional spatial coordinates being expressed relative to a location of the fisheye image sensor when the image frame was captured (Yathirajam page 11 lines 19-30, Yathirajam teaches in case of a rectangular bounding box, its position may be given by a center position of the rectangle or a corner position of the rectangle or another defined position of the rectangle. In this case, the size of the bounding box may be given by a width and/or height of the rectangle or by equivalent quantities. In particular, the intrinsic parameters define, how a point in the three- dimensional environment is mapped to a pixel in the two-dimensional image plane or sensor plane of the environmental sensor modality. For example, in case of a visible range camera or a thermal camera, the intrinsic parameters may comprise coordinates of a center of projection, commonly denoted as cx, cy, for example, and focal lengths, commonly denoted as fx, fy, for example. The intrinsic parameters may for example be described in terms of a corresponding camera matrix. The distortion parameters of an environmental sensor modality describe the distortion of a respective image generated by the environmental sensor modality, in particular compared to an undistorted two-dimensional representation of the environment. The distortion parameters define, in particular, how a pixel position (u', v') of the respective image is transformed to a corresponding pixel position (u, v) in an undistorted image, also denoted as rectified image. The distortion parameters may be given by a respective model for the environmental sensor modality, such as a pin-hole camera model or a fisheye camera model, et cetera); and
determining control instructions for a vehicle based on the bounding surfaces (Yathirajam page 13 lines 26-31, Yathirajam teaches the method further comprises generating at least one control signal for guiding the vehicle at least in part automatically depending on the result of the at least one visual perception task).
Yathirajam fails to explicitly teach a region containing the one or more objects is defined by three-dimensional spatial coordinates.
However, Beijbom teaches a set of measurements includes a plurality of data points representing a plurality of objects in a three-dimensional (3D) space around the vehicle. Each of the plurality of data points is a set of 3D spatial coordinates. The location module 408 determines the AV location by calculating the location using data from the sensor 121 and data from the database module 410 (e.g., geographic data) (Beijbom [0004], [0087]).
Therefore, taking the teachings of Yathirajam and Beijbom as a whole, it would have been obvious to one or ordinary skill in the art before the effective filing date of the application to defined the location of the object by three-dimensional spatial coordinates relative to the camera in order to properly and more accurately locate the object in the 3D spatial environment especially when considering the camera mounted on the vehicle is fixed.
The combination fails to teach wherein the bounding surfaces comprise at least one curved bounding surface; determining, in the fisheye coordinate space, locations for one or more objects depicted within the image frame based on the fisheye image data without modifying the fisheye image data;
However, Ren teaches at a high level, objects (e.g., salient objects) may be detected from sensor data (e.g., fisheye images) representing an environment surrounding an ego-object such as a vehicle. 3D object detection is performed (e.g., by processing sensor data) and/or a representation of detected 3D objects (e.g., 3D cuboids in rig coordinates) is accessed. For example, 2D bounding boxes and/or 3D cuboids may be predicted from image data (e.g., fisheye images). The 3D object detector 1620 may include or trigger one or more machine learning models (e.g., neural networks) to predict 3D cuboids from the input images 1610, corresponding LiDAR or RADAR detections, and/or other sensor data, and distances and directions between the ego-object (e.g., vehicle center) to detected objects (e.g., closest point, 3D cuboids corner(s), center) may be computed from the predicted 3D cuboids (Ren [0063], [0076], [0143], [0179] [0187], [0222], figs. 15, 16, 27-28).
Therefore, taking the teachings of Yathirajam, Beijbom and Ren as a whole it would have been obvious to one of ordinary skill in the art before the effective filing date of the application to determine the object location in the fish-eyes image(s) bounding curve surface and using a cuboid that represents the coordinate of the location of the object in the fish-eye image(s), in order to avoid additional steps of converting or transforming the fisheye image(s) to world coordinate space or other coordinate and save power to increase the efficiency of the system and the speed of presenting the image for viewing.
Claims 2-5, 7-8, 10-13, 15-16, 18-21, 23-24, 26-29 and 31 are rejected under 35 U.S.C. 103 as being unpatentable over Yathirajam et al. (WO 2024062025 A1, hereinafter “Yathirajam”) in view of Beijbom et al. (CN 111160561 B, hereinafter “Beijbom”), in view of Ren et al. (US 20230319218 A1, hereinafter “Ren”) and in view of Oya (US 20240087100 A1, hereinafter “Oya”).
Regarding claims 2, 10, 18 and 26:
Yathirajam in view of Beijbom and in view of Ren fails to teach wherein the three-dimensional spatial coordinates are three-dimensional polar coordinates representing portions of a viewing area of the fisheye image sensor
However, Oya teaches FIG. 1 is a schematic diagram illustrating projection in imaging a three-dimensional space represented by a polar coordinate system through an optical system 103 as a fisheye lens (Oya [0017]-[0018], [0022]-[0023]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the application to obtain the image frame as fisheye image data from the fisheye image sensor, the fisheye image data being in a fisheye coordinate space, in order to transform the location of each object or bounding box to the corresponding image space or world coordinate and/or to facilitate further calculation and to take advantage of using polar coordinate because some calculations are better or easier when performed in polar coordinate and some are better when performed in cartesian (Oya [0022]-[0023
Regarding claims 3, 11, 19 and 27:
Yathirajam in view of Beijbom, in view Ren and in view of Oya teaches wherein, for each respective bounding surface of the bounding surfaces, the three-dimensional polar coordinates include a center coordinate for the respective bounding surface, depth of the respective bounding surface, a first latitude dimension of the respective bounding surface, a first longitude dimension of the respective bounding surface, a second latitude dimension of the respective bounding surface, and a second longitude dimension of the respective bounding surface (Oya [0019]-[0022], [0032], [0133]-[0139]).
Regarding claims 4, 12, 20 and 28:
Yathirajam in view of Beijbom, in view of Ren and in view of Oya teaches wherein the three-dimensional polar coordinates are converted into cartesian coordinates, and wherein the control instructions are determined based on the cartesian coordinates (Beijbom 0075]-[0076], where X and Y coordinate can be used which is a simple conversion to one of ordinary skill in the art to determine the control of the vehicle on the map (see us patent (12033482 B1) for example).
Regarding claims 5, 13, 21 and 29:
Yathirajam in view of Beijbom, in view of Ren and in view of Oya teaches wherein an encoder model is configured to determine the locations for the one or more objects and a decoder model is configured to determine the bounding surfaces for the one or more objects (Yathirajam page 3; Beijbom [0040], [0137]-[0138], [0150]).
Regarding claims 7, 15 and 23:
Yathirajam in view of Beijbom, in view of Ren and in view of Oya teaches further comprising training a model based on the bounding surfaces (Yathirajam page 1 line 5-14, page 6 line 17-35; Oya [0033]).
Regarding claims 8, 16 and 24:
Yathirajam in view of Beijbom in view of Ren and in view of Oya teaches wherein the locations are determined to identify pixels within the image frame that correspond to the one or more objects (Yathirajam page 11 lines 15-24).
Regarding claim 31:
Yathirajam in view of Beijbom in view of Ren and in view of Oya teaches wherein determining the bounding surfaces in the fisheye coordinate space includes determining the bounding surfaces using the fisheye image data without rectifying the fisheye image data ((Beijbom [0092]-[0093]; Oya [0005], [0018], [0022]-[0023], claim 1).
Regarding claim 32:
Yathirajam in view of Beijbom in view of Ren and in view of Oya teaches wherein each respective bounding surface is a bounding box defined by a center coordinate, a depth, a first latitude dimension, a first longitude dimension, a second latitude dimension, and a second longitude dimension (Oya [0019]-[0022], [0032], [0133]-[0139]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEDNEL CADEAU whose telephone number is (571)270-7843. The examiner can normally be reached Mon-Fri 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chieh Fan can be reached at 571-272-3042. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WEDNEL CADEAU/Primary Examiner, Art Unit 2632 August 27, 2026