Prosecution Insights
Last updated: August 17, 2026
Application No. 18/920,258

Object Recognition Apparatus and Object Recognition Method

Non-Final OA §103
Filed
Oct 18, 2024
Priority
Apr 03, 2024 — RE 10-2024-0045436
Examiner
THOMAS, SOUMYA
Art Unit
Tech Center
Assignee
Kia Corporation
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
3 granted / 4 resolved
+15.0% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
26 currently pending
Career history
28
Total Applications
across all art units

Statute-Specific Performance

§101
8.7%
-31.3% vs TC avg
§103
71.7%
+31.7% vs TC avg
§102
7.6%
-32.4% vs TC avg
§112
7.6%
-32.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1-3 5, 12-14, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ko (US Pub No 20240005632), hereinafter Ko in view of Cha et al. (US Pat. No 11210533), hereinafter Cha, and further in view of Yang et al. (US Pub No 20240249538), hereinafter Yang. As to Claim 1, Ko teaches an object recognition apparatus of a vehicle (see Fig. 21, autonomous driving system 100), the object recognition apparatus comprising: a camera (see Fig. 21, sensors 103, and see paragraph [0126], “Such an object detection device 2140 includes a camera module. The controller 2120 extracts object information from an external image captured by the camera module and allows the controller 2120 to process the information thereabout” ); and a processor (see Fig. 21, image pre-processor 105 and AI processor 109), wherein the processor is configured to: obtain, via the camera, at least one image of an object external to the vehicle (see paragraph [0052], “The image acquiring unit 11 acquires a 2D image obtained by capturing the non-ego vehicle from a camera mounted in an ego vehicle in driving”); determine, based on the at least one image, a camera object box comprising a first plurality of line segments, wherein the camera object box is a two-dimensional rectangular box, and wherein the camera object box surrounds an object image that represents the object (see paragraph [0055], “The bounding box detecting unit 12 detects a 2D bounding box for an object (for example, non-ego vehicle or other vehicle) detected from the 2D image captured by the camera. A shape of the 2D bounding box for a non-ego vehicle detected according to one exemplary embodiment of the present disclosure may be a rectangular shape”); determine, based on information about the first plurality of line segments of the camera object box (see paragraph [0057], “The 3D information reconstructing unit 13 detects a 3D coordinate corresponding to at least one 3D vertex of a virtual 3D bounding box enclosing the non-ego vehicle in the 3D space, from a coordinate corresponding to at least one vertex of the 2D bounding box detected by the bounding box detecting unit 12”), a first point where a portion, of the object, closest to the vehicle is projected onto a ground from an outer contour of the object image (see paragraph [0088], “In step 1136, the 3D information reconstructing unit 13 acquires the 3D coordinate M2 from the 2D intersection point m2 by means of the unprojection conversion”), a second point where a portion, of the object, furthest from the vehicle in a longitudinal direction is projected onto the ground from the outer contour of the object image (see paragraph [0091], “ The 3D coordinate M3 may be acquired from the 2D intersection point m3 by means of the unprojection conversion as illustrated in FIG. 17 . The 3D coordinate M3 acquired as described above becomes a 3D coordinate of a second 3D vertex A of FIG. 5”), and a third point where a portion, of the object, furthest from the vehicle in a lateral direction is projected onto the ground from the outer contour of the object image (see paragraph [0088], “ The 3D coordinate M4 may be acquired from the 2D intersection point m4 by means of the unprojection conversion as illustrated in FIG. 17”, and see Fig. 17 shown below); PNG media_image1.png 519 659 media_image1.png Greyscale determining a length of the object, a width of the object, or a heading of the object (see paragraph [0102], “For example, the length, the width, and the height of the 3D bounding box correspond to the length, the width, and the height of the non-ego vehicle so that a type of vehicle (a sedan, SUV, a Van, a bus, a small truck, or a large truck) may be classified using the length, the width, and the height of the 3D bounding box as vehicle specifications”); track, based on at least one of the length, the width, or the heading of the object, a position of the object (see paragraph [0169], “The information may include a category or a class of a subject which is identifiable through an image. The information may include a position, a width, a height and/or size of a visual object corresponding to the subject in the image”; and control, based on the tracked position of the object, the vehicle (see paragraph [0105], “In various exemplary embodiments, the network interface 113 performs communication with the electronic device in the vehicle to transmit autonomous driving route information and/or autonomous driving control instructions for autonomous driving of the vehicle to the internal block configurations”). Ko fails to teach determining based on a top-view perspective of the object, top-view points respectively corresponding to the first point, the second point, and the third point; determine, based on a second plurality of line segments connecting the top-view points, at least one of a length of the object, a width of the object, or a heading of the object; track, based on at least one of the length, the width, or the heading of the object, a position of the object; and control, based on the tracked position of the object, the vehicle. However, in an analogous art, Cha teaches a method for predicting a trajectory of a vehicle (see abstract) comprising obtaining a camera object box comprising a first plurality of line segments, wherein the camera object box is a two-dimensional rectangular box, and wherein the camera object box surrounds an object image that represents the object (see Col. 1, lines 49-54, “receiving a first image having a plurality of target objects representing a plurality of target vehicles on a road from a camera mounted on the ego vehicle, generating a plurality of preliminary bounding boxes surrounding the plurality of target objects using a first convolutional neural network”, and see Fig 7, 2D, rectangular bounding box 210-IB surrounding a vehicle); obtaining a first, second and third point of an object (see Col. 9, lines 52-54, “Using the first to third sub-bounding boxes 210-FLBB, 210-RLBB and 210-RRBB, the processor 122 operates to generate three tire contact information of the target vehicle”); obtaining top-view points respectively corresponding to the first point, the second point, and the third point (see Col. 11, lines 52-59, “With reference to step S210 of FIG. 12 and FIG. 13A, the processor 122 operates to generate a top-down view image TV01 of the input image IMG01, mapping the contact information of the target vehicle from the input image IMG01 onto the top-down view image TV01 using an inverse perspective matrix. The inverse perspective matrix corresponds to a projective transformation configured to generate the bird's eye view image”, and see Fig. 13, shown, below, the top-view image TV01 generated from IMG01), PNG media_image2.png 269 489 media_image2.png Greyscale determining, based on a second plurality of line segments connecting the top-view points, at least one of a length of the object, a width of the object, or a heading of the object (see Col. 11, lines 28-36, “With reference to step S220 of FIG. 12, FIG. 13A and FIG. 14, the processor 122 operates to determine whether the input image IMG01 has tire contact information of two rear tires of a target vehicle (S221), and to calculate, in response to determining that the target vehicle has the tire contact information (e.g., tire contacts RL and RR on the input image 01) of the two rear tires, a first moving direction /u1 from the two corresponding points RL′ and RR′ on the top down view image (S224)”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the top-view projection taught by Cha with the object detection apparatus taught by Ko. The motivation for doing so would be to accurately reliable to determine the trajectory of an extrinsic object. Cha teaches in Col. 1, lines 20-25, “With at least one of those sensors, the ego vehicle may identify a neighboring vehicle, predict a trajectory thereof and control itself to avoid collision to the neighboring vehicle. Accordingly, safe driving of the ego vehicle may demand reliable and fast prediction of a trajectory of the neighboring vehicle.” Ko in view of Cha fails to teach that determine that the first, second, and third points are derived based on inputting information about the first plurality of line segments of the camera object box into a model that is trained through machine learning. However, in an analogous art, Yang teaches an object detection system (see abstract), which comprises, obtaining a camera object box comprising a first plurality of line segments (see paragraph [0020], “In operation 102, a 2D bounding box is computed for an object in an image of a scene”) determine, based on inputting information about the first plurality of line segments of the camera object box into a model that is trained through machine learning (see paragraph [0022], “In an embodiment, the 2D bounding box may be defined by a position of the box (e.g. center point or top-left anchor point) and a size of the box (i.e. height/width)”, and see paragraph [0024], “In operation 104, the 2D bounding box is processed, using a neural network, to predict a 3D bounding box for the object in the scene”) a first, second, and third point (see paragraph [0024], “In operation 104, the 2D bounding box is processed, using a neural network, to predict a 3D bounding box for the object in the scene. With respect to the present description, the 3D bounding box refers to a 3D cuboid with orientation”, where the Examiner has interpreted the bottom 3 vertices of the cuboid 3D bounding box to be the first, second, and third point). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the neural network taught by Yang with the teachings of Ko and Cha in order to obtain the first, second, and third points. The motivation for doing so would use the neural network to obtain increased accuracy. Yang teaches in paragraph [0060], “Deep learning is a technique that models the neural learning process of the human brain, continually learning, continually getting smarter, and delivering more accurate results more quickly over time.” Thus, it would have been obvious to combine the neural network taught by Yang with the teachings of Ko and Cha in order to obtain the invention as claimed in Claim 1. As to Claim 2, Ko in view of Cha teaches wherein the top-viewpoints comprise: a first top-view point corresponding to the first point, a second top-view point corresponding to the second point, and a third top-view point corresponding to the third point (see Cha, Col. 11, lines 52-59 “With reference to step S210 of FIG. 12 and FIG. 13A, the processor 122 operates to generate a top-down view image TV01 of the input image IMG01, mapping the contact information of the target vehicle from the input image IMG01 onto the top-down view image TV01 using an inverse perspective matrix. The inverse perspective matrix corresponds to a projective transformation configured to generate the bird's eye view image”, and see Fig. 13A, where a top-view image of the three points is shown). Ko in view of Cha fails to teach wherein the information about the first plurality of line segments comprise at least one of: a longitudinal position of a midpoint of a line segment, wherein the line segment is closest, among the first plurality of line segments, to the ground, a lateral position of the midpoint, a width of the camera object box, a height of the camera object box, an area of the camera object box, or a ratio of the width to the height. Ko teaches that vertex corresponding to the line segments may be used (see paragraph [0057], “The 3D information reconstructing unit 13 detects a 3D coordinate corresponding to at least one 3D vertex of a virtual 3D bounding box enclosing the non-ego vehicle in the 3D space, from a coordinate corresponding to at least one vertex of the 2D bounding box detected by the bounding box detecting unit 12”), but fails to teach inputting a midpoint, length, or width of a camera object box are used to obtain the first, second, and third point. However, in an analogous art, Yang teaches obtaining a camera object box comprising a plurality of line segments (see paragraph [0023], “The 2D bounding box refers to a 2D shape (e.g. rectangle) that encloses the object in the image. In an embodiment, the 2D bounding box may be defined by a position of the box (e.g. center point or top-left anchor point) and a size of the box (i.e. height/width).), and using a width and height information of the camera object box to obtain a first, second, and third point (see paragraph [0024], “In operation 104, the 2D bounding box is processed, using a neural network, to predict a 3D bounding box for the object in the scene”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed inventio to combine the neural network taught by Yang with the teachings of Ko and Cha The motivation for doing so would use the neural network to obtain increased accuracy (see Yang, paragraph [0060]). Thus, it would have been obvious to combine the neural network taught by Yang with the teaching of Ko and Cha in order to obtain the invention as claimed in Claim 1. As to Claim 3, Ko in view of Yang fails to teach the processor is configured to determine at least one of the length, the width, or the heading, further based on at least one of: a longitudinal position of the first top-view point, a lateral position of the first top-view point, a longitudinal position of the second top-view point, a lateral position of the second top-view point, a longitudinal position of the third top-view point, or a lateral position of the third top-view point. However, Cha teaches the length, the width, or the heading, may be determined based on at least one of: a longitudinal position of the first top-view point, a lateral position of the first top-view point, a longitudinal position of the second top-view point, a lateral position of the second top-view point, a longitudinal position of the third top-view point, or a lateral position of the third top-view point (see Col. 11, lines 28-36, “With reference to step S220 of FIG. 12, FIG. 13A and FIG. 14, the processor 122 operates to determine whether the input image IMG01 has tire contact information of two rear tires of a target vehicle (S221), and to calculate, in response to determining that the target vehicle has the tire contact information (e.g., tire contacts RL and RR on the input image 01) of the two rear tires, a first moving direction /u1 from the two corresponding points RL′ and RR′ on the top down view image (S224). The first moving direction /u1 may correspond to a direction perpendicular to a line connecting the two corresponding points RL′ and RR′.”, where the examiner has interpreted the ‘moving direction’ as the heading, and see Fig. 13A, where the moving direction u1 is shown). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the top-view projection taught by Cha with the object detection apparatus taught by Ko. The motivation for doing so would be to accurately reliable to determine the trajectory of an extrinsic object (see Cha, Col. 1, lines 20-25). Thus, it would have been obvious to combine the top-view projection taught by Cha with the teachings of Ko in order to obtain the invention as claimed in Claim 3. As to Claim 5, Ko in view of Yang and Cha wherein the processor is further configured to: display a top-view image representing a position, relative to the vehicle, of the object, wherein the top-view image is based on a longitudinal position and a lateral position of the first top-view point, a longitudinal position and a lateral position of the second top-view point, and a longitudinal position and a lateral position of the third top-view point (see Cha, Col. 11, lines 52-59, “With reference to step S210 of FIG. 12 and FIG. 13A, the processor 122 operates to generate a top-down view image TV01 of the input image IMG01, mapping the contact information of the target vehicle from the input image IMG01 onto the top-down view image TV01 using an inverse perspective matrix”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the top-view projection taught by Cha with the object detection apparatus taught by Ko and Yang. The motivation for doing so would be to accurately reliable to determine the trajectory of an extrinsic object (see Cha, Col. 1, lines 20-25). Thus, it would have been obvious to combine the top-view projection taught by Cha with the teachings of Ko and Yang in order to obtain the invention as claimed in Claim 3. As to Claim 12, Ko in view of Yang and Cha teaches an object recognition method (see Ko, paragraph [0007], “a method for measuring a distance between vehicles”), comprising the same limitations recited in Claim 1. Thus, the rejection and rationale are analogous to that of Claim 1. As to Claim 13, Claim 13 claims the same limitation claimed as Claim 2 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 2. As to Claim 14, Claim 14 claims the same limitation claimed as Claim 3 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 3. As to Claim 16, Claim 16 claims the same limitation claimed as Claim 5 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 5. Claim(s) 4 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Ko (US Pub No 20240005632), hereinafter Ko in view of Cha et al. (US Pat. No 11210533), hereinafter Cha, further in view of Yang et al. (US Pub No 20240249538), hereinafter Yang, and further in view of Nishimura et al. (US Pub No 20130182908 ), hereinafter Nishimura. As to Claim 4, Ko in view of Cha and Yang teaches obtaining a heading based on longitudinal and lateral positions of the first top-view point and longitudinal and lateral positions of the second top-view point (see Cha,. 11, lines 28-36). However, Ko in view of Cha and Yang fails to explicitly teach the processor is configured to determine at least one of the length and the width, by determining the width of the object by determining, based on longitudinal and lateral positions of the first top-view point and longitudinal and lateral positions of the second top-view point, a length of a first side of the object; and determining the length of the object by determining, based on the longitudinal and lateral positions of the first top-view point and longitudinal and lateral positions of the third top-view point, a length of a second side of the object. However, in an analogous art, Nishimura teaches obtaining a device for top view of a vehicle(see abstract, “There is provided a vehicle type identification device”, and see Fig. 4, top-view shown) , And using the determining the width of the object by determining, based on longitudinal and lateral positions of the first top-view point and longitudinal and lateral positions of the second top-view point (see paragraph [0008], “a vehicle width estimation section which estimates a vehicle width indicating a width of the vehicle in a real space based on the ground point, the first endpoint”), a length of a first side of the object; and determining the length of the object by determining, based on the longitudinal and lateral positions of the first top-view point and longitudinal and lateral positions of the third top-view point, a length of a second side of the object (see paragraph [0011], “ vehicle length estimation section which estimates a vehicle length indicating a length of the vehicle in the real space based on the ground point, the first endpoint, and the second endpoint which have been detected by the detection section”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the height and width taught by Nishimura with the teachings of Ko, Cha, and Yang. The motivation for doing so would be to increase the accuracy of vehicle identification. Nishimura teaches in paragraph [0088], “According to the first embodiment of the present invention, a vehicle type is identified by taking into account the size of the vehicle V, and therefore, accuracy of the vehicle type identification can be enhanced.” Thus, it would have been obvious to combine the teachings of Nishimura with the teachings of Ko, Cha, and Yang in order to obtain the invention as claimed in Claim 4. As to Claim 15, Claim 15 claims the same limitation claimed as Claim 4 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 4. Claim(s) 6, 8-10, 17, and 19-21 are rejected under 35 U.S.C. 103 as being unpatentable over Ko (US Pub No 20240005632), hereinafter Ko in view of Cha et al. (US Pat. No 11210533), hereinafter Cha, further in view of Yang et al. (US Pub No 20240249538), hereinafter Yang, and further in view of Araki et al. (US Pub No 20230245323), hereinafter Araki. As to Claim 6, Ko in view of Cha and Yang fails to explicitly teach repeatedly performing, for each frame of the at least one image, processes of: the obtaining of the at least one image, the determining of the camera object box, the determining of the first point, the second point, and the third point, and the determining of at least one of the length, the width, or the heading; and tracking the position of the object based on at least one of the length, the width, or the heading in each frame of the at least one image. However, in an analogous art, Araki teaches an object tracking device (see abstract), which comprises: determining a camera object box (see paragraph [0062], “The guidance system 112 obtains an image 210 of the environment surrounding the operating vehicle from a frontal view. In the example shown in FIG. 2, the image 210 contains a vehicle on the road. The guidance system 112 obtains a region-of-interest 212 of the image containing the vehicle”, and see Fig. 12, bounding box 212 surrounding the vehicle) determining a first point corresponding to an object (see paragraph [0035], “The position of an object is recognized, for example, as a position on absolute coordinates with a representative point”, and see first point on Fig. 10), determining at least one point corresponding to the object in bird’s eye view (see paragraph [0071], “In the processing of step S100, the area predictor 138 converts, for example, a coordinate system (a camera coordinate system) of a camera image at a front viewing angle into a coordinate system (a vehicle coordinate system) that is based on the position of the host vehicle M viewed from above. determining the length, width, and heading of the object (see paragraph [0038], “The recognizer 120 may analyze the image captured by the camera 10, and recognize the direction of a vehicle body of another…The direction of the vehicle body is, for example, a yaw angle of another vehicle (an angle of the vehicle body with respect to a line connecting the centers of a lane in a traveling direction of another vehicle)”), and see heading of object shown in Fig. 10), and tracking the position of an object in multiple image frames (see paragraph [0070], “The area predictor 138 obtains the amount of change in the position and size of a bounding box between frames on the basis of the position and size of the bounding box BX(t) recognized by the recognizer 120 and the position and size of a bounding box BX(t−1) recognized in an image frame at a past time (t−1).”, and see paragraph [0071], “FIG. 11 is a flowchart which shows an example of area setting processing by the area predictor 138. In the example of FIG. 11 , the area predictor 138 projects and converts a camera image (for example, the image IM20 in FIG. 10 ) acquired by the image acquirer 110 into a bird's-eye view”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of claimed invention to combine the object tracking across multiple frames taught by Araki with the first, second, and third point taught by Ko and Cha. The motivation for doing so would be to increase tracking accuracy. Araki teaches in paragraph [0072], “By recognizing an object in the next frame in an attention area set in this manner, a possibility that a tracking target object (the motorcycle B) is included in the attention area increases, so that the tracking accuracy can be further improved” Thus, it would have been obvious to combine the object tracking taught by Araki with the object detection method taught by Ko and Cha in order to obtain the invention as claimed in Claim 6. As to Claim 8, Ko in view of Cha fails to teach determining the camera object box by determining, from each of a plurality of frames of the at least one image, the camera object box, wherein the first point, the second point, and the third point are determined from each of the camera object boxes in the plurality of frames, and wherein the processor is further configured to: determine, based on a difference between a first position in a first frame of the plurality of frames and a second position in a second frame of the plurality of frames, a speed of the object in the second frame, wherein the first position is at least one of the first point, the second point, or the third point in the first frame, wherein the second position is at least one of the first point, the second point, or the third point in the second frame, and wherein the second frame occurs later than the first frame in the at least one image. However, Araki teaches obtaining a camera object box for a plurality of frames (see paragraph [0069], “ FIG. 10 is a schematic diagram for describing the setting of an image area and the tracking processing. In an example of FIG. 10, a frame IM20 of a camera image at a current time (t) and a bounding box BX(t) including the motorcycle B at the current time (t) are shown”) obtaining a first point from each camera box (see paragraph [0035], “The position of an object is recognized, for example, as a position on absolute coordinates with a representative point”, and see first point on Fig. 10), determine, based on a difference between a first position in a first frame of the plurality of frames and a second position in a second frame of the plurality of frames, a speed of the object in the second frame (see paragraph [0064], “For example, the area predictor 138 estimates a position and a speed of the motorcycle B after a time point of recognition on the basis of an amount of change in the position of the motorcycle B in the past prior to the time point of recognition of the motorcycle B by the recognizer 120”) wherein the first position is at least one of the first point, and wherein the second position is at least one of the first point, in the second frame and wherein the second frame occurs later than the first frame in the at least one image (see paragraph, [0070], “The area predictor 138 obtains the amount of change in the position and size of a bounding box between frames on the basis of the position and size of the bounding box BX(t) recognized by the recognizer 120 and the position and size of a bounding box BX(t−1) recognized in an image frame at a past time (t−1).”, PNG media_image3.png 542 938 media_image3.png Greyscale Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of claimed invention to combine the object tracking across multiple frames taught by Araki with the first, second, and third point taught by Ko and Cha. The motivation for doing so would be to increase tracking accuracy (see, Araki paragraph [0072]). Thus, it would have been obvious to combine the object tracking taught by Araki with the object detection method taught by Ko and Cha in order to obtain the invention as claimed in Claim 8. As to Claim 9, Ko in view of Cha fails to teach determining, from each of a plurality of frames of the at least one image, the camera object box, wherein the first point, the second point, and the third point are determined from each of the camera object boxes in the plurality of frames, and wherein the processor is further configured to: determine, based on a difference between a first position in a first frame of the plurality of frames and a second position in a second frame of the plurality of frames, a traveling direction of the object, wherein the first position is at least one of the first point, the second point, or the third point in the first frame, wherein the second position is at least one of the first point, the second point, or the third point in the second frame, and wherein the second frame occurs later than the first frame in the at least one image. However, Araki teaches obtaining a camera object box for a plurality of frames (see paragraph [0069], “ FIG. 10 is a schematic diagram for describing the setting of an image area and the tracking processing. In an example of FIG. 10, a frame IM20 of a camera image at a current time (t) and a bounding box BX(t) including the motorcycle B at the current time (t) are shown”) determining a position of a first point from a camera object box (see paragraph [0035], “ The position of an object is recognized, for example, as a position on absolute coordinates with a representative point”, and see first point on Fig. 10), determine, based on a difference between a first position in a first frame of the plurality of frames and a second position in a second frame of the plurality of frames, a traveling direction of the object (see paragraph [0038], “The recognizer 120 may analyze the image captured by the camera 10, and recognize the direction of a vehicle body of another vehicle with respect to a front direction of the host vehicle M or an extending direction of the lane, a width of the vehicle, a position and a direction of wheels of the another vehicle, and the like on the basis of feature information (for example, edge information, color information, information such as a shape and a size of the object) obtained from results of the analysis. The direction of the vehicle body is, for example, a yaw angle of another vehicle (an angle of the vehicle body with respect to a line connecting the centers of a lane in a traveling direction of another vehicle)”), wherein the first position is at least one of the first point in the first frame and wherein the second position is at least one of the first point in the second frame and wherein the second frame occurs later than the first frame in the at least one image (see Fig. 10, shown below, where first point is tracked across multiple frames, at times t-1 and t, and traveling direction is indicated by the black arrows connecting the first points). PNG media_image4.png 542 938 media_image4.png Greyscale Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of claimed invention to combine the object tracking across multiple frames taught by Araki with the first, second, and third point taught by Ko and Cha. The motivation for doing so would be to increase tracking accuracy (see, Araki paragraph [0072]). Thus, it would have been obvious to combine the object tracking taught by Araki with the object detection method taught by Ko and Cha in order to obtain the invention as claimed in Claim 9. As to Claim 10, Ko in view of Cha fails to teach determining the camera object box by determining, from each of a plurality of frames of the at least one image, the camera object box, wherein the first point, the second point, and the third point are determined from each of the camera object boxes in the plurality of frames, and wherein the processor is further configured to: determine a speed of the object by comparing the first point in a first frame of the plurality of frames with the first point in a second frame of the plurality of frames, wherein the second frame occurs later than the first frame, wherein a first portion, of the object, corresponding to the first point in the first frame coincides with a second portion, of the object, corresponding to the first point in the second frame. However, Araki teaches obtaining a camera object box for a plurality of frames (see paragraph [0069], “ FIG. 10 is a schematic diagram for describing the setting of an image area and the tracking processing. In an example of FIG. 10, a frame IM20 of a camera image at a current time (t) and a bounding box BX(t) including the motorcycle B at the current time (t) are shown”) obtaining a first point from each camera box (see paragraph [0035], “The position of an object is recognized, for example, as a position on absolute coordinates with a representative point”, and see first point on Fig. 10), determine a speed of the object by comparing the first point in a first frame of the plurality of frames with the first point in a second frame of the plurality of frames, wherein the second frame occurs later than the first frame (see paragraph [0064], “For example, the area predictor 138 estimates a position and a speed of the motorcycle B after a time point of recognition on the basis of an amount of change in the position of the motorcycle B in the past prior to the time point of recognition of the motorcycle B by the recognizer 120”, and see Fig. 10, where first point is tracked across multiple sequential frames) wherein a first portion, of the object, corresponding to the first point in the first frame coincides with a second portion, of the object, corresponding to the first point in the second frame. (see Fig. 10, where the same portion of the object is tracked across frames). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of claimed invention to combine the object tracking across multiple frames taught by Araki with the first, second, and third point taught by Ko and Cha. The motivation for doing so would be to increase tracking accuracy (see, Araki paragraph [0072]). Thus, it would have been obvious to combine the object tracking taught by Araki with the object detection method taught by Ko and Cha in order to obtain the invention as claimed in Claim 10. As to Claim 17, Claim 17 claims the same limitation claimed as Claim 6 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 6. As to Claim 19, Claim 19 claims the same limitation claimed as Claim 8 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 8. As to Claim 20, Claim 20 claims the same limitation claimed as Claim 9 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 9. As to Claim 21, Claim 21 claims the same limitation claimed as Claim 10 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 10. Claims 7and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Ko (US Pub No 20240005632), hereinafter Ko in view of Cha et al. (US Pat. No 11210533), hereinafter Cha, further in view of Yang et al. (US Pub No 20240249538), hereinafter Yang, and further in view of Park et al. (US Pat No 20220300746), hereinafter Park. As to Claim 7, Ko in view of Cha and Yang teaches the object recognition apparatus comprises a light detection and ranging (LIDAR) device (see Ko, paragraph [0106], “In some exemplary embodiment, sensors 103 include a RADAR, a light detection and ranging (LiDAR) and/or ultrasonic sensors in addition to the image sensor”). However Ko in view of Cha and Yang fails to teach the processor is further configured to: obtain, via the LIDAR device, a LIDAR object box that surrounds the object image, wherein the LIDAR object box is a three-dimensional hexahedron box, wherein the LIDAR object box comprises four top vertices and four bottom vertices, wherein the four bottom vertices of the LIDAR object box are closer, to the ground, than the four top vertices of the LIDAR object box, wherein the four bottom vertices of the LIDAR object box comprise: a first vertex that is closest, among the four bottom vertices, to the vehicle, and a second vertex and a third vertex that are on both sides of the first vertex, wherein the second vertex is closer, between the second vertex and the third vertex, to the vehicle; and train the model based on: inputting, as training data for the first point, a point obtained by projecting, into the at least one image of the object, the first vertex; inputting, as training data for the second point, a point obtained by projecting, into the at least one image of the object, the second vertex; and inputting, as training data for the third point, a point obtained by projecting, into the at least one image of the object, the third vertex. However, in an analogous art, Park teaches a method for training an object recognition model (see abstract) comprising a LIDAR device (see paragraph [0033], “Located on the road 11 is a vehicle 12 that includes a LIDAR sensor 14 and a camera sensor 16”), wherein the device is configured to obtain a LIDAR object box that surrounds the object image (see paragraph [0036], “The 3D bounding boxes used as ground truths to train the 3D object detection model are generally based on point cloud information generated by a LIDAR sensor, such as the LIDAR sensor”), wherein the LIDAR object box is a three-dimensional hexahedron box, wherein the LIDAR object box comprises four top vertices and four bottom vertices, wherein the four bottom vertices of the LIDAR object box are closer, to the ground, than the four top vertices of the LIDAR object box (see paragraph [0036], “Here, points of the point cloud 30 have been utilized to identify objects within the point cloud 30. In this example, objects within the point cloud 30 have been identified by 3D bounding boxes 32A-32E and 34A-34B. The 3D bounding boxes 32A-32E have been identified as vehicles”, and see Fig. 2A, where a 3D box comprising 4 top vertices and 4 bottom vertices) wherein the four bottom vertices of the LIDAR object box comprise: a first vertex that is closest, among the four bottom vertices, to the vehicle, and a second vertex and a third vertex that are on both sides of the first vertex, wherein the second vertex is closer, between the second vertex and the third vertex, to the vehicle; (see paragraph [0036], see Fig. 2A, wherein the 3D bounding boxes comprise four bottom vertices, with and the four bottom vertices comprise a first vertex closest to the ego vehicle and a second and third vertices) and train the model based on: inputting, as training data for the first point, a point obtained by projecting, into the at least one image of the object, the first vertex; inputting, as training data for the second point, a point obtained by projecting, into the at least one image of the object, the second vertex; and inputting, as training data for the third point, a point obtained by projecting, into the at least one image of the object, the third vertex (see paragraph [0044], “As such, the 3D bounding boxes 144A, 144D, 144F, and 144H, which form the training set 145, will be utilized to train the monocular 3D object detection model 170. In this example, the training of the monocular 3D object detection model 170”, and the 3D bounding box comprises the first, second, and third vertex, and see paragraph [0061], “Here, the monocular 3D object detection model 170 receives the image 142 and outputs predicted 3D bounding boxes 544A, 544F, and 544H forming the set 544. The processor(s) 110 uses a loss function 212 to determine a loss between the predicted 3D bounding boxes 544A, 544F, and 544H and the 3D bounding boxes 144A, 144D, 144F, and 144H that act as ground truths”, wherein the predicted bounding boxes comprise a first, second, and third point). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the LIDAR data and training method taught by Park with the object recognition apparatus taught by Ko, Cha, and Yang. The motivation for doing so would be to improve the accuracy of the predicted locations of points. Park teaches in paragraphs [0004-0006], “The neural network models have been trained to identify objects located within the image in a 3D space and generate appropriate 3D bounding boxes around these images…To generate the 3D location of an object, some annotations are based on point cloud information captured from a light detection and ranging (LIDAR) sensor. While these training sets may provide useful data for generating annotations, they suffer from drawbacks”, and see paragraph [0032], “3D bounding boxes that have corresponding 2D bounding boxes from the second subset should be of higher quality and should not suffer as much from parallax and/or synchronization issues. The 3D bounding boxes that form the training set can then be used to train a monocular 3D object detection model”). Thus, it would have been obvious to combine the LIDAR bounding box and training taught by Park with the teachings of Ko, Cha, and Yang in order to obtain the invention as claimed in Claim 7. As to Claim 18, Claim 18 claims the same limitation claimed as Claim 7 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 7. Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Ko (US Pub No 20240005632), hereinafter Ko, in view of Cha et al. (US Pat. No 11210533), hereinafter Cha, further in view of Yang et al. (US Pub No 20240249538), hereinafter Yang, and further in view of Kim et al. (US Pat No 10402724), hereinafter Kim. As to Claim 11, Ko in view of Cha and Yang fails to teach wherein the processor is further configured to: model, based on performing regression , a relationship between: input data comprising information about the camera object box, and output data comprising the first point, the second point, and the third point. However in an analogous art, Kim teaches a method for acquiring a 3D bounding box of a vehicle (see abstract), which comprises obtaining a camera bounding box (see Col. 5, lines 11-14, “A first part 201 in the CNN is configured to acquire the 2D bounding box in the training image”) and model , based on performing regression input data comprising information about the camera object box, and output data comprising the first point, the second point, and the third point (see Col. 7, lines 17-24, “the processor 120 performs or supports another device to perform a process of calculating respective displacements of the vertices of the pseudo-3D box from vertices of the 2D bounding box by using the regression analysis”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the regression taught by Kim with the teachings of Ko, Cha, and Yang. The motivation for doing so would be to increase the accuracy of the final three points. Kim teaches in Col. 2, lines 4-6 “Thus, the present invention proposes a new method for removing such redundant computation and improving the accuracy of detection”. Thus, it would have been obvious to combine the regression taught by Kim with the teachings of Ko, Cha, and Yang in order to obtain the invention as claimed in Claim 11. As to Claim 22, Claim 22 claims the same limitation claimed as Claim 11 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 11. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kwon et al. (US Pub No 20190294177) teaches a method for tracking virtual vehicles which comprises obtaining a bounding box, generating a 3D bounding box from the 2D bounding box, and then generating a bird’s-eye view of the 3D bounding box. Wang et al. (US Pat No 12293543) teaches a method of training an object recognition apparatus which comprises training using LIDAR data and annotations. Velankar et al. (US Pub No 20230032669), hereinafter Velankar teaches a method of generating a 3D bounding box of an object from a 2D bounding box of an object. Dorn et al. (DE Pub No 02022112317-A1) teaches a method of obtaining 3D data from a 2D bounding box surrounding a vehicle. Wang et al. (CN Pub No 117789156-A) teaches a method of obtaining 3D data from a 2D bounding box surrounding a vehicle. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOUMYA THOMAS whose telephone number is (571)272-8639. The examiner can normally be reached M-F 8:30-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.T./ Examiner, Art Unit 2664 /JENNIFER MEHMOOD/ Supervisory Patent Examiner, Art Unit 2664
Read full office action

Prosecution Timeline

Oct 18, 2024
Application Filed
Jul 14, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12614064
NOVEL NEUROMORPHIC VISION SYSTEM
3y 5m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
99%
With Interview (+33.3%)
2y 8m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month