DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. KR10-2024-0061714, filed on 05/10/2024.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 11/20/2024 has been considered by the examiner.
Status of Claims
Claims 1-20 are currently pending in this application.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 5, 6-7, 11-12, 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Paik et al. (US 10,885,667 B2) (hereinafter, “Paik”) in view of Lee (US 2022/0275677 A1).
Regarding claim 1, Paik discloses [a method of controlling operation of a vehicle], the method comprising:
acquiring an image [from a camera of the vehicle] (Column 5 [lines 12-15] “generating a multi-ellipsoid based three-dimensional (3D) human model using perspective features of a plurality of two-dimensional (2D) images obtained by the multiple cameras”);
extracting, [based on semantic segmentation of the image]:
a first foot pixel coordinate (foot position coordinates in Column 10 [lines 5-7] equates to a first foot pixel coordinate) corresponding to a foot of a person detected in the acquired image (Column 10 [lines 5-7] "Step 1 (step S121) is to extract valid foot and head data to calculate homology from foot to head"; Column 10 [lines 62-65] "When the foot position coordinates on the ground plane are given, the corresponding head position in the image plane be may determined using Hhf. H=Hfh is defined as the homology from foot to head"; Column 20 [lines 53-58] "In addition, the object occlusion detection device 100 may determine a position of about 20% from the bottom as the foot position, rather than determining the lowest position, in determining a vertically lower position as the foot position in the object region."); and
a head pixel coordinate (head position in Column 20 [lines 48-53] equates to a head pixel coordinate) corresponding to a head of the person (Column 20 [lines 48-53] " In step 115, the object occlusion detection device 100 detects a head position and a foot position using the extracted object region, individually... the object occlusion detection device 100 may determine the vertically highest position—the vertex—as the head position in the object region.");
generating, based on transforming a coordinate system of the acquired image to a transformed coordinate system, a transformed view image (Column 8 [lines 12-18] " In order to normalize the object information, camera parameters are estimated using automatic scene calibration, and a projection matrix is estimated using the camera parameters obtained by scene calibration. After the normalized object information is obtained, the object in the two-dimensional image is projected onto the three-dimensional world coordinate using the projection matrix."; Equation 39 and Column 22 [lines 49-53] “Where xf denotes the foot position on the 2D image, P denotes the projection matrix, and X denotes the coordinates of the inversely-projected xf. The coordinates of the inversely-projected X are normalized by the Z-axis value to detect the foot position in the 3D space”);
determining, based on the transformed view image, a second foot pixel coordinate (foot position in the 3D space in Column 22 equates to second foot pixel coordinate), in the transformed view coordinate system, corresponding to the first foot pixel coordinate (foot position in the 2D image in Column 22 equates to first foot pixel coordinate) (Column 22 [lines 40-45] "The foot position of the object with respect to the reference plane (ground plane) in the 3D space may be calculated using the foot position of the object in the 2D image. To detect the foot position in the 3D space, the foot position in the 2D image may be projected inversely onto the reference plane in the 3D space using the projection matrix."; Column 22 [lines 51-53] “The coordinates of the inversely-projected X are normalized by the Z-axis value to detect the foot position in the 3D space as in Equation 40.”);
estimating, based on the first foot pixel coordinate, a distance between the vehicle and the detected person (Column 16 [lines 9-13] “In order to extract physical information of an object from the three-dimensional world coordinates, the foot position on the ground plane
X
~
f
=
H
-
1
x
~
f
(i.e. first foot pixel coordinate) is required to be calculated using Equation 1.”; Column 22 [lines 63-64] “The depth of the object may be predicted by calculating the distance between the object and the camera.”; Column 23 [lines 4-7] “When the object is sufficiently far from the camera, the depth of the foot position of the object may be estimated as a Y-axis coordinate because a camera pan angle is zero and the center point is on the ground plane.”);
estimating, based on the distance and the head pixel coordinate, a height of the detected person (Column 16 [lines 20-25 “Where P represents a projection matrix, and Ho represents an object height. Using Equation 29, Ho may be calculated from y as follows:
PNG
media_image1.png
100
434
media_image1.png
Greyscale
”); and
[controlling, based on the estimated height, an operation of the vehicle].
However, Paik fails to teach a method of controlling operation of a vehicle, [an image] from a camera of the vehicle, semantic segmentation of the image, and controlling, based on the estimated height, an operation of the vehicle.
Lee teaches a method of controlling operation of a vehicle (Paragraph [0036] “the present invention may detect an opening angle, mapped to the estimated user height information, as a target opening angle from the lookup table with reference to the lookup table and may adjust the opening amount of the tailgate based on the detected target opening angle.”), [an image] from a camera of the vehicle (Paragraph [0044] "input a preprocessing image, obtained by preprocessing a rear camera image 11 input from a rear camera 10 by frame units, to the semantic segmentation image generating unit 120."), semantic segmentation of the image (Paragraph [0055]"The semantic segmentation image generating unit 120 may be implemented as a software module, a hardware module, or a combination thereof and may generate a semantic segmentation image corresponding to a preprocessing image input from the preprocessing unit 110. "), and controlling, based on the estimated height, an operation of the vehicle (Paragraph [0036] “the present invention may detect an opening angle, mapped to the estimated user height information, as a target opening angle from the lookup table with reference to the lookup table and may adjust the opening amount of the tailgate based on the detected target opening angle.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Paik’s reference to include a method of controlling operation of a vehicle, [an image] from a camera of the vehicle, semantic segmentation of the image, and controlling, based on the estimated height, an operation of the vehicle taught by Lee’s reference. The motivation for doing so would have been to automatically adjust the opening amount of a tailgate based on the estimated height of the user to increase the convenience of an electrical tailgate as suggested by Lee (see Lee, Paragraph [0032]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lee with Paik to obtain the invention specified in claim 1.
Regarding claim 2, which claim 1 is incorporated, Paik discloses determining, based on a value of the height being within a predetermined range, that the detected person is a child ("Column 8 [lines 59 - 63] "the average heights for a child, a juvenile, and an adult was set to 100 cm, 140 cm, and 180 cm, respectively for application to real human models. The ratio of head, body and leg is set to 2:4:4."; Column 9 [lines 36-38] "an ellipsoid-based human model with the minimum matching error ei is selected for three human models including a child, a juvenile, and an adult."), or
determining, based on the value being outside of the predetermined range, that the detected person is not a child.
Regarding claim 5 which claim 1 is incorporated, Paik discloses wherein: the extracting the head pixel coordinate comprise: setting a second temporary region (object region in Column 20 [lines 58-59] equates to a temporary region) to be a region corresponding to the head of the person detected in the acquired image (Column 20 [lines 58-59] “In step 115, the object occlusion detection device 100 detects a head position and a foot position using the extracted object region, individually.”; and
extracting a coordinate of an uppermost pixel (the vertex in Column 20 [lines 48-53] equates to an uppermost pixel), of one or more pixels in the second temporary region, as the head pixel coordinate (Column 20 [lines 48-53] " In step 115, the object occlusion detection device 100 detects a head position and a foot position using the extracted object region, individually... the object occlusion detection device 100 may determine the vertically highest position—the vertex—as the head position in the object region.").
Regarding claim 6, which claim 1 is incorporated, Paik discloses wherein: the generating of the transformed view image comprises: obtaining intrinsic and extrinsic parameters of the camera (Column 3 [lines 59-65] "Internal parameters and external parameters of the camera may be estimated using the detected vanishing line and the vanishing points, and the internal parameters may include a focal length, a principal point and an aspect ratio, and the external parameters include a panning angle, a tilting angle, a rolling angle, a camera height with respect to the z-axis, transformation in x-axis and y-axis directions.");
converting, via a homography matrix based on the intrinsic and extrinsic parameters, each point, of a plurality of points in the acquired image, into a point in the transformed view image (Column 8 [lines 12-18] " In order to normalize the object information, camera parameters are estimated using automatic scene calibration, and a projection matrix is estimated using the camera parameters obtained by scene calibration. After the normalized object information is obtained, the object in the two-dimensional image is projected onto the three-dimensional world coordinate using the projection matrix."; Column 8 [lines 35-37] "Where H=[p 1p2p3]T is a 3×3 homography matrix, pi (i=1, 2, 3) is the first three columns of a 3×4 projection matrix P calculated by estimating the camera parameters.");
generating, based on the converting, [a look-up table] (Column 2 [lines 60-65] “performing scene calibration based on the three-dimensional human model to normalize object information of the object included in the two-dimensional images, and generating normalized metadata of the object from the two-dimensional images on which the scene calibration is performed.”); and
generating, [based on the look-up table] and the acquired image, the transformed view image (Column 8 [lines 12-18] "In order to normalize the object information, camera parameters are estimated using automatic scene calibration, and a projection matrix is estimated using the camera parameters obtained by scene calibration. After the normalized object information is obtained, the object in the two-dimensional image is projected onto the three-dimensional world coordinate using the projection matrix.”; Column 22 [lines 32-39] "the coordinates of the 2D image are projected onto the reference plane in the 3D space using the projection matrix. The foot position of the object is placed on the ground plane, the camera height is calculated as the distance between the ground plane and the camera, and the ground plane is regarded as the XY plane, so the ground plane may be used as the reference plane.").
However, Paik fails to teach a look-up table.
Lee teaches a look-up table (Paragraph [0084] “Based on an arm length proportional to a height of a person, the lookup table 152 may store a plurality of opening angles which are learned so that a closing button of a tailgate is disposed at a statistical height”)
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Paik’s reference to include a look-up table taught by Lee’s reference. The motivation for doing so would have been to set the closing button at a target height with which a person can stably reach as suggested by Lee (see Lee, Paragraph [0084]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lee with Paik to obtain the invention specified in claim 6.
Regarding claim 7, which claim 6 is incorporated, Paik discloses acquiring, based on the homography matrix, the second foot pixel coordinate (foot position in the 3D space in Column 22 equates to second foot pixel coordinate) corresponding to the first foot pixel coordinate (foot position in the 2D space in Column 22 equates to first foot pixel coordinate) (Column 22 [lines 40-45] "The foot position of the object with respect to the reference plane (ground plane) in the 3D space may be calculated using the foot position of the object in the 2D image. To detect the foot position in the 3D space, the foot position in the 2D image may be projected inversely onto the reference plane in the 3D space using the projection matrix."; Column 22 [lines 51-53] “The coordinates of the inversely-projected X are normalized by the Z-axis value to detect the foot position in the 3D space as in Equation 40.”).
Regarding claim 11, Paik discloses [a device of a vehicle], wherein the device comprises one or more processors configured to executes program codes loaded on one or more memory devices (Column 24 [lines 51-56] “above-described technical features may be implemented in the form of program instructions that can be executed through various computer means and recorded in a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures”), wherein:
the program codes, when executed by the one or more processors, configure the device to (Column 25 [lines 1-5] “high-level language code that can be executed by a computer using an interpreter or the like. The hardware device may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.”):
acquire an image [from a camera of the vehicle] (Column 5 [lines 12-15] “generating a multi-ellipsoid based three-dimensional (3D) human model using perspective features of a plurality of two-dimensional (2D) images obtained by the multiple cameras”),
extract, [based on semantic segmentation of the image]:
a first foot pixel coordinate (foot position coordinates in Column 10 [lines 5-7] equates to a first foot pixel coordinate) corresponding to a foot of a person detected in the image (Column 10 [lines 5-7] "Step 1 (step S121) is to extract valid foot and head data to calculate homology from foot to head"; Column 10 [lines 62-65] "When the foot position coordinates on the ground plane are given, the corresponding head position in the image plane be may determined using Hhf. H=Hfh is defined as the homology from foot to head"; Column 20 [lines 53-58] "In addition, the object occlusion detection device 100 may determine a position of about 20% from the bottom as the foot position, rather than determining the lowest position, in determining a vertically lower position as the foot position in the object region."); and
a head pixel coordinate (head position in Column 20 [lines 48-53] equates to a head pixel coordinate) corresponding to a head of the person (Column 20 [lines 48-53] " In step 115, the object occlusion detection device 100 detects a head position and a foot position using the extracted object region, individually... the object occlusion detection device 100 may determine the vertically highest position—the vertex—as the head position in the object region.");
generate, based on transforming a coordinate system of the acquired image to a transformed view coordinate system, a transformed view image (Column 8 [lines 12-18] " In order to normalize the object information, camera parameters are estimated using automatic scene calibration, and a projection matrix is estimated using the camera parameters obtained by scene calibration. After the normalized object information is obtained, the object in the two-dimensional image is projected onto the three-dimensional world coordinate using the projection matrix."; Equation 39 and Column 22 [lines 49-53] “Where xf denotes the foot position on the 2D image, P denotes the projection matrix, and X denotes the coordinates of the inversely-projected xf. The coordinates of the inversely-projected X are normalized by the Z-axis value to detect the foot position in the 3D space”);
determine, based on the transformed view image, a second foot pixel coordinate (foot position in the 3D space in Column 22 equates to second foot pixel coordinate), in the transformed view coordinate system, corresponding to the first foot pixel coordinate (foot position in the 2D image in Column 22 equates to first foot pixel coordinate) (Column 22 [lines 40-45] "The foot position of the object with respect to the reference plane (ground plane) in the 3D space may be calculated using the foot position of the object in the 2D image. To detect the foot position in the 3D space, the foot position in the 2D image may be projected inversely onto the reference plane in the 3D space using the projection matrix."; Column 22 [lines 51-53] “The coordinates of the inversely-projected X are normalized by the Z-axis value to detect the foot position in the 3D space as in Equation 40.”);
estimate, based on the first foot pixel coordinate, a distance between the vehicle and the detected person (Column 16 [lines 9-13] “In order to extract physical information of an object from the three-dimensional world coordinates, the foot position on the ground plane
X
~
f
=
H
-
1
x
~
f
(i.e. first foot pixel coordinate) is required to be calculated using Equation 1.”; Column 22 [lines 63-64] “The depth of the object may be predicted by calculating the distance between the object and the camera.”; Column 23 [lines 4-7] “When the object is sufficiently far from the camera, the depth of the foot position of the object may be estimated as a Y-axis coordinate because a camera pan angle is zero and the center point is on the ground plane.”);
estimate, based on the distance and the head pixel coordinate, a height of the detected person (Column 16 [lines 20-25 “Where P represents a projection matrix, and Ho represents an object height. Using Equation 29, Ho may be calculated from y as follows:
PNG
media_image1.png
100
434
media_image1.png
Greyscale
”); and
[control, based on the estimated height, an operation of the vehicle].
However, Paik fails to teach a device of a vehicle, [an image] from a camera of the vehicle, semantic segmentation of the image, and control, based on the estimated height, an operation of the vehicle.
Lee teaches a device of a vehicle, [an image] from a camera of the vehicle (Paragraph [0044] "input a preprocessing image, obtained by preprocessing a rear camera image 11 input from a rear camera 10 by frame units, to the semantic segmentation image generating unit 120."), semantic segmentation of the image (Paragraph [0055] "The semantic segmentation image generating unit 120 may be implemented as a software module, a hardware module, or a combination thereof and may generate a semantic segmentation image corresponding to a preprocessing image input from the preprocessing unit 110. "), and control, based on the estimated height, an operation of the vehicle (Paragraph [0036] “the present invention may detect an opening angle, mapped to the estimated user height information, as a target opening angle from the lookup table with reference to the lookup table and may adjust the opening amount of the tailgate based on the detected target opening angle.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Paik’s reference to include a device of a vehicle, [an image] from a camera of the vehicle, semantic segmentation of the image, and control, based on the estimated height, an operation of the vehicle taught by Lee’s reference. The motivation for doing so would have been to automatically adjust the opening amount of a tailgate based on the estimated height of the user to increase the convenience of an electrical tailgate as suggested by Lee (see Lee, Paragraph [0032]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Lee with Paik to obtain the invention specified in claim 11.
Regarding claim 12 (drawn to a device), claim 12 is rejected the same as claim 2 and the arguments similar to that presented above for claim 2 are equally applicable to the claim 12, and all the other limitations similar to claim 2 are not repeated herein, but incorporated by reference.
Regarding claim 15 (drawn to a device), claim 15 is rejected the same as claim 5 and the arguments similar to that presented above for claim 5 are equally applicable to the claim 15, and all the other limitations similar to claim 5 are not repeated herein, but incorporated by reference.
Regarding claim 16 (drawn to a device), claim 16 is rejected the same as claim 6 and the arguments similar to that presented above for claim 6 are equally applicable to the claim 16, and all the other limitations similar to claim 6 are not repeated herein, but incorporated by reference.
Regarding claim 17 (drawn to a device), claim 17 is rejected the same as claim 7 and the arguments similar to that presented above for claim 7 are equally applicable to the claim 17, and all the other limitations similar to claim 7 are not repeated herein, but incorporated by reference.
Claims 3-4, and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Paik et al. (US 10885667 B2) (hereinafter, “Paik”) in view of Lee (US 2022/0275677 A1), and further in view of Fujimoto (US 2007/0127778 A1).
Regarding claim 3, which claim 2 is incorporated, Paik discloses [displaying, on the transformed view image]:
the second foot pixel coordinate (Column 22 [lines 40-45] "The foot position of the object with respect to the reference plane (ground plane) in the 3D space may be calculated using the foot position of the object in the 2D image. To detect the foot position in the 3D space, the foot position in the 2D image may be projected inversely onto the reference plane in the 3D space using the projection matrix."); and
a result of determining whether the detected person is a child (Column 18 [lines 5-12] "FIGS. 20A-20C show results of an object search experiment using size queries including a child (small, FIG. 20A), a juvenile (medium, FIG. 20B), and an adult (large, FIG. 20C). FIG. 20A shows that the normalized metadata generation method according to an embodiment of the present invention successfully searches for a child less than 110 cm, and FIGS. 20B and 20C show results similar to those of with respect to the juvenile and the adult.").
However, Paik and Lee both fail to teach displaying, on the transformed view image.
Fujimoto teaches displaying, on the transformed view image (Paragraph [0069] “the object-attribute determining unit 107 determines in which region(s) of the specified ZX-plane that the coordinate-transformed top points RT1 to RT11 and the coordinate-transformed bottom points RB1 to RB11 belong, determines if an object is two-dimensional or three-dimensional, and determines positional information with regard to two-dimensional objects disposed on the road surface.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Paik in view of Lee to include displaying, on the transformed view image taught by Fujimoto’s reference. The motivation for doing so would have been to determine the position information of an object on a road surface as suggested by Fujimoto (see Fujimoto, Paragraph [0069]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Fujimoto with Paik and Lee to obtain the invention specified in claim 3.
Regarding claim 4, which claim 1 is incorporated, Paik discloses the extracting the first foot pixel coordinate comprises (Column 10 [lines 5-7] "Step 1 (step S121) is to extract valid foot and head data to calculate homology from foot to head”; Column 10 [lines 62-65] "When the foot position coordinates on the ground plane are given, the corresponding head position in the image plane be may determined using Hhf. H=Hfh is defined as the homology from foot to head"):
setting a first temporary region (object region in Column 20 [lines 58-59] equates to a temporary region) to be a region corresponding to the foot of the person detected in the acquired image (Column 20 [lines 40-42] "In step 110, an object occlusion detection device 100 extracts an object region using a background model after inputting a current frame to the background model.”); and
extracting [a coordinate of a lowermost pixel], of one or more pixels in the first temporary region, as the first foot pixel coordinate (Column 20 [lines 48-58] " In step 115, the object occlusion detection device 100 detects a head position and a foot position using the extracted object region, individually... In addition, the object occlusion detection device 100 may determine a position of about 20% from the bottom as the foot position, rather than determining the lowest position, in determining a vertically lower position as the foot position in the object region.”).
However, Paik and Lee both fail to teach a coordinate of a lowermost pixel
Fujimoto teaches a coordinate of a lowermost pixel (Paragraph [0109] “The object detecting system further includes the grouping unit 106 for grouping the pixels when pixels that are adjacent in the vertical direction of the image also have velocity differences that are within a predetermined range. The object-attribute determining unit 107 performs…on the basis of the three-dimensional position coordinates of the uppermost pixel and the lowermost pixel of the pixels grouped by the grouping unit 106.”)
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Paik in view of Lee to include a coordinate of a lowermost pixel taught by Fujimoto’s reference. The motivation for doing so would have been to determine the position information of an object on a road surface as suggested by Fujimoto (see Fujimoto, Paragraph [0069]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Fujimoto with Paik and Lee to obtain the invention specified in claim 4.
Regarding claim 13 (drawn to a device), claim 13 is rejected the same as claim 3 and the arguments similar to that presented above for claim 3 are equally applicable to the claim 13, and all the other limitations similar to claim 3 are not repeated herein, but incorporated by reference.
Regarding claim 14 (drawn to a device), claim 14 is rejected the same as claim 4 and the arguments similar to that presented above for claim 4 are equally applicable to the claim 14, and all the other limitations similar to claim 4 are not repeated herein, but incorporated by reference.
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Paik et al. (US 10885667 B2) (hereinafter, “Paik”) in view of Lee (US 2022/0275677 A1), and further in view of Kai Ruhl (Coordinates: Bundler/VisualSFM, Matlab Calibration Toolbox, and OpenGL, 2014-05 (https://www.land-of-kain.de/docs/coords/)) (hereinafter, “Ruhl”).
Regarding claim 10, which claim 1 is incorporated, Paik discloses wherein the estimating the height comprises: estimating, as the height (Column 16 [lines 20-25] “Where P represents a projection matrix, and Ho represents an object height. Using Equation 29, Ho may be calculated from y as follows:
PNG
media_image1.png
100
434
media_image1.png
Greyscale
”), [a value of Z of coordinates (X, Y, Z) obtained by converting] the head pixel coordinate (Column 20 [lines 48-53] " In step 115, the object occlusion detection device 100 detects a head position and a foot position using the extracted object region, individually... the object occlusion detection device 100 may determine the vertically highest position—the vertex—as the head position in the object region.") [(x, y) into a world coordinate system by using Equation 6:
Equation 6
X
Y
Z
1
=
s
[
R
|
t
]
-
1
K
-
1
x
y
1
where, K denotes an intrinsic parameter matrix of the camera, [R|t] denotes an extrinsic parameter matrix of the camera, and s denotes] the distance (Column 16 [lines 9-13] “In order to extract physical information of an object from the three-dimensional world coordinates, the foot position on the ground plane
X
~
f
=
H
-
1
x
~
f
(i.e. first foot pixel coordinate) is required to be calculated using Equation 1.”; Column 22 [lines 63-64] “The depth of the object may be predicted by calculating the distance between the object and the camera.”; Column 23 [lines 4-7] “When the object is sufficiently far from the camera, the depth of the foot position of the object may be estimated as a Y-axis coordinate because a camera pan angle is zero and the center point is on the ground plane.”).
However, Paik and Lee both fail to teach [estimating, as the height], a value of Z of coordinates (X, Y, Z) obtained by converting [the head pixel coordinate] (x, y) into a world coordinate system by using Equation 6:
Equation 6
X
Y
Z
1
=
s
[
R
|
t
]
-
1
K
-
1
x
y
1
where, K denotes an intrinsic parameter matrix of the camera, [R|t] denotes an extrinsic parameter matrix of the camera, and s [denotes the distance].
Ruhl teaches [estimating, as the height], a value of Z of coordinates (X, Y, Z) obtained by converting [the head pixel coordinate] (x, y) into a world coordinate system by using Equation 6 (Page 1 Section 01 Camera Projection “if you have a world space (wld) coordinate [X Y Z] (arbitrary scale, arbitrary orientation), you turn it into a homogeneous coordinate by making it into [X Y Z 1] (if it was not 1, all other components would count as if you multiplied it by that number).”; Figure 1)
PNG
media_image2.png
152
536
media_image2.png
Greyscale
.:
Equation 6
X
Y
Z
1
=
s
[
R
|
t
]
-
1
K
-
1
x
y
1
where, K denotes an intrinsic parameter matrix of the camera, [R|t] denotes an extrinsic parameter matrix of the camera (Page 1 Section 01 Camera Projection “camera projection requires a perspective division by z, a 3x3 intrinsic matrix K (sometimes C), and a 3x4 extrinsic matrix Rt put together from a 3x3 rotation matrix R and a translation 3-vector t (not multiplied! Rt is not R·t).”; Page 5 Conclusion “Projection (wld2cam) has three steps: Rt, perspective division, and K. Back-projection (cam2wld) has 3 steps in reverse: K-1, perspective multiplication, and Rt-1. Bundler converts RHS (wld) to LHS (cam)”), and s [denotes the distance] (End of Page 4 Section 04 OpenGL “the world coordinate system (wld) is right-handed: it originated from CAD, where on a sheet of paper, x is right, y up, and z upwards out of the paper. Therefore, in the OpenGL pipeline a z-swap is performed”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Paik in view of Lee to include [estimating, as the height], a value of Z of coordinates (X, Y, Z) obtained by converting [the head pixel coordinate] (x, y) into a world coordinate system by using Equation 6:
Equation 6
X
Y
Z
1
=
s
[
R
|
t
]
-
1
K
-
1
x
y
1
where, K denotes an intrinsic parameter matrix of the camera, [R|t] denotes an extrinsic parameter matrix of the camera, and s [denotes the distance] taught by Ruhl’s reference. The motivation for doing so would have been because camera projection requires a perspective division by z (the depth) as suggested by Ruhl (see Ruhl, Page 1 Section 01 Camera Projection).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more that predictable results. Therefore, it would have been obvious to combine Ruhl with Paik and Lee to obtain the invention specified in claim 10.
Regarding claim 20 (drawn to a device), claim 20 is rejected the same as claim 10 and the arguments similar to that presented above for claim 10 are equally applicable to the claim 20, and all the other limitations similar to claim 10 are not repeated herein, but incorporated by reference.
Allowable Subject Matter
Claims 8-9 and 18-19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 8-9 and 18-19 contain subject matter that is not disclosed or made obvious in the cited art.
In regards to claim 8, when considering claim 8 as a whole, prior art fails to disclose or render obvious, alone or in combination:
“[…] setting the first foot pixel coordinate and a coordinate of the camera in a normalized coordinate system;
determining corresponding coordinates in a world coordinate system, wherein the corresponding coordinates correspond to the first foot pixel coordinate and the coordinate of the camera in the normalized coordinate system; and
estimating, based on the corresponding coordinates in the world coordinate, a distance between the vehicle and the detected person.”.
In regards to claim 9, when considering claim 9 as a whole, prior art fails to disclose or render obvious, alone or in combination:
“[…]
Equation 1
d
=
(
C
P
'
)
2
+
(
P
P
'
)
2
where d denotes the distance, and C'P' is calculated by using Equation 2:
Equation 2
C
'
P
'
=
C
P
'
*
tan
(
π
2
+
θ
t
i
l
t
-
tan
-
1
v
)
where, CC' denotes a height of the camera, θtilt denotes a tilt angle of the camera, v denotes a y coordinate of the first foot pixel coordinate, and PP' is calculated by using Equation 3:
Equation 3
P
P
'
=
u
*
C
P
'
C
p
'
where, u denotes an x coordinate of the first foot pixel coordinate, CP' is calculated by using Equation 4, and Cp' is calculated by using Equation 5:
Equation 4
C
P
'
=
(
C
C
'
)
2
+
(
C
'
P
'
)
2
Equation 5
C
p
'
=
1
+
v
2
where, CC' denotes the height of the camera, v denotes the y coordinate of the first foot pixel coordinate, and C'P' is calculated by using Equation 2.”
In regard to claims 18-19: Claims 18-19 mirror claims 8-9 in a device and similarly considering claims 18-19 as a whole, prior art of record fails to disclose or render obvious the subject matter in the claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Zhang et al. (US 2018/0116556 A1) discloses a height measurement method based on monocular machine vision taken from a camera attached to a robot. The height is measured by picking up two-dimensional identifier from the head to the feet of a person under measurement.
Perrault (US 2023/0136084 A1) discloses a method for calibrating parameters of an imaging device by selecting a plurality of pairs of points wherein each pair of points comprise a head point and a foot point associated with the person.
Sasatani et al. (US 2014/0348382 A1) discloses a people counting device that detect a person’s head and estimates foot coordinates from an image to determine whether the foot coordinates fall within a detection region.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to UROOJ FATIMA whose telephone number is (571)272-2096. The examiner can normally be reached M-F 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/UROOJ FATIMA/Examiner, Art Unit 2676
/Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676