DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/21/2026 has been entered.
Claims 1, 3, 5 and 6 are currently pending in U.S. Patent Application No. 18/552,722 and an Office action on the merits follows.
Response to Arguments/Remarks
Applicant’s 05/21/2026 remarks accompanying the Request for Continued Examination have been fully considered but they are not persuasive. Independent claim 1 as amended incorporates the subject matter of now cancelled claims 2 and 7, previously rejected as being unpatentable over Urushido et al. (US 2022/031400 A1) in view of Ha (US 2006/0177124 A1). Applicant’s remarks reference the proposed combination and motivation statement previously presented at page 12 of the Final Office Action and assert that Ha is not in the same context/ field-of-endeavor as the claimed invention.
PNG
media_image1.png
128
1012
media_image1.png
Greyscale
Urushido is relied upon for a baseline selection, and the Examiner pre-emptively identified how Ha is analogous art under at least the reasonable-pertinence theory, for a same problem that is edge detection – even if it is argued that Ha concerns a different field of endeavor.
PNG
media_image2.png
446
1164
media_image2.png
Greyscale
Ha’s field of endeavor is not disqualifying, and Ha evidences the manner in which relying on directional frequency information is known, and arguably even routine, in the context of edge detection broadly – to include such an edge detection that may then be used in any number of subsequent determinations not limited to any context in Ha. POSITA would recognize that stereo-matching for distance/depth determinations requires a point/feature match/ correspondence between images (otherwise depth cannot be accurately resolved) – and that a dominance of a horizontal direction for spatial frequency components is indicative of vertically oriented lines/edges – best handled by a horizontal baseline (inverse situation similarly, a scene (or object(s) of interest) with clear/discernable horizontally oriented edges (dominant spatial frequency in vertical direction) – e.g. powerlines (Urushido Fig. 4) – is best handled by a vertical baseline). Urushido is relied upon for these teachings, and Ha simply evidences the manner in which directional frequency information can be readily used to ascertain the directions/ orientations of edges in an image. Applicant’s remarks appear to concede that Ha teaches measuring equivalent directional frequencies, but that this is mooted by the fact that Ha doesn’t use that information for a baseline selection (further asserting that even a selection between a ‘top-down’ vs. ‘side-by-side’ format is not a baseline selection equivalent). This is not dispositive because Urushido is relied upon for a selection on the basis of scene/object edges, and Ha need only evidence the obvious nature of relying on directional frequency information for edge detection, and no more (see MPEP 2145 arguing non-analogous art with reference to MPEP § 2141.01(a)).
PNG
media_image3.png
316
1044
media_image3.png
Greyscale
In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Applicant’s remarks at pages 6-7 further assert that Urushido’s ‘framework’ for evaluating stereo images differs from one implied/required for the case of the claimed invention – because Urushido at [0047-0048] discloses the manner in which the selected baseline is ideally perpendicular to an extending direction of an object of interest (so as to most accurately resolve that objects depth/distance). Examiner respectfully disagrees with any assertion that this disclosure in Urushido is ‘unrelated to’ that image selection as recited (i.e. selecting which pair of images to rely upon for resolving object depth among other processing), because this selection is for exactly that purpose – to determine which pair of images is best considered. Nor does the Examiner understand Applicant to assert that an implementation of Applicant’s method as claimed would/necessarily result(s) in the opposite outcome – i.e. selecting a baseline (and a corresponding image pair for subsequent depth/distance determinations therefrom) so as to ensure that the selected baseline is substantially parallel to salient/distinguishing/dominant edges of an object – since such a selection, as is known to POSITA and further evidenced by Urushido, would likely impair a system’s ability to most accurately resolve object depth/distance (absent some other obvious consideration e.g. one of the three cameras is completely occluded/malfunctioning, etc.).
No recited claim elements appear distinguished from/outside the teachings of the prior art of record, even when considered in combination – it is known/recognized that select baselines better resolve object depth given known/recognized object dimensions/orientation, and it is additionally known/ recognized that spatial frequency information can be readily used to detect edges (and object orientation accordingly). Applicant’s remarks point to Applicant’s Specification at e.g. [0023] in asserting that the limitations in question “allow for exemplary non-limiting embodiments which have technical effects not realized in the cited art” (naming none specifically), however Examiner would assert that [0023] suggests no more than what is already known to POSITA, as evidenced at least by Urushido and references of record more broadly, that a baseline selection may facilitate depth/distance resolution because it may ensure a greater degree of point/feature correspondences/matches between the resultant/associated image pair.
Additional search and consideration identifies relevant literature evidencing the same.
Please see e.g. Kallwies et al. “Effective Combination of Vertical and Horizontal Stereo Vision” (2018), reproduced in part below, among that/those additionally cited literature:
PNG
media_image4.png
368
676
media_image4.png
Greyscale
PNG
media_image5.png
406
656
media_image5.png
Greyscale
PNG
media_image6.png
504
654
media_image6.png
Greyscale
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
1. Claims 1, 3 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Urushido et al. (US 2022/031400 A1) in view of Ha (US 2006/0177124 A1) and Kallwies et al. “Effective Combination of Vertical and Horizontal Stereo Vision” (2018).
As to claim 1, Urushido discloses an object distance detecting device which detects a distance to a target object (Abs “an object detection unit that detects the object based on the depth estimated by the depth estimation unit and reliability of the depth”) around a vehicle ([0052] “Note that, in the present embodiment, a description will be given of an example where the information processing apparatus 1 is mounted on the unmanned moving body. However, besides this, the information processing apparatus 1 may be mounted on an autonomous mobile robot, a vehicle, a portable terminal, or the like”), the object distance detecting device comprising
at least three imaging cameras which image a same target object (Figures 6-8, apparatus 1 comprises 10 further comprising three imaging sections 10a, 10b, and 10c, Fig. 18, [0052] “As illustrated in FIG. 6, the information processing apparatus 1 includes three imaging units (cameras) 10a, 10b, and 10c, a control unit 20, and a storage unit 30”, [0054] “As illustrated in FIG. 7 for example, the stereo camera system 10 is a trinocular camera system including three imaging units 10a, 10b, and 10c. The stereo camera system 10 is attached to, for example, a lower portion of the unmanned moving body with a support member 12 interposed therebetween. The imaging units 10a, 10b, and 10c are arranged in a V shape. That is, the first and second stereo cameras 11a and 11b are arranged so that a direction of a baseline length of the first stereo camera 11a and a direction of a baseline length of the second stereo camera 11b are perpendicular to each other”, etc.,), wherein the object detecting device is configured to:
specify a region, in which the target object exists, on a basis of an image acquired from at least one of the imaging cameras (detection section 22, Fig. 12, processing steps Pr2-Pr3, [0059] “The edge detection unit 22 detects the edge of the object from a monocular image (RGB image) captured by any of the imaging units 10a, 10b and 10c, and generates an edge image (see FIG. 10 to be described later)”; Examiner identifies edge regions corresponding to one or more objects as detected by 22 to be those most equivalent to the specified regions, since an edge is detected prior to and for that processing of Pr4 and Pr5 – see also the Response to Remarks above);
select one base line direction among a plurality of base line directions defined by any two imaging sections among the at least three imaging cameras on a basis of object edge orientation, and select an image of the region acquired from each of the two imaging cameras defining the selected base line direction ([0080] “For example, a case is considered where edges in the horizontal direction and edges in the vertical direction are detected by the edge detection unit 22, for example, as illustrated in FIG. 12. In this case, a second probability distribution corresponding to the first stereo camera 11a in which the direction of the baseline length is the horizontal direction as illustrated in FIG. 13 becomes such a distribution in which the highest probability overlaps the edges in the vertical direction perpendicular to the baseline length as illustrated in FIG. 14. Meanwhile, a second probability distribution corresponding to the second stereo camera 11b in which the direction of the baseline length is the vertical direction as illustrated in FIG. 15 becomes such a distribution in which the highest probability overlaps the edges in the horizontal direction perpendicular to the baseline length as illustrated in FIG. 16”, [0081-0082]; Examiner notes the ‘selected’ images are those corresponding to the baseline, i.e. that of either 11a, or 11b, associated with the highest reliability (based on the object(s), edge lines, and corresponding edge line angle of directions); [0008] “the reliability being determined in accordance with an angle of a direction of an edge line of the object with respect to the directions of the baseline lengths of the plurality of stereo cameras”, [0006] “In the object detection using the stereo camera, for example, on the basis of a parallax of an object seen from right and left cameras, a distance between the camera and the object is measured. However, when the object as a measuring target extends in a direction of a baseline length of the stereo camera, there is a problem that it is difficult to measure the distance”);
detect a distance to the target object existing in the region on a basis of the image selected ([0047] “In object detection using such a stereo camera system 110 as described above, by using a method such as triangulation for example, a distance to an object (hereinafter, the distance will be referred to as a "depth") is estimated on the basis of a parallax of the object seen from the left and right imaging units 110a and 110b”, [0048-0049], etc.,); and
generate a parallax image of the region from the image selected (Urushido determining depth reliability/probability distributions (parallax image equivalents) for both 11a (horizontal baseline, Fig. 13, cameras 10a and 10b), and 11b (vertical baseline, Fig. 15 (cameras 10b and 10c)); see e.g. Fig. 16 and voting for ‘second stereo camera’ (11b) of Fig. 15; [0079-0080]), and
detect the distance to the target object on a basis of the parallax image ([0047] “In object detection using such a stereo camera system 110 as described above, by using a method such as triangulation for example, a distance to an object (hereinafter, the distance will be referred to as a "depth") is estimated on the basis of a parallax of the object seen from the left and right imaging units 110a and 110b”, [0048-0049], [0058] “From the captured images captured by the first and second stereo cameras 11a and 11b, the depth estimation unit 21 estimates a depth of the object included in the captured images. On the basis of a parallax of the object seen from the imaging units 10a and 10b and a parallax of the object seen from the imaging units 10b and 10c, the depth estimation unit 21 estimates the depth by using, for example, a known method such as triangulation”, etc.,).
Urushido further discloses the device as configured to: obtain a spatial frequency component (Urushido’s the “extending direction of the object”) in a vertical direction (Fig. 2, wherein extending direction of the object exists in the vertical direction, e.g. for an object such as Fig. 5, wherein the tall building is characterized by a primarily vertical direction of spatial components/edges/features) and a spatial frequency component (“extending direction of the object”) in a horizontal direction of the image of the region (Fig. 3, e.g. that instance of Fig. 4 wherein the electrical wires have an extending direction/ more/a higher frequency/count of components in the horizontal direction; these interpretations are not inconsistent with Applicant’s disclosure e.g. pgpub [0038-0039] – it should be noted that the “extending direction of the object” is e.g. that axis in which the object has longest/most edges – which corresponds to a greater number of high frequency components in the direction that is perpendicular/orthogonal to such an edge/line – to illustrate Urushido Fig. 16 is understood to illustrate an extending direction that is primarily horizontal (and has a greater number of high frequency components in the vertical direction accordingly) and results in selecting for a vertical baseline (11b Fig. 15) accordingly), and
select the base line direction on a basis of the obtained spatial frequency component in the vertical direction and the obtained spatial frequency component in the horizontal direction (Urushido selects for a baseline that is most orthogonal relative to the extending direction of a target object, since such a baseline enables more accurate disparity (and depth accordingly given their relationship) determination(s)) – having a higher “reliability of the depth”, [0080], [0006], [0047] “In this depth estimation, when a direction of a baseline length that indicates a distance between the center of the imaging unit 110a and the center of the imaging unit 110b and an extending direction of the object as a measuring target are not parallel to each other but intersect each other as illustrated in FIG. 2 for example, the depth can be estimated appropriately since it is easy to grasp a correlation between such an object reflected in the video of the right camera and such an object reflected in the video of the left camera”, [0048], etc.).
Ha further evidences the obvious nature of determining line/edge directions on the basis of measuring directivity (horizontal vs vertical) of high frequency components (Fig. 6 620, 630, Fig. 9 S920, [0011] “The measuring the directivity of the high frequency components comprises measuring high frequency components of the selected image to determine if the selected image has more high frequency components in the horizontal direction or the vertical direction”, Figures 5A, high frequency in horizontal direction indicative of vertically oriented lines/edges, Fig. 5B high frequency in vertical direction corresponding to horizontally oriented lines; While Ha utilizes measuring directivity of frequency components to determine which components are more prevalent and thereby minimize loss of image quality for any subsequent compression/decompression, Ha evidences the manner in which POSITA would look to such a measuring, with a reasonable expectation of success, when attempting to solve that same problem of determining the/a dominant orientation for edges of one or more objects, and Ha is analogous art accordingly (see MPEP 2141.01(a) and at least reasonable-pertinence theory)).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date, to modify the system and method of Urushido to further comprise measuring a directivity of high frequency components when determining that ‘extending direction of the object’ as taught/suggested by Ha, the motivation as similarly taught/ suggested therein such a means for determining the extending direction of the object would constitute no more than readily implemented filtering operations (see MPEP 2143 Rationale (C) – Ha evidencing how measuring directivity of frequency components serves as a known edge detection technique), robust to noise and efficiently performed in a frequency domain, etc., for the purposes of edge detection and subsequent processing, while characterized by a reasonable expectation of success.
Urushida as modified by Ha, wherein the edge based baseline selection of Urushida further comprises an identification of edges on the basis of spatial frequency components as taught/suggested by Ha, further teaches/suggests select[ing] the vertical direction as the base line direction when a number of the spatial frequency components in the vertical direction obtained by the target object region specifying section is large, and selects the horizontal direction as the base line direction when a number of the spatial frequency components in the horizontal direction is large (Urushido as modified selects the same by consequence of that selection identified above for the case of those limitations previously recited for the case of now cancelled/incorporated claim 2 – as previously identified, sharp edges are characterized by high frequency components (higher the frequency, sharper the edge/transition) and high ‘horizontal frequency’ corresponds to sharp vertical edges (a rapid transition when traversing horizontally →, signals a strong vertical edge |), e.g. Fig. 16 has longest/more edges/lines that are horizontal, and so a higher/‘large’ number of ‘vertical’ high frequency components, and the vertical baseline is selected; while large is arguably Relative Terminology as identified in MPEP 2173.05(b), the instant claims are not understood to be indefinite because POSITA would understand the abovementioned relationships, Clearone, Inc. v. Shure Acquisition Holdings, Inc., 35 F.4th 1345, 1349, 2022 USPQ2d 509 (Fed. Cir. 2022) (similar to the manner in which how 202 operates specifically need not be disclosed in the Specification) see also Ha, Figures 5A and 5B and that modification/ motivation as proposed above and discussed in the associated remarks).
Under any assumption that given Ha’s disclosed/preferred context, Ha’s use of directional frequency information as taught/suggested therein cannot be applied to the context of a baseline/camera off-set selection, Kallwies further evidences the obvious nature of the same, i.e. a baseline selection ensuring that the selected baseline/offset is in the same direction as the significant gradients for structures/objects present in a scene/ environment/ resultant imagery (Abs “Actually, the depth of structures with significant gradients in just one direction can only be measured using cameras with an offset in the same direction”; see also remarks above, Fig. 8, etc.,).
It would have been obvious to a person of ordinary skill in the art, before the effective filing date, to further modify the system and method of Urushido in view of Ha, so as to select for baseline/offset additionally and/or alternatively based at least in part on associated directional frequency information, as a known/recognized characteristic of associated edges within an image/scene, as taught/suggested by Kallwies, the motivation as similarly taught/suggested therein that for particular instances (e.g. poles, powerlines, etc.,), such a selection may serve as the most reliable means for ensuring success in the context of subsequent depth/disparity (known inverse relationship) processing.
As to claim 3, Urushido as modified by Ha and Kallwies teaches/suggests the device of claim 1.
Urushido further teaches/suggests the device configured to:
generate a plurality of parallax images from the image acquired from each of the at least three imaging cameras, select the parallax image generated from the image acquired from each of the two imaging cameras defining the selected base line direction (Urushido determining depth reliability/probability distributions (parallax image equivalents) for both 11a (horizontal baseline, Fig. 13, cameras 10a and 10b), and 11b (vertical baseline, Fig. 15 (cameras 10b and 10c)); see e.g. Fig. 16 and voting for ‘second stereo camera’ (11b) of Fig. 15; [0079-0080]), and
detect the distance to the target object on a basis of the parallax image selected by the image selection section ([0047-0049], [0058] “From the captured images captured by the first and second stereo cameras 11a and 11b, the depth estimation unit 21 estimates a depth of the object included in the captured images. On the basis of a parallax of the object seen from the imaging units 10a and 10b and a parallax of the object seen from the imaging units 10b and 10c, the depth estimation unit 21 estimates the depth by using, for example, a known method such as triangulation”, etc.,).
As to claim 5, Urushido as modified by Ha and Kallwies teaches/suggests the device of claim 1.
Urushido further teaches/suggests the device configured to:
acquire vehicle information of at least one of motion state information or position information of the vehicle (Fig. 18 control 20A receives information from e.g. IMU 40, [0093] “Moreover, a control unit 20A of the information processing apparatus 1A includes a position/attitude estimation unit 24 in addition to the respective constituents of the above-described control unit 20”, Fig. 22 S11, [0094] “The inertial measurement unit 40 is composed of an inertial measurement unit (IMU) including, for example, a three-axis acceleration sensor, a three-axis gyro sensor, and the like, and outputs acquired sensor information to the position/attitude estimation unit 24 of the control unit 20A. The position/attitude estimation unit 24 detects a position and attitude (for example, an orientation, an inclination, and the like) of an unmanned moving body, on which the information processing apparatus 1A is mounted, on the basis of the captured images captured by the imaging units 10a, 10b, and 10c and the sensor information input from the inertial measurement unit 40. Note that a method for detecting the position and attitude of the unmanned moving body is not limited to a method using the above-described IMU”, [0100] “First, the position/attitude estimation unit 24 of the control unit 20A estimates the position and attitude of the subject machine (Step S11)”) and, obtain a spatial frequency component in a vertical direction and a spatial frequency component in a horizontal direction of image data of the region by weighting based on the vehicle information (see claims above, Urushido Fig. 21, [0095-0096], [0097] “Referring to FIG. 21, a description will be given below of an example of the deformation of the second probability distribution, in which the changes of the position and attitude of the subject machine are considered. When the second probability distribution is approximated by a two-dimensional normal distribution of a periphery of the edge as illustrated in FIG. 21 for example, this normal distribution can be represented by values of an x average, a y average, an edge horizontal dispersion, an edge vertical dispersion, an inclination, a size of the entire distribution, and the like. These values are changed in accordance with an angle of the edge and the variations of the position and attitude of the subject machine”, [0101]), and
select the base line direction on a basis of the spatial frequency component in the vertical direction and the spatial frequency component in the horizontal direction obtained by weighting (Fig. 22 S12-13, [0101], etc.,).
2. Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Urushido et al. (US 2022/031400 A1) in view of Ha (US 2006/0177124 A1), Kallwies et al. “Effective Combination of Vertical and Horizontal Stereo Vision” (2018), and Fujiwara (US 2014/0247344 A1).
As to claim 6, Urushido as modified by Ha and Kallwies teaches/suggests the device of claim 1.
Urushido suggests the device as configured to in a case where there are a plurality of regions where a target object exists, select the base line direction for one or more subsets of the regions (Urushido at least suggests selecting a baseline on the basis of target object portions, that may be characterized by a different spatial frequency component/extending direction relative to other object portions – e.g. while the building as a whole may be characterized by a vertical extending direction (and a large number of high frequency components in a horizontal direction), if the vehicle is primarily concerned with just that ‘top portion of a building’ B of Fig. 5 ([0048]), the target object portion may be characterized by an extending direction that is horizontal (as vehicle is clearing the building and a corresponding FOV concerns primarily region B, but does not image/capture the lower half of the building as the vehicle/UAV approaches and is close to the building top portion)).
Urushido however fails to disclose selecting multiple and different baselines, for different portions of any same object – stated differently Urushido appears to only explicitly disclose one target object at a time, and if considering an object portion as the target object, does not then additionally explicitly determine a baseline best suited for object portions that are not the target object portion. Urushido however appears modifiable in this respect, in view of a motivation to consider a plurality of objects in a scene at a time and allowing for a subsequent prioritization of an object of interest, and teachings of Urushido as applied to target object portions are readily extended to instances involving a plurality of objects and respective baseline selections accordingly.
Fujiwara further evidences the obvious nature of dividing an image into a plurality of grid/block portions which are then individually analyzed, reading on a case where there are a plurality of regions where a target object exists, and performing analysis for each region (Fig. 3, regions A1-A9, [0066-0067], etc., calculating evaluation values related to a focus state (characterized by edge sharpness/degree of high frequency components)). POSITA would further recognize such a partitioning may minimize the impact of any outlying information if present in only a few/minority of the corresponding regions.
It would have been obvious to a person of ordinary skill in the art, before the effective filing date, to modify the system and method of Urushido as proposed for the case of claim 1, to further comprise partitioning images into distinct sub-regions and analyzing each individually as taught/suggested by Fujiwara and Ha, and/or in addition to determining an optimum baseline for one target object, repeat such a determination for a plurality of objects and/or object portions as suggested by Urushido, the motivation as similarly taught/suggested therein such an analysis of a plurality of regions would enable handling multiple target objects and/or target object portions of higher priority.
Inquiry
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IAN L LEMIEUX whose telephone number is (571)270-5796. The examiner can normally be reached Mon - Fri 9:00 - 6:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached on 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IAN L LEMIEUX/Primary Examiner, Art Unit 2669