DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . It is responsive to the submission dated 02/14/2025. Claims 1-16 are presented for examination. Claims 1, 11 and 16 are independent claims.
Information Disclosure Statement
2. The information disclosure statements (IDSs) submitted on 02/14/2025 are in compliance with the provisions of 37 CFR 1.97 and are being considered by the Examiner.
Claim Rejections - 35 USC § 102
3. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
4. Claims 1-3, 6, and 11-13 are rejected under 35 U.S.C. 102(a)(a1) as being anticipated by Mann et al. (US 20200098135).
Considering claim 1, Mann discloses a method of controlling an electronic device for determining a relative position of at least one object in an image (e.g., Mann discloses a method for determining a geographical location and orientation of a vehicle travelling through a road network comprises obtaining a sequence of images (500) reflecting the environment of the road network on which the vehicle is travelling, wherein each of the images has an associated camera location at which the image was recorded. See paras. 14-15. In addition, Mann discloses: using a ground mesh to generate an orthorectified image of the ground-level features of the road network, wherein the orthorectified road image may comprise a bird's eye mosaic including a top down view of the area within which the vehicle is travelling in which features extracted from the images are projected onto the ground mesh and blended together …. so as to project the image onto the ground from a number of different perspectives. See paras. 95-96), the method comprising:
obtaining at least one semantic parameter associated with the image and
segmenting the at least one object based on the at least one semantic parameter to generate a first segmented object and a second segmented object (e.g., Mann discloses: in order to automatically detect or extract objects appearing in the images, a step of semantic segmentation may be performed in order to classify the elements of the images according to one or more of a plurality of different “object classes”…. in embodiments, at least some of the images (e.g. at least the key frames) may be processed in order to allocate an object class or list of classes (and associated probabilities) for each pixel in the images. See para. 58. In addition, Mann discloses performing the semantic segmentation using a stereo point cloud to segment the images into ground level pixels, walls/housing, traffic sign poles, and so on. See para. 62. Mann further teaches: Image Segmentation is performed to classify the objects appearing in the images, e.g. so that the classified objects can be extracted, and used, by the other processing modules in the flow. Thus, a step of “vehicle environment” semantic segmentation may be performed that uses as input the obtained image data, and processes each of the images, on a pixel by pixel basis, to assign an object class vector for each pixel, the object class vector containing a score (or likelihood value) for each of a plurality of classes. See para. 208);
identifying a camera level of the electronic device (e.g., Mann discloses: using a ground mesh to generate an orthorectified image of the ground-level features of the road network (see para. 95), wherein the orthorectified road image may comprise a bird's eye mosaic including a top down view of the area within which the vehicle is travelling in which features extracted from the images are projected onto the ground mesh and blended together …. so as to project the image onto the ground from a number of different perspectives…. and capturing linear features along a reference-line across a surface for use in a map database, wherein a height map of the ground mesh can then be generated to allow the height to be sampled at any arbitrary point. See para. 96);
applying a ground mesh to the image based on the camera level (e.g., Mann discloses a “ground mesh” is generated from a sequence of images. For instance, a “ground mesh” image may be generated containing (only) the ground-level features in the area that the vehicle is travelling. The ground mesh may be generated using the object class(es) obtained from the vehicle environment semantic segmentation, e.g. by extracting or using any pixels that have been allocated ground-level object classes ….. The ground mesh may in turn be used to generate an orthorectified image of the ground-level features of the road network. See paras. 94-95. In addition, Man discloses capturing linear features along a reference-line across a surface for use in a map database, wherein a height map of the ground mesh can then be generated to allow the height to be sampled at any arbitrary point. See para. 96);
determining placements of the first segmented object and the second segmented object based on the at least one semantic parameter associated with the first segmented object, the second segmented object, and the ground mesh (e.g., Mann discloses the object class or classes determined from the semantic segmentation can then be used to create a landmark observation feature. See para. 63. The images are processed in order to detect (and extract) one or more landmark object features for inclusion into the local map representation. In general, a landmark object feature may comprise any feature that is indicative or characteristic of the environment of the road network and that may be suitably and desirably incorporated into the local map representation, e.g. to facilitate the matching and/or aligning of the local map representation with a reference map section. For instance, the image content will generally include whatever objects are within the field of view of the camera or cameras used to obtain the images and any of these objects may in principle be extracted from the images, as desired. However, typical landmark objects that may suitably be extracted and incorporated into the local map representation may include features such as buildings, traffic signs, traffic lights, billboards and so on. See para. 65. Additionally, Man discloses the local map representation comprises a ground mesh and orthorectified road image together with both of the lane observations and landmark observations as described above. For instance, the lane marking objects and/or landmarks may suitable be embedded into the ground mesh and/or orthorectified road image, appropriately, to build up the local map representation. See para. 117. The method may further comprise using the local map representation for extracting a secondary set of features and using secondary set of features for the matching and aligning. See paras. 127 and 169);
determining a relative position of the first segmented object with respect to the second segmented object based on the determined placement of the first segmented object and the second segmented object (e.g., Mann discloses: After the vehicle environment semantic segmentation is performed on the images, as described above, and one or more regions of interest are determined as potentially containing a landmark object on the basis of the first semantic segmentation, a second or further step of semantic segmentation (or object detection and classification) may be performed specifically on the determined regions of interest. That is, a further specific classification step may be performed on any regions of interest determined to contain a landmark in order to further refine the landmark classification. See para. 67. For each landmark objects that are detected in one or more of the images, a landmark observation feature may then be created for inclusion within the local map representation. That is, when the local map representation is generated, a feature representing the detected landmark may be included into the local map representation. A landmark observation feature typically comprises: a landmark location; a landmark orientation; and a landmark shape. Thus, in embodiments, the techniques may comprise processing the sequence of images to detect one or more landmark objects appearing in one or more of the images, and generating for each detected landmark object a landmark observation for inclusion into the local map representation. See para. 69. In addition, Mann discloses using the estimated camera poses for each detected image within a landmark to aggregate the information from the plurality of images together, e.g. so as to be able to determine the position and orientation of a landmark relative to the vehicle, so that the landmark can be incorporated into the local map representation. See para. 70); and
displaying the at least one object on a screen of the electronic device based on the determined relative position (e.g., Mann discloses: a bird's eye mosaic georeferenced image of the road may be generated containing a 2D top view of the trip in which the images are projected onto the ground mesh and blended/weighted together, such that the pixel value of each pixel in the image represents the colour of the location in environment detected from the images used to generate the image. See para. 229).
Aa such, it is submitted that the Mann reference meets all the conditions for controlling an electronic device for determining a relative position of at least one object in an image, as recited in claim 1.
As per claim 2, Mann discloses determining at least one optimal location for at least one virtual object in the image based on the determined relative position of the first segmented object with respect to the second segmented object; and displaying the at least one object with the at least one virtual object on the screen of the electronic device based on the determined at least one optimal location (e.g., Using the known absolute positions and orientations of each key frame and its associated virtual camera (provided by the visual odometry estimation), all 2D image information, i.e. pixel colour, segmentation information along with the high-level feature positions, is projected onto the 3D ground geometry. A 2D orthophoto of this patch is then generated by projecting the data again on a virtual orthographic camera which looks perpendicularly down at the ground, thus yielding a bird's eye view of the scenery. See para. 227, wherein the 2D orthophoto of the patch generated by projecting the data again on a virtual orthographic camera corresponds to the determined optimal location for a virtual object in the image).
As per claim 3, Mann discloses the at least one semantic parameter comprises at least one of: the at least one object within the image, an edge of the at least one object, a ground corner point of the at least one object, a boundary of the at least one object, or a ground intersection edge of the at least one object. See paras. 67-70, 127-130 and 169.
As per claim 6, Mann discloses the applying the ground mesh based on the camera level comprises: applying the ground mesh covering an area of the image below the camera level. See paras. 94-96 and 229-230.
The subject-matter of independent claim 11 is analogous to and similar in scope to that of claim 1, and the rationale raised above to reject the later also apply, mutatis mutandis, to the former. In addition, Mann discloses an electronic device for determining a relative position of at least one object in an image (See paras. 14-15), comprising: a display and at least one processor (see paras. 229-230 and claim 15 of Mann); and memory storing instructions to be executed by the at least one processor (see paras. 256-264 and claim 19 of Mann).
Claim 12 is rejected under the same rationale as claim 2.
Claim 13 is rejected under the same rationale as claim 3.
Allowable Subject Matter
5. Claims 4-5, 7-10 and 14-15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, because the prior art of record fail to teach the method of claim 1, wherein the determining the placements of the first segmented object and the second segmented object based on the at least one semantic parameter associated with the first segmented object, the second segmented object, and the ground mesh comprises: determining ground corner points of the first segmented object and the second segmented object based on ground intersection edges of the first segmented object and the second segmented object; determining distances of the determined ground corner points to the camera level based on the ground mesh and the camera level; classifying the determined ground corner points as at least one of a near-ground corner point, a mid-ground corner point, or a far-ground corner point; and determining the placements of the first segmented object and the second segmented object based on the determined distances and the classified ground corner points (as recited in claims 4 and 14); wherein the determining the relative position of the first segmented object with respect to the second segmented object based on the determined placements of the first segmented object and the second segmented object comprises at least one of: comparing a distance of a near-ground corner point of the first segmented object with a distance of a near-ground corner point of the second segmented object; or comparing a distance of a far-ground corner point of the first segmented object with a distance of a far-ground corner point of the second segmented object (as recited in claims 5 and 15); and the method of claim 1, further comprising: locating corner points of the first segmented object and the second segmented object; determining intersection points of the located corner points with the ground mesh; calculating distances of each intersection point to the camera level; and determining relative positions of the first segmented object and the second segmented objected based on the calculated distances (as recited in claim 10).
6. Claim 16, after further consideration and search, is deemed to contain allowable subject matters, and are allowed over the prior art, because, while the prior art of record discloses substantial features of the claimed invention, as exemplified above with respect to the Mann reference, the prior art of record fails to particularly teach determining a relative position of at least one object in an image, by identifying a near object and a far object in the image; determining a first ground point of the near object and a second ground point of the far object based on ground intersection edges of the near object and the far object; determining placements of the near object and the far object based on the first ground point, the second ground point, and the ground mesh;
determining the relative position of the near object with respect to the far object based on the determined placement of the near object and the far object; and control the display to display the near object and the far object based on the determined relative position.
Conclusion
7. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Citraro et al. (US 20210225034) discloses Data processing systems are disclosed for determining semantic and person keypoints for an environment and an image and matching the keypoints for the image to the keypoints for the environment. A homography is generated based on the keypoint matching and decomposed into a matrix. Camera parameters are then determined from the matrix. A plurality of random camera poses can be generated and used to project keypoints for an environment using image keypoints. The projected keypoints can be compared to the actual keypoints for the environment to determine an error and weighting for each of the random camera poses.
8. Any inquiry concerning this communication or earlier communications from the examiner should be directed to WESNER SAJOUS whose telephone number is (571)272-7791. The examiner can normally be reached on M-F 10:00 TO 7:30 (ET).
Examiner interviews are available via telephone and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice or email the Examiner directly at wesner.sajous@uspto.gov.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said Broome can be reached on 571-272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WESNER SAJOUS/Primary Examiner, Art Unit 2612
WS
07/16/2026