Prosecution Insights
Last updated: August 18, 2026
Application No. 18/658,019

DEVICE AND METHOD WITH DETECTION OF OBJECTS FROM IMAGES

Final Rejection §103§112
Filed
May 08, 2024
Priority
Dec 08, 2023 — RE 10-2023-0177788
Examiner
KUDO, KEN
Art Unit
2671
Tech Center
2600 — Communications
Assignee
Samsung Electronics Co., Ltd.
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
38 currently pending
Career history
35
Total Applications
across all art units

Statute-Specific Performance

§101
16.1%
-23.9% vs TC avg
§103
51.6%
+11.6% vs TC avg
§102
8.1%
-31.9% vs TC avg
§112
23.4%
-16.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed on June 11, 2026 has been entered. Claims 1, 3–13, 15–16, and 19–20 are currently pending. Claims 1, 3, 8, 11, 13, and 15–16 have been amended. Response to Arguments Applicant’s arguments filed June 11, 2026 have been fully considered. Applicant’s arguments are persuasive in part with respect to the previous rejections under 35 U.S.C. § 102(a)(1) over Tsafrir, as explained below. Applicant argues that amended independent claims 1, 11, and 13, and amended dependent claim 16, now recite limitations not expressly disclosed by Tsafrir. In particular, amended claim 1 requires determining to collect an image as part of a training dataset when an accuracy level of a detected static-object region is less than a threshold accuracy level. Amended claim 11 requires selecting a ground-truth static-object region based on a degree of overlap between a pose-based projection of a three-dimensional occupancy map and a corresponding region in a training image. Amended claim 16 further requires generating permutations of a sensor pose by changing the sensor location and orientation, projecting the occupancy map according to the respective pose permutations, and selecting a ground-truth region from the resulting candidate regions based on comparisons with the detected region. Upon reconsideration, the Examiner agrees that Tsafrir does not expressly disclose the newly added low-accuracy threshold criterion, the expressly recited degree-of-overlap selection criterion, or the complete candidate-pose-permutation procedure now required by the amended claims. Accordingly, the previous rejections of claims 1–20 under 35 U.S.C. § 102(a)(1) as anticipated by Tsafrir have been withdrawn. However, Applicant’s arguments are not persuasive to the extent they suggest that Tsafrir fails to teach the underlying image-acquisition, sensor-pose, static-object-detection, three-dimensional-scene projection, projected-bounding-box comparison, ground-truth annotation, and training-dataset framework. As explained in the detailed rejections below, Tsafrir continues to teach those underlying limitations. New grounds of rejection under 35 U.S.C. § 103 are therefore made in this Office Action. Furthermore, it is noted that Applicant's amendments introduce new informalities and indefiniteness issues under 35 U.S.C. § 112(b). These new issues are addressed in the detailed Objections and 35 U.S.C. § 112(b) rejections below. Applicant’s amendments to claims 1, 3, 8, 11, 13, and 15–16 have been fully considered. The newly added limitations have been addressed by the updated rejections utilizing Tsafrir together with newly cited Je, Romero, and Armagan. Because the necessity to apply the newly cited references and to make the new § 112(b) rejections was directly necessitated by Applicant’s substantive amendments, this action is properly made final in accordance with MPEP § 706.07(a). Based on these facts, this action is made FINAL. Claim Objections Claim 1 is objected to because of the following informalities: Amended claim 1 bases the accuracy determination on "agreement between the static object region and a reference static object region", while amended claim 13 (the device claim, meant to mirror claim 1) uses "similarity" instead. The specification uses "comparison", "degree of registration", and "agree" interchangeably at [0075], so both terms are individually supported, but using two different words for what should be parallel limitations across the method and apparatus claims invites a claim-differentiation argument that the scopes are not identical, and could draw an indefiniteness inquiry into whether "agreement" and "similarity" are meant to be coextensive. Claim 11 is objected to because of the following informalities: claim 11 recites “the ground truth static object region is selected as such based on --” is awkward. It should read: “the ground truth static object region is selected based on--”. This is a grammatical informality. Appropriate correction is required. Claim 13 is objected to because of the following informalities: claim 13 appears to recites “wherein ground truth static object region is associated with the static object…”. The second occurrence should read: “wherein the ground truth static object region is associated with the static object…”. The antecedent is still reasonably apparent, so this is a grammatical informality. Appropriate correction is required. Claim 16 is objected to because of the following informalities: claim 16 introduces the sensor pose as comprising “a position and an orientation” but subsequently refers to “a corresponding change to the location and the orientation”, which is inconsistent wording. Appropriate correction is required. Claim Rejections - 35 USC § 112(b) The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3, 5-9, and 15 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 3, the claim is internally inconsistent with the parent claim from which it depends, rendering the scope of the claim unclear. Parent claim 1 explicitly requires “determining an accuracy level of the detected static object region based on agreement between the static object region and a reference static object region.”. Dependent claim 3, however, recites “determining the accuracy level of the detected static object region based on a ratio of a number of valid detection states and a number of invalid detection states”. It is unclear whether the "ratio" of valid to invalid states in claim 3 is intended to be the specific mathematical mechanism for calculating the "agreement" required by claim 1, or if claim 3 is improperly attempting to replace the "agreement" limitation of the parent claim with an entirely different alternative calculation. This contradiction renders the metes and bounds of claim 3 indefinite. The Examiner continue to interpret this limitation as to the same context with last Office Action. Regarding claims 5, the claims depend from claim 1 and recite “localization information of the sensor”. Claim 1, as amended, recites “a sensor pose of the sensor”. It is unclear whether “localization information of the sensor” in claims 5 is intended to refer back to “a sensor pose of the sensor” recited in claim 1, or whether “localization information” is intended to introduce a separate, distinct concept from “sensor pose” without antecedent basis in the claims (as intended in the last Office Action). Clarification and/or correction is required. For similar reasons, “candidate localization information”, “obtaining candidate static object regions of the respective units of candidate localization information”, and “the corresponding candidate localization information” in claim 5 are indefinite for setting a clear antecedent. Claims 6-7 depend from claim 5 and inherits the indefiniteness. Applicant is directed to the correction made in claim 16, wherein analogous language directed to “candidate calibration parameter sets” and “candidate localization information” was amended to instead recite “permutations of the sensor pose”, consistent with the “sensor pose” terminology introduced in claim 13. No corresponding correction was made to claims 5-7 to maintain consistency with the “sensor pose” terminology introduced in amended claim 1. Claim 8 recites the limitation "the determining the image" and "the selecting the ground truth static object region." in claim. There is insufficient antecedent basis for this limitation in the claim. Claim 1, from which claim 8 depends, does not recite a step of "determining the image", claim 1 recites "determining to collect the image as part of a training dataset". Claim 1 also does not recite any step of "selecting" a ground truth static object region, claim 1 recites "determining a ground truth static object region". It is unclear whether "the determining the image" refers to "determining to collect the image" of claim 1, and whether "the selecting the ground truth static object region" refers to "the determining a ground truth static object region" of claim 1, or whether these terms are intended to introduce new, undefined steps. Clarification and/or correction is required. Claim 9 depends from claim 8 and inherits the indefiniteness. Claim 15 recites the limitation "a partial reference region" in claim. There is insufficient antecedent basis for this limitation in the claim. The claim initially recites determining a reference static object region "which comprises reference sub-regions respectively corresponding to the static objects". Subsequently, the claim requires making a determination based on "a partial region and a partial reference region." It is unclear whether the newly introduced "partial reference region" is intended to refer to one of the previously recited "reference sub-regions" or if it constitutes an entirely distinct limitation. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3–13, 15–16, and 19–20 are rejected under 35 U.S.C. §103 as being unpatentable over Tsafrir in view of Je (Je et al., US 2021/0334652 A1, 2021). Regarding claim 1, Tsafrir teaches a method performed by an electronic device (in [0049]: Tsafrir describes systems, methods, a computing device, and/or code instructions … executable by one or more hardware processor(s) for automatically creating an annotated training dataset for training an ML model), the method comprising: obtaining an image of a scene captured by a sensor, wherein a sensor pose of the sensor corresponds to the obtaining of the image by the sensor; ( Tsafrir, [0049], [0064]: Tsafrir teaches capturing a sequence of digital images of a three-dimensional scene using one or more sensors. Tsafrir further teaches applying visual simultaneous localization and mapping (SLAM) to the captured images to generate registered images and a camera path including, for each registered image, a corresponding camera position and camera orientation. The corresponding camera position and orientation constitute the sensor pose associated with obtaining the respective image. ) detecting, in the image, a static object region corresponding to a static object of the image, wherein the detecting is performed by applying an object detection model to the image; ( Tsafrir, in [0007], [0049] & [FIG. 1, Step 103], static objects and moving objects are differentiated during detection using a stability score indicative of a likelihood that each point is associated with a static object of the scene, further teaches detecting and classifying and annotating at least one static object in the scene of the image; in [0058], objects captured in the images are each classified and bounded by a bounding box; [FIG. 4A, step 402]: project a bounding box of the object onto an image of the sequence) and [FIG. 4B, step 451]: project bounding box onto an image; ) determining an accuracy level of the detected static object region based on agreement between the static object region and a reference static object region that is a projection of a 3D occupancy map of the scene according to the sensor pose; ( Tsafrir, [0064], [0112], [0124]: Tsafrir teaches generating a static three-dimensional stacked-scene representation using registered images and a SLAM-derived camera path comprising camera positions and orientations; projecting an object or an object bounding box from the three-dimensional stacked scene onto an image; and comparing the expected projected bounding box with an identified or detected bounding box in the image. Tsafrir uses the difference between the expected and identified bounding boxes as an error parameter to verify or modify the annotation and/or the camera position and orientation. The disclosed difference/ error constitutes a quantitative measure of agreement and thus an accuracy level of the detected object region relative to the pose-based projected reference region. ) determining whether to collect the image as part of a training dataset based on an accuracy level of the detected static object region ( Tsafrir, [FIG. 4B], step 452 teaches annotating the object in the image with a confidence scores; wherein [0114]-[0115] explains the confidence score is “indicative of a likelihood the object is annotated correctly"; then step 453 identifying a highest confidence score, and step 454 to determining using the image associated with the highest confidence score as the one; this corresponds to determining whether to collect / use an image based on an accuracy level [confidence] of the detected / annotated object region [bounding box of the object]. Tsafrir further teaches creating annotated dataset(s) for training an ML model [0007], [0049], [0086] ) determining a ground truth static object region for the static object of the image from 3D occupancy information of the static object with respect to the image collected as part of the training dataset. ( Tsafrir, [0007], [0049], [0115], [0144–0147], and [Fig. 4B]: Tsafrir teaches detecting a static object in a static three-dimensional stacked scene, projecting the object’s bounding box from the three-dimensional scene onto an image, using the projected bounding box as the ground-truth annotation for the static object, and storing the image and corresponding ground-truth annotation as a record in an annotated training dataset. ) Tsafrir teaches selecting an image for an annotated training dataset based on a confidence/ accuracy level associated with the detected static-object region, but fails to expressly disclose selecting the image when the confidence/ accuracy level is less than a threshold, where Je teaches: determining to collect the image as part of a training dataset based on the accuracy level of the detected static object region being less than a threshold accuracy level; and ( Je, [0022], [0061–0064], [0070], [0078]: Je teaches selecting a certain image frame in which an object has a confidence score, included in the object-detection information, that is equal to or lower than a preset value; storing the selected frame as a specific frame useful for training a perception network; and sampling the stored frame to generate training data.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Tsafrir’s confidence-based training-image selection method by applying Je’s known low-confidence selection criterion, such that an image is selected and stored for training when the object-detection confidence is equal to or lower than a preset value. Both references concern selecting informative vehicle images for training an object-detection or perception model, and the modification would predictably collect low-confidence or hard-example images that are useful for improving training efficiency, while each technique performs its known function. Regarding claim 3, Tsafrir [as modified by Je] teaches the method of claim 1, wherein the detecting of the static object region comprises detecting sub-regions of the image respectively corresponding to static objects of the image, including the static object, ( Tsafrir, [FIG. 1, steps 101–103], teaches producing a sequence of registered images, producing a stacked scene, and using the stacked scene to “detect and annotate (i.e., classify and identify a bounding box)” one or more objects in the sequence of images, including “one or more static objects,” which corresponds to detecting image sub-regions (bounding boxes) for respective static objects ) determining reference static object sub-regions of the static objects from the detected sub-regions, and ( Tsafrir, [FIG. 4A, steps 401–406], teaches identifying an object in the stacked scene, projecting the object’s bounding box from the stacked scene onto an image in the sequence and annotating the object according to the projected bounding box, and further projecting the bounding box to another image and annotating that instance, thereby providing a projected / reference sub-region corresponding to each detected sub-region ) the determining of the accuracy level of the detected static object region comprises: determining detection states for the respective static objects, wherein each static object's detection state is determined based on its detected sub-region and its reference sub-region, and wherein static object's detection state indicates whether the detection of the static object is valid or invalid; and ( [0115] & [0129]: Tsafrir teaches comparing a reference / expected bounding box (computed from the stacked scene) to an identified bounding box in another image, and using the difference between them as an error parameter to modify classification of bounding box and/or camera pose / orientation, which corresponds to determining whether the detected sub-region matches the reference sub-region (i.e., correct/ valid vs incorrect/ invalid detection) ) determining the accuracy level of the detected static object region based on a ratio of a number of valid detection states and a number of invalid detection states. ( [0115]: Tsafrir teaches computing a plurality of confidence scores across iterations and selecting the highest confidence score / associated image for annotation, which is a confidence-based selection mechanism as a ratio of a number of valid detection states and a number of invalid detection states. ) Regarding claim 4, Tsafrir [as modified by Je] teaches the method of claim 1, further comprising: displaying the detected static object region, wherein the determining of whether to collect the image is based on detecting a user input requesting a collection of the image. ( [0107]: Tsafrir discloses computing device includes a display to view the annotated dataset, the detected static object region; also includes a mechanism designed for a user to enter data (e.g., enter manual annotation) to collect the image. ) Regarding claim 5, Tsafrir [as modified by Je] teaches the method of claim 1, wherein the determining of the ground truth static object region comprises: obtaining, for candidate calibration parameter sets, respectively corresponding units of candidate localization information, each obtained by calibrating localization information of the sensor using a corresponding candidate calibration parameter set; ( [0064]: Tsafrir teaches aligning / calibrating a sequence of images for consistent representation of points of a captured 3D scene, then applying SLAM to determine a camera path comprising camera positions and camera orientations, as a localization information [location + orientation as are described in the Applicant's specification] associated with registered images; and such calibration is in terms that include, one of more of: aspect ratio, scale, focal point, view point, resolution, scan patterns, and frame rate... ) obtaining candidate static object regions of the respective units of candidate localization information, wherein each candidate static object region is obtained from a view of a three-dimensional space occupied by the static object as viewed from a viewpoint direction of the corresponding candidate localization information; and ( [0061], [0064]: Tsafrir teaches using SLAM to “calculate the position and orientation of a sensor with respect to its surroundings” i.e., localization information defining the sensor viewpoint and viewing direction; then, providing consistent 3D-scene representation for particular viewpoints constitutes obtaining static object regions from a view of a three-dimensional space occupied by the static object as viewed from a viewpoint direction of corresponding localization information) determining, from among the candidate static object regions, the ground truth static object region based on comparison between each candidate static object region with the static object region. ( [0129]: teaches computing an expected bounding box in another image and comparing it to an identified bounding box, and using the difference as an error parameter to modify annotations and/or camera position / orientation, thereby selecting / refining the ground-truth region based on comparisons between candidate regions and the object region. ) Regarding claim 6, Tsafrir [as modified by Je] teaches the method of claim 5, wherein the determining of the ground truth static object region comprises: determining, for each candidate static object region, a similarity level between the static object region and a corresponding candidate static object region based on a number of pixels classified into the same class in the static object region and in a corresponding candidate static object region; and ( Tsafrir, [0005–0006]. [0015], [0064], teaches pixel-level class identification for static content by “identifying … a set of static pixels by projecting the static three-dimensional stacked scene onto … the image” and “extracting the set of static pixels … to create a static image,” and further teaches “classifying and annotating” the static object based on that static / pixel set; [0069], Tsafrir provides pixels classified into a same class (static-pixel class) within an object region produced by projection into a three-dimensional representation of a scene; and that the stacked scene comprises, for each pixel of the sequence of registered images; [0118]: Tsafrir teaches that one or more stability scores of the stacked scene are modified according to static objects detected / annotated in the images; then in [0166], Tsafrir determines a likelihood score (similarity level) based on the pixel/point-level static classification / stability information associated with the object region. ) determining, to be the ground truth static object region, from among the candidate static object regions, a candidate static object region with a maximum similarity level. ( [0166]: Tsafrir further teaches leaving only pixels with a high likelihood of being stable and belonging to ground truth static objects. ) Regarding claim 7, Tsafrir [as modified by Je] teaches the method of claim 5, wherein each candidate calibration parameter set comprises a position delta and an orientation delta. ( [0064]: Tsafrir further discloses the principle of SLAM is to use the location of visual features (for example corners) between consecutive images to calculate the position and orientation of a sensor with respect to its surroundings in the three-dimensional scene [a position delta and an orientation delta as are described in the Applicant's specification]. ) Regarding claim 8, Tsafrir [as modified by Je] teaches the method of claim 1, wherein the sensor is mounted on a moving object, and ( [0006]: Tsafrir teaches “a target image captured by a camera located on a moving vehicle” is fed into an ML model. ) the obtaining the image, the detecting the static object region, the determining to collect the image, the selecting the ground truth static object region, and detecting an object are performed while the moving object travels, wherein the object is the static object or another static object. ( Tsafrir, [0009], [0064], [0114–0115], and [Fig. 4B]; Je, [0002], [0015], [0049], [0059–0064], and [0070]: Tsafrir teaches capturing sequences of images by sensors moving through a scene, detecting a static object and another static object in the captured images, projecting a bounding box of the static object from a static three-dimensional scene onto an image, annotating the projected object region, and selecting an image and its corresponding annotation for use in a training dataset. Je teaches an on-vehicle active-learning process in which, while an autonomous vehicle is driven, driving-video frames are acquired from a camera mounted on the vehicle, objects are detected in the frames by a deep-learning object detector, a frame having object-detection confidence equal to or lower than a preset value is selected, and the selected frame is stored as useful training data. Accordingly, Tsafrir [as modified by Je] teaches performing the image acquisition, static-object detection, training-image collection, and detection of the static object or another object while the moving object travels. ) Regarding claim 9, Tsafrir [as modified by Je] teaches the method of claim 8, further comprising controlling driving of the moving object based on a result obtained by the detecting of the object. ( [0158-0159]: Tsafrir teaches that the outcome (indication of the target object) is used to generate instructions for controlling the vehicle, and further teaches generating instructions for automatically maneuvering the vehicle to avoid collision with a detected target object; Tsafrir also provides examples in consistent with vehicle control actions such as triggering automatic braking or maneuvering to avoid collision based on detection/recognition outcomes [0049]. ) Regarding claim 10, Tsafrir [as modified by Je] teaches the method of claim 1, wherein the method further comprises: determining a loss value for adaptive learning based on a difference between the static object region and the ground truth static object region; and ( [0124]: Tsafrir teaches the expected bounding box is compared to an identified bounding box of the object, where the difference is used as an error parameter to the classification process, creating records for training. ) updating a parameter of the object detection model based on the determined loss value for adaptive learning. ( [0147]: Tsafrir, FIG. 9 teaches creating a training dataset (914) comprises plurality of record and training an ML model (916) using that dataset; wherein [0085], Tsafrir teaches a central server based implementation in which one or more servers receive images / sensor data from client terminals and perform training, i.e., performing adaptive learning/training that updates the ML model. ) Regarding claims 13, 15, and 19–20 the rationale provided for claims 1 and 3–10 is incorporated herein. In addition, the electronic device of claims 13, 15, and 19–20 corresponds to the method of claims 1 and 3–10, and performs the steps disclosed herein. Therefore, the claims are all rejected. Claims 11–12 are rejected under 35 U.S.C. §103 as being unpatentable over Tsafrir in view of Romero (Romero et al., US 2023/0109712 A1, 2023). Regarding claim 11, Tsafrir teaches a method performed by an electronic device, the method comprising: obtaining an image of a scene captured by a sensor, wherein a sensor pose of the sensor corresponds to the obtaining of the image by the sensor; ( Tsafrir, [0049], [0064]: Tsafrir teaches capturing a sequence of digital images of a three-dimensional scene using one or more sensors. Tsafrir further teaches applying visual simultaneous localization and mapping (SLAM) to the captured images to generate registered images and a camera path including, for each registered image, a corresponding camera position and camera orientation. The corresponding camera position and orientation constitute the sensor pose associated with obtaining the respective image. ) detecting, in the image, a static object region corresponding to a static object of the image, wherein the detecting is performed by applying an object detection model to the image; ( Tsafrir, in [0007], [0049] & [FIG. 1, Step 103], static objects and moving objects are differentiated during detection using a stability score indicative of a likelihood that each point is associated with a static object of the scene, further teaches detecting and classifying and annotating at least one static object in the scene of the image; in [0058], objects captured in the images are each classified and bounded by a bounding box; [FIG. 4A, step 402]: project a bounding box of the object onto an image of the sequence) and [FIG. 4B, step 451]: project bounding box onto an image; ) wherein the object detection model is an adaptively learned model, based on a training dataset, which comprises a training image with respect to a training static object, and a ground truth static object region mapped to the training image, and ( Tsafrir expressly teaches training ML models on images annotated with a ground truth label (e.g., bounding box) [0006], and training a machine learning model on a training dataset comprising records that include images and ground-truth label indications [0007] & [0019], as part of creating an annotated training dataset for training an ML model [0049].) wherein the ground truth static object region is selected as such based on a difference between a projection, according to the sensor pose, of a 3D occupancy map of the scene and a region in the training image corresponding to the training static object ( Tsafrir, [0064], [0112], [0124], and [0144–0147]: Tsafrir teaches projecting an expected bounding box from a static three-dimensional (3D) stacked scene according to the camera pose, comparing the expected bounding box with an identified bounding box in the image, and using a difference between the boxes to verify or modify the annotation. It would have been obvious to use the annotation verified or modified through Tsafrir’s comparison as the selected ground-truth static-object region for the corresponding training image, because Tsafrir expressly uses the resulting annotated images as ground-truth-labeled records for training an ML model. ) Tsafrir teaches projecting a static-object bounding box from a 3D scene according to the camera pose, comparing the projected expected bounding box with an identified bounding box in an image, and using their spatial difference to verify or modify the annotation, but fails to expressly disclose selecting the ground-truth static-object region based specifically on a degree of overlap between the projected and identified regions, where Romero teaches: wherein the ground truth static object region is selected as such based on a degree of overlap between a projection, according to the sensor pose, of a 3D occupancy map of the scene and a region in the training image corresponding to the training static object ( Romero, [0029]–[0031], [0041–0044], [Algorithm 1], and [Figs. 1–2]: Romero teaches obtaining three-dimensional object bounding boxes from point-cloud information representing a vehicle scene, projecting the three-dimensional bounding boxes into a camera image using sensor-calibration information, calculating an Intersection-over-Union value between each projected three-dimensional bounding box and corresponding two-dimensional object-detection boxes, identifying the maximum IoU, and retaining the three-dimensional object result when the maximum IoU exceeds a threshold. Accordingly, Romero teaches selecting an object region based on a degree of overlap between a calibrated projection of three-dimensional scene information and a corresponding region detected in the image. ) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Tsafrir’s projected-bounding-box comparison by applying Romero’s IoU-based overlap criterion, because both references compare a three-dimensional-derived projected object region with a corresponding two-dimensional detected image region. Using Romero’s known IoU measure would predictably quantify the spatial agreement already evaluated by Tsafrir and permit more reliable selection, verification, or refinement of the ground-truth static-object region. Regarding claim 12, Tsafrir [as modified by Romero] teaches a non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 11. ( Tsafrir, in [0049], discloses embodiments as “systems, methods, a computing device, and/or code instructions, "stored on a memory and executable by one or more hardware processors" for automatically creating an annotated training dataset for training an ML model. ) Claim 16 is rejected under 35 U.S.C. §103 as being unpatentable over Tsafrir [as modified by Je] in view of Armagan (Armagan et al., "Semantic segmentation for 3D localization in urban environments". IEEE, 2017). Regarding claim 16, Tsafrir [as modified by Je] teaches the electronic device of claim 13, wherein the one or more processors are configured to: Tsafrir [as modified by Je] teaches generating a pose-based projected static-object region and comparing the projected region with an object region detected in the image, but fails to expressly disclose where Armaganteaches: obtain, permutations of the sensor pose, the sensor pose comprising a position and an orientation, wherein each permutation of the sensor pose is determined by a corresponding change to the location and the orientation; ( [Sec. V, “3D Localization”], [Eq. (2)], [Sec. VI, “Evaluation”], and [Fig. 1]: Armagan teaches obtaining an initial camera pose from device sensors and sampling a plurality of candidate poses around the initial sensor pose. Armagan samples candidate camera locations over a two-dimensional region in one-meter increments and candidate camera orientations over a range centered on the sensor-provided orientation. Accordingly, each candidate pose constitutes a permutation of the sensor pose determined by a corresponding change in camera location and orientation. ) obtain candidate static object regions of the respective permutations of the sensor pose, wherein each candidate static object region is obtained by projecting the occupancy map from a corresponding permutation of the sensor pose; and ( [Sec. V, Eqs. (2)–(3)], and [Figs. 1 and 3]: Armagan teaches rendering a 2.5D map of surrounding static buildings under each sampled camera pose. Each rendering projects the mapped building façades and vertical edges into the image plane according to the corresponding candidate camera location and orientation, thereby producing respective candidate static-object regions for the candidate sensor poses. ) determine, from among the candidate static object regions, the ground truth static object region based on comparison of each candidate static object region with the static object region. ( [Secs. III and V, Eqs. (1)–(3)], and [Figs. 1 and 3]: Armagan teaches semantically segmenting the input image to obtain probability maps for static building façades and edges and, for each candidate pose, comparing the corresponding rendered map regions with the segmented image regions by calculating a pixelwise log-likelihood. Armagan retains the candidate pose having the largest log-likelihood, thereby selecting the candidate projected static-object region having the greatest agreement with the static-object region detected in the image. When incorporated into Tsafrir’s annotated-training-dataset process, the selected candidate projected region is used as the ground-truth static-object region. ) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to further modify Tsafrir [as modified by Je] by applying Armagan’s candidate-pose search technique, because Tsafrir [as modified by Je] recognizes that differences between projected and detected object regions may result from camera-position or camera-orientation error, while Armagan teaches varying the sensor position and orientation, projecting the mapped scene under each candidate pose, and selecting the best-matching projection. The modification would predictably improve the accuracy of the selected ground-truth static-object region. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. KEN KUDO Examiner Art Unit 2671 /KEN KUDO/Examiner, Art Unit 2671 /VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

May 08, 2024
Application Filed
Mar 11, 2026
Non-Final Rejection mailed — §103, §112
May 06, 2026
Examiner Interview Summary
May 06, 2026
Applicant Interview (Telephonic)
Jun 11, 2026
Response Filed
Aug 04, 2026
Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month