Prosecution Insights
Last updated: October 01, 2026
Application No. 19/040,535

METHOD AND APPARATUS WITH THREE-DIMENSIONAL OBJECT DETECTION

Non-Final OA §102§103§112
Filed
Jan 29, 2025
Priority
May 16, 2024 — RE 10-2024-0064025 +1 more
Examiner
MOTSINGER, SEAN T
Art Unit
Tech Center
Assignee
Korea University Research and Business Foundation
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 3m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
547 granted / 697 resolved
+18.5% vs TC avg
Moderate +12% lift
Without
With
+11.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
30 currently pending
Career history
718
Total Applications
across all art units

Statute-Specific Performance

§101
14.2%
-25.8% vs TC avg
§103
40.7%
+0.7% vs TC avg
§102
19.0%
-21.0% vs TC avg
§112
19.2%
-20.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 697 resolved cases

Office Action

§102 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 3-5 and 13-15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Re clams 3 this claim contains the language “intrinsic/extrinsic parameters of a camera”, the examiner note that it is unclear what this is intended to mean does this mean both intrinsic and extrinsic parameters are need or one of intrinsic or extrinsic parameters are needed. Claims 4-5 are rejected because they depend from claim 3 and therefore contain the same features. Re clams 4 this claim contains the language “intrinsic/extrinsic parameters of a camera”, the examiner note that it is unclear what this is intended to mean does this mean both intrinsic and extrinsic parameters are need or one of intrinsic or extrinsic parameters are needed. Claim 5 is rejected because it depends from claim 4 and therefore contains the same features. Re clams 13 this claim contains the language “intrinsic/extrinsic parameters of a camera”, the examiner note that it is unclear what this is intended to mean does this mean both intrinsic and extrinsic parameters are need or one of intrinsic or extrinsic parameters are needed. Claims 14-15 are rejected because they depend from claim 13 and therefore contain the same features. Re clams 14 this claim contains the language “intrinsic/extrinsic parameters of a camera”, the examiner note that it is unclear what this is intended to mean does this mean both intrinsic and extrinsic parameters are need or one of intrinsic or extrinsic parameters are needed. Claim 15 is rejected because it depends from claim 14 and therefore contains the same features. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-9 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chang et al “Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection” arXiv:2410.22461 29 October 2024. (Note that this paper lists some similar inventors to the present application however “Donghyun Kim” is an author but not a joint inventor of the present application.) Applicant cannot rely upon the certified copy of the foreign priority application to overcome this rejection because a translation of said application has not been made of record in accordance with 37 CFR 1.55. When an English language translation of a non-English language foreign application is required, the translation must be that of the certified copy (of the foreign application as filed) submitted together with a statement that the translation of the certified copy is accurate. See MPEP §§ 215 and 216. Re claim 1 Chang discloses ”A method of detecting a three-dimensional (3D) object, the method comprising: extracting two-dimensional (2D) image features from images using an image backbone (see figure 3 note that Multiview input images are input into a X6 features used in the view formers also see section 3.1 “ encode 2D visual features alongside the 3D spatial environment into a bird’s eye view (BEV) representation” note these are 2D visual features); extracting a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features (see figure 3 view transformer note that the depth net extracts the 3d depth ) by using a view transformer configured to perform domain generalization (see abstract and figure 3 note that domain generalization is performed during the process); extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder (see section 3.1 “encode 2D visual features alongside the 3D spatial environment into a bird’s eye view (BEV) representation” see also figure 3 note that 2d feature “X6” and the output of the depth net are multiplied and inserted into a BEV pool to form “X1” features where is then input into the BEV back bone and BEV neck which correspond to the BEV encoder.); and predicting a position of the object and a class of the object from the BEV feature by using a detection head (see abstract note that object detection is performed see figure 3 note that scale orientation and translation are determined which could correspond to the position and the class see also section 3.1 “Subsequently, Detector Head modules D supervises BEV features with 3D labels Y in a three-dimensional manner”). Re claim 2 Chang discloses wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features and by inputting, into a BEV pool, an outer product of the depth output of the DepthNet and the 2D image features ( see figure 3 “(see section 3.1 “encode 2D visual features alongside the 3D spatial environment into a bird’s eye view (BEV) representation” see also figure 3 note that 2d feature “X6” and the output of the depth net are multiplied [e.g. outer product] and inserted into a BEV pool to form “X1” features where is then input into the BEV back bone and BEV neck which correspond to the BEV encoder.”) Re claim 3 Chang discloses wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images (see section 3.3 approach whole section “Concretely, we take two advantages leveraging Lp in narrow occluded regions; First, Lp effectively mitigates the triangular misalignment. Second, Lp potentially supports insufficiently scaled Lov. Ultimately, we alleviate perspective view gaps by directly constraining the corresponding depth and the photometric matching between adjacent views.” And “ to achieve generalized BEV extraction, we directly constrain depth estimation network from adjacent overlap regions between multi-view cameras.” Note that in the section transformation matrices are used in the overlap region based on camera parameters K and T this mitigates misalignment for the depth estimation network, note the depth is generalized i.e. normalized) Re claim 4 Chang discloses wherein cameras, including the camera, provide the respective images, and wherein the relative depth normalization method comprises calculating a transformation matrix through which geometric transformation is performed between adjacent pairs of the cameras from the intrinsic/extrinsic parameters and the camera (see section 3.3 approach “To achieve generalized BEV extraction, we directly constrain depth estimation network from adjacent overlap regions between multi-view cameras. Also, we advocate that multi-frame image inputs substantially complement geometric understanding in dynamic scenes with speedy translation and rotation shifts. To this end, we formulate corresponding depth D∗ leveraging spatial and temporal adjacent views. First, we calculate overlap transformation matrices Ti→j” note that transformation matrices are calculated for the overlap region during generalized the depth estimation) Re claim 5 Chang discloses wherein the relative depth normalization method obtains a relative depth after projecting an image feature onto an adjacent image feature by using the depth prediction information and the transformation matrix and minimizing a relative depth loss based on a depth loss function (see section 3.3 approach section in particular equation 3 note that the difference between estimated depth projected using the transformation matrix is used to calculate Lov which penalized unmatched depths between Dj and Di-to-j[projected using the transformation matrix].) Re claim 6 Chang discloses wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method (see section 3.3 approch section especially “Ultimately, we alleviate perspective view gaps by directly constraining the corresponding depth and the photometric matching between adjacent views” note that alignment is optimized using photometric matching). Re claim 7 Chang discloses wherein the image backbone, the view transformer, the BEV encoder, and/or the detection head comprise respective domain adaptation adapters (see section 3.4 label efficient domain adaptation and Appendix C Searching adapter structures “We explore various modules and structures to find a suitable adapter architecture.Tab.9,10 show which structures and locations affects the model’s performance. For adapter locations, performance is optimal when adapters are attached to all modules.” Note that domain adapter are attached to all modules note that A is built parallel to operation block). Re claim 8 Chang discloses wherein each domain adaptation adapter is added in parallel to an operation block to enable fine-tuning on parameters (see figure 3 label efficient DA portion and caption note that the adapter is used in a fine tuning process in parallel to the operational block, See section 3.4 note that the adapter is parallelly built). Re claim 9 Chang disclose wherein each domain adaptation adapter is configured to perform a skip connection in which features input to the view transformer, the BEV encoder, and/or the detection head are received, operated, and summed to update a gradient (see section 3.4 “Secondly, we fuse each outputs from B, and Adapter by exploiting skip-connections that directly link between the down sampling and up sampling paths. By doing so, these extensible modules allow to capture high-resolution spatial details while reducing network and computational complexity.” Note that that the adapters provide a skip connection see figure 3 and equation 6 note that the input to the operational block is operated on by the adapter and then summed with the output of the operational block). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3 7, 8 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al “Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View.” Computer Vision and Pattern Recognition March 2023 arXiv:2303.01686 in view of Huang et al “BEVDet: High-Performance Multi-Camera 3D Object Detection in Bird-Eye-View” arXiv:2112.11790 December 2021. Re claim 1 Wang discloses A method of detecting a three-dimensional (3D) object, the method comprising: extracting two-dimensional (2D) image features from images using an image backbone (see figure 2 Backbone element which generates domain agnostic features from the 2d imagers also see section 3.2 last paragraph “In MV3D-Det, P(X,K,E) indicates the feature distribution of 2D images projected into 3D space, which is determined by the depth estimation and 2D image feature jointly”); extracting a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features (see figure 3 depth distribution note that a depth net takes in the domain agnostic features and used them to determine the depth distribution) by using a view transformer configured to perform domain generalization;(section 3.2. note that intrinsics decoupled depth estimation is used to transform into a depth distribution see figure 3 caption not this is done with a strategy for domain generalization) extracting a bird's eye view (BEV) feature (see section 3.2 note that the features are projected into a BEV space) a detection head (see figure 3 detection head element) Wang does not expressly disclose” extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder; predicting a position of the object and a class of the object from the BEV feature by using a detection head. Huang discloses in a similar BEV object detection environment discloses extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder ( see figure 2 note that the 3d depth data is input into a BEV encoder see also section 2.2 first paragraph “a BEV encoder for further encoding the feature in BEV”); predicting a position of the object and a class of the object from the BEV feature by using a detection head (see figure 2 note that BEV features are input into a detection head to perform 3D detection see also section 3.1 note that a detection head is used, see section 2.1 note that “ object detection, aiming at both classification labels and the instance bounding boxes of all pre-defined objects”). The motivation to combine the BEV encoder “perceives some pivotal cues like depth, scale, orientation, and speed, which is difficult for predicting in image view” (see section 3.1 Bev encoder) and “ 3D object detection in automatic pilot aims at the position, scale, orientation, and speed of movable objects like pedestrians, vehicles, barriers, and so on” (section 3.1 Heads) and “Benefiting from explicitly encoding feature in BEV, BEVDet is genius in perceiving the targets’ translation, orientation, and velocity.” section 1 last paragraph. One of ordinary skill in the art could have easily added the Bev encoder and projection head of Huang to the teaching of Wang to reach the aforementioned advantage. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Wang and Huang to reach the aforementioned advantages. Re claim 2 Wang discloses wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features (see figure 3 note that features from the 2d images are input into the depth net) and by inputting, into a BEV pool (see figure 3 voxel pooling), an outer product of the depth output of the Depth Net and the 2D image feature (see figure 3 note that the product of the 2d features and the depth distribution are multiplied and input into the voxel pooling) Re claim 3 Wang further discloses wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images (see section 3.3.1 and section 3.3.2 page 5 first paragraph note “As showninFig.2, when two cameras with various intrinsic parameters (i.e. focal lengths) shoot the same object at the same distance, the imaging size of the object, which is determined by the intrinsic parameters, can be quite different. If a model is only optimized on the dataset collected from a specific camera, it can be difficult for the model to predict an identical depth for the object pictured from another camera, and thus it is the cause of inaccurate depth prediction when domain shifts, similar in [20]. Furthermore, we empirically find that the estimated depth has been entangled with the intrinsic parameters of the camera, and results in non-compliance with the intuition of “Everything looks small in the distance and big on the contrary”. Hence, we attempt to decouple the estimated depth from the intrinsic parameters” note that the process for computing depth is trained such that it is invariant [i.e. normalized] to the camera parameters this is to remove errors caused by the camera position of the images being captured being different from the one the model trains on). Re claim 7 Wang further discloses image backbone, the view transformer (see figure 3 portion i and associated caption as well as section 3.3.1 note that intrinsic decoupled depth estimation is performed and intrinsic parameters are used to adapt to the domain during generation of the depth distribution), the BEV encoder, and/or the detection head comprise respective domain adaptation adapters (see figure 3 portion i and associated caption as well as section 3.3.1 note that intrinsic decoupled depth estimation is performed and intrinsic parameters are used to adapt to the domain during generation of the depth distribution). Re claim 8 Wang further discloses wherein each domain adaptation adapter is added in parallel to an operation block to enable fine-tuning on parameters (see figure 3 portion i and associated caption as well as section 3.3.1 note that intrinsic decoupled depth estimation is performed and intrinsic parameters are used to adapt to the domain during generation of the depth distribution this tunes the depth distribution to the particular camera parameters; note that this runs in parallel to some of the operation blocks). Re claim 10 Wang discloses further comprising augmenting the 3D feature map by performing a generalization method of decoupling-based image depth estimation (see figure 3 and associated caption see section 3.3.1 note that depth prediction is intrinsics decoupled to perform domain generalization). Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al “Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View.” Computer Vision and Pattern Recognition March 2023 arXiv:2303.01686 in view of Huang et al “BEVDet: High-Performance Multi-Camera 3D Object Detection in Bird-Eye-View” arXiv:2112.11790 December 2021 in view of Rho et al “ORA3D: Overlap Region Aware Multi-view 3D Object Detection” ARXiv 2207.00865v4 2023. Re claim 6 Wang and Huang do not expressly disclose wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method. ORA3d discloses wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method (see section 3.1. “Learning Depth Cue by Multi-view Stereo Matching. When considering the multi-view camera system, adjacent cameras have a strong association. We regard this association comes from overlap regions and can be extended to geometric guides. To interactively supervise the network, we train the Stereo Disparity Estimation head, which reconstructs a dense disparity map with overlap region pairs of neighboring cameras” The examiner notes that disparity between overlap regions is detected see figure 2 caption “1) Stereo Matching Network for Weak Depth Supervision, where our depth estimation head is trained to predict a dense depth map Of the overlap region” note that the disparity is used to predict a depth map for overlapped images.) . The motivation to combine is “ Moreover, objects in the overlap region are often largely occluded or suffer from deformation due to camera distortion, causing a domain shift. To mitigate this issue, we propose using the following two main modules: (1) Stereo Disparity Estimation for Weak Depth Supervision and (2) Adversarial Overlap Region Discriminator. The former utilizes the traditional stereo disparity estimation method to obtain reliable disparity information from the overlap region” (see abstract). One of ordinary skill in the art could have easily modified the teachings of Wang and Huang to incorporate the teachings of Roh to reach the aforementioned advantage. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Wang Huang and Roh to reach the aforementioned advantage. Claim(s) 11-13 17-18 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al “Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View.” Computer Vision and Pattern Recognition March 2023 arXiv:2303.01686 in view of Huang et al “BEVDet: High-Performance Multi-Camera 3D Object Detection in Bird-Eye-View” arXiv:2112.11790 December 2021 in view of Nagori US2024/0354888. Re claim 11 Wang discloses extract two-dimensional (2D) image features from images using an image backbone (See figure2 Backbone element which generates domain agnostic features from the 2d imagers also see section 3.2 last paragraph “In MV3D-Det, P(X,K,E) indicates the feature distribution of 2D images projected into 3D space, which is determined by the depth estimation and 2D image feature jointly”); extract a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features by using a view transformer (see figure 3 depth distribution note that a depth net takes in the domain agnostic features and used them to determine the depth distribution) (section 3.2. note that intrinsics decoupled depth estimation is used to transform into a depth distribution see figure 3 caption not this is done with a strategy for domain generalization) extract a bird's eye view (BEV) feature from the 3D feature map (see section 3.2 note that the features are projected into a BEV space, see also figure 3) a detection head (see figure 3 detection head element) Wang does not expressly disclose” extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder; predicting a position of the object and a class of the object from the BEV feature by using a detection head. An electronic device comprising: a memory storing instructions; and one or more processors, wherein the instructions, when performed by the one or more processors, cause the one or more processors Huang discloses in a similar BEV object detection environment discloses extracting a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder ( see figure 2 note that the 3d depth data is input into a BEV encoder see also section 2.2 first paragraph “a BEV encoder for further encoding the feature in BEV”); predicting a position of the object and a class of the object from the BEV feature by using a detection head (see figure 2 note that BEV features are input into a detection head to perform 3D detection see also section 3.1 note that a detection head is used, see section 2.1 note that “ object detection, aiming at both classification labels and the instance bounding boxes of all pre-defined objects”). The motivation to combine the BEV encoder “perceives some pivotal cues like depth, scale, orientation, and speed, which is difficult for predicting in image view” (see section 3.1 Bev encoder) and “ 3D object detection in automatic pilot aims at the position, scale, orientation, and speed of movable objects like pedestrians, vehicles, barriers, and so on” (section 3.1 Heads) and “Benefiting from explicitly encoding feature in BEV, BEVDet is genius in perceiving the targets’ translation, orientation, and velocity.” section 1 last paragraph. One of ordinary skill in the art could have easily added the Bev encoder and projection head of Huang to the teaching of Wang to reach the aforementioned advantage. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Wang and Huang to reach the aforementioned advantages. While Wang and Huang are clearly intended to be performed on a computer they do not expressly disclose an electronic device comprising: a memory storing instructions; and one or more processors, wherein the instructions, when performed by the one or more processors, cause the one or more processors. In a similar field of endeavor Nagori US2024/035488 discloses an electronic device comprising: a memory storing instructions; and one or more processors, wherein the instructions, when performed by the one or more processors, cause the one or more processors. (see paragraph 51-53 note that object detection is performed using a computer processor and a memory storing instructions). One of ordinary skill in the art could have easily used a processor to perform the method of Wang and Huang and the results would have be predictable. The processor and memory perform the same function separately as they do in combination as the these elements merely used as tools to implement the function. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Wang and Huang with Nagori. Re claim 12 Wang discloses wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features (see figure 3 note that features from the 2d images are input into the depth net) and by inputting, into a BEV pool (see figure 3 voxel pooling), an output of the DepthNet and the 2D image features (see figure 3 note that the product of the 2d features and the depth distribution are multiplied and input into the voxel pooling) Re claim 13 Wang further discloses wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images (see section 3.3.1 and section 3.3.2 page 5 first paragraph note “As showninFig.2, when two cameras with various intrinsic parameters (i.e. focal lengths) shoot the same object at the same distance, the imaging size of the object, which is determined by the intrinsic parameters, can be quite different. If a model is only optimized on the dataset collected from a specific camera, it can be difficult for the model to predict an identical depth for the object pictured from another camera, and thus it is the cause of inaccurate depth prediction when domain shifts, similar in [20]. Furthermore, we empirically find that the estimated depth has been entangled with the intrinsic parameters of the camera, and results in non-compliance with the intuition of “Everything looks small in the distance and big on the contrary”. Hence, we attempt to decouple the estimated depth from the intrinsic parameters” note that the process for computing depth is trained such that it is invariant [i.e. normalized] to the camera parameters this is to remove errors caused by the camera position of the images being captured being different from the one the model trains on). Re claim 17 Wang further discloses wherein the image backbone, the view transformer (see figure 3 portion i and associated caption as well as section 3.3.1 note that intrinsic decoupled depth estimation is performed and intrinsic parameters are used to adapt to the domain during generation of the depth distribution), the BEV encoder, and/or the detection head have respective domain adaptation adapters (see figure 3 portion i and associated caption as well as section 3.3.1 note that intrinsic decoupled depth estimation is performed and intrinsic parameters are used to adapt to the domain during generation of the depth distribution). Re claim 18 Wang further discloses wherein the domain adaptation adapters temporarily supplant layers in the image backbone, the view transformer, the BEV encoder, and/or the detection head, respectively (see figure 3 portion i and associated caption as well as section 3.3.1 note that intrinsic decoupled depth estimation is performed and intrinsic parameters are used to adapt to the domain during generation of the depth distribution this tunes the depth distribution to the particular camera parameters; note that this runs in parallel to some of the operation blocks and supplants the invariant depth distribution with one compensated with the specific camera parameters of the system). Re claim 20 Wang discloses further wherein the 3D feature map is augmented by performing a generalization method of decoupling-based image depth estimation (see figure 3 and associated caption see section 3.3.1 note that depth prediction is intrinsics decoupled to perform domain generalization). Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al “Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View.” Computer Vision and Pattern Recognition March 2023 arXiv:2303.01686 in view of Huang et al “BEVDet: High-Performance Multi-Camera 3D Object Detection in Bird-Eye-View” arXiv:2112.11790 December 2021 in view of Nagori US2024/0354888 in view of Rho et al “ORA3D: Overlap Region Aware Multi-view 3D Object Detection” ARXiv 2207.00865v4 2023. Re claim 16 Wang and Huang do not expressly disclose wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method. ORA3d discloses wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method (see section 3.1. “Learning Depth Cue by Multi-view Stereo Matching. When considering the multi-view camera system, adjacent cameras have a strong association. We regard this association comes from overlap regions and can be extended to geometric guides. To interactively supervise the network, we train the Stereo Disparity Estimation head, which reconstructs a dense disparity map with overlap region pairs of neighboring cameras” The examiner notes that disparity between overlap regions is detected see figure 2 caption “1) Stereo Matching Network for Weak Depth Supervision, where our depth estimation head is trained to predict a dense depth map Of the overlap region” note that the disparity is used to predict a depth map for overlapped images.). The motivation to combine is “ Moreover, objects in the overlap region are often largely occluded or suffer from deformation due to camera distortion, causing a domain shift. To mitigate this issue, we propose using the following two main modules: (1) Stereo Disparity Estimation for Weak Depth Supervision and (2) Adversarial Overlap Region Discriminator. The former utilizes the traditional stereo disparity estimation method to obtain reliable disparity information from the overlap region” (see abstract). One of ordinary skill in the art could have easily modified the teachings of Wang and Huang to incorporate the teachings of Roh to reach the aforementioned advantage. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Wang Huang Nagori and Roh to reach the aforementioned advantage. Claim(s)11-17 and19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang et al “Unified Domain Generalization and Adaptation for Multi-View 3D Object Detection” arXiv:2410.22461 29 October 2024 in view of Nagori US2024/0354888. Re claim 11 Chang discloses ”A method of detecting a three-dimensional (3D) object, the method comprising: extract two-dimensional (2D) image features from images using an image backbone (see figure 3 note that Multiview input images are input into a X6 features used in the view formers also see section 3.1 “ encode 2D visual features alongside the 3D spatial environment into a bird’s eye view (BEV) representation” note these are 2D visual features); extract a 3D feature map, the 3D feature map reflecting depth prediction information, from the 2D image features (see figure 3 view transformer note that the depth net extracts the 3d depth ) by using a view transformer (see abstract and figure 3 note that domain generalization is performed during the process of view transformation); extract a bird's eye view (BEV) feature from the 3D feature map by using a BEV encoder (see section 3.1 “encode 2D visual features alongside the 3D spatial environment into a bird’s eye view (BEV) representation” see also figure 3 note that 2d feature “X6” and the output of the depth net are multiplied and inserted into a BEV pool to form “X1” features where is then input into the BEV back bone and BEV neck which correspond to the BEV encoder.); predict a position of the object and a class of the object from the BEV feature by using a detection head (see abstract note that object detection is performed see figure 3 note that scale orientation and translation are determined which could correspond to the position and the class see also section 3.1 “Subsequently, Detector Head modules D supervises BEV features with 3D labels Y in a three-dimensional manner”). While Chang is clearly intended to be performed on a computer Chang does not expressly disclose an electronic device comprising: a memory storing instructions; and one or more processors, wherein the instructions, when performed by the one or more processors, cause the one or more processors. In a similar field of endeavor Nagori US2024/035488 discloses an electronic device comprising: a memory storing instructions; and one or more processors, wherein the instructions, when performed by the one or more processors, cause the one or more processors. (see paragraph 51-53 note that object detection is performed using a computer processor and a memory storing instructions). One of ordinary skill in the art could have easily used a processor to perform the method of Wang and Huang and the results would have be predictable. The processor and memory perform the same function separately as they do in combination as the these elements merely used as tools to implement the function. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Wang and Huang with Nagori. Re claim 12 Chang discloses wherein the 3D feature map is extracted by a DepthNet predicting a depth output from the 2D image features and by inputting, into a BEV pool, an output of the DepthNet and the 2D image features ( see figure 3 “(see section 3.1 “encode 2D visual features alongside the 3D spatial environment into a bird’s eye view (BEV) representation” see also figure 3 note that 2d feature “X6” and the output of the depth net are multiplied [e.g. outer product] and inserted into a BEV pool to form “X1” features where is then input into the BEV back bone and BEV neck which correspond to the BEV encoder.”) Re claim 13 Chang discloses wherein the view transformer is configured to perform a relative depth normalization method that minimizes depth and position prediction errors caused by a difference in intrinsic/extrinsic parameters of a camera that provided one of the images (see section 3.3 approach whole section “Concretely, we take two advantages leveraging Lp in narrow occluded regions; First, Lp effectively mitigates the triangular misalignment. Second, Lp potentially supports insufficiently scaled Lov. Ultimately, we alleviate perspective view gaps by directly constraining the corresponding depth and the photometric matching between adjacent views.” And “ to achieve generalized BEV extraction, we directly constrain depth estimation network from adjacent overlap regions between multi-view cameras.” Note that in the section transformation matrices are used in the overlap region based on camera parameters K and T this mitigates misalignment for the depth estimation network, note the depth is generalized i.e. normalized) Re claim 14 Chang discloses wherein cameras, including the camera, provide the respective images, and wherein the relative depth normalization method comprises calculating a transformation matrix through which geometric transformation is performed between adjacent pairs of the cameras from the intrinsic/extrinsic parameters and the camera (see section 3.3 approach “To achieve generalized BEV extraction, we directly constrain depth estimation network from adjacent overlap regions between multi-view cameras. Also, we advocate that multi-frame image inputs substantially complement geometric understanding in dynamic scenes with speedy translation and rotation shifts. To this end, we formulate corresponding depth D∗ leveraging spatial and temporal adjacent views. First, we calculate overlap transformation matrices Ti→j” note that transformation matrices are calculated for the overlap region during generalized the depth estimation) Re claim 15 Chang discloses wherein the relative depth normalization method obtains a relative depth after projecting an image feature onto an adjacent image feature by using the depth prediction information and the transformation matrix and minimizing a relative depth loss based on a depth loss function (see section 3.3 approach section in particular equation 3 note that the difference between estimated depth projected using the transformation matrix is used to calculate Lov which penalized unmatched depths between Dj and Di-to-j[projected using the transformation matrix].) Re claim 16 Chang discloses wherein the view transformer is configured to perform a photometric matching method using depth prediction to optimize alignment between an image and an adjacent image, based on the photometric matching method (see section 3.3 approch section especially “Ultimately, we alleviate perspective view gaps by directly constraining the corresponding depth and the photometric matching between adjacent views” note that alignment is optimized using photometric matching). Re claim 17 Chang discloses wherein the image backbone, the view transformer, the BEV encoder, and/or the detection head have respective domain adaptation adapters (see section 3.4 label efficient domain adaptation and Appendix C Searching adapter structures “We explore various modules and structures to find a suitable adapter architecture.Tab.9,10 show which structures and locations affects the model’s performance. For adapter locations, performance is optimal when adapters are attached to all modules.” Note that domain adapter are attached to all modules note that A is built parallel to operation block). Re claim 19 Chang disclose wherein each domain adaptation adapter is configured to perform a skip connection in which features input to the view transformer, the BEV encoder, and/or the detection head are received, operated, and summed to update a gradient (see section 3.4 “Secondly, we fuse each outputs from B, and Adapter by exploiting skip-connections that directly link between the down sampling and up sampling paths. By doing so, these extensible modules allow to capture high-resolution spatial details while reducing network and computational complexity.” Note that that the adapters provide a skip connection see figure 3 and equation 6 note that the input to the operational block is operated on by the adapter and then summed with the output of the operational block). Cited Art The following is a recitation of art considered relevant but not cited in a rejection above. Achim US 12579725 B1 discloses: Determining an absolute depth estimate of a vehicle object in a 2D image captured by a camera includes receiving the 2D image captured by the camera. It further includes determining, at least in part by using a segmentation model, a plurality of pixels corresponding to the vehicle object in the 2D image. It further includes using a prediction model to determine, for each pixel in the plurality of pixels determined at least in part by using the segmentation model, a corresponding camera invariant distance value comprising a predicted height of a cross-section of the vehicle object that is visible within a given pixel. It further includes determining the absolute depth estimate of the vehicle object in the 2D image based on: a focal length of the camera that captured the 2D image; and an aggregation of the predicted heights determined for each of the pixels in the plurality of pixels. (see abstract) Kum et al US 12651445 B2 discloses generate a 3-D feature map that is used to perceive an environment through the fusion of a camera and a radar sensor, the method of perceiving a 3-D environment may include extracting a two-dimensional (2-D) feature map from an image obtained by the camera and transforming the 2-D feature map into a feature map in a 3-D space by using first distance information extracted from the image and second distance information measured by the radar sensor. (see abstract) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN T MOTSINGER whose telephone number is (571)270-1237. The examiner can normally be reached 9AM-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SEAN T MOTSINGER/Primary Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Jan 29, 2025
Application Filed
Sep 09, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743901
INFORMATION PROCESSING DEVICE, INFORMATION PROESSING METHOD, AND NON-TRANSITORY COMPUTER READABLE MEDIUM
4y 0m to grant Granted Sep 22, 2026
Patent 12737864
INSPECTION SYSTEM
2y 6m to grant Granted Sep 15, 2026
Patent 12731425
METHOD OF CLASSIFYING A DOCUMENT FOR A STRAIGHT-THROUGH PROCESSING
2y 11m to grant Granted Sep 08, 2026
Patent 12725437
IMAGE PROCESSING APPARATUS AND IMAGE PROCESSING METHOD
3y 3m to grant Granted Sep 01, 2026
Patent 12725255
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND IMAGE PROCESSING PROGRAM
2y 5m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
90%
With Interview (+11.9%)
2y 11m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 697 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month