Prosecution Insights
Last updated: October 02, 2026
Application No. 19/003,086

Three-dimensional (3D) Object Detection Method, Apparatus, Controller, Vehicle, and Medium

Non-Final OA §101§102§103§112
Filed
Dec 27, 2024
Priority
Dec 29, 2023 — CN 2023 1184 8053.4
Examiner
PHAM, NHUT HUY
Art Unit
Tech Center
Assignee
Robert Bosch GmbH
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
64 granted / 78 resolved
+22.1% vs TC avg
Strong +24% interview lift
Without
With
+23.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
15 currently pending
Career history
91
Total Applications
across all art units

Statute-Specific Performance

§101
8.9%
-31.1% vs TC avg
§103
60.3%
+20.3% vs TC avg
§102
14.2%
-25.8% vs TC avg
§112
14.5%
-25.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 78 resolved cases

Office Action

§101 §102 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION The United States Patent & Trademark Office appreciates the application that is submitted by the inventor/assignee. The United States Patent & Trademark Office reviewed the following application and has made the following comments below. Priority This application claims benefit of foreign priority under 35 U.S.C. 119(a)-(d) of: CN2023 1184 8053.4, filed in China on 12/29/2023. Copies of certified papers required by 37 CFR 1.55 have been retrieved. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “an obtaining unit”, “a generation unit” and “a detection unit” in claim 9. The interpretation for these limitations is a computing system and equivalent thereof. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 5-7 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The examiner strongly suggested that appropriate corrections be made to clarify the claim scope. With respect to Claim 5, the claim recites the following, each of which renders the claim indefinite: “the feature diagram” on line 5 (unclear antecedent basis). Claims 6-7 are also rejected due to their dependence on rejected independent claim 5. USC § 101 Consideration Regarding Claim 12, the Examiner has reviewed Applicant’s specification, paragraph 75 (“The computer-readable storage medium may be a tangible device that maintains and stores instructions used to instruct execution devices … The computer-readable storage medium used herein is not to be construed as transient signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires”) and determined that the claim does not contain transitory signal. Thus, a 101 rejection is not necessary. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-2 and 9-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. When reviewing independent claim 1, and based upon consideration of all of the relevant factors with respect to the claim as a whole, claim(s) 1 are held to claim an abstract idea without reciting elements that amount to significantly more than the abstract idea and is/are therefore rejected as ineligible subject matter under 35 U.S.C. 101. The Examiner will analyze Claim 1, and similar rationale applies to independent claim/s 15. The rationale, under MPEP § 2106, for this finding is explained below: The claimed invention (1) must be directed to one of the four statutory categories, and (2) must not be wholly directed to subject matter encompassing a judicially recognized exception, as defined below. The following two step analysis is used to evaluate these criteria. Step 1: Is the claim directed to one of the four patent-eligible subject matter categories: process, machine, manufacture, or composition of matter? When examining the claim under 35 U.S.C. 101, the Examiner interprets that the claims is related to a process since the claim is directed to a method. Step 2a, Prong 1: Does the claim wholly embrace a judicially recognized exception, which includes laws of nature, physical phenomena, and abstract ideas, or is it a particular practical application of a judicial exception? YES, the claims are directed toward a mental process (i.e., abstract idea). With regard to STEP 2A (PRONG 1), the guidelines provide three groupings of subject matter that are considered abstract ideas: Mathematical concepts – mathematical relationships, mathematical formulas or equations, mathematical calculations; Certain methods of organizing human activity – fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions); and Mental processes – concepts that are practicably performed in the human mind (including an observation, evaluation, judgment, opinion). The method in claim 1 comprise a mental process that can be practicably performed in the human mind therefore, an abstract idea. Claim 1 recites: detecting a 3D object in the target 3D scene based on the 2D image and the depth data (a human can review RGB image and radar image of a scene, detecting a 3D object in the scene by drawing an outline for an image region associated with an object, using a pen and paper, as a mental process as an abstract idea); These limitations, as drafted, is a simple process that, under their broadest reasonable interpretation, covers performance of the limitations in the mind or by a human. The Examiner notes that under MPEP 2106.04(a)(2)(III), the courts consider a mental process (thinking) that “can be performed in the human mind, or by a human using a pen and paper" to be an abstract idea. CyberSource Corp. v. Retail Decisions, Inc., 654 F.3d 1366, 1372, 99 USPQ2d 1690, 1695 (Fed. Cir. 2011). As the Federal Circuit explained, "methods which can be performed mentally, or which are the equivalent of human mental work, are unpatentable abstract ideas the ‘basic tools of scientific and technological work’ that are open to all.’" 654 F.3d at 1371, 99 USPQ2d at 1694 (citing Gottschalk v. Benson, 409 U.S. 63, 175 USPQ 673 (1972)). See also Mayo Collaborative Servs. v. Prometheus Labs. Inc., 566 U.S. 66, 71, 101 USPQ2d 1961, 1965 ("‘[M]ental processes[] and abstract intellectual concepts are not patentable, as they are the basic tools of scientific and technological work’" (quoting Benson, 409 U.S. at 67, 175 USPQ at 675)); Parker v. Flook, 437 U.S. 584, 589, 198 USPQ 193, 197 (1978) (same). The courts do not distinguish between mental processes that are performed entirely in the human mind and mental processes that require a human to use a physical aid (e.g., pen and paper or a slide rule) to perform the claim limitation. See, e.g., Benson, 409 U.S. at 67, 65, 175 USPQ at 674-75, 674 (noting that the claimed "conversion of [binary-coded decimal] numerals to pure binary numerals can be done mentally," i.e., "as a person would do it by head and hand."); Synopsys, Inc. v. Mentor Graphics Corp., 839 F.3d 1138, 1139, 120 USPQ2d 1473, 1474 (Fed. Cir. 2016) (holding that claims to a mental process of "translating a functional description of a logic circuit into a hardware component description of the logic circuit" are directed to an abstract idea, because the claims "read on an individual performing the claimed steps mentally or with pencil and paper"). Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer. As the Federal Circuit has explained, "[c]ourts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind." Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015). See also Intellectual Ventures I LLC v. Symantec Corp., 838 F.3d 1307, 1318, 120 USPQ2d 1353, 1360 (Fed. Cir. 2016) (‘‘[W]ith the exception of generic computer-implemented steps, there is nothing in the claims themselves that foreclose them from being performed by a human, mentally or with pen and paper.’’); Mortgage Grader, Inc. v. First Choice Loan Servs. Inc., 811 F.3d 1314, 1324, 117 USPQ2d 1693, 1699 (Fed. Cir. 2016) (holding that computer-implemented method for "anonymous loan shopping" was an abstract idea because it could be "performed by humans without a computer"). Because both product and process claims may recite a "mental process", the phrase "mental processes" should be understood as referring to the type of abstract idea, and not to the statutory category of the claim. The courts have identified numerous product claims as reciting mental process-type abstract ideas, for instance the product claims to computer systems and computer-readable media in Versata Dev. Group. v. SAP Am., Inc., 793 F.3d 1306, 115 USPQ2d 1681 (Fed. Cir. 2015). As such, a person could identify an object in RGB image and radar image of a scene. The mere nominal recitation that the various steps are being executed by a device/in a device (e.g. processing unit) does not take the limitations out of the mental process grouping. Thus, the claims recite a mental process. If a claim limitation, under its broadest reasonable interpretation, covers performance of a mental step which could be performed with a simple tool such as a pen and paper, then it falls within the “mental steps” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2a, Prong 2: Does the claim recite additional elements that integrate the judicial exception into a practical application? NO, the claims do not recite additional elements that integrate the judicial exception into a practical application. With regard to STEP 2A (prong 2), whether the claim recites additional elements that integrate the judicial exception into a practical application, the guidelines provide the following exemplary considerations that are indicative that an additional element (or combination of elements) may have integrated the judicial exception into a practical application: an additional element reflects an improvement in the functioning of a computer, or an improvement to other technology or technical field; an additional element that applies or uses a judicial exception to affect a particular treatment or prophylaxis for a disease or medical condition; an additional element implements a judicial exception with, or uses a judicial exception in conjunction with, a particular machine or manufacture that is integral to the claim; an additional element effects a transformation or reduction of a particular article to a different state or thing; and an additional element applies or uses the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception. While the guidelines further state that the exemplary considerations are not an exhaustive list and that there may be other examples of integrating the exception into a practical application, the guidelines also list examples in which a judicial exception has not been integrated into a practical application: an additional element merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea; an additional element adds insignificant extra-solution activity to the judicial exception; and an additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use. Claims 1-2 and 9-12 do not recite any of the exemplary considerations that are indicative of an abstract idea having been integrated into a practical application. The steps “detecting a 3D object in the target 3D scene based on the 2D image and the depth data using a predetermined neural network model”, amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application. Please See MPEP 2106.05(f). The steps “obtaining a two-dimensional (2D) image of a target 3D scene;” and “obtaining depth data corresponding to the 2D image;” merely constitutes activity involving data gathering and data outputting. Such extra-solution activity does not integrate the abstract idea into a practical application. Please see MPEP §2106.05(g). Claim 2 recites “2D object detection neural network” and “3D object-detection task head neural network model” amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)). Claim 9 recites “An apparatus”, “an obtaining unit”, “a generation unit” and “a detection unit” amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)). Claim 10 recites: “A controller”, “at least one processor”, and “a memory, coupled to the at least one processor, and having instructions stored thereon …” amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)). Claim 12 recites: “A computer-readable storage medium having stored thereon computer-executable instructions …” amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)). These limitations are recited at a high level of generality (i.e. as a general action or change being taken based on the results of the acquiring step) and amounts to mere post solution actions, which is a form of insignificant extra-solution activity. Further, the claims are claimed generically and are operating in their ordinary capacity such that they do not use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. Accordingly, even in combination, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. Step 2b: If a judicial exception into a practical application is not recited in the claim, the Examiner must interpret if the claim recites additional elements that amount to significantly more than the judicial exception. With regard to STEP 2B, whether the claims recite additional elements that provide significantly more than the recited judicial exception, the guidelines specify that the pre-guideline procedure is still in effect. Specifically, that examiners should continue to consider whether an additional element or combination of elements: adds a specific limitation or combination of limitations that are not well-understood, routine, conventional activity in the field, which is indicative that an inventive concept may be present; or simply appends well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, which is indicative that an inventive concept may not be present. With regard to (2b) the Guidance provided the following examples of limitations that may be enough to qualify as “significantly more" when recited in a claim with a judicial exception: Improvement to another technology or technical field Improvement to functioning of computer itself and/or applying the judicial exception with, or by use of, a particular machine Effecting a transformation or reduction of a particular article to a different state or thing. Adding a specific limitation other that what is well understood, routine and conventional in the field, or adding unconventional steps that confine the claim to a particular useful application Meaningful limitation beyond generally linking the use of an abstract idea to a particular technological environment. The Guidance further set forth limitations that were found not to be enough to qualify as “significantly more” when recited in a claim with a judicial exception include: Adding words to “apply it” (or an equivalent) with the judicial exception or mere instructions to implement abstract ideas on a computer Simply appending well-understood, routine and conventional activities previously known to the industry specified at a high level of generality to the judicial exception, e.g. a claim to an abstract idea requiring no more than a generic Computer to perform generic computer functions that are well -understood, routine and conventional activities previously known to the industry. Adding insignificant extra-solution activity to the judicial exception, e.g. mere data gathering in conjunction with a law of nature or abstract idea Generally linking the use of the judicial exception to a particular technological environment or field of use. Claims 1-2 and 9-12 do not recite any additional elements that are not well-understood, routine or conventional. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The above identified additional computer components, using instructions to apply the judicial exception, are merely generic computer components that are well-known, routine, and conventional as is evidenced by Bancorp Services v. Sun Life (Fed. Cir. 2012) and Alice Corp. v. CLS Bank (2014). Thus, since claims 1 and 9 are: (a) directed toward an abstract idea, (b) do not recite additional elements that integrate the judicial exception into a practical application, and (c) do not recite additional elements that amount to significantly more than the judicial exception, claims 1 and 15 are not eligible subject matter under 35 U.S.C 101. Similar analysis is made for the dependent claims 2 and 10-12 and the dependent claims are similarly identified as: being directed towards an abstract idea, not reciting additional elements that integrate the judicial exception into a practical application, and not reciting additional elements that amount to significantly more than the judicial exception. Regarding Claims 3-8, The examiner finds that the disclosed improvement to the technological field of camera-radar fusion for object detection is set forth in the originally filed specification for example in pages 8-9, and requires the 3D object detection limitations to realize the improvement, with these limitations being recited in various scope in claims 3-8. Accordingly, 3-8 are found to be subject matter eligible under 35 U.S.C. 101 and are not rejected therefor. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1 and 9-12 are rejected under 35 U.S.C. 102 (a) (1) & (2) as being anticipated by Goel et al. (US20230033177A1, hereinafter Goel). CLAIM 1 Regarding Claim 1, Goel teaches a method for three-dimensional (3D) object detection (Goel, Abstract: “Techniques are discussed herein for generating three-dimensional (3D) representations of an environment based on two-dimensional (2D) image data, and using the 3D representations to perform 3D object detection and other 3D analyses of the environment”; see method of claim 6), comprising: obtaining a two-dimensional (2D) image of a target 3D scene (Goel, ¶ [0018]: “ the image-based object detector 102 (e.g., within a vehicle) receives image data 106 associated with an environment. In some examples, vehicle sensors may be configured to capture the image data 106 via one or more cameras (or image sensors) on the vehicle, including but not limited to red-green-blue (RGB) cameras … as the vehicle traverses through an environment, the image sensors can capture image data 106 associated with the environment, and provide the image data to the image-based object detector 102” Goel teaches capturing 2D image of surrounding environment of a vehicle); obtaining depth data corresponding to the 2D image (Goel, ¶ [0020]: “At operation 108, the image-based object detector 102 receives depth data 110 associated with the image data …”); and detecting a 3D object in the target 3D scene (Goel, ¶ [0027-0030]: “perform various different 3D object detection techniques in different examples. Different 3D object detection techniques may be performed based on various different inputs. For example, scene data associated with an environment may be stored as one or a combination of a point cloud, a multi-channel image, a 3D grid, and/or various other 2D or 3D representations/views of the environment. The scene data associated with an environment may be provided as input to various 3D object detection techniques, including techniques configured to receive 3D scene data representations such as point clouds” Goel teaches performing 3D object detection for surrounding environment, based on input of 3D point cloud data generated from image data and associated depth data) based on the 2D image and the depth data (Goel, ¶ [0023-0025]: “generate one or more 3D representations, based on the image data 106 and the associated depth data 110. In various examples, the 3D representation(s) generated by the image-based object detector 102 may include 3D point clouds and/or 3D grids” Goel teaches generating 3D point cloud data based on image data and associated depth data) using a predetermined (***The Examiner interprets a predetermined neural network model is a pre-trained neural network model) neural network model (Goel, ¶ [0027, 0043, 0045 and 0055]: “image-based object detector 102 may perform one or more 3D object detection techniques using a 3D point cloud representation generated … 3D object detection techniques performed based on a point cloud may include executing neural network 118 (e.g., a 3D convolutional neural network (CNN)) configured to perform object detection techniques based on a 3D point cloud provided as input to the neural network 118 … the neural network 118 may be designed or trained to receive lidar or radar point clouds” Goel teaches using a trained 3D CNN for 3D object detection). CLAIM 9 Regarding Claim 19, Goel teaches an apparatus for three-dimensional (3D) object detection (Goel, Abstract: “Techniques are discussed herein for generating three-dimensional (3D) representations of an environment based on two-dimensional (2D) image data, and using the 3D representations to perform 3D object detection and other 3D analyses of the environment”; ¶ [0062-0063]: “vehicle computing device”), comprising: an obtaining unit (Goel, ¶ [0063]: “vehicle computing device(s) 1004 can include one or more processors 1016 and memory … an image-based object detector 1024, one or more trained models/networks …”) configured to obtain a two-dimensional (2D) image of a target 3D scene (Goel, ¶ [0018]: “ the image-based object detector 102 (e.g., within a vehicle) receives image data 106 associated with an environment. In some examples, vehicle sensors may be configured to capture the image data 106 via one or more cameras (or image sensors) on the vehicle, including but not limited to red-green-blue (RGB) cameras … as the vehicle traverses through an environment, the image sensors can capture image data 106 associated with the environment, and provide the image data to the image-based object detector 102” Goel teaches capturing 2D image of surrounding environment of a vehicle); a generation unit (Goel, ¶ [0063]: “vehicle computing device(s) 1004 can include one or more processors 1016 and memory … an image-based object detector 1024, one or more trained models/networks …”) configured to obtain depth data corresponding to the 2D image (Goel, ¶ [0020]: “At operation 108, the image-based object detector 102 receives depth data 110 associated with the image data …”); and a detection unit (Goel, ¶ [0063]: “vehicle computing device(s) 1004 can include one or more processors 1016 and memory … an image-based object detector 1024, one or more trained models/networks …”) configured to detect a 3D object in the target 3D scene (Goel, ¶ [0027-0030]: “perform various different 3D object detection techniques in different examples. Different 3D object detection techniques may be performed based on various different inputs. For example, scene data associated with an environment may be stored as one or a combination of a point cloud, a multi-channel image, a 3D grid, and/or various other 2D or 3D representations/views of the environment. The scene data associated with an environment may be provided as input to various 3D object detection techniques, including techniques configured to receive 3D scene data representations such as point clouds” Goel teaches performing 3D object detection for surrounding environment, based on input of 3D point cloud data generated from image data and associated depth data) based on the 2D image and the depth data (Goel, ¶ [0023-0025]: “generate one or more 3D representations, based on the image data 106 and the associated depth data 110. In various examples, the 3D representation(s) generated by the image-based object detector 102 may include 3D point clouds and/or 3D grids” Goel teaches generating 3D point cloud data based on image data and associated depth data) using a predetermined (***The Examiner interprets a predetermined neural network model is a pre-trained neural network model) neural network model (Goel, ¶ [0027, 0043, 0045 and 0055]: “image-based object detector 102 may perform one or more 3D object detection techniques using a 3D point cloud representation generated … 3D object detection techniques performed based on a point cloud may include executing neural network 118 (e.g., a 3D convolutional neural network (CNN)) configured to perform object detection techniques based on a 3D point cloud provided as input to the neural network 118 … the neural network 118 may be designed or trained to receive lidar or radar point clouds” Goel teaches using a trained 3D CNN for 3D object detection). CLAIM 10 Regarding Claim 10, Goel teaches the method of Claim 1. In addition, Goel teaches a controller (Goel, ¶ [0069]: “ the vehicle computing device(s) 1004 can include one or more system controllers 1030, which can be configured to control steering, propulsion, braking, safety, emitters, communication, and other systems of the vehicle 1002”), comprising: at least one processor; and a memory, coupled to the at least one processor (Goel, ¶ [0063]: “The vehicle computing device(s) 1004 can include one or more processors 1016 and memory 1018 communicatively coupled with the one or more processors”), and having instructions stored thereon that, when executed by the at least one processor, cause the controller to perform the method according to claim 1. (Goel, ¶ [0084]: “Memory 1018 and 1042 are examples of non-transitory computer-readable media. The memory 1018 and 1042 can store an operating system and one or more software applications, instructions, programs, and/or data to implement the methods described herein and the functions attributed to the various systems”, see the rejection of claim 1) CLAIM 11 Regarding Claim 11, Goel teaches the controller of claim 10. In addition, Goel teaches A vehicle (Goel, ¶ [0062]: “the system 1000 can include a vehicle 1002, which can correspond to an autonomous or semi-autonomous vehicle configured to perform object perception and prediction functionality, route planning and/or optimization”) comprising the controller of claim 10 (Goel, ¶ [0062]: “the vehicle 1002 can include vehicle computing device”, ¶ [0069]: “ the vehicle computing device(s) 1004 can include one or more system controllers 1030, which can be configured to control steering, propulsion, braking, safety, emitters, communication, and other systems of the vehicle 1002”) and the predetermined neural network model. (Goel, ¶ [0061]: “ The vehicle 1002 may include components configured to perform image-based 3D object detection”, Goel, ¶ [0063]: “vehicle computing device(s) 1004 can include one or more processors 1016 and memory … an image-based object detector 1024, one or more trained models/networks …”) CLAIM 12 Regarding Claim 12, Goel teaches the method of claim 1. In addition, Goel teaches a computer-readable storage medium (Goel, ¶ [0063]: “The vehicle computing device(s) 1004 can include one or more processors 1016 and memory 1018 communicatively coupled with the one or more processors”) having stored thereon computer-executable instructions, wherein the computer-executable instructions are executed by the processor to perform the method according to claim 1. (Goel, ¶ [0084]: “Memory 1018 and 1042 are examples of non-transitory computer-readable media. The memory 1018 and 1042 can store an operating system and one or more software applications, instructions, programs, and/or data to implement the methods described herein and the functions attributed to the various systems”, see the rejection of claim 1.) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. CLAIM 2 Claim(s) 2-3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Goel in view of Sengupta et al. (Sengupta, Arindam, Atsushi Yoshizawa, and Siyang Cao. "Automatic radar-camera dataset generation for sensor-fusion applications." IEEE, published 2022, hereinafter Sengupta). In regards to Claim 2, Goel teaches the method of Claim 1. In addition, Goel teaches the predetermined neural network model comprises a first neural network model (Goel, ¶ [0020 and 0052]: “0020: the image-based object detector 102 may perform initial 2D object detection techniques on the image data 106 (e.g., algorithms, trained neural networks or machine-learned models), to identify one or more regions of interest in the image data containing objects such as other vehicles, bicycles, pedestrians, traffic signs/signals, etc; 0052: the image-based object detector 102 may execute one or more 2D object detection and/or instance segmentation algorithms (or models and/or network) on the 2D image data, to identify the presence of various objects of interest within the environment. An object of interest may include, for example, another vehicle, bicycle, pedestrian, etc” Goel teaches a neural network object detector 102 may include a sub-network to perform 2D object detection) and a second neural network model (Goel, ¶ [0027, 0043, 0045 and 0055]: “image-based object detector 102 may perform one or more 3D object detection techniques using a 3D point cloud representation generated … 3D object detection techniques performed based on a point cloud may include executing neural network 118 (e.g., a 3D convolutional neural network (CNN)) configured to perform object detection techniques based on a 3D point cloud provided as input to the neural network 118 … the neural network 118 may be designed or trained to receive lidar or radar point clouds” Goel teaches using a trained 3D CNN for 3D object detection), the first neural network model is obtained based on a 2D object-detection neural network model (Goel, ¶ [0020 and 0052]. Goel teaches trained neural network, model for 2D object detection, see above), the second neural network model being a 3D object-detection task head neural network model (Goel, ¶ [0027, 0043, 0045 and 0055]. Goel teaches a trained 3D CNN for 3D object detection, see above), Goel does not explicitly disclose the 2D object detection neural network model is a neural network model having a 2D anchor for single-stage 2D object detection. Sengupta is in the same field of art of camera-radar fusion for object detection. Further, Sengupta teaches the 2D object detection neural network model is a neural network model having a 2D anchor for single-stage 2D object detection. (Sengupta, Abstract: “This paper presents a novel approach that leverages YOLOv3 based highly accurate object detection from camera to automatically label point cloud data obtained from a co-calibrated radar sensor to generate labeled radar-image and radar-only data-sets to aid learning algorithms for different applications” The Examiner notes YOLOv3 is single stage and use anchor boxes, see a Google Overview of YOLOv3 below; see the document Redmon for original paper of YOLOv3 in the Pertinent Arts section) PNG media_image1.png 299 788 media_image1.png Greyscale Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Goel by incorporating single shot detector YOLOv3 that is taught by Sengupta, to make a system to perform 2D object detection by using a single shot detector network; thus, one of ordinary skilled in the art would be motivated to combine the references since among its several aspects, the present invention recognizes there is a need to accuracy, efficiency and reduce overhead (Sengupta, page 2882, section V: “ The proposed approach is >97% accurate, efficient and extremely easy to implement …”). Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. CLAIM 3 Regarding Claim 3, the combination of Goel and Sengupta teaches the method of claim 2. In addition, the combination of Goel and Pang teaches obtaining the first neural network model (Sengupta, page 2881, left col, 2nd paragraph: “The model was trained using an Adam optimizer with the objective to minimize the sparse categorical cross-entropy. The SFFR based approach yielded a test accuracy of 95% for the vehicle class and >99% for the pedestrian class. We also explored the advantage offered by the SFFR over individual sensors data” Sengupta teaches training his SFFR model) based on the 2D object detection neural network model (Sengupta, pages 2878-2879, section C. Sensor-Fusion Data-Set Generation, Algorithm 1. Sengupta’s model is based on the 2D object detection neural network YOLOv3); obtaining predetermined depth data (Sengupta, pages 2878-2879, section C. Sensor-Fusion Data-Set Generation, Algorithm 1, see input Datarad which denotes the input radar data; pages 2877-2878, sections A and B, see FIG. 2, the radar data is first collected by a setup of 2 mmWave radars); clustering the predetermined depth data to cluster the predetermined depth data as a plurality of clusters (Sengupta, pages 2878-2879, section C. Sensor-Fusion Data-Set Generation: “In each of the radar frames, DBSCAN is performed on the 3-D PCL data to segregate multiple target clusters.”; see Algorithm, lines 5-6. Sengupta teaches clustering radar points into clusters using DBSCAN (Density-Based Spatial Clustering of Applications with Noise)); and obtaining the first neural network model (Sengupta, page 2881, left col, 2nd paragraph: “The model was trained using an Adam optimizer with the objective to minimize the sparse categorical cross-entropy. The SFFR based approach yielded a test accuracy of 95% for the vehicle class and >99% for the pedestrian class. We also explored the advantage offered by the SFFR over individual sensors data” Sengupta teaches training his SFFR model) by setting the 2D anchor into each cluster of the plurality of clusters to obtain a model anchor for each cluster (Sengupta, pages 2878-2879, section C. Sensor-Fusion Data-Set Generation: “The pixel indices of the YOLO centroids and the projected radar cluster centroids are then subjected to the Hungarian Algorithm for intra-frame radar-to-image association. For every associated YOLO-cluster pair (i) the image-region inside the bounding box is cropped and reshaped to a 64×64×3 png file (CropImg), and (ii) the X,Y,Z, Doppler and SNR of all the points from the radar cluster, are saved to disk as Numpy arrays, with the YOLO class as the label”, see Algorithm 1. Sengupta teaches associating bounding boxes from YOLO with the projected radar cluster, to obtained a YOLO-cluster pair, see annotated FIG. 3 below, bottom-right picture), wherein the setting is such that for each cluster, the 2D anchor is associated with a center value of the predetermined depth data in the cluster to obtain a model anchor for the cluster. (Sengupta, pages 2878-2879, section C. Sensor-Fusion Data-Set Generation, Algorithm 1, see above. Sengupta teaches each cluster’s centroid, a central value of each radar cluster, is associated with the corresponding YOLO bounding box; A YOLO-radar centroid pair is obtained as the result of the association, see annotated FIG. 3 below, bottom-right picture) PNG media_image2.png 568 776 media_image2.png Greyscale Allowable Subject Matter Claims 4-8 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Pertinent Arts The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Gupta et al. (US-20250095345-A1) which is directed to two-stage three-dimensional (3D) object detection, which includes a distinct fusion of radar data and camera data. The radar data includes a four-dimensional (4D) millimeter-wave (MMW) radar point cloud, and the camera data includes a high-resolution image in the two-dimensional space (2D). Thereafter, a 3D ROI proposal is fused with 2D image data generating a 2D proposal projection. The 2D proposal projection comprises proposals that predict the position of objects in the high-resolution image. Redmon et al. (Redmon, Joseph, and Ali Farhadi. "Yolov3: An incremental improvement." arXiv) which is directed to the 2D object detection network YOLOv3 used in Sengupta. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NHUT HUY (JEREMY) PHAM whose telephone number is (703)756-5797. The examiner can normally be reached Mo - Fr. 8:30am - 6pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O'Neal Mistry can be reached on (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NHUT HUY PHAM/Examiner, Art Unit 2674 /ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674
Read full office action

Prosecution Timeline

Dec 27, 2024
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743877
TRAINING DATASET AUGMENTATION METHOD AND SYSTEM FOR TRAINING DEEP LEARNING NETWORK
2y 8m to grant Granted Sep 22, 2026
Patent 12744876
CROSS-VIEW ATTENTION FOR VISUAL PERCEPTION TASKS USING MULTIPLE CAMERA INPUTS
3y 0m to grant Granted Sep 22, 2026
Patent 12737844
METHOD, APPARATUS AND SOFTWARE PROGRAM FOR INCREASING RESOLUTION IN MICROSCOPY
3y 10m to grant Granted Sep 15, 2026
Patent 12737883
PORTABLE DEVICE FOR ENUMERATION AND SPECIATION OF FOOD ANIMAL PARASITES
2y 7m to grant Granted Sep 15, 2026
Patent 12731236
METHOD OF MEASURING STRUCTURE DISPLACEMENT BASED ON FUSION OF ASYNCHRONOUS VISION MEASUREMENT DATA OF NATURAL TARGET AND ACCELERATION DATA OF STRUCTURE, AND SYSTEM FOR THE SAME
2y 11m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
99%
With Interview (+23.5%)
2y 10m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 78 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month