DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of Applicant’s claim of priority from the Korean Patent Application No. 10-2023-0178980, filed on December 11, 2023. It is noted, however, that applicant has not filed a certified copy of the prior application as required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statements (“IDS”) filed on 03/31/2025 and 08/12/2024 has been reviewed and the listed references have been considered.
Drawings
The 8-page drawings have been considered and placed on record in the file.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: “depth sensor” in claims 1, 11, 12, 14, 15 and 20 and “imaging device” in claims 10-12 and 19.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, these are being interpreted to cover the corresponding structures described in the applicant’s drawings: schematics depicted in Fig. 5B, and applicant’s specification: ¶0026: “The depth sensor may include a light detection and ranging (lidar)”; and ¶0070: “an image obtained from an image device (e.g., a camera with a color, grayscale, or monochrome visual and/or non-visual spectrum image sensor” as performing the claimed functions, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 3, 4, 7, 8, 10, 11, 13 and 14 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kim et al. (3D Dual-Fusion: Dual-Domain Dual-Query Camera-LiDAR Fusion for 3D Object Detection).
Regarding Claim 1, Kim teaches, An object detection method, (Kim, page 1, col. 1, Abstract: “technique to achieve robust 3D object detection”) the method comprising: extracting voxel features of valid cells from voxel data corresponding to a point cloud (Kim, page 1, col. 2, ¶01: “Voxel-based encoding is a widely used LiDAR encoding method, that voxelizes LiDAR points in 3D space and encodes the points in each voxel”) obtained using a depth sensor; and generating object detection data (Kim, page 1, col. 1, ¶01: “Both cameras and light detection and ranging (LiDAR) sensors provide useful information for 3D object detection”) using a transformer-based model (Kim, page 2 col. 2, section II, ¶01: “Transformer models have been used to encode LiDAR point clouds”) for object detection (Kim, page 1, col. 2, ¶01: “uses both voxel features and camera features together to perform 3D object detection”) based on the extracted voxel features, (Kim, page 3, col. 1, ¶02: “extracts features from data collected by two sensors”) wherein the valid cells comprise cells among a plurality of cells, included in the voxel data that comprises point data, (Kim, page 2 col. 2, section II, ¶01: “The voxel-based backbone partitions the LiDAR points using a voxel or a pillar structure and encodes the points of each grid element”) and wherein a cell among the plurality of cells that does not comprise corresponding point data is not a valid cell. (Kim, page 2, col. 2, ¶02: “We assign dual-queries only for the non empty voxels and the corresponding pixels of the camera features projected from those voxels. Because only a small fraction of voxels are non-empty, the number of queries used for dual-domain interactive feature fusion is much smaller than the size of the entire voxel”).
Regarding claim 3, Kim teaches, The object detection method of claim 1, wherein the transformer-based model comprises a transformer decoder, (Kim, page 4, Fig. 3: “DDA decodes dual-queries to perform dual-domain interactive feature fusion”) and wherein the extracted voxel features are applied to the transformer decoder, that is configured to perform cross attention. (Kim, page 5, col.1, ¶02: “The dual-query cross attention decodes the dual-queries qv and qc over multiple attention layers”).
Regarding claim 4, Kim teaches, The object detection method of claim 1, wherein the transformer-based model comprises a transformer encoder (Kim, page 1, col.2, ¶01: “detect objects based on the features obtained from the encoder”) and a transformer decoder, (Kim, page 4, Fig. 3: “DDA decodes dual-queries to perform dual-domain interactive feature fusion”) and wherein the extracted voxel features are applied to the transformer encoder, that is configured to perform cross attention. (Kim, page 1, col. 2, Fig. 1: “Transformer first performs 3D local self-attention to encode the voxel domain features. Then, it applies dual-query cross attention”).
Regarding claim 7, Kim in view of Li teaches, The object detection method of claim 6, wherein the corresponding point in each of the valid cells is a corresponding center point in each of the valid cells. (Kim, page 4, col.1, ¶01: “c-queries qc = {qc,1, qc,2, ..., qc, Q} are assigned to the image pixels indicated by projecting the center points of the non-empty voxels into the camera domain”).
Regarding claim 8, Kim in view of Li teaches, The object detection method of claim 7, wherein the generating of the position embeddings of the valid cells comprises performing the respective positional encodings of the 3D coordinates (Kim, page 4, col. 2, ¶01: “FFN(·) denotes a position-wise feed forward network, and PE(xi,xj) is the positional encoding function that encodes the difference of two 3D coordinates”) of the corresponding center points by applying, for each of the valid cells, (Kim, page 4, col. 2, ¶03: “a 3D reference point p3D,q = (x,y,z) at the center point of the qth non-empty voxel”) a lowest weight to coordinate of a z-axis of the 3D coordinates of the corresponding center point, which includes coordinates at an x-axis, a y-axis, and the z-axis of the corresponding center point. (Kim, page 6, col. 2, ¶01: “3D Dual-Fusion with CenterPoint baseline is called 3D Dual-Fusion (C) and 3D Dual-Fusion with TransFusion baseline is called 3D Dual-Fusion (T). The range of the point clouds was within [−54,54]×[−54,54]×[−5,3]m on (x,y,z) axis”).
Regarding claim 10, Kim teaches, The object detection method of claim 1, wherein the generating of the object detection data comprises: extracting image features from an image obtained from an imaging device; (Kim, page 3, col. 1, ¶02: “Feature-level fusion extracts features from data collected by two sensors and aggregates them through a domain transformation”) generating respective fusion features corresponding to the valid cells based on the voxel features of the valid cells and corresponding image features among the extracted image features; (Kim, page 2, col. 1, ¶02: “Feature-level fusion methods [9]–[15] are designed to extract semantic features separately from the camera image and LiDAR data and aggregate them in the voxel domain”) and generating the object detection data from the transformer-based model based on the fusion features. (Kim, page 1, col. 1, ¶01: “Camera-LiDAR sensor fusion aims to combine the complementary information provided by the two sensing modalities to achieve robust 3D object detection”).
Regarding claim 11, Kim teaches, The object detection method of claim 10, wherein the generating of the object detection data further comprises: identifying which image features correspond to which valid cells based on a transformation matrix corresponding to the depth sensor and the imaging device. (Kim, page 2, col. 1, ¶02: “extract semantic features separately from the camera image and LiDAR data and aggregate them in the voxel domain. MMF[9], ContFuse [10], PointAugmenting [11], EPNet [12], and 3D-CVF [13] transformed the image features into a voxel domain using a calibration matrix and conducted the element-wise feature fusion”).
Regarding claim 13, Kim teaches, The object detection method of claim 1, wherein the voxel data comprises the plurality of cells obtained by dividing a space corresponding to the point cloud in a grid unit of a predetermined volume. (Kim, page 3, col. 2, ¶02: “We voxelize the LiDAR points using the voxel grid structure of the width Wv, length Lv and height Hv”).
Regarding claim 14, Kim teaches, The object detection method of claim 1, wherein the depth sensor comprises a light detection and ranging (lidar). (Kim, page 1, col. 1, ¶01: “light detection and ranging (LiDAR) sensors provide useful information for 3D object detection”).
Claim 20 is rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li et al. (US 2024/0087222 A1).
Regarding claim 20, An electronic device, (Li, ¶0143: “a client device may be any appropriate electronic and/or computing devices”) comprising: one or more processors configured to execute instructions; and a memory storing the instructions, wherein execution of the instructions configures the one or more processors to: (Li, ¶0059: “the apparatus 200 includes one or more processors 202 in communication with a memory 204. The memory 204 may store parameters 206 to implement the framework”) generate object detection data using a transformer-based model, (Li, ¶0048: “a transformer-based architecture that divides the depth range into bins and adaptively estimates the depth center of each bin based on a monocular image”) configured for object detection, (Li, ¶0064: “The LiDAR sensor 230 may use laser range finding techniques to detect the depth of objects visible to the image sensor”) through provision of tokens that include three-dimensional (3D) space voxel features extracted from 3D point cloud data generated using a depth sensor. (Li, ¶0072: “3D feature map is generated using a second transformer. Initial voxel features are generated by combining the updated query proposals with corresponding mask tokens. The initial voxel features are then refined using the second transformer to generate the 3D feature map (i.e., refined voxel features), which implements a self-attention mechanism”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 2, 5, 6 and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (3D Dual-Fusion: Dual-Domain Dual-Query Camera-LiDAR Fusion for 3D Object Detection) in view of Li et al. (US 2024/0087222 A1).
Regarding claim 2, Kim teaches, The object detection method of claim 1, wherein the voxel features are three-dimensional (3D) features, (Kim, page 6, col.2, ¶01: “The point clouds were voxelized with a voxel sizes of 0.075×0.075×0.2m”). However, Kim does not explicitly teach, wherein the generating of the object detection data includes inputting tokens, which include the voxel features, to the transformer-based model.
In an analogous field of endeavor, Li teaches, wherein the generating of the object detection data includes inputting tokens, which include the voxel features, (Li, ¶0045: “a mask token 106 are then used to populate a voxel feature map, which is further refined by a second transformer that implements a deformable self-attention mechanism to generate a dense voxel feature map”) to the transformer-based model. (Li, ¶0044: “The query proposals Q.sub.p 104 estimate the location of objects that appear in the images 102 within a sparse binary voxel grid”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Kim using the teachings of Li to introduce inputting object feature data to a transformer in form of tokens. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of detecting the object locations in the data. Therefore, it would have been obvious to combine the analogous arts Kim and Li to obtain the invention in claim 2.
Regarding claim 5, Kim teaches, The object detection method of claim 1, wherein the extracting of the voxel features comprises extracting a respective voxel feature of the valid cells, and (Kim, page 7, col.2, ¶01: “fusion method that brings the camera features associated with the non-empty voxels and fuses them with LiDAR voxel features”). However, Kim does not explicitly teach, wherein the generating of the object detection data includes inputting a token corresponding to each of the respective voxel features to the transformer-based model.
In an analogous field of endeavor, Li teaches, wherein the generating of the object detection data includes inputting a token corresponding to each of the respective voxel features (Li, ¶0045: “a mask token 106 are then used to populate a voxel feature map, which is further refined by a second transformer that implements a deformable self-attention mechanism to generate a dense voxel feature map”) to the transformer-based model. (Li, ¶0044: “The query proposals Q.sub.p 104 estimate the location of objects that appear in the images 102 within a sparse binary voxel grid”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Kim using the teachings of Li to introduce inputting respective voxel features into a transformer. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of detecting the respective object locations. Therefore, it would have been obvious to combine the analogous arts Kim and Li to obtain the invention in claim 5.
Regarding claim 6, Kim teaches, The object detection method of claim 1, further comprising: generating position embeddings of the valid cells by performing respective positional encodings (Kim, page 5, col.1, ¶01: “depth-aware positional embedding is then added to both queries”) of three-dimensional (3D) coordinates of a corresponding point in each of the valid cells, and (Kim, page 4, col.2, ¶01: “positional encoding function that encodes the difference of two 3D coordinates”). However, Kim does not explicitly teach, wherein the generating of the object detection data comprises: generating the object detection data based on the position embeddings and the extracted voxel features.
In an analogous field of endeavor, Li teaches, wherein the generating of the object detection data (Li, ¶0044: “transformer-based framework 100, which processes the images 102 using a depth estimation network to generate a set of query proposals”) comprises: generating the object detection data based on the position embeddings and the extracted voxel features. (Li, ¶0044: “The query proposals Q.sub.p 104 estimate the location of objects that appear in the images 102 within a sparse binary voxel grid”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Kim using the teachings of Li to introduce generating positional information of an object. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of locating an object using positional data in a voxel grid. Therefore, it would have been obvious to combine the analogous arts Kim and Li to obtain the invention in claim 6.
Regarding claim 15, it recites am apparatus with elements corresponding to the steps of the method recited in claim 1. Therefore, the recited elements of apparatus claim 15 are mapped to the proposed combination in the same manner as the corresponding steps in method claim 1. However, Kim does not explicitly teach, An apparatus, comprising: one or more processors configured to execute instructions; and a memory storing the instructions, wherein execution of the instructions configures the one or more processors to.
In an analogous field of endeavor, Li teaches, An apparatus, comprising: one or more processors configured to execute instructions; and a memory storing the instructions, wherein execution of the instructions configures the one or more processors to (Li, ¶0059: “the apparatus 200 includes one or more processors 202 in communication with a memory 204. The memory 204 may store parameters 206 to implement the framework”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Kim using the teachings of Li to introduce an apparatus including a processor and a memory. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of automatically performing the object detection method on the apparatus. Therefore, it would have been obvious to combine the analogous arts Kim and Li to obtain the invention in claim 15.
Regarding claim 16, it recites am apparatus with elements corresponding to the steps of the method recited in claim 3. Therefore, the recited elements of apparatus claim 16 are mapped to the proposed combination in the same manner as the corresponding steps in method claim 3.
Regarding claim 17, it recites am apparatus with elements corresponding to the steps of the method recited in claim 4. Therefore, the recited elements of apparatus claim 17 are mapped to the proposed combination in the same manner as the corresponding steps in method claim 4.
Regarding claim 18, it recites am apparatus with elements corresponding to the steps of the method recited in claim 6. Therefore, the recited elements of apparatus claim 18 are mapped to the proposed combination in the same manner as the corresponding steps in method claim 6. Additionally, the rationale and motivation to combine Kim and Li presented in rejection of claim 6, apply to this claim.
Regarding claim 19, it recites am apparatus with elements corresponding to the steps of the method recited in claim 10. Therefore, the recited elements of apparatus claim 19 are mapped to the proposed combination in the same manner as the corresponding steps in method claim 10.
Claims 9 is rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (3D Dual-Fusion: Dual-Domain Dual-Query Camera-LiDAR Fusion for 3D Object Detection) in view of Li et al. (US 2024/0087222 A1) and in further view of Zhao (US 2022/0122321 A1).
Regarding claim 9, Kim in view of Li teaches, The object detection method of claim 6, wherein the generating of the position embeddings of the valid cells comprises. However, the combination of Kim and Li does not explicitly teach, generating the position embeddings of the valid cells by performing the respective positional encodings of the 3D coordinates of the corresponding point in the valid cells based on a sign function.
In another analogous field of endeavor, Zhao teaches, generating the position embeddings of the valid cells by performing the respective positional encodings of the 3D coordinates of the corresponding point in the valid cells based on a sign function. (Zhao, ¶0030: “a sign distance function (SDF): SDF(x)=z(x)−D(u), where u represents a pixel point of an image and X represents a 3D coordinate point in the TSDF model corresponding to u. If there is a dynamic object in the original image, the SDF value of the original image data differs greatly from a SDF value stored in a corresponding point in the TSDF model”).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Kim in view of Li using the teachings of Zhao to introduce a sign distance function. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of detecting an object in the image data based on a change in sign distance function value. Therefore, it would have been obvious to combine the analogous arts Kim, Li and Zhao to obtain the invention in claim 9.
Claims 12 is rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (3D Dual-Fusion: Dual-Domain Dual-Query Camera-LiDAR Fusion for 3D Object Detection) in view of Wang et al. (US 2024/0127534 A1).
Regarding claim 12, Kim teaches, The object detection method of claim 11. However, Kim does not explicitly teach, wherein the transformation matrix corresponding to the depth sensor and the imaging device is determined based on a relative position relationship between the depth sensor and the imaging device.
In an analogous field of endeavor, Wang teaches, wherein the transformation matrix corresponding to the depth sensor and the imaging device is determined (Wang, ¶0081: “calibration data 710 includes transformation matrices (K, R, t) between point clouds generated by LiDAR sensors (e.g., LiDAR sensors 202b of FIG. 2) and camera images captured by cameras”) based on a relative position relationship between the depth sensor and the imaging device. (Wang, ¶0081: “LiDAR sensors and cameras are installed at different positions in AV. The transformation matrices provide information indicating relative positions of LiDAR sensors and cameras).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify Kim using the teachings of Wang to introduce a positional transformation matrix. A person skilled in the art would be motivated to combine the known elements as described above and achieve the predictable result of determining the positional relationship between a camera and a lidar sensor. Therefore, it would have been obvious to combine the analogous arts Kim and Wang to obtain the invention in claim 12.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MEHRAZUL ISLAM whose telephone number is (571)270-0489. The examiner can normally be reached Monday-Friday: 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Saini Amandeep can be reached on (571) 272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MEHRAZUL ISLAM/Examiner, Art Unit 2662
/AMANDEEP SAINI/Supervisory Patent Examiner, Art Unit 2662