DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 9/26/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fuxman et al (US20210264203) in view of Balachandran et al (US20250029355).
Regarding claim 1, Fuxman teaches a processor-implemented method, the method comprising:
text-guided training (figs. 1-2; para. [0040]) using a pre-trained text-guided model (122 in fig. 1; 204 in fig. 2; para. [0032]) and an image feature extractor based on one or more text inputs and one or more image inputs corresponding to the one or more text inputs (124 in fig. 1; 206 in fig. 2; para. [0035]);
wherein the text-guided training comprises:
outputting one or more text-image features that are used to train the object detection model by using the text-guided model and the image feature extractor (fig. 2; para. [0009], [0036]); and
Fuxman fails to teach light detection and ranging (LiDAR)-guided training using a point cloud encoder and a bird's-eye view (BEV) encoder; and
training an object detection model based on a result of the text-guided training and a result of the LiDAR-guided training,
However, Fuxman does teach training an object detection model (para. [0003]). It would be obvious to train the model based on the text-guided training to provide high predictive accuracy and generalizability to unseen data (para. [0003]) as well as provide significant improvement in the field of image classification (para. [0006]).
However Balachandran teaches light detection and ranging (LiDAR)-guided training using a point cloud encoder (408 in fig. 4; para. [0078) and a bird's-eye view (BEV) encoder (404 in fig. 4; para. [0079]-[0080]); and
training an object detection model based on a result of the LiDAR-guided training (428 in fig. 4).
Therefore taking the combine teachings of Fuxman and Balachandran as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Balachandran into the method of Fuxman. The motivation to combine Balachandran and Fuxman would be to improve object recognition (para. [0005] of Balachandran).
Regarding claim 13, the claim recites similar subject matter as claim 1 and is rejected for the same reasons as stated above.
Claim(s) 2 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fuxman et al (US20210264203) and Balachandran et al (US20250029355) in view of Harikumar et al (US20220156992).
Regarding claim 2, the modified method of Fuxman fails to teach a method wherein the outputting the one or more text-image features comprises:
outputting camera-variant information by performing semantic information encoding on the one or more text inputs through a text encoder and a first projection layer model, which are comprised in the text-guided model; and
generating the one or more text-image features by adding one or more image features extracted by the image feature extractor to the camera-variant information.
However Harikumar teaches outputting camera-variant information by performing semantic information encoding on the one or more text inputs (para. [0005], generating, by a text embedding model, a text embedding of a text query, where the text embedding and the learned image representation of the target image are in a same embedding space; para. [0023], [0029]) through a text encoder (208 in fig. 2) and a first projection layer model, which are comprised in the text-guided model (para. [0030], After converting the text to the cross-lingual embedding, fully connected layers of the deep learning architecture use the cross-lingual embedding to generate a text embedding in an image space); and
generating the one or more text-image features by adding one or more image features extracted by the image feature extractor to the camera-variant information (para. [0005], [0018]).
Therefore taking the combine teachings of Fuxman and Balachandran with Harikumar as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Harikumar into the method of Fuxman and Balachandran. The motivation to combine Balachandran, Harikumar and Fuxman would be to improve segmentation accuracy (para. [0019] of Harikumar).
Regarding claim 14, the claim recites similar subject matter as claim 2 and is rejected for the same reasons as stated above.
Claim(s) 3 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fuxman et al (US20210264203) and Balachandran et al (US20250029355) in view of Borse et al (US20240153249).
Regarding claim 3, the modified method of Fuxman fails to teach a method wherein the LiDAR-guided training comprises:
performing contrastive training on LiDAR BEVs obtained from the point cloud encoder and image BEVs generated based on the BEV encoder.
However Borse teaches performing contrastive training (para. [0100]) on LiDAR BEVs obtained from a point cloud encoder (para. [0090]) and image BEVs generated based on a BEV encoder (para. [0091]).
Therefore taking the combine teachings of Fuxman and Balachandran with Borse as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Borse into the method of Fuxman and Balachandran. The motivation to combine Balachandran, Borse and Fuxman would be to improve the detection of objects in a scene (para. [0092] of Borse).
Regarding claim 15, the claim recites similar subject matter as claim 3 and is rejected for the same reasons as stated above.
Claim(s) 4 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fuxman et al (US20210264203), Balachandran et al (US20250029355) and Borse et al (US20240153249) in view of Redford et al (US20240104913).
Regarding claim 4, the modified method of Fuxman fails to teach a method wherein the contrastive training comprises:
training a cross-correlation of the LiDAR BEVs and the image BEVs based on a second loss function.
However Redford teaches training a cross-correlation of the LiDAR BEVs and the image BEVs based on a second loss function (para. [0036], [0014], [0079], [0081]).
Therefore taking the combine teachings of Fuxman, Balachandran and Borse with Redford as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Redford into the method of Fuxman, Borse and Balachandran. The motivation to combine Redford, Balachandran, Borse and Fuxman would be to extract useful features from sensor data (para. [0001] of Redford).
Regarding claim 16, the claim recites similar subject matter as claim 4 and is rejected for the same reasons as stated above.
Claim(s) 5 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fuxman et al (US20210264203), Balachandran et al (US20250029355) and Borse et al (US20240153249) in view of Jung et al (US20230326051).
Regarding claim 5, the modified method of Fuxman teaches a method wherein the training the object detection model comprises:
updating one or more of the image feature extractor (para. [0094] of Borse), a depth extractor, the BEV encoder (para. [0091] of Borse), and a detection head (para. [0091] of Borse), based on one of the one or more text-image features (para. [0036] of Fuxman) or a result of the contrastive training.
The modified method of Fuxman fails to teach updating a depth extractor. However Jung teaches updating a depth extractor (para. [0048]-[0050]).
Therefore taking the combine teachings of Fuxman, Balachandran and Borse with Jung as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Jung into the method of Fuxman, Borse and Balachandran. The motivation to combine Jung, Balachandran, Borse and Fuxman would be to more accurately estimate depth information of an input image (para. [0036] of Jung).
Regarding claim 17, the claim recites similar subject matter as claim 5 and is rejected for the same reasons as stated above.
Claim(s) 8-10, 18 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fuxman et al (US20210264203) and Balachandran et al (US20250029355) in view of Minderer et al (US20240161459).
Regarding claim 8, the modified method of Fuxman fails to teach a method further comprising:
text-guided model training to obtain the pre-trained text guided model, the text-guided model training including updating a second projection layer model to a first projection layer model by using a text encoder, the second projection layer model, and an image encoder.
However Minderer teaches text-guided model training to obtain the pre-trained text guided model (para. 0031]), the text-guided model training including updating a second projection layer model to a first projection layer model (para. [0011], [0065], [0096]) by using a text encoder (para. [0011]), the second projection layer model (para. [0019]), and an image encoder (para. [0011]).
Therefore taking the combine teachings of Fuxman and Balachandran with Minderer as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Minderer into the method of Fuxman and Balachandran. The motivation to combine Balachandran, Minderer and Fuxman would be to allow the detection of objects that are difficult to describe through text yet easy to capture in an image (para. [0029] of Minderer).
Regarding claim 9, the modified method of Fuxman teaches a method wherein the text-guided model training comprises:
extracting a text feature for training, the text feature comprising camera-variant information for training, from a training text input by using the text encoder (para. [0009], [0093] of Minderer) and the second projection layer model (para. [0019] of Minderer) and projecting the extracted text feature for training onto a shared embedding space (para. [0031], [0097] of Minderer);
extracting an image feature for training from a training image input by using the image encoder (fig. 5 of Minderer; para. [0010], [0095] of Minderer) and projecting the extracted image feature for training onto the shared embedding space (para. [0031], [0097] of Minderer); and
updating the second projection layer model to the first projection layer model (para. [0012], [0104] of Minderer).
Regarding claim 10, the modified method of Fuxman teaches a method wherein the updating the second projection layer model comprises:
performing contrastive alignment training on the text feature for training and the image feature for training in the shared embedding space (para. [0009], [0031], [0097] of Minderer); and
training the second projection layer model to the first projection layer model by using a result of the contrastive alignment training and a first loss function (para. [0012], [0104], [0112] of Minderer).
Regarding claim 18, the claim recites similar subject matter as claims 1, 8, and 9 and is rejected for the same reasons as stated above.
Regarding claim 19, the claim recites similar subject matter as claim 10 and is rejected for the same reasons as stated above.
Claim(s) 11 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fuxman et al (US20210264203), Balachandran et al (US20250029355) and Minderer et al (US20240161459) in view of Liu et al ("Feature-level camera style transfer for person re-identification." Applied Sciences 12.14 (2022): pages 1-18, retrieved from the Internet on 5/27/2026).
Regarding claim 11, the modified method of Fuxman fails to teach a method wherein the first loss function comprises:
a camera classifier configured to suppress unclear geometric noise in the text feature for training.
However Liu teaches a camera classifier (fig. 2) configured to suppress unclear geometric noise (page 13, As we can see in the figure, the baseline model tends to be affected by camera style-related factors, e.g., posture, background, illumination, etc. The proposed CST can generate a large amount of camera style transferred features to alleviate this problem so that the model can overcome the effect of camera style variance) in a feature for training.
Therefore taking the combine teachings of Fuxman, Balachandran and Minderer with Liu as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the steps of Liu into the method of Fuxman, Minderer and Balachandran. The motivation to combine Balachandran, Liu, Minderer and Fuxman would be to perform more accurate matching (page 13 of Liu).
Regarding claim 20, the claim recites similar subject matter as claim 11 and is rejected for the same reasons as stated above.
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minderer et al (US20240161459) in view of Jung et al (US20230326051) and Park et al (US20240020953).
Regarding claim 12, Minderer teaches a non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, configure the one or more processors to perform (claim 16) a text-guided object detection model (para. [0031], [0097]), the text guided object detection model comprising:
a pre-trained text-guided model configured to generate camera-variant information from a text-image pair (para. [0031], [0097]); and
an image feature extractor configured to extract an image feature from the text-image pair (para. [0101]).
Minderer fails to teach a depth extractor configured to extract depth information based on the image feature;
a BEV encoder configured to generate an image bird's-eye view (BEV) based on the camera-variant information, the image feature, and the depth information; and
a detection head configured to perform object detection based on the image BEV.
Jung teaches a depth extractor configured to extract depth information based on the image feature (para. [0036], [0047]).
Therefore taking the combine teachings of Minderer and Jung as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the features of Jung into the apparatus of Minderer. The motivation to combine Minderer and Jung would be to more accurately estimate depth information of an input image (para. [0036] of Jung).
Park teaches a BEV encoder configured to generate an image bird's-eye view (BEV) (para. [0029]) based on camera-variant information, an image feature, and depth information (para. [0039]. Furthermore, it would be obvious to use the depth information of Jung as sensor data); and
a detection head configured to perform object detection based on the image BEV para. [0029]).
Therefore taking the combine teachings of modified Minderer and Park as a whole, it would have been obvious to one of ordinary skill in the art at the time the invention was filed to incorporate the features of Park into the apparatus of modified Minderer. The motivation to combine modified Minderer and Park would be to provide robustness to camera dropout (para. [0023] of Park).
Allowable Subject Matter
Claims 6-7 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Related Art
Chen et al (US20250104409) – see para. [0053], [0056], [0061]
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEON VIET Q NGUYEN whose telephone number is (571)270-1185. The examiner can normally be reached Mon-Fri 11AM-7PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gregory Morse can be reached at 571-272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LEON VIET Q NGUYEN/ Primary Examiner, Art Unit 2663