DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-4, 6-9, 13, 15-20, 24-25, 28-29 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Gu et al.: "Zero-Shot Detection via Vision and Language Knowledge Distillation", arxiv.org, Submitted on 28 Apr 2021 [retrieved on 7/29/2026]. Retrieved from the internet <https://arxiv.org/abs/2104.13921v1>, hereinafter Gu.
Regarding claim 1, Gu teaches A computer-implemented method of training a detector head for object detection of a training object category based on a frozen vision and language model (VLM), comprising: receiving, by a computing device, the frozen VLM pre-trained on a plurality of image-text pairs; (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17])." Pg. 1, Col. 2, Para. 2, see "Recently, Radford et al. [33] train a joint vision and language model using 400 million image and text pairs and demonstrate impressive zero-shot recognition abilities on over 30 computer vision datasets. Despite the great success on learning image-level representations for zero-shot classification, learning object-level representations for zero-shot object detection is still challenging. In this work, we consider the idea of borrowing the knowledge from a pre-trained zero-shot image classification model to enable zero-shot object detection." Pg. 3, Col. 2, Para. 2 see "We achieve zero-shot detection by leveraging an off-the-shelf pre-trained zero-shot image classification model (e.g. CLIP [33]). The model has a text encoder T (·) and an image encoder V(·), which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2." Pg. 5, Col. 2, Para. 1 see "We train the model from scratch for 180,000 iterations... For all our experiments involving a zero-shot classification model, we use the pre-trained CLIP model that is publicly available."). determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in an image; (Pg. 4, Section 3.3 see "In ViLD, we learn region embeddings in a two-stage detector to represent each proposal. We define region embeddings as R(φ(I),r), where φ(·) is the backbone model and R is a lightweight model to generate region embeddings for each proposal r. Specifically, we take outputs of the layer before the detection classifier as region embeddings. Our goal is to train the region embeddings such that they can be classified with the text embeddings encoded by T (·)." Pg. 4, Section 3.4 see "We then introduce ViLD-image, which aims to align region embeddings R(φ(I),r) to image embeddings V(crop(I,r)), introduced in Section 3.2. The goal is to distill the knowledge in the teacher image encoder V into the student detector. We extract proposals r offline from the training images using a region proposal network pre-trained on base categories." Pg. 3, Fig. 2 see "Then, ViLD uses the text embeddings as the region classifier (ViLD-text) and minimizes the distance of the region embedding to the image embedding for each proposal (ViLD-image). During inference, text embeddings of novel categories are used to enable zero-shot detection."). generating, by a pre-trained text encoder of the frozen VLM, a text embedding of the training object category; (Pg. 4, Col. 2, Para. 2 see " For training, we generate the text embeddings T (CB) by feeding the text prompts of base categories (CB), e.g., “a photo of {category} in the scene”, into the text encoder." Pg. 2, Col. 1, Para. 2 see " In ViLD-text, we obtain the text embeddings by feeding the category text prompts into the pretrained text encoder. We then replace the trainable detection classifier with the fixed text embeddings." Pg. 3, Fig. 2 see "First, the category text embeddings and the image embeddings of cropped object proposals are computed using the text and image encoders in the classification model."). predicting, by the detector head and based on the detection region embedding and the text embedding of the training object category, an object from a target object vocabulary associated with the training object category; (Abstract see "Our method aligns the region embeddings in the detector to the text and image embeddings inferred by the pre-trained model. We use the text embeddings as the detection classifier, obtained by feeding category names into the pre-trained text encoder." Pg. 2, Col. 1, Para. 2 see " We then replace the trainable detection classifier with the fixed text embeddings. Similar approaches have been used in prior zero-shot detection work [3, 35]. However, these existing methods only use text embeddings learned from a language corpus, e.g., [32]. In contrast, we find text embeddings learned jointly with visual data, e.g., [33], can better represent the visual similarity between text prompts." Pg. 4, Col. 2, Para. 2 see "We compute the cosine similarity between each region embedding R(φ(I),r) and all category embeddings, including T (CB) and ebg. Then we apply softmax activation with a temperature τ to these cosine similarities to compute the cross entropy loss. Let sim(a,b) = a b/( a b), ti denote elements in T (CB), yr denote the class label of the region r, and LCE denote the cross entropy loss. The loss function for ViLD-text can be written as: er = R(φ(I),r)z(r) = sim(er,ebg), sim(er,t1), ··· , sim(er,t|CB|)LViLD-text = LCE softmaxz(r)/τ ,yr ."). and providing, by the computing device, the pre-trained frozen VLM and the trained detector head. (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17])." Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset." Pg. 5, Col. 1, Para. 1 see "The difference between ViLD-text and ViLD-image is only in the training, where ViLD-text and ViLD-image are trained with LViLD-text and LViLD-image respectively. During inference, ViLD-image, ViLD-text and ViLD share the same model architecture for zero-shot detection.").
Regarding claim 2, Gu teaches The computer-implemented method of claim 1. wherein the predicting of the object comprises: determining, by the detector head, one or more detection scores for the one or more regions of interest, wherein the one or more detection scores for the one or more regions of interest is indicative of the predicted object. (Pg. 5, Col. 1, Para. 3 see "We use a trained ViLD-text detector to obtain a set of candidate regions and their category confidence scores. We then filter out background regions and apply NMS to obtain the top-k proposals. We use pi,ViLD-text to denote the confidence scores for each proposal r.").
Regarding claim 3, Gu teaches The computer-implemented method of claim 2. wherein the predicting of the object comprises training the detector head to predict one or more object detection boxes and associated masks corresponding to the one or more regions of interest, and wherein the one or more detection scores are associated with the one or more predicted object detection boxes. (Pg. 3, Col. 2, Para. 3 see "We modify a two-stage object detector (e.g.,Mask R-CNN [17]) to detect object proposals with bounding boxes and masks. We replace class-specific localization modules, i.e., the second stage bounding box regression and mask prediction layers, with class-agnostic modules for general object proposals. For each region of interest, these modules only predict a single bounding box and a single mask for all possible categories, instead of one prediction for each category." Pg. 5, Col. 1, Para. 3 see "We use a trained ViLD-text detector to obtain a set of candidate regions and their category confidence scores. We then filter out background regions and apply NMS to obtain the top-k proposals. We use pi,ViLD-text to denote the confidence scores for each proposal r." Pg. 13, Col. 2, Para. 6 see "in object detection, it is important for the higher-quality boxes of the same object to have higher scores. In Figure 12(c), we simply rescore by taking the geometric mean of CLIP probabilities and the Mask R-CNN score of the bounding box.").
Regarding claim 4, Gu teaches The computer-implemented method of claim 1. wherein the training of the detector head is based on one or more of a box region loss, a box classification loss, or a mask classification loss. (Pg. 4, Col. 2, Para. 2 see "We compute the cosine similarity between each region embedding R(φ(I),r) and all category embeddings, including T (CB) and ebg. Then we apply softmax activation with a temperature τ to these cosine similarities to compute the cross entropy loss. Let sim(a,b) = a b/( a b), ti denote elements in T (CB), yr denote the class label of the region r, and LCE denote the cross entropy loss. The loss function for ViLD-text can be written as: er = R(φ(I),r)z(r) = sim(er,ebg), sim(er,t1), ··· , sim(er,t|CB|)LViLD-text = LCE softmaxz(r)/τ ,yr ." Pg. 4, Col. 2, Para. 1 see "The training loss of ViLD is simply a wegithed sum of both objectives: LViLD = LViLD-text + w · LViLD-image,").
Regarding claim 6, Gu teaches The computer-implemented method of claim 1. wherein the detector head is a neural network. (Abstract see "We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17]).").
Regarding claim 7, Gu teaches The computer-implemented method of claim 6. wherein the detector head is one of a Mask R-CNN or a Faster R-CNN. (Abstract see "We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17]).").
Regarding claim 8, Gu teaches The computer-implemented method of claim 1. wherein the detector head further comprises a feature pyramid network. (Pg. 5, Col. 1, Para. 5 see "We benchmark on the Mask RCNN [17] with ResNet [18] FPN [27] backbone and use the same settings for all models unless explicitly specified.").
Regarding claim 9, Gu teaches The computer-implemented method of claim 1. wherein the pre-trained text encoder and the pre-trained image encoder of the frozen VLM are jointly trained based on contrastive learning. (Pg. 3, Col. 2, Para. 2 see " The model has a text encoder T (·) and an image encoder V(·),which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2.").
Regarding claim 13, Gu teaches The computer-implemented method of claim 1. further comprising: maintaining an image normalization scheme of the pre-trained frozen VLM to enable open vocabulary object detection. (Pg. 4, Fig. 3 see "A projection layer and a normalization layer are introduced to adjust the normand dimension of region embeddings in order to be compatible with the fixed text embeddings.").
Regarding claim 15, Gu teaches A computer-implemented method of applying a trained detector head for object detection of a training object category based on a frozen vision and language model (VLM), comprising: receiving, by a computing device, an input image; (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17])." Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset." Pg. 1, Col. 2, Para. 2, see "Recently, Radford et al. [33] train a joint vision and language model using 400 million image and text pairs and demonstrate impressive zero-shot recognition abilities on over 30 computer vision datasets. Despite the great success on learning image-level representations for zero-shot classification, learning object-level representations for zero-shot object detection is still challenging. In this work, we consider the idea of borrowing the knowledge from a pre-trained zero-shot image classification model to enable zero-shot object detection." Pg. 3, Col. 2, Para. 2 see "We achieve zero-shot detection by leveraging an off-the-shelf pre-trained zero-shot image classification model (e.g. CLIP [33]). The model has a text encoder T (·) and an image encoder V(·), which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2."). applying a trained neural network for object detection, wherein the neural network comprises the frozen VLM pre-trained on a plurality of image-text pairs, and the trained detector head associated with the pre-trained frozen VLM and pre-trained on the training object category; (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17]). Our method aligns the region embeddings in the detector to the text and image embeddings inferred by the pre-trained model. We use the text embeddings as the detection classifier, obtained by feeding category names into the pre-trained text encoder." Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset." Pg. 1, Col. 2, Para. 2, see "Recently, Radford et al. [33] train a joint vision and language model using 400 million image and text pairs and demonstrate impressive zero-shot recognition abilities on over 30 computer vision datasets. Despite the great success on learning image-level representations for zero-shot classification, learning object-level representations for zero-shot object detection is still challenging. In this work, we consider the idea of borrowing the knowledge from a pre-trained zero-shot image classification model to enable zero-shot object detection." Pg. 3, Col. 2, Para. 2 see "We achieve zero-shot detection by leveraging an off-the-shelf pre-trained zero-shot image classification model (e.g. CLIP [33]). The model has a text encoder T (·) and an image encoder V(·), which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2."). determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in the input image; (Pg. 4, Section 3.3 see "In ViLD, we learn region embeddings in a two-stage detector to represent each proposal. We define region embeddings as R(φ(I),r), where φ(·) is the backbone model and R is a lightweight model to generate region embeddings for each proposal r. Specifically, we take outputs of the layer before the detection classifier as region embeddings. Our goal is to train the region embeddings such that they can be classified with the text embeddings encoded by T (·)." Pg. 4, Section 3.4 see "We then introduce ViLD-image, which aims to align region embeddings R(φ(I),r) to image embeddings V(crop(I,r)), introduced in Section 3.2. The goal is to distill the knowledge in the teacher image encoder V into the student detector. We extract proposals r offline from the training images using a region proposal network pre-trained on base categories." Pg. 3, Fig. 2 see "Then, ViLD uses the text embeddings as the region classifier (ViLD-text) and minimizes the distance of the region embedding to the image embedding for each proposal (ViLD-image). During inference, text embeddings of novel categories are used to enable zero-shot detection."). predicting, by the detector head and based on the detection region embedding and a text embedding of the training object category, an object from a target object vocabulary associated with the training object category; (Abstract see "Our method aligns the region embeddings in the detector to the text and image embeddings inferred by the pre-trained model. We use the text embeddings as the detection classifier, obtained by feeding category names into the pre-trained text encoder." Pg. 2, Col. 1, Para. 2 see " We then replace the trainable detection classifier with the fixed text embeddings. Similar approaches have been used in prior zero-shot detection work [3, 35]. However, these existing methods only use text embeddings learned from a language corpus, e.g., [32]. In contrast, we find text embeddings learned jointly with visual data, e.g., [33], can better represent the visual similarity between text prompts." Pg. 4, Col. 2, Para. 2 see "We compute the cosine similarity between each region embedding R(φ(I),r) and all category embeddings, including T (CB) and ebg. Then we apply softmax activation with a temperature τ to these cosine similarities to compute the cross entropy loss. Let sim(a,b) = a b/( a b), ti denote elements in T (CB), yr denote the class label of the region r, and LCE denote the cross entropy loss. The loss function for ViLD-text can be written as: er = R(φ(I),r)z(r) = sim(er,ebg), sim(er,t1), ··· , sim(er,t|CB|)LViLD-text = LCE softmaxz(r)/τ ,yr ."). and providing, by the computing device, the input image with the object from the target object vocabulary. (Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset.").
Regarding claim 16, Gu teaches The computer-implemented method of claim 15. wherein the predicting of the object comprises: determining, by the trained detector head, one or more detection scores for the one or more regions of interest, wherein the one or more detection scores for the one or more regions of interest is indicative of the predicted object. (Pg. 5, Col. 1, Para. 3 see "We use a trained ViLD-text detector to obtain a set of candidate regions and their category confidence scores. We then filter out background regions and apply NMS to obtain the top-k proposals. We use pi,ViLD-text to denote the confidence scores for each proposal r.").
Regarding claim 17, Gu teaches The computer-implemented method of claim 16. wherein the predicting of the object comprises predicting one or more object detection boxes and associated masks corresponding to the one or more regions of interest, and wherein the one or more detection scores are associated with the one or more predicted object detection boxes. (Pg. 3, Col. 2, Para. 3 see "We modify a two-stage object detector (e.g.,Mask R-CNN [17]) to detect object proposals with bounding boxes and masks. We replace class-specific localization modules, i.e., the second stage bounding box regression and mask prediction layers, with class-agnostic modules for general object proposals. For each region of interest, these modules only predict a single bounding box and a single mask for all possible categories, instead of one prediction for each category." Pg. 5, Col. 1, Para. 3 see "We use a trained ViLD-text detector to obtain a set of candidate regions and their category confidence scores. We then filter out background regions and apply NMS to obtain the top-k proposals. We use pi,ViLD-text to denote the confidence scores for each proposal r." Pg. 13, Col. 2, Para. 6 see "in object detection, it is important for the higher-quality boxes of the same object to have higher scores. In Figure 12(c), we simply rescore by taking the geometric mean of CLIP probabilities and the Mask R-CNN score of the bounding box.").
Regarding claim 18, Gu teaches The computer-implemented method of claim 15. further comprising: providing, by the computing device, the predicted object from the target object vocabulary associated with the training object category. (Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset.").
Regarding claim 19, Gu teaches The computer-implemented method of claim 15. further comprising: receiving, by the pre-trained text encoder, an inference object category different from the training object category; (Pg. 4, Col. 2, Para. 3 see "During inference, we include novel categories (CN) and generate T (CB ∪ CN) (sometimes T (CN) only) for zero-shot detection (Figure 2). Our hope is that the model learned from labeled CB can generalize to novel categories CN."). augmenting, by the trained detector head, the text embedding of the training object category with an additional embedding of the inference object category, and wherein the predicting of the object comprises predicting, based on the augmented text embedding and the detection region embedding, an additional object from an augmented target object vocabulary associated with the training object category and the inference object category. (Pg. 1, Fig. 1 see "An example of our zero-shot object detection with free-form text classifiers. After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset." Pg. 7, Col. 2, Para. 3 see "On-the-fly interactive object detection: We tap the potential of ViLD by using free-form text to interactively recognize fine-grained categories and attributes. After obtaining detection results on base categories, we extract the region embedding and compute its cosine similarity with a small set of on-the-fly free-form texts describing attributes and/or fine-grained categories; we apply a softmax with temperature on top of the similarities.").
Regarding claim 20, Gu teaches The computer-implemented method of claim 19. wherein the predicting of the additional object comprises: determining, by the trained detector head, one or more augmented detection scores for the one or more regions of interest, wherein the one or more augmented detection scores for the one or more regions of interest is indicative of the predicted additional object. (Pg. 5, Col. 1, Para. 3 see "We use a trained ViLD-text detector to obtain a set of candidate regions and their category confidence scores. We then filter out background regions and apply NMS to obtain the top-k proposals. We use pi,ViLD-text to denote the confidence scores for each proposal r. We then feed crop(I,r) to the zero-shot classification model to obtain confidence scores pi,cls. ").
Regarding claim 24, Gu teaches The computer-implemented method of claim 15. wherein the detector head is a neural network. (Abstract see "We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17]).").
Regarding claim 25, Gu teaches The computer-implemented method of claim 24. wherein the detector head is one of a Mask R-CNN or a Faster R-CNN. (Abstract see "We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17]).").
Regarding claim 28, Gu teaches The computer-implemented method of claim 15. wherein the detector head further comprises a feature pyramid network. (Pg. 5, Col. 1, Para. 5 see "We benchmark on the Mask RCNN [17] with ResNet [18] FPN [27] backbone and use the same settings for all models unless explicitly specified.").
Regarding claim 29, Gu teaches The computer-implemented method of claim 15. the pre- trained text encoder and the pre-trained image encoder of the frozen VLM having been jointly trained based on contrastive learning. (Pg. 3, Col. 2, Para. 2 see " The model has a text encoder T (·) and an image encoder V(·),which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2.").
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Gu et al.: "Zero-Shot Detection via Vision and Language Knowledge Distillation", arxiv.org, Submitted on 28 Apr 2021 [retrieved on 7/29/2026]. Retrieved from the internet <https://arxiv.org/abs/2104.13921v1>, hereinafter Gu, in view of Radford et al.: "Learning Transferable Visual Models From Natural Language Supervision", arxiv.org, Submitted on 26 Feb 2021 [retrieved on 7/29/2026]. Retrieved from the internet <https://arxiv.org/abs/2103.00020>, hereinafter Radford.
Regarding claim 10, Gu teaches The computer-implemented method of claim 1.
While Gu teaches a pre-trained image encoder, Gu does not teach wherein the pre-trained image encoder comprises a (i) feature extractor to generate the image representation for the image, and (ii) a feature pooling layer.
However, Radford teaches wherein the pre-trained image encoder comprises a (i) feature extractor to generate the image representation for the image, and (ii) a feature pooling layer. (Pg. 4, Col. 2, Para. 3 see "We consider two different architectures for the image encoder. For the first, we use ResNet-50 as the base architecture for the image encoder… We also replace the global average pooling layer with an attention pooling mechanism. The attention pooling is implemented as a single layer of “transformer-style” multi-head QKV attention.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gu to incorporate the teachings of Radford to include a feature extractor to generate the image representation for the image and a feature pooling layer. Doing so would predictably save money and time by avoiding retraining an entirely new model to extract features from images and compacting region embeddings into a vector that can be directly compared to text embeddings of object categories.
Regarding claim 11, Gu in view of Radford teaches The computer-implemented method of claim 10.
While Gu teaches using ResNet-50 architecture, Gu does not teach wherein the pre-trained image encoder comprises the feature extractor comprising a ResNet-50 architecture.
However, Radford teaches wherein the feature extractor comprises a ResNet-50 architecture. (Pg. 4, Col. 2, Para. 3 see "We consider two different architectures for the image encoder. For the first, we use ResNet-50 as the base architecture for the image encoder." Pg. 5, Col. 2, Para. 2 see "We train a series of 5 ResNets and 3 Vision Transformers. For the ResNets we train a ResNet-50, a ResNet-101, and then 3 more which follow EfficientNet-style model scaling and use approximately 4x, 16x, and 64x the compute of a ResNet-50.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gu and Radford to incorporate the teachings of Radford to use a ResNet-50 architecture to extract features. Doing so would predictably increase reliability and reduce cost by using a high performing and relatively cheap architecture to extract multi-scale visual details from images.
Regarding claim 12, Gu in view of Radford teaches The computer-implemented method of claim 10.
While Gu teaches an image encoder, Gu does not teach wherein the feature pooling layer is an attention layer of the image encoder.
However, Radford teaches wherein the feature pooling layer is an attention layer of the image encoder. (Pg. 4, Col. 2, Para. 3 see "We also replace the global average pooling layer with an attention pooling mechanism. The attention pooling is implemented as a single layer of “transformer-style” multi-head QKV attention.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gu and Radford to incorporate the teachings of Radford to use a feature pooling attention layer. Doing so would predictably increase accuracy by allowing the model to focus on the most important parts of the image.
Claim 26 is rejected under 35 U.S.C. 103 as being unpatentable over Gu et al.: "Zero-Shot Detection via Vision and Language Knowledge Distillation", arxiv.org, Submitted on 28 Apr 2021 [retrieved on 7/29/2026]. Retrieved from the internet <https://arxiv.org/abs/2104.13921v1>, hereinafter Gu, in view of Xie et al.: "Zero-shot Object Detection Through Vision-Language Embedding Alignment", arxiv.org, Submitted on 24 Sep 2021 [retrieved on 7/29/2026]. Retrieved from the internet <https://arxiv.org/abs/2109.12066>, hereinafter Xie.
Regarding claim 26, Gu teaches The computer-implemented method of claim 24.
While Gu teaches training a detector head, Gu does not teach the detector head having been trained to perform one-stage object detection.
However, Xie teaches the detector head having been trained to perform one-stage object detection. (Pg. 1, Col. 2, Para. 3, see "we propose a method for adapting a one stage detector to perform the ZSD task through aligning detector semantic outputs to embeddings from a trained vision-language model." Pg. 1, Col. 2, Para. 4 see "Specifically, we train ZSD-YOLO, our one stage zero-shot detection model [16] that aligns detector semantic outputs to embeddings from a contrastively trained vision-language model CLIP [30].").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gu to incorporate the teachings of Xie to train the detector to perfrom one-stage object detection. Doing so would predictably increase the speed of object detection and use few computational resources by reducing the amount of processing needed to determine an object detection.
Claims 30, 35 are rejected under 35 U.S.C. 103 as being unpatentable over Gu et al.: "Zero-Shot Detection via Vision and Language Knowledge Distillation", arxiv.org, Submitted on 28 Apr 2021 [retrieved on 7/29/2026]. Retrieved from the internet <https://arxiv.org/abs/2104.13921v1>, hereinafter Gu, in view of Gu-Jiuxiang et al. (US 20220147838 A1), hereinafter Gu-Jiuxiang.
Regarding claim 30, Gu teaches training a detector head for object detection of a training object category based on a frozen vision and language model (VLM), receiving, by a computing device, the frozen VLM pre-trained on a plurality of image-text pairs; (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17])." Pg. 1, Col. 2, Para. 2, see "Recently, Radford et al. [33] train a joint vision and language model using 400 million image and text pairs and demonstrate impressive zero-shot recognition abilities on over 30 computer vision datasets. Despite the great success on learning image-level representations for zero-shot classification, learning object-level representations for zero-shot object detection is still challenging. In this work, we consider the idea of borrowing the knowledge from a pre-trained zero-shot image classification model to enable zero-shot object detection." Pg. 3, Col. 2, Para. 2 see "We achieve zero-shot detection by leveraging an off-the-shelf pre-trained zero-shot image classification model (e.g. CLIP [33]). The model has a text encoder T (·) and an image encoder V(·), which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2." Pg. 5, Col. 2, Para. 1 see "We train the model from scratch for 180,000 iterations... For all our experiments involving a zero-shot classification model, we use the pre-trained CLIP model that is publicly available."). determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in an image; (Pg. 4, Section 3.3 see "In ViLD, we learn region embeddings in a two-stage detector to represent each proposal. We define region embeddings as R(φ(I),r), where φ(·) is the backbone model and R is a lightweight model to generate region embeddings for each proposal r. Specifically, we take outputs of the layer before the detection classifier as region embeddings. Our goal is to train the region embeddings such that they can be classified with the text embeddings encoded by T (·)." Pg. 4, Section 3.4 see "We then introduce ViLD-image, which aims to align region embeddings R(φ(I),r) to image embeddings V(crop(I,r)), introduced in Section 3.2. The goal is to distill the knowledge in the teacher image encoder V into the student detector. We extract proposals r offline from the training images using a region proposal network pre-trained on base categories." Pg. 3, Fig. 2 see "Then, ViLD uses the text embeddings as the region classifier (ViLD-text) and minimizes the distance of the region embedding to the image embedding for each proposal (ViLD-image). During inference, text embeddings of novel categories are used to enable zero-shot detection."). generating, by a pre-trained text encoder of the frozen VLM, a text embedding of the training object category; (Pg. 4, Col. 2, Para. 2 see " For training, we generate the text embeddings T (CB) by feeding the text prompts of base categories (CB), e.g., “a photo of {category} in the scene”, into the text encoder." Pg. 2, Col. 1, Para. 2 see " In ViLD-text, we obtain the text embeddings by feeding the category text prompts into the pretrained text encoder. We then replace the trainable detection classifier with the fixed text embeddings." Pg. 3, Fig. 2 see "First, the category text embeddings and the image embeddings of cropped object proposals are computed using the text and image encoders in the classification model."). predicting, by the detector head and based on the detection region embedding and the text embedding of the training object category, an object from a target object vocabulary associated with the training object category; (Abstract see "Our method aligns the region embeddings in the detector to the text and image embeddings inferred by the pre-trained model. We use the text embeddings as the detection classifier, obtained by feeding category names into the pre-trained text encoder." Pg. 2, Col. 1, Para. 2 see " We then replace the trainable detection classifier with the fixed text embeddings. Similar approaches have been used in prior zero-shot detection work [3, 35]. However, these existing methods only use text embeddings learned from a language corpus, e.g., [32]. In contrast, we find text embeddings learned jointly with visual data, e.g., [33], can better represent the visual similarity between text prompts." Pg. 4, Col. 2, Para. 2 see "We compute the cosine similarity between each region embedding R(φ(I),r) and all category embeddings, including T (CB) and ebg. Then we apply softmax activation with a temperature τ to these cosine similarities to compute the cross entropy loss. Let sim(a,b) = a b/( a b), ti denote elements in T (CB), yr denote the class label of the region r, and LCE denote the cross entropy loss. The loss function for ViLD-text can be written as: er = R(φ(I),r)z(r) = sim(er,ebg), sim(er,t1), ··· , sim(er,t|CB|)LViLD-text = LCE softmaxz(r)/τ ,yr ."). and providing, by the computing device, the pre-trained frozen VLM and the trained detector head. (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17])." Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset." Pg. 5, Col. 1, Para. 1 see "The difference between ViLD-text and ViLD-image is only in the training, where ViLD-text and ViLD-image are trained with LViLD-text and LViLD-image respectively. During inference, ViLD-image, ViLD-text and ViLD share the same model architecture for zero-shot detection.").
While Gu teaches a computer implemented method, Gu does not teach A computing device for comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out functions comprising:.
However, Gu-Jiuxiang teaches A computing device for (Abstract see "Methods and systems disclosed herein relate generally to systems and methods for generating visual relationship graphs that identify relationships between objects depicted in an image." Para. 139 see "the computing system 1300 includes a processing device 1302 that executes the VL modeling application 102, a memory that stores various data computed or used by the VL modeling application 102." Para. 140 see "a processing device 1302 communicatively coupled to one or more memory devices 1304. The processing device 1302 executes computer-executable program code stored in a memory device 1304, accesses information stored in the memory device 1304."). comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out functions comprising: (Abstract see "Methods and systems disclosed herein relate generally to systems and methods for generating visual relationship graphs that identify relationships between objects depicted in an image." Para. 139 see "the computing system 1300 includes a processing device 1302 that executes the VL modeling application 102, a memory that stores various data computed or used by the VL modeling application 102." Para. 140 see "a processing device 1302 communicatively coupled to one or more memory devices 1304. The processing device 1302 executes computer-executable program code stored in a memory device 1304, accesses information stored in the memory device 1304.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gu to incorporate the teachings of Gu-Jiuxiang to train the object detector on a computing device with processors and computer memory. Doing so would predictably save time by running calculations on a computing device as opposed to a human performing calculations with pen and paper.
Regarding claim 35, Gu teaches applying a trained detector head for object detection of a training object category based on a frozen vision and language model (VLM), receiving, by a computing device, an input image; (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17])." Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset." Pg. 1, Col. 2, Para. 2, see "Recently, Radford et al. [33] train a joint vision and language model using 400 million image and text pairs and demonstrate impressive zero-shot recognition abilities on over 30 computer vision datasets. Despite the great success on learning image-level representations for zero-shot classification, learning object-level representations for zero-shot object detection is still challenging. In this work, we consider the idea of borrowing the knowledge from a pre-trained zero-shot image classification model to enable zero-shot object detection." Pg. 3, Col. 2, Para. 2 see "We achieve zero-shot detection by leveraging an off-the-shelf pre-trained zero-shot image classification model (e.g. CLIP [33]). The model has a text encoder T (·) and an image encoder V(·), which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2."). applying a trained neural network for object detection, wherein the neural network comprises the frozen VLM pre-trained on a plurality of image-text pairs, and the trained detector head associated with the pre-trained frozen VLM and pre-trained on the training object category; (Abstract see "We propose ViLD, a training method via Vision and Language knowledge Distillation. We distill the knowledge from a pre-trained zero-shot image classification model (e.g., CLIP [33]) into a two-stage detector (e.g., Mask R-CNN [17]). Our method aligns the region embeddings in the detector to the text and image embeddings inferred by the pre-trained model. We use the text embeddings as the detection classifier, obtained by feeding category names into the pre-trained text encoder." Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset." Pg. 1, Col. 2, Para. 2, see "Recently, Radford et al. [33] train a joint vision and language model using 400 million image and text pairs and demonstrate impressive zero-shot recognition abilities on over 30 computer vision datasets. Despite the great success on learning image-level representations for zero-shot classification, learning object-level representations for zero-shot object detection is still challenging. In this work, we consider the idea of borrowing the knowledge from a pre-trained zero-shot image classification model to enable zero-shot object detection." Pg. 3, Col. 2, Para. 2 see "We achieve zero-shot detection by leveraging an off-the-shelf pre-trained zero-shot image classification model (e.g. CLIP [33]). The model has a text encoder T (·) and an image encoder V(·), which are pre-trained by joint image-text contrastive learning. In ViLD, we do not update these encoders during training, as shown in Figure 2."). determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in the input image; (Pg. 4, Section 3.3 see "In ViLD, we learn region embeddings in a two-stage detector to represent each proposal. We define region embeddings as R(φ(I),r), where φ(·) is the backbone model and R is a lightweight model to generate region embeddings for each proposal r. Specifically, we take outputs of the layer before the detection classifier as region embeddings. Our goal is to train the region embeddings such that they can be classified with the text embeddings encoded by T (·)." Pg. 4, Section 3.4 see "We then introduce ViLD-image, which aims to align region embeddings R(φ(I),r) to image embeddings V(crop(I,r)), introduced in Section 3.2. The goal is to distill the knowledge in the teacher image encoder V into the student detector. We extract proposals r offline from the training images using a region proposal network pre-trained on base categories." Pg. 3, Fig. 2 see "Then, ViLD uses the text embeddings as the region classifier (ViLD-text) and minimizes the distance of the region embedding to the image embedding for each proposal (ViLD-image). During inference, text embeddings of novel categories are used to enable zero-shot detection."). predicting, by the detector head and based on the detection region embedding and a text embedding of the training object category, an object from a target object vocabulary associated with the training object category; (Abstract see "Our method aligns the region embeddings in the detector to the text and image embeddings inferred by the pre-trained model. We use the text embeddings as the detection classifier, obtained by feeding category names into the pre-trained text encoder." Pg. 2, Col. 1, Para. 2 see " We then replace the trainable detection classifier with the fixed text embeddings. Similar approaches have been used in prior zero-shot detection work [3, 35]. However, these existing methods only use text embeddings learned from a language corpus, e.g., [32]. In contrast, we find text embeddings learned jointly with visual data, e.g., [33], can better represent the visual similarity between text prompts." Pg. 4, Col. 2, Para. 2 see "We compute the cosine similarity between each region embedding R(φ(I),r) and all category embeddings, including T (CB) and ebg. Then we apply softmax activation with a temperature τ to these cosine similarities to compute the cross entropy loss. Let sim(a,b) = a b/( a b), ti denote elements in T (CB), yr denote the class label of the region r, and LCE denote the cross entropy loss. The loss function for ViLD-text can be written as: er = R(φ(I),r)z(r) = sim(er,ebg), sim(er,t1), ··· , sim(er,t|CB|)LViLD-text = LCE softmaxz(r)/τ ,yr ."). and providing, by the computing device, the input image with the object from the target object vocabulary. (Pg. 1, Fig. 1 see "After training our zero-shot detector on base categories (purple), we can use the text embeddings of novel categories (pink) to detect novel object categories that do not exist in the training dataset.").
While Gu teaches a computer implemented method, Gu does not teach A computing device for comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out functions comprising:.
However, Gu-Jiuxiang teaches A computing device for (Abstract see "Methods and systems disclosed herein relate generally to systems and methods for generating visual relationship graphs that identify relationships between objects depicted in an image." Para. 139 see "the computing system 1300 includes a processing device 1302 that executes the VL modeling application 102, a memory that stores various data computed or used by the VL modeling application 102." Para. 140 see "a processing device 1302 communicatively coupled to one or more memory devices 1304. The processing device 1302 executes computer-executable program code stored in a memory device 1304, accesses information stored in the memory device 1304."). comprising: one or more processors; and data storage, wherein the data storage has stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing device to carry out functions comprising: (Abstract see "Methods and systems disclosed herein relate generally to systems and methods for generating visual relationship graphs that identify relationships between objects depicted in an image." Para. 139 see "the computing system 1300 includes a processing device 1302 that executes the VL modeling application 102, a memory that stores various data computed or used by the VL modeling application 102." Para. 140 see "a processing device 1302 communicatively coupled to one or more memory devices 1304. The processing device 1302 executes computer-executable program code stored in a memory device 1304, accesses information stored in the memory device 1304.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Gu to incorporate the teachings of Gu-Jiuxiang to run inference through the object detector on a computing device with processors and computer memory. Doing so would predictably save time by running calculations on a computing device as opposed to a human performing calculations with pen and paper.
Allowable Subject Matter
Claim(s) 5, 14, 21-23, 27 is/are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 5, none of Gu et al., Radford et al., Xie et al., or Gu-Jiuxiang et al. teach a two-stage detector head where the first stage determines the detection region embedding and the second stage determines the text embedding.
Regarding claim 14, none of Gu et al., Radford et al., Xie et al., or Gu-Jiuxiang et al. teach a detector head being trained on a collection of object categories from the same dataset as the one used to train the frozen VLM with image-text pairs. Nor do they teach substituting that same collection of object categories with a different collection of object categories from another dataset during testing.
Regarding claim 21, none of Gu et al., Radford et al., Xie et al., or Gu-Jiuxiang et al. teach computing scores for regions of interest with the feature pooling layer based on region embeddings and augmented text embeddings.
Regarding claim 22, none of Gu et al., Radford et al., Xie et al., or Gu-Jiuxiang et al. teach claim 21. Gu et al. teaches computing scores using a geometric mean.
Regarding claim 23, none of Gu et al., Radford et al., Xie et al., or Gu-Jiuxiang et al. teach claim 22. Radford et al. teaches a feature pooling attention layer.
Regarding claim 27, none of Gu et al., Radford et al., Xie et al., or Gu-Jiuxiang et al. teach a two-stage detector head where the first stage determines the detection region embedding and the second stage determines the text embedding.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Lin et al. (US 11256918 B2) discloses implementations of object detection in images, object detectors are trained using heterogeneous training datasets.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER VAUGHN whose telephone number is (571) 272-5253. The examiner can normally be reached M-F 11am-7pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER MEHMOOD can be reached on (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEXANDER VAUGHN/Examiner, Art Unit 2675
/XIAO LIU/Primary Examiner, Art Unit 2664