DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending regarding this application.
Priority
The present application claims foreign priority benefits from KR10-2024-0013805 filed on
01/30/2024 and KR10-2025-0009180 filed on 01/22/2025. The certified copies of the priority documents were electronically retrieved on 03/06/2025.
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/23/2025 is considered and attached.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 1, line 17 recites “from another domain”. Furthermore, the only other instance of the word “domain” is in the preamble wherein the claim recites “category-based domain learning” in line 1. As such, it is unclear as to what the term “another” refers to in this context, as there is no other “domain” to which “another domain” can be in reference. Applicant’s specification discusses this subject matter in para. [0015], [0016], and [0025]. However, none of these sections clarify how there can exist “another domain” when no first domain is introduced in the claim. As such, claim 1 is rejected for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Additionally, regarding claim 1, claim 1 recites “an object in an image from another domain” in line 17. Claim 1 additionally recites “an image with strong augmentation” in line 11 and “an image with weak augmentation […] for which an object is to be detected” in lines 5-6. As such, it is unclear whether the object and image as recited in line 17 are equivalent to or distinct from the image as recited in lines 5 and 11, and the object as recited in line 6. Applicant discusses the object/image as recited in line 17 of claim 1 in para. [0015], [0016], and [0025] of applicant’s specification. However, none of these sections clarify whether the object and image as recited in line 17 of claim 1 are equivalent to the image as recited in lines 5 and 11, and the object as recited in line 6 of claim 1. Similar analysis may be applied to the phrase “the detecting the object in an image from another domain” as recited in claim 10 and corresponding claim 20. As such, claims 1, 10, and 20 are rejected for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Similar analysis can be applied to independent claim 11. Claim 11 is similarly rejected.
Claims 2-10 and 12-20 are rejected due to their dependency upon rejected claims 1 and 11.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 5, 11, 12, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Vu et al. (U.S. Publication No. 2024/0161477 A1), hereinafter Vu in view of Li et al. (U.S. Publication No. 2023/0154167 A1), hereinafter Li and Luo et al. (“Exploiting Negative Learning for Implicit Pseudo Label Rectification in Source-Free Domain Adaptive Semantic Segmentation”), hereinafter Luo.
Regarding claim 1, Vu teaches a multi-domain object detection method through category-based domain learning, performed by a computing device including a processor and a memory (Vu teaches “a method for performing cross-domain training of an object detection machine learning model” in para. [0007] wherein the method may be implemented by a computing device including a processor and a memory as shown in para. [0059]), the method comprising:
inputting, by the processor, an image with weak augmentation applied to a target image, for which an object is to be detected, to the teacher model (Vu teaches “the weakly augmented target images 113 generated by weak augmentation module 112 are processed by teacher machine learning model 116” as shown in para. [0029]. See FIG. 2 wherein the unlabeled target image 104 is input into the weak augmentation 112 and eventually input into the teacher model);
determining, by the processor, whether a pseudo label generated by the teacher model is below a preset threshold (Vu teaches “to further control and reduce pseudo labeling noise, in some implementations, a confidence threshold (e.g. 0.8) may be used to remove low-quality, false-positive bounding boxes” in para. [0035]);
inputting, by the processor, an image with strong augmentation applied to the target image to the student model (Vu teaches “strong augmentations such as color jittering, grayscaling, blurring, and/or cutout, may be applied by strong augmentation module 114 to the input(s) processed by student machine learning model 116′ to encourage robustness against appearance changes” in para. [0035]. See FIG. 2 wherein the unlabeled target image 104 is input into the strong augmentation 114 and eventually input into the student model);
calculating, by the processor, an unsupervised loss by comparing a first prediction generated by the student model with the pseudo label (Vu teaches obtaining an unsupervised distillation loss 126 generated by the student model based on the pseudo labels generated by the teacher MLM 116 and the predictions output by the student model as shown in para. [0032] and FIG. 2. Here, since the pseudo labels output by the teacher model 116 are used to train the student model 116’ and the unsupervised distillation loss 126 is determined based on the output of the teacher model, it is determined that the unsupervised loss is calculated by comparing the prediction of the student model with the pseudo label generated by the teacher model);
updating, by the processor, the teacher model using an exponential moving average (EMA) predetermined in the student model (Vu teaches “weights of the teacher machine learning model 116 may be iteratively updated based on an EMA of weights of the object detection model [student model]” as shown in para. [0032] and FIG. 2); and
detecting, by the processor, an object in an image from another domain using the teacher model (Vu teaches “during the adaptation phase, and leveraging the pseudo labels 124 generated by teacher machine learning model 116, mixer 106 (particularly iterations 106C and 106D) may perform intra-domain mixing with both labeled source (source-source) and pseudo labeled target data (target-target)” as shown in para. [0046] and FIG. 2, wherein “the mixed-domain teacher machine learning model 116 may be considered “holistic” in the sense that it leverages the joint mixing of both intra-domain and inter-domain data to simultaneously tackle the challenges of missing object ground truths, bias toward source domain, and pseudo label noise” as shown in para. [0043]. Here, the teacher model detects an object in the target domain 104, wherein the “target domain” is interpreted as equivalent to the claimed “another domain”).
While the warmup phase (Vu, see FIG. 1) may be interpreted as equivalent to a pre-training phase, Vu fails to teach generating, by the processor, a teacher model and a student model from a pre-trained model; and when the pseudo label is determined to be below the preset threshold, performing, by the processor, negative learning for a class corresponding to the pseudo label.
However, Li teaches generating, by the processor, a teacher model and a student model from a pre-trained model (Li teaches a method for cross domain detection method with strong data augmentation (see para. wherein the student and teacher model are initialized with a pretrained model as shown in para. [0059])).
Vu and Li are both considered to be analogous to the claimed invention because they are in the same field of using a teacher student network framework to train an object detection model using strongly augmented images. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu to incorporate the teachings of Li and include “generating, by the processor, a teacher model and a student model from a pre-trained model”. The motivation for doing so would have been to “improve the cross-domain detection performance and outperforms many existing domain adaptive detectors that use various sophisticated techniques to align the domains using images from both domains”, as suggested by Li in para. [0019]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu with Li to obtain the invention specified in the above claim limitation.
While Vu teaches applying a threshold to a pseudo label to remove low-quality, false-positive bounding boxes (see para. [0035]) and Li teaches selecting confidence objects as pseudo ground truth objects by “thresholding the classification probability and measuring the repeatability when mapped to the corresponding strongly views” (see para. [0070]), Vu and Li fail to teach when the pseudo label is determined to be below the preset threshold, performing, by the processor, negative learning for a class corresponding to the pseudo label.
However, Luo teaches when the pseudo label is determined to be below the preset threshold, performing, by the processor, negative learning for a class corresponding to the pseudo label (Luo teaches “Figure 2 illustrates the concept of negative learning. The marked pixel has a less confident pseudo label, which is noisy (labeled as sidewalk instead of road)” wherein “negative learning uses its complementary label” as shown in the subsection “Negative Learning” within the section “Preliminaries”. The negative learning in this instance takes place for the class corresponding to the pseudo label as shown above. Luo additionally teaches that “a pre-defined threshold (e.g., 0.6 in our experiments) is utilized to generate binary invalid masks, which indicates the distribution of potential noisy regions” as shown in the subsection “Noise-Aware Pseudo Label Learning” within the section “Method”).
Vu, Li, and Luo are all considered to be analogous to the claimed invention because they are in the same field of using pseudo labels to train an object detection model. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li) to incorporate the teachings of Luo and include “when the pseudo label is determined to be below the preset threshold, performing, by the processor, negative learning for a class corresponding to the pseudo label”. The motivation for doing so would have been that “negative learning can help avoid the impact of noisy labels via providing correct information even from wrong labels”, as suggested by Luo in the subsection “Negative Learning” within the section “Preliminaries”. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu and Li with Luo to obtain the invention specified in claim 1.
Regarding claim 2, Vu, Li, and Luo teach the method of claim 1, further comprising:
passing, by the processor, a feature map generated by the student model to a first discriminator, and transmitting, by the processor, the feature map to a first head that generates the first prediction (Vu teaches “discriminator 134 may be used to process the feature embedding of object detection model/student machine learning model 116′ in order to classify whether the feature embedding is from a source domain image (e.g., 102) or target domain image (e.g., 104)” in para. [0037]. This classification is interpreted as equivalent to the claimed first prediction. While the first prediction here is not equivalent to the mapping of the first prediction as recited in claim 1, it would have been obvious to one of ordinary skill in the art to combine the teachings of Vu regarding the first prediction as recited in the above mapping of claim 2, and the associated adversarial loss (as shown in FIG. 2 and para. [0033]) with the teachings of the Vu regarding the teaching of the unsupervised loss as taught in the mapping of claim 1, and the motivation for such a combination would have been that, “by performing adversarial training between object detection (student) machine learning model 116′ and discriminator machine learning model 134 using adversarial domain discovery loss 136, object detection (student) machine learning model 116′ is encouraged to learn domain invariant features” as suggested by Vu in para. [0033]).
Regarding claim 5, Vu, Li, and Luo teach the method of claim 1,
wherein the first prediction comprises a class prediction value and a bounding box prediction value (Vu teaches that “both RPN 120′ and ROI layer 122′ may perform bounding box regression and classification tasks, such as binary classification for RPM (e.g., “object” or “no object”) and/or multi-class classification” as shown in para. [0027], wherein RPN 120’ and ROI Layer 122’ are within the student model 116’ as shown in FIG. 2).
Regarding claim 11, Vu teaches a multi-domain object detection apparatus for performing object detection through category-based domain learning, the apparatus comprising a memory storing computer-executable instructions, and at least one processor configured to access the memory and execute the instructions, wherein the instructions comprise (Vu teaches “a method for performing cross-domain training of an object detection machine learning model” in para. [0007] wherein the method may be implemented by a computing device including a processor and a memory as shown in para. [0059]):
inputting an image with weak augmentation applied to a target image, for which an object is to be detected, to the teacher model (Vu teaches “the weakly augmented target images 113 generated by weak augmentation module 112 are processed by teacher machine learning model 116” as shown in para. [0029]. See FIG. 2 wherein the unlabeled target image 104 is input into the weak augmentation 112 and eventually input into the teacher model);
determining whether a pseudo label generated by the teacher model is below a preset threshold (Vu teaches “to further control and reduce pseudo labeling noise, in some implementations, a confidence threshold (e.g. 0.8) may be used to remove low-quality, false-positive bounding boxes” in para. [0035]);
inputting an image with strong augmentation applied to the target image to the student model (Vu teaches “strong augmentations such as color jittering, grayscaling, blurring, and/or cutout, may be applied by strong augmentation module 114 to the input(s) processed by student machine learning model 116′ to encourage robustness against appearance changes” in para. [0035]. See FIG. 2 wherein the unlabeled target image 104 is input into the strong augmentation 114 and eventually input into the student model);
calculating an unsupervised loss by comparing a first prediction generated by the student model with the pseudo label (Vu teaches obtaining an unsupervised distillation loss 126 generated by the student model based on the pseudo labels generated by the teacher MLM 116 and the predictions output by the student model as shown in para. [0032] and FIG. 2. Here, since the pseudo labels output by the teacher model 116 are used to train the student model 116’ and the unsupervised distillation loss 126 is determined based on the output of the teacher model, it is determined that the unsupervised loss is calculated by comparing the prediction of the student model with the pseudo label generated by the teacher model);
updating the teacher model using an exponential moving average (EMA) predetermined in the student model (Vu teaches “weights of the teacher machine learning model 116 may be iteratively updated based on an EMA of weights of the object detection model [student model]” as shown in para. [0032] and FIG. 2); and
detecting an object in an image from another domain using the teacher model (Vu teaches “during the adaptation phase, and leveraging the pseudo labels 124 generated by teacher machine learning model 116, mixer 106 (particularly iterations 106C and 106D) may perform intra-domain mixing with both labeled source (source-source) and pseudo labeled target data (target-target)” as shown in para. [0046] and FIG. 2, wherein “the mixed-domain teacher machine learning model 116 may be considered “holistic” in the sense that it leverages the joint mixing of both intra-domain and inter-domain data to simultaneously tackle the challenges of missing object ground truths, bias toward source domain, and pseudo label noise” as shown in para. [0043]. Here, the teacher model detects an object in the target domain 104, wherein the “target domain” is interpreted as equivalent to the claimed “another domain”).
While the warmup phase (Vu, see FIG. 1) may be interpreted as equivalent to a pre-training phase, Vu fails to teach generating a teacher model and a student model from a pre-trained model; and when the pseudo label is determined to be below the preset threshold, performing negative learning for a class corresponding to the pseudo label.
However, Li teaches generating a teacher model and a student model from a pre-trained model (Li teaches a method for cross domain detection method with strong data augmentation (see para. wherein the student and teacher model are initialized with a pretrained model as shown in para. [0059])).
Vu and Li are both considered to be analogous to the claimed invention because they are in the same field of using a teacher student network framework to train an object detection model using strongly augmented images. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu to incorporate the teachings of Li and include “generating a teacher model and a student model from a pre-trained model”. The motivation for doing so would have been to “improve the cross-domain detection performance and outperforms many existing domain adaptive detectors that use various sophisticated techniques to align the domains using images from both domains”, as suggested by Li in para. [0019]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu with Li to obtain the invention specified in the above claim limitation.
While Vu teaches applying a threshold to a pseudo label to remove low-quality, false-positive bounding boxes (see para. [0035]) and Li teaches selecting confidence objects as pseudo ground truth objects by “thresholding the classification probability and measuring the repeatability when mapped to the corresponding strongly views” (see para. [0070]), Vu and Li fail to teach when the pseudo label is determined to be below the preset threshold, performing, by the processor, negative learning for a class corresponding to the pseudo label.
However, Luo teaches when the pseudo label is determined to be below the preset threshold, performing, by the processor, negative learning for a class corresponding to the pseudo label (Luo teaches “Figure 2 illustrates the concept of negative learning. The marked pixel has a less confident pseudo label, which is noisy (labeled as sidewalk instead of road)” wherein “negative learning uses its complementary label” as shown in the subsection “Negative Learning” within the section “Preliminaries”. The negative learning in this instance takes place for the class corresponding to the pseudo label as shown above. Luo additionally teaches that “a pre-defined threshold (e.g., 0.6 in our experiments) is utilized to generate binary invalid masks, which indicates the distribution of potential noisy regions” as shown in the subsection “Noise-Aware Pseudo Label Learning” within the section “Method”).
Vu, Li, and Luo are all considered to be analogous to the claimed invention because they are in the same field of using pseudo labels to train an object detection model. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li) to incorporate the teachings of Luo and include “when the pseudo label is determined to be below the preset threshold, performing, by the processor, negative learning for a class corresponding to the pseudo label”. The motivation for doing so would have been that “negative learning can help avoid the impact of noisy labels via providing correct information even from wrong labels”, as suggested by Luo in the subsection “Negative Learning” within the section “Preliminaries”. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu and Li with Luo to obtain the invention specified in claim 11.
Regarding claim 12, Vu, Li, and Luo teach the apparatus of claim 11, further comprising:
passing a feature map generated by the student model to a first discriminator, and transmitting the feature map to a first head that generates the first prediction (Vu teaches “discriminator 134 may be used to process the feature embedding of object detection model/student machine learning model 116′ in order to classify whether the feature embedding is from a source domain image (e.g., 102) or target domain image (e.g., 104)” in para. [0037]. This classification is interpreted as equivalent to the claimed first prediction. While the first prediction here is not equivalent to the mapping of the first prediction as recited in claim 1, it would have been obvious to one of ordinary skill in the art to combine the teachings of Vu regarding the first prediction as recited in the above mapping of claim 2, and the associated adversarial loss (as shown in FIG. 2 and para. [0033]) with the teachings of the Vu regarding the teaching of the unsupervised loss as taught in the mapping of claim 1, and the motivation for such a combination would have been that, “by performing adversarial training between object detection (student) machine learning model 116′ and discriminator machine learning model 134 using adversarial domain discovery loss 136, object detection (student) machine learning model 116′ is encouraged to learn domain invariant features” as suggested by Vu in para. [0033]).
Regarding claim 15, Vu, Li, and Luo teach the apparatus of claim 11,
wherein the first prediction comprises a class prediction value and a bounding box prediction value (Vu teaches that “both RPN 120′ and ROI layer 122′ may perform bounding box regression and classification tasks, such as binary classification for RPM (e.g., “object” or “no object”) and/or multi-class classification” as shown in para. [0027], wherein RPN 120’ and ROI Layer 122’ are within the student model 116’ as shown in FIG. 2).
Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Vu et al. (U.S. Publication No. 2024/0161477 A1), hereinafter Vu in view of Li et al. (U.S. Publication No. 2023/0154167 A1), hereinafter Li, Luo et al. (“Exploiting Negative Learning for Implicit Pseudo Label Rectification in Source-Free Domain Adaptive Semantic Segmentation”), hereinafter Luo, and Liu et al. (U.S. Publication No. 2021/0407086 A1), hereinafter Liu.
Regarding claim 6, Vu, Li, and Luo teach the method of claim 1.
Vu, Li, and Luo fail to teach performing, by the processor, pre-training based on a pre-configured second discriminator.
However, Liu teaches performing, by the processor, pre-training based on a pre-configured second discriminator (Liu teaches “retrain[ing] the pre-trained image segmentation model according to a loss function of the pre-trained image segmentation model, an adversarial loss function of the first discriminator, and an adversarial loss function of the second discriminator, such iterative loop training being performed until converging to obtain a trained image segmentation model” as shown in para. [0058]. Here, Liu teaches pre-training an image segmentation model based on a second discriminator, wherein the discriminator is already preconfigured to judge the result of the generator and generate a loss value).
Vu, Li, Luo, and Liu are all considered to be analogous to the claimed invention because they are in the same field of using networks to detect objects in different domains. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li and Luo) to incorporate the teachings of Liu and include “performing, by the processor, pre-training based on a pre-configured second discriminator”. The motivation for doing so would have been that “through such an adversarial learning process, the segmentation accuracy of the image segmentation model is improved” and “the trained image segmentation model can reduce a difference between the source domain image and the target domain image, and reduce an error in segmentation of a target domain by the trained image segmentation model, to further enable image visual information outputted by the target domain image in an output space to be more accurate”, as suggested by Liu in para. [0060] and para. [0062], respectively. See also para. [0061]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu, Li, and Luo with Liu to obtain the invention specified in claim 6.
Regarding claim 16, Vu, Li, and Luo teach the apparatus of claim 11.
Vu, Li, and Luo fail to teach performing pre-training based on a pre-configured second discriminator.
However, Liu teaches performing pre-training based on a pre-configured second discriminator (Liu teaches “retrain[ing] the pre-trained image segmentation model according to a loss function of the pre-trained image segmentation model, an adversarial loss function of the first discriminator, and an adversarial loss function of the second discriminator, such iterative loop training being performed until converging to obtain a trained image segmentation model” as shown in para. [0058]. Here, Liu teaches pre-training an image segmentation model based on a second discriminator, wherein the discriminator is already preconfigured to judge the result of the generator and generate a loss value).
Vu, Li, Luo, and Liu are all considered to be analogous to the claimed invention because they are in the same field of using networks to detect objects in different domains. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li and Luo) to incorporate the teachings of Liu and include “performing pre-training based on a pre-configured second discriminator”. The motivation for doing so would have been that “through such an adversarial learning process, the segmentation accuracy of the image segmentation model is improved” and “the trained image segmentation model can reduce a difference between the source domain image and the target domain image, and reduce an error in segmentation of a target domain by the trained image segmentation model, to further enable image visual information outputted by the target domain image in an output space to be more accurate”, as suggested by Liu in para. [0060] and para. [0062], respectively. See also para. [0061]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu, Li, and Luo with Liu to obtain the invention specified in claim 16.
Claims 8 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Vu et al. (U.S. Publication No. 2024/0161477 A1), hereinafter Vu in view of Li et al. (U.S. Publication No. 2023/0154167 A1), hereinafter Li, Luo et al. (“Exploiting Negative Learning for Implicit Pseudo Label Rectification in Source-Free Domain Adaptive Semantic Segmentation”), hereinafter Luo, Liu et al. (U.S. Publication No. 2021/0407086 A1), hereinafter Liu, and Ren et al. (CN 117351192 A, see attached English translation for citations), hereinafter Ren.
Regarding claim 8, Vu, Li, Luo, and Liu teach the method of claim 6.
While Liu teaches “retrain[ing] the pre-trained image segmentation model according to a loss function of the pre-trained image segmentation model, an adversarial loss function of the first discriminator, and an adversarial loss function of the second discriminator, such iterative loop training being performed until converging to obtain a trained image segmentation model” in para. [0058], Vu, Li, Luo, and Liu fails to specifically teach wherein the performing the pre-training comprises repeating the pre-training for a predetermined number of iterations (emphasis added).
However, Ren teaches wherein the performing the pre-training comprises repeating the pre-training for a predetermined number of iterations (Ren teaches “in an alternative embodiment, the reaching the model training convergence condition may be that the number of training iterations reaches a preset number of training iterations” in para. [0116]. Here, Ren’s teaching can be combined with Liu’s teaching of pre-training a model for a number of iterations until convergence to teach the above claim limitation).
Vu, Li, Luo, Liu, and Ren are all considered to be analogous to the claimed invention because they are in the same field of using networks to detect objects. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li, Luo, and Liu) to incorporate the teachings of Liu and include “wherein the performing the pre-training comprises repeating the pre-training for a predetermined number of iteration”. The motivation for doing so would have been that “the unknown object appearing in the input potential image can be positioned by the network by utilizing the regression coordinates and the positioning confidence score of the region of the output prediction frame”, as suggested by Ren in para. [0113]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu, Li, Luo and Liu with Ren to obtain the invention specified in claim 8.
Regarding claim 18, Vu, Li, Luo, and Liu teach the apparatus of claim 16.
While Liu teaches “retrain[ing] the pre-trained image segmentation model according to a loss function of the pre-trained image segmentation model, an adversarial loss function of the first discriminator, and an adversarial loss function of the second discriminator, such iterative loop training being performed until converging to obtain a trained image segmentation model” in para. [0058], Vu, Li, Luo, and Liu fails to specifically teach wherein the performing the pre-training comprises repeating the pre-training for a predetermined number of iterations (emphasis added).
However, Ren teaches wherein the performing the pre-training comprises repeating the pre-training for a predetermined number of iterations (Ren teaches “in an alternative embodiment, the reaching the model training convergence condition may be that the number of training iterations reaches a preset number of training iterations” in para. [0116]. Here, Ren’s teaching can be combined with Liu’s teaching of pre-training a model for a number of iterations until convergence to teach the above claim limitation).
Vu, Li, Luo, Liu, and Ren are all considered to be analogous to the claimed invention because they are in the same field of using networks to detect objects. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li, Luo, and Liu) to incorporate the teachings of Liu and include “wherein the performing the pre-training comprises repeating the pre-training for a predetermined number of iteration”. The motivation for doing so would have been that “the unknown object appearing in the input potential image can be positioned by the network by utilizing the regression coordinates and the positioning confidence score of the region of the output prediction frame”, as suggested by Ren in para. [0113]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu, Li, Luo and Liu with Ren to obtain the invention specified in claim 18.
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Vu et al. (U.S. Publication No. 2024/0161477 A1), hereinafter Vu in view of Li et al. (U.S. Publication No. 2023/0154167 A1), hereinafter Li, Luo et al. (“Exploiting Negative Learning for Implicit Pseudo Label Rectification in Source-Free Domain Adaptive Semantic Segmentation”), hereinafter Luo, and Ligocki et al. (“Fully Automated DCNN-Based Thermal Images Annotation Using Neural Network Pretrained on RGB Data”), hereinafter Ligocki.
Regarding claim 10, Vu, Li, and Luo teach the method of claim 1.
Vu, Li, and Luo fail to teach wherein the detecting the object in an image from another domain using the teacher model comprises: detecting the object in an infrared (IR) domain related to IR images, a thermal imaging domain related to thermal images, or a Light Detection And Ranging (LiDAR) domain related to LiDAR images using the teacher model trained in an RGB domain related to RGB images.
However, Ligocki teaches wherein the detecting the object in an image from another domain using the teacher model comprises: detecting the object in an infrared (IR) domain related to IR images, a thermal imaging domain related to thermal images, or a Light Detection And Ranging (LiDAR) domain related to LiDAR images using the teacher model trained in an RGB domain related to RGB images (Ligocki teaches a method of training a teacher network using rgb images, and using said teacher network to detect objects in a thermal image domain as shown in Figure 10, wherein the teacher network was trained on color images as shown in Section 4.4 Reducing No. of Classes).
Vu, Li, Luo, and Ligocki are all considered to be analogous to the claimed invention because they are in the same field of using networks to detect objects in different domains. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li and Luo) to incorporate the teachings of Ligocki and include “wherein the detecting the object in an image from another domain using the teacher model comprises: detecting the object in an infrared (IR) domain related to IR images, a thermal imaging domain related to thermal images, or a Light Detection And Ranging (LiDAR) domain related to LiDAR images using the teacher model trained in an RGB domain related to RGB images”. The motivation for doing so would have been to be “able to map detections from RGB into the thermal image” and to “present a technique to leverage the usability of [RGB object detection datasets [which] are quite well-covered in science, and some projects provide even millions of RGB annotated images] and harness them in IR image processing”, as suggested by Ligocki in Section 4.3 and Section 1.1, respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu, Li, and Luo with Ligocki to obtain the invention specified in claim 10.
Regarding claim 20, Vu, Li, and Luo teach the apparatus of claim 11.
Vu, Li, and Luo fail to teach wherein the detecting the object in an image from another domain using the teacher model comprises: detecting the object in an infrared (IR) domain related to IR images, a thermal imaging domain related to thermal images, or a Light Detection And Ranging (LiDAR) domain related to LiDAR images using the teacher model trained in an RGB domain related to RGB images.
However, Ligocki teaches wherein the detecting the object in an image from another domain using the teacher model comprises: detecting the object in an infrared (IR) domain related to IR images, a thermal imaging domain related to thermal images, or a Light Detection And Ranging (LiDAR) domain related to LiDAR images using the teacher model trained in an RGB domain related to RGB images (Ligocki teaches a method of training a teacher network using rgb images, and using said teacher network to detect objects in a thermal image domain as shown in Figure 10, wherein the teacher network was trained on color images as shown in Section 4.4 Reducing No. of Classes).
Vu, Li, Luo, and Ligocki are all considered to be analogous to the claimed invention because they are in the same field of using networks to detect objects in different domains. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Vu (as modified by Li and Luo) to incorporate the teachings of Ligocki and include “wherein the detecting the object in an image from another domain using the teacher model comprises: detecting the object in an infrared (IR) domain related to IR images, a thermal imaging domain related to thermal images, or a Light Detection And Ranging (LiDAR) domain related to LiDAR images using the teacher model trained in an RGB domain related to RGB images”. The motivation for doing so would have been to be “able to map detections from RGB into the thermal image” and to “present a technique to leverage the usability of [RGB object detection datasets [which] are quite well-covered in science, and some projects provide even millions of RGB annotated images] and harness them in IR image processing”, as suggested by Ligocki in Section 4.3 and Section 1.1, respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Vu, Li, and Luo with Ligocki to obtain the invention specified in claim 20.
Allowable Subject Matter
Claims 3, 4, 7, 9, 13, 14, 17, and 19 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter.
The best prior art of record is Vu, Li, Luo, Ligocki, Liu, Ren, and Yu et al. (U.S. Publication No. 2022/0261593 A1), hereinafter Yu. Prior art applied alone or in combination with fails to anticipate or render obvious claims 3, 4, 7, 9, 13, 14, 17, and 19.
Claim 3
Regarding Claim 3, Vu, Li, and Luo teach the method of claim 1. Luo further teaches wherein the determining whether the pseudo label generated by the teacher model is below a preset threshold comprises: the pseudo label comprises a plurality of pseudo labels; determining, by the processor, whether a first pseudo label
Vu further teaches wherein the determining whether the pseudo label generated by the teacher model is below a preset threshold comprises: the pseudo label comprises a plurality of pseudo labels; determining, by the processor, whether a first pseudo label
However, neither Vu, nor Li, nor Luo, nor Ligocki, nor Yu, nor Liu, nor Ren, nor the combination, teaches that the performing the negative learning comprises: selecting, by the processor, k pseudo labels, where k is a natural number, from the plurality of pseudo labels, excluding the first pseudo label, when the first pseudo label is determined to be below the preset threshold; and performing, by the processor, the negative learning for classes corresponding to the first pseudo label and the k pseudo labels, in combination with the other elements of the claim.
Similar analysis is applicable to claim 13.
Claim 4
Regarding Claim 4, Vu, Li, and Luo teach the method of claim 1.
Luo further teaches using negative learning for implicit pseudo label rectification.
However, neither Vu, nor Li, nor Luo, nor Ligocki, nor Yu, nor Liu, nor Ren, nor the combination, teaches performing, by the processor, the negative learning based on a negative learning loss according to:
PNG
media_image1.png
68
332
media_image1.png
Greyscale
wherein B is a batch size, C is an object category class, II is an indicator function,
p
c
(
i
)
is a probability that a sample does not belong to class c,
q
c
(
i
)
is a probability predicted by a model for class c of i-th sample, Rank is ranking sorted in descending order based on confidence scores, and k is the top k ranks calculated adaptively, in combination with the other elements of the claim.
Similar analysis is applicable to claim 14.
Claim 7
Regarding Claim 7, Vu, Li, and Luo teach the method of claim 6. Vu and Li teach inputting, by the processor, a pre-configured dataset to a backbone to generate a feature map.
Vu further teaches passing, by the processor, the feature map to the pre-configured
However, neither Vu, nor Li, nor Luo, nor Ligocki, nor Yu, nor Liu, nor Ren, nor the combination, teaches that the second discriminator, second head, and calculating, by the processor, a supervised loss by comparing the second prediction with ground truth, and updating weights through backpropagation, in combination with the other elements of the claim.
Similar analysis is applicable to claim 17.
Claims 9 and 19 include allowable subject matter by virtue of being dependent upon claims 7 and 17 respectively.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Li et al. (“Cross-Domain Adaptive Teacher for Object Detection”) teaches determining Pseudo-labels from the target-domain teacher model and comparing them with the output of the student model to generate an unsupervised loss.
Any inquiry concerning this communication or earlier communications from the examiner
should be directed to KYLA G ALLEN whose telephone number is (703)756-5315. The examiner can
normally be reached M-F 7:30am - 4:30pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a
USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use
the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor,
John Villecco can be reached on (571) 272-7319. The fax phone number for the organization where this
application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from
Patent Center. Unpublished application information in Patent Center is available to registered users. To
file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit
https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and
https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional
questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like
assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or
571-272-1000.
/Kyla Guan-Ping Tiao Allen/
Examiner, Art Unit 2661
/AARON W CARTER/Primary Examiner, Art Unit 2661