DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Applicants Amendments filed on July 14, 2026, has been entered and made of record.
Currently pending Claim(s): 1-20
Independent Claim(s): 1, 8, 15
Amended Claim(s): 1, 5, 8, 15
Canceled Claim(s): 3, 6, 14, 18
Specification Objections
In view of Applicant’s amendments to paragraphs [0004-0006], the previous objection to the Specification is withdrawn.
Claim Objections
In view of Applicant’s amendments to Claims 1, 8, and 15, the previous objections to Claims 1, 8, and 15 are withdrawn.
Claim Rejections – 35 U.S.C. 101
In view of Applicant’s amendments to Claims 15, the previous 101 rejection of Claim 15-17 and 19-20 are withdrawn.
Response to Arguments
This office action is responsive to the Applicant’s Arguments/Remarks Made in an Amendment
received on July 14, 2026.
In view of amendments filed on, the Applicant has amended independent Claim 1 to recite the additional limitation of “by at least conditioning a transformer decoder on specific dataset semantics by applying dataset-specific query embeddings to object queries, wherein the object queries identify objects within the received images and the dataset- specific query embeddings are the same dimensionality as the object queries; generating predictions for segmentation masks and classes for the received images using the trained transformer-based segmentation model”. Independent claims 8 and 15 have been amended in a similar manner. The applicant has also incorporated the limitations of Claim 6 into Claim 1.
Originally, (from the claim set dated July 14, 2026), the Examiner rejected Claims 1 and 2 over Zhou (Q. Zhou, et al., “LMSeg: Language-Guided Multi-Dataset Segmentation”, International Conference on Learning Representations (ICLR), 2023, pp. 1-12). Claim 6 was rejected over Zhou in view of Weber (M. Weber et al., "Single-Shot Panoptic Segmentation," 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 2020, pp. 8476-8483).
As discussed in the following paragraphs, the Zhou does not render obvious the newly amended claims. However, the Applicant’s amendment necessitated the new grounds of rejection presented in this Office action. Upon conducting a new search, the Examiner argues that the newly amended Claim(s) 1, 8, and 15 are unpatentable over Zhou, Meng, and Webber, since Meng teaches the limitation of applying dataset-specific query embeddings to object queries.
In view of Applicant’s Arguments/Remarks filed on July 14, 2026, with respect to the claims, the Applicant explained (on Remarks pg. 12, paragraph 2) that Zhou fails to teach the limitation of ‘training a transformer based segmentation model by at least by at least conditioning a transformer decoder on specific dataset semantics by applying dataset-specific query embeddings to object queries, wherein the object queries identify objects within the received images and the dataset- specific query embeddings are the same dimensionality as the object queries’. The Applicant explained that the pixel embeddings of Zhou are ‘structurally and functionally different from the object queries’, and that the pixel embeddings are not used to ‘condition a transformer decoder’. The Examiner agrees. Zhou fails to teach that each pixel embedding is used to identify objects, nor does Zhou teach conditioning a transformer with the pixel embeddings.
The Applicant then argued (on Remarks pg. 15, paragraph 8) that Weber fails to teach the limitations of Claim 6. The Applicant explained that Weber discloses a panoptic head with a distinct way of dealing with different-class overlaps, and a distinct way of dealing with same class overlaps, and thus, Weber’s panoptic head cannot teach resolving annotations of different classes.
The Examiner respectfully disagrees. Weber teaches that the panoptic head can be combined with the method for solving intra-class overlaps (see pg.4, Section III, Subsection B, “Hence, we pro pose three different policies to overcome this shortcoming. Combined with the panoptic head, we can solve both kinds of overlaps”). Thus, the combined algorithm can address different classes using the highest-confidence policy and smallest first policy. Thus, it is clear that Weber teaches the limitation of wherein “the panoptic inference algorithm resolves conflicting annotations from the multiple datasets by allowing a smaller mask to override a larger mask if both have confidences above a certain threshold and the smaller mask is fully contained within the larger mask and is of a different class.” Therefore, for the reasons cited above, the Examiner maintains Weber.
Thus, the Applicant’s amendment necessitated the new grounds of rejection presented in Office Action, and Claim 1 is rejected under 35 USC 103 as being unpatentable over Zhou, Meng, and Weber. Claims 8 and 15 are rejected over Zhou and Meng. Therefore, the rejections to the dependent claims are maintained.
Claim Objections
Claims 16-17 and 19-20 objected to because of the following informalities:
Claim 16-17 and 19-20 are directed towards a ‘computer program product’. However, Claim 15 is directed towards a “non-transitory computer program product”. The Examiner suggests amending Claims 16-17 and 19-20 to recite “The non-transitory computer program product of claim..”
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-2, 4-5, and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou et al. (Q. Zhou, et al., “LMSeg: Language-Guided Multi-Dataset Segmentation”, International Conference on Learning Representations (ICLR), 2023, pp. 1-12), hereinafter Zhou, further in view of Meng et al. (L. Meng, et al., “Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language Embedding”, arXiv, 2023), hereinafter Meng, and further in view of in view of Weber et al (M. Weber et al., "Single-Shot Panoptic Segmentation," 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 2020, pp. 8476-8483), hereinafter Weber.
As to Claim 1, Zhou teaches a computer-implemented method for multi-dataset panoptic segmentation (see pg. 1, Abstract, “In this paper, we investigate the multi-dataset segmentation and propose a scalable Language-guided Multi-dataset Segmentation framework, dubbed LMSeg, which supports both semantic and panoptic segmentation”),
comprising: processing received images from multiple datasets (see pg. 4, Figure 3, multiple images from Datasets 1-N are input into the model),
to extract multi-scale features (see pg.4, Section 3, Subsection 3.1, “The image is first preprocessed using the proposed dataset-aware augmentation strategy, and then image features are extracted through the image encoder and pixel decoder”),
using a backbone network (see pg. 5, Section 3, Subsection 3.2, “The image encoder can use arbitrary backbone models, not limited to the ResNet (He et al., 2016) as we use in this work”),
each of the multiple datasets including a unique label space (see pg., 4, Figure 3, where each dataset has their own label space);
generating text-embeddings for class names from the unique label space for each of the multiple datasets (see pg. 4, Section 3, Subsection 3.1, “The class names are mapped to text embeddings by the pre-trained text encoder”);
integrating the text-embeddings with visual features extracted from the received images (see pg. 4, Section 3, Subsection 3.1, “The category-guided decoding module bridges the text embeddings and the image features”),
to create a unified semantic space (see pg. 2, Section 1, “we introduce a pre-trained text encoder to automatically map the category identification to a unified representation”);
training a transformer-based segmentation at least conditioning a transformer decoder on specific dataset semantics (see pg. 6, Section 3, Subsection 3.4, Through prediction redirection, we can arbitrarily specify the categories that the model needs to predict so that we can use the original annotations of each dataset….We propose a category-guided decoding module to dynamically adapt to classes to be predicted by the model, as shown in Figure 4. The decoder module follows the standard architecture of the transformer, using multi-head self- and cross-attention mechanisms and an FFN module to transform N segment queries _query. The self-attention to query embeddings enables the model to make global inferences for all masks using pairwise relationships between them.”, and see Table 1, where the model training data is shown);
generating predictions for segmentations masks and classes for the received images using the trained transformer-based segmentation model (see pg.4, Section 3, Subsection 3.1, “The LMSeg is decomposed of an encoder-decoder pixel feature extractor, a pre-trained text encoder, a Transformer decoder with category guided module” and see pg. 5, Section 3.4, “Finally, we obtain each binary mask prediction”);
and generating a unified panoptic segmentation map from the predicted segmentation masks and classes by performing inference using a panoptic inference algorithm (see pg. 6, Section 3.4, “The self-attention to query embeddings enables the model to make global inferences for all masks using pairwise relationships between them” and see pg. 9, Section 5, “we propose a language-guided multi-dataset segmentation framework that supports both semantic and panoptic segmentation”).
Zhou fails to explicitly teach that the transformer model is trained by at least conditioning a transformer decoder on specific dataset semantics by applying dataset-specific query embeddings to object queries, wherein the object queries identify objects within the received images and the dataset- specific query embeddings are the same dimensionality as the object queries.
However, in an analogous art, Meng teaches a method for dataset aware object detection (see pg. 1, Abstract, “Combining multiple datasets enables performance boost on many computer vision tasks. But similar trend has not been witnessed in object detection when combining multiple datasets due to two inconsistencies among detection datasets: taxonomy difference and domain gap. In this paper, we address these challenges by a new design (named Detection Hub) that is dataset-aware and category-aligned"),
which comprises training a transformer model (see pg. 3, Section 3, Subsection 3. 1, “Given a set of N learnable queries v ∈ RN×d, an end-to-end detector utilizes a transformer T to generate N corresponding predictions”, and see pg. 4, Figure 2., shown below)
PNG
media_image1.png
500
958
media_image1.png
Greyscale
Figure 2 of Meng
by applying dataset-specific query embeddings to object queries (see pg. 4, Section 3, Subsection 3.3, “As described above, to fully unleash the power of large amount of data in different datasets, our method proposes to use the dataset-specific language embedding to adapt the
object queries so that the model can learn to adapt its behavior for each dataset”),
wherein the object queries identify objects within the received images (see pg. 3, Section 3, Subsection 3.1, “End-to-end detectors [2, 3, 37] utilize object queries to encode the content and position statistics over the training dataset and drive the detector to predict desired objects”), wherein the training dataset comprises images),
and the dataset- specific query embeddings are the same dimensionality as the object queries (see pgs. 4-5, Section 3, Subsection 3.3, “In detail, given a dataset D and its language embedding E, we can achieve the query adaptation by simply performing a cross-attention between the
learnable queries Q and E,, Conceptually, through cross-attention, we encode the dataset-specific language embedding into the adapted queries
Q
D
, which is used as object query as Fig 2 and makes our detector dataset-aware”, and see Formula 7, shown below).
PNG
media_image2.png
44
460
media_image2.png
Greyscale
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the dataset-specific query embeddings and object queries taught by Meng with the panoptic segmentation method taught by Zhou. The motivation for doing so would be to mitigate dataset inconsistency (see pg. 1, Abstract, “In this paper, we address these challenges by a new design (named Detection Hub) that is dataset-aware and category-aligned. It not only mitigates the dataset inconsistency but also provides coherent guidance for the detector to learn across multiple datasets. In particular, the dataset-aware design is achieved by learning a dataset embedding that is used to adapt object queries as well as convolutional kernels in detection heads”).
Both Zhou and Meng fail to explicitly teach wherein the panoptic inference algorithm resolves conflicting annotations from the multiple datasets by allowing a smaller mask to override a larger mask if both have confidences above a certain threshold and the smaller mask is fully contained within the larger mask and is of a different class.
However, in an analogous art Weber teaches a panoptic inference algorithm that resolves overlapping masks (see pg.4, Section III, Subsection B, “Combined with the panoptic head, we can solve both kinds of overlaps”),
that teaches prioritizing smaller masks contained in larger masks (see Section III, Subsection B. Overlap Resolution, pg. 4, “Hence, sorting overlapping instances in increasing order of size prevents large instances from overshadowing smaller ones.” )
if both confidences above a certain threshold (see Section IV, Subsection B, page 5, “We choose the highest-confidence policy as this strategy is used in previous work to resolve overlaps in Mask R-CNN. Additionally, we apply a confidence threshold of 0.4 to the outputs of our detector”)
and if the masks are of different class (see Section V, page 7, “Still, our panoptic head can resolve inter-and intra-class overlaps by combining semantic segmentation, object detection and instance center prediction”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the panoptic inference algorithm taught by Weber with the teachings of Zhou and Meng. The motivation for doing so would be to correct false segmentation outputs. Weber teaches in Section I, pg. 1, “This novel approach can resolve both intra- and inter-class overlaps to obtain instance segmentation and is able correct false predictions of individual components.”). Thus, it would have been obvious to combine the inference algorithm taught by Weber with the panoptic segmentation method taught by Zhou and Meng in order to obtain the invention as claimed in Claim 1.
As to Claim 2, Zhou teaches that the backbone network comprises a convolutional neural network or a transformer network that processes the images to extract multi-scale visual features (see pg. 5, Section 3.2, “The image encoder can use arbitrary backbone models, not limited to the ResNet (He et al., 2016) as we use in this work”, where ResNet is a well-known convolutional neural network in the art).
As to Claim 4, Zhou teaches that the text-embeddings are generated using a pre-trained vision-and-language model (see Section 3, pg. 3, “The text encoder takes the class names of a dataset as input and outputs the corresponding text embeddings… The implementation of the text encoder follows CLIP... During the parameters of the text encoder are initialized with a pre-trained CLIP model”, where CLIP is a well-known vision-and-language model in the art),
that maps category names of different datasets into a single consistent space where semantic relations are preserved (see pg.3, Figure 2, “visualization of category embeddings for several semantic segmentation datasets using CLIP’s text encoder for feature extraction. The text embedding space of CLIP is suitable as a unified taxonomy, with semantically similar categories holding closer text embeddings”).
As to Claim 5, Zhou teaches that the pre-trained vision-and-language model is a Contrastive Language-Image Pre-training (CLIP) model for generating text-embeddings that map category names from different datasets into a consistent semantic space model (see Section 3.2, pg. 4, “The text encoder takes the class names of a dataset as input and outputs the corresponding text embeddings… The implementation of the text encoder follows CLIP”).
As to Claim 7, Zhou teaches wherein the transformer-based segmentation model is trained to handle overlapping label spaces from the multiple datasets (see pg.3, Figure 2, overlapping label spaces from multiple datasets shown, see Section 1, pg. 2, “To this end, we introduce a category-guided decoding (CGD) module to guide the model to predict involved labels for the specified taxonomy”)
enhancing a robustness and semantic understanding of the transformer-based segmentation model (see pg.8, Section 4.2.1, “This experimental results show that compared with single-dataset training, multi-dataset training can obtain a single robust model capable of recognizing more categories and improving the average performance on multiple datasets”).
Claims 8-10, 12-13, 15-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou et al. (Q. Zhou, et al., “LMSeg: Language-Guided Multi-Dataset Segmentation”, International Conference on Learning Representations (ICLR), 2023, pp. 1-12), hereinafter Zhou, in view of Meng et al. (L. Meng, et al., “Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language Embedding”, arXiv, 2023), hereinafter Meng, and further in view of Xu et al. (US Pub No 20240153093), hereinafter Xu.
As to Claim 8, Zhou teaches a model to process received images from multiple datasets (see pg. 4, Figure 3, multiple images from Datasets 1-N are input into the model),
to extract multi-scale features (see pg. 4, Section 3, Subsection 3.1),
using a backbone network (see pg. 5, Section 3, Subsection 3.2),
each of the multiple datasets including a unique label space (see pg. 4, Figure 3);
generate text-embeddings for class names from the unique label space for each of the multiple datasets (see pg. 4, Section 3, Subsection 3.1),
integrate the text-embeddings with visual features extracted from the received images (see pg. 4, Section 3, Subsection 3.1),
train a transformer-based segmentation model by at least conditioning a transformer decoder on specific dataset semantics (see pg. 6, Section 3, Subsection 3.4, and see Table 1);
generate predictions for segmentation masks and classes for the received images using the trained transformer-based segmentation model (see pg.4, Section 3, Subsection 3.1); and
and generate a unified panoptic segmentation map from the predicted segmentation masks and classes by performing inference using a panoptic inference algorithm (see pg. 6, Section 3.4).
Zhou fails to explicitly teach that the transformer model is trained by at least conditioning a transformer decoder on specific dataset semantics by applying dataset-specific query embeddings to object queries, wherein the object queries identify objects within the received images and the dataset- specific query embeddings are the same dimensionality as the object queries.
However, in an analogous art, Meng teaches a method for dataset aware object detection (see pg. 1, Abstract, “Combining multiple datasets enables performance boost on many computer vision tasks. But similar trend has not been witnessed in object detection when combining multiple datasets due to two inconsistencies among detection datasets: taxonomy difference and domain gap. In this paper, we address these challenges by a new design (named Detection Hub) that is dataset-aware and category-aligned"),
which comprises training a transformer model (see pg. 3, Section 3, Subsection 3. 1, “Given a set of N learnable queries v ∈ RN×d, an end-to-end detector utilizes a transformer T to generate N corresponding predictions”, and see pg. 4, Figure 2., shown below)
by applying dataset-specific query embeddings to object queries (see pg. 4, Section 3, Subsection 3.3, “As described above, to fully unleash the power of large amount of data in different datasets, our method proposes to use the dataset-specific language embedding to adapt the
object queries so that the model can learn to adapt its behavior for each dataset”),
wherein the object queries identify objects within the received images (see pg. 3, Section 3, Subsection 3.1, “End-to-end detectors [2, 3, 37] utilize object queries to encode the content and position statistics over the training dataset and drive the detector to predict desired objects”), wherein the training dataset comprises images),
and the dataset- specific query embeddings are the same dimensionality as the object queries (see pgs. 4-5, Section 3, Subsection 3.3, “In detail, given a dataset D and its language embedding E, we can achieve the query adaptation by simply performing a cross-attention between the
learnable queries Q and E,, Conceptually, through cross-attention, we encode the dataset-specific language embedding into the adapted queries
Q
D
, which is used as object query as Fig 2 and makes our detector dataset-aware”, and see Formula 7, where both the object query Q is the same dimensionality as the learnable dataset specific query embedding E).
PNG
media_image2.png
44
460
media_image2.png
Greyscale
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the dataset-specific query embeddings and object queries taught by Meng with the panoptic segmentation method taught by Zhou. The motivation for doing so would be to mitigate dataset inconsistency (see pg. 1, Abstract).
Both Zhou and Meng fail to explicitly teach a processor device and memory storing instructions. However, in an analogous art, Xu teaches a system for panoptic segmentation (see paragraph [0008], “FIG. 2A illustrates a block diagram of a diffusion-based panoptic segmentation system”),
that includes a processor device (see Fig. 4, parallel processing unit 400)
and a memory (see Fig. 4, Memory 404) comprising instructions for panoptic segmentation (see paragraph [0098], “The computer-storage media may include both volatile and nonvolatile media and/or removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the memory and processor taught by Xu with the multi-dataset panoptic segmentation method taught by Zhou and Meng. The motivation for doing so would be to integrate the panoptic segmentation method into a variety of devices. Xu teaches in paragraph [0085], “The PPU 400 may be included in a desktop computer, a laptop computer, a tablet computer, servers, supercomputers, a smart-phone (e.g., a wireless, hand-held device), personal digital assistant (PDA), a digital camera, a vehicle, a head mounted display, a hand-held electronic device, and the like. In an embodiment, the PPU 400 is embodied on a single semiconductor substrate”. Thus, it would have been obvious to combine the processor and memory taught by Xu with the panoptic segmentation taught by Zhou and Meng in order to obtain the invention as claimed in Claim 8.
As to Claim 9, Zhou in view of Meng and Xu teaches that the text-embeddings are generated using a pre-trained vision-and-language model (see Zhou, Section 3, pg. 3, “The text encoder takes the class names of a dataset as input and outputs the corresponding text embeddings… The implementation of the text encoder follows CLIP... During the parameters of the text encoder are initialized with a pre-trained CLIP model”, where CLIP is a well-known vision-and-language model in the art),
that maps category names of different datasets into a single consistent space where semantic relations are preserved (see pg.3, Figure 2, “visualization of category embeddings for several semantic segmentation datasets using CLIP’s text encoder for feature extraction. The text embedding space of CLIP is suitable as a unified taxonomy, with semantically similar categories holding closer text embeddings”).
As to Claim 10, Zhou in view of Meng and Xu teaches that the pre-trained vision-and-language model is Contrastive Language-Image Pre-training (CLIP) model for generating text-embeddings that map category names from different datasets into a consistent semantic space model (see Zhou, Section 3.2, pg. 4, “The text encoder takes the class names of a dataset as input and outputs the corresponding text embeddings… The implementation of the text encoder follows CLIP”).
As to Claim 12, Zhou in view of Meng and Xu teaches a memory (see Xu, Fig. 4, Memory 404) storing instructions that enables the system to adapt to varying label spaces from the datasets by employing language-based embeddings thereby allowing the system to process images from new or unseen datasets without retraining (see Zhou, pg. 1, Abstract, “However, there are two major issues with multi-dataset segmentation: (i) the inconsistent taxonomy demands manual reconciliation to construct a unified taxonomy; (ii) the inflexible one-hot common taxonomy causes time consuming model retraining and defective supervision of unlabeled categories. In this paper, we investigate the multi-dataset segmentation and propose a scalable Language-guided Multi-dataset Segmentation framework, dubbed LMSeg, which supports both semantic and panoptic segmentation. Specifically, we intro duce a pre-trained text encoder to map the category names to a text embedding space as a unified taxonomy, instead of using inflexible one-hot label…Furthermore, we adopt a dataset-aware augmentation strategy that assigns each dataset a specific image augmentation pipeline, which can suit the properties of images from different datasets”).
As to Claim 13, Zhou in view of Meng and Xu teaches wherein the transformer-based segmentation model is trained to handle overlapping label spaces from the multiple datasets (see Zhou, pg.3, Figure 2, overlapping label spaces from multiple datasets shown, see Section 1, pg. 2, “To this end, we introduce a category-guided decoding (CGD) module to guide the model to predict involved labels for the specified taxonomy”) ,
enhancing a robustness and semantic understanding of the transformer-based segmentation model (see Zhou, pg. 8, Section 4.2.1, “This experimental results show that compared with single-dataset training, multi-dataset training can obtain a single robust model capable of recognizing more categories and improving the average performance on multiple datasets.” )
As to Claim 15, Zhou teaches a model to process received images from multiple datasets (see pg. 4, Figure 3, multiple images from Datasets 1-N are input into the model),
to extract multi-scale features (see pg. 4, Section 3, Subsection 3.1),
using a backbone network (see pg. 5, Section 3, Subsection 3.2),
each of the multiple datasets including a unique label space (see pg. 4, Figure 3);
generate text-embeddings for class names from the unique label space for each of the multiple datasets (see pg. 4, Section 3, Subsection 3.1),
integrate the text-embeddings with visual features extracted from the received images (see pg. 4, Section 3, Subsection 3.1),
train a transformer-based segmentation model by at least conditioning a transformer decoder on specific dataset semantics (see pg. 6, Section 3, Subsection 3.4, and see Table 1);
generate predictions for segmentation masks and classes for the received images using the trained transformer-based segmentation model (see pg.4, Section 3, Subsection 3.1); and
and generate a unified panoptic segmentation map from the predicted segmentation masks and classes by performing inference using a panoptic inference algorithm (see pg. 6, Section 3.4).
Zhou fails to explicitly teach that the transformer model is trained by at least conditioning a transformer decoder on specific dataset semantics by applying dataset-specific query embeddings to object queries, wherein the object queries identify objects within the received images and the dataset- specific query embeddings are the same dimensionality as the object queries.
However, in an analogous art, Meng teaches a method for dataset aware object detection (see pg. 1, Abstract, “Combining multiple datasets enables performance boost on many computer vision tasks. But similar trend has not been witnessed in object detection when combining multiple datasets due to two inconsistencies among detection datasets: taxonomy difference and domain gap. In this paper, we address these challenges by a new design (named Detection Hub) that is dataset-aware and category-aligned"),
which comprises training a transformer model (see pg. 3, Section 3, Subsection 3. 1, “Given a set of N learnable queries v ∈ RN×d, an end-to-end detector utilizes a transformer T to generate N corresponding predictions”, and see pg. 4, Figure 2., shown below)
by applying dataset-specific query embeddings to object queries (see pg. 4, Section 3, Subsection 3.3, “As described above, to fully unleash the power of large amount of data in different datasets, our method proposes to use the dataset-specific language embedding to adapt the
object queries so that the model can learn to adapt its behavior for each dataset”),
wherein the object queries identify objects within the received images (see pg. 3, Section 3, Subsection 3.1, “End-to-end detectors [2, 3, 37] utilize object queries to encode the content and position statistics over the training dataset and drive the detector to predict desired objects”), wherein the training dataset comprises images),
and the dataset- specific query embeddings are the same dimensionality as the object queries (see pgs. 4-5, Section 3, Subsection 3.3, “In detail, given a dataset D and its language embedding E, we can achieve the query adaptation by simply performing a cross-attention between the
learnable queries Q and E,, Conceptually, through cross-attention, we encode the dataset-specific language embedding into the adapted queries
Q
D
, which is used as object query as Fig 2 and makes our detector dataset-aware”, and see Formula 7, shown below, where both the object query Q is the same dimensionality as the learnable dataset specific query embedding E).
PNG
media_image2.png
44
460
media_image2.png
Greyscale
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the dataset-specific query embeddings and object queries taught by Meng with the panoptic segmentation method taught by Zhou. The motivation for doing so would be to mitigate dataset inconsistency (see pg. 1, Abstract).
Both Zhou and Meng fails to explicitly teach a computer program product for multi-dataset panoptic segmentation, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a hardware processor. However, in an analogous art, Xu teaches a system for panoptic segmentation (see paragraph [0008], “FIG. 2A illustrates a block diagram of a diffusion-based panoptic segmentation system”),
that includes a computer program product for multi-dataset panoptic segmentation, the computer program product comprising a computer readable storage medium having program instructions embodied therewith (see paragraph [0098], “The computer-storage media may include both volatile and nonvolatile media and/or removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, and/or other data types. For example, the main memory 540 may store computer-readable instructions (e.g., that represent a program(s) and/or a program element(s), such as an operating system”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the memory and processor taught by Xu with the multi-dataset panoptic segmentation method taught by Zhou and Meng. The motivation for doing so would be to integrate the panoptic segmentation method into a variety of devices (see Xu, paragraph [0085]). Thus, it would have been obvious to combine the computer readable storage medium taught by Xu with the panoptic segmentation taught by Zhou and Meng in order to obtain the invention as claimed in Claim 15.
As to Claim 16, Claim 16 claims the same limitation claimed as Claim 9 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 9.
As to Claim 17, Claim 17 claims the same limitation claimed as Claim 10 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 10.
As to Claim 19, Claim 19 claims the same limitation claimed as Claim 12 and is dependent on a similarly rejected independent claim. Therefore, the rejection and rationale are similar to that of Claim 12.
Claims 11 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou et al. (Q. Zhou, et al., “LMSeg: Language-Guided Multi-Dataset Segmentation”, International Conference on Learning Representations (ICLR), 2023, pp. 1-12), hereinafter Zhou, further in view of Meng et al. (L. Meng, et al., “Detection Hub: Unifying Object Detection Datasets via Query Adaptation on Language Embedding”, arXiv, 2023), hereinafter Men, in view of Xu et al. (US Pub No 20240153093), hereinafter Xu, and further in view of Weber et al (M. Weber et al., "Single-Shot Panoptic Segmentation," 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 2020, pp. 8476-8483), hereinafter Weber.
As to Claim 11, Zhou in view of Meng and Xu fails to teach resolving conflicting annotations by prioritizing smaller masks over larger masks based on containment and class difference.
However, in an analogous art, Weber teaches a method of panoptic segmentation that resolves overlapping masks based on size (see Section III, Subsection B. Overlap Resolution, pg. 4, “Hence, sorting overlapping instances in increasing order of size prevents large instances from overshadowing smaller ones.” ), containment (see Section III, Subsection B. Overlap Resolution, pg. 4, “The overlap is resolved by adding the predicted offset to the pixels’ location and computing the L2 distance to all the centers of boxes containing that pixel”), and class difference (see Section V, page 7, “Still, our panoptic head can resolve inter-and intra-class overlaps by combining semantic segmentation, object detection and instance center prediction”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the panoptic inference algorithm taught by Weber with the teachings of Zhou, Meng, and Xu. The motivation for doing so would be to correct false segmentation outputs. Weber teaches in Section I, pg. 1, “This novel approach can resolve both intra- and inter-class overlaps to obtain instance segmentation and is able correct false predictions of individual components.”).
Thus, it would have been obvious to combine the inference algorithm taught by Weber with the panoptic segmentation system taught by Zhou and Xu in order to obtain the invention as claimed in Claim 11.
As to Claim 20, Zhou in view of Meng and Xu teaches evaluating the trained model's performance using metrics that assess the model performance of overlapping label spaces (see Zhou, pg. 7, Tables 2 and 3, where the panoptic segmentation accuracy (PQ) is used as a metric to evaluate the model’s performance), enhancing its utility in diverse application scenarios (see pg. 1, Section 1, “Image Segmentation has been a longstanding challenge in computer vision and plays a pivotal role in a wide variety of applications ranging from autonomous driving (Levinson et al., 2011; Maurer et al., 2016) to remote sensing image analysis (Ghassemian, 2016). Building a general and inclusive segmentation model is meaningful to real-world applications.”)
Zhou in view of Meng and Xu fails to explicitly teach that the panoptic inference algorithm includes steps for sequentially placing segmentation masks based on their confidence scores and sizes to effectively manage overlapping masks, ensuring accurate segmentation outcomes.
However, in an analogous art, Weber teaches a panoptic inference algorithm that includes steps for sequentially placing segmentation masks (see page 4, Section IIII, Subsection B, “Hence, we pro pose three different policies to overcome this shortcoming. Combined with the panoptic head, we can solve both kinds of overlaps, thus making the use of instance segmentation subnetworks optional”),
Based on confidence scores (see pg., 4, Section III, Subsection B, “Highest-Confidence Policy: This policy is similar to the way previous work handles overlapping instances (of any class) predicted by Mask R-CNN. Assuming the confidence score correlates strongly to the probability of the existence of an object, sorting the overlapping bounding boxes in decreasing order of confidence leads to a higher score”) and sizes (see pg. 4 Section III, Subsection B. Overlap Resolution, pg. 4, “Hence, sorting overlapping instances in increasing order of size prevents large instances from overshadowing smaller ones” ),
to effectively manage overlapping masks, ensuring accurate segmentation outcomes (see pg. 4, Section III, Subsection B., “With this policy, very likely predictions could hide false positives”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the panoptic inference algorithm taught by Weber with the panoptic segmentation system taught by Zhou, Meng, and Xu. The motivation for doing so would be to correct false segmentation outputs (see Weber Section I, pg. 1). Thus, it would have been obvious to combine the panoptic inference algorithm taught by Weber with the panoptic segmentation system taught by Zhou, Meng, and Xu in order to obtain the invention as claimed in Claim 20.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Smock et al. (US Pub No 20220335240) teaches identifying smaller masks of a different class than an overlapping larger mask.
Lee et al. (US Pub No 20240151814) teaches a method for eliminating overlapping masks of different class using a confidence threshold.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOUMYA THOMAS whose telephone number is (571)272-8639. The examiner can normally be reached M-F 8:30-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.T./Examiner, Art Unit 2664
/JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664