DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Notice to Applicants
2. This communication is in response to the application filed on 09/26/2024.
3. Claims 1-20 are pending.
4. Limitations appearing inside {} are intended to indicate the limitations not taught by said prior art(s)/combinations.
Information Disclosure Statement
5. The information disclosure statements (IDS) submitted on 11/27/2024 has been considered by the examiner.
Claim Rejections - 35 USC § 101
6. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
7. Claims 1-7, and 10-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis below follows Subject Matter Eligibility Test (See flowchart in MPEP 2106).
Step 1: Is the claim to a process, machine, manufacture or composition of matter? YES.
Step2A, Prong 1: Does the claim recite an abstract idea, law of nature, or natural phenomenon? YES. Claim 1 recites “A computer-implemented method comprising (1): generating, by at least one processor, one or more group masks corresponding to user-tagged groups of objects in a vector image by executing a search on a vector hierarchy of the vector image (2); determining, by the at least one processor and from the one or more group masks, a group mask comprising a semantically relevant set of objects based on semantic information from one or more segmentation masks comprising a plurality of semantic segmentations generated utilizing one or more segmentation neural networks (3); and extracting, by the at least one processor and from the group mask, one or more masks corresponding to the semantically relevant set of objects. (4)” [emphasis added].
Limitation (2), (3), and (4) recite a mental process, directed toward the mental process grouping of abstract ideas (MPEP 2106.04(a)(2)). Specifically, a person, can generate one or more group masks corresponding to user-tagged groups of objects in a vector image by executing a search on a vector hierarchy of the vector image (e.g., if a vector hierarchy includes “animals”, “cats”, and “dogs”, highlight all animals, cats, and dogs in an image blue, since cats and dogs are both animals), determine, from the one or more group masks, a group mask comprising a semantically relevant set of objects based on semantic information from one or more segmentation masks comprising a plurality of semantic segmentations and extract, from the group mask one or more masks corresponding to semantically relevant set of objects (e.g., cats and dogs are both animals, therefore, belong to same group mask and would be highlighted blue, but they are also not the same species, which can be determined by semantic information such as shape and size, therefore, highlight cats red and dogs green). Specifically, a person could perform all such actions using a pen and paper.
Step2A, Prong 2: Does the claim recite an additional elements that integrate the judicial exception into a practical application? NO. Limitation (1) recites additional limitation “A computer-implemented method”, which constitutes a generic computer recitation (MPEP 2106.05(f)). Limitation (2) recites additional limitation “…by at least one processor…”, which constitutes a generic computer recitation and mere instructions to apply the judicial exception to a computer environment (MPEP 2106.05(f)). Limitation (3) recites additional limitations “…by at least one processor…” and “…from one or more segmentation masks comprising a plurality of semantic segmentations generated utilizing one or more segmentation neural networks”. The first additional limitation of (3) constitutes a generic computer recitation and mere instructions to apply the judicial exception to a computer environment (MPEP 2106.05(f)), and the second limitation constitutes mere instruction to apply the judicial exception to the particular field of semantic segmentation using neural networks (MPEP 2106.05(f)) and/or mere data gathering (MPEP 2106.05(g)). Limitation (4) recites “by the at least one processor”, which constitutes a generic computer recitation and mere instructions to apply the judicial exception to a computer environment (MPEP 2106.05(f)).
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? NO. The claim’s additional elements, as stated in Prong 2, do not amount to significantly more than the judicial exception. Limitations (1), (2), (3), and (4) lack sufficient structure to amount to significantly more than the judicial exception as described in Step2A. Furthermore, segmentation using neural networks is well-understood and routine within the industry. Therefore, limitations (1), (2), (3), and (4) use well-understood, routine, and conventional activities previously known to the industry, specified at a high level of generality, to accomplish the judicial exception.
Claim 2 recites additional limitations “generating the one or more segmentation masks by generating the plurality of semantic segmentations utilizing a plurality of separate segmentation neural networks.”, which constitute mere instruction to apply the judicial exception to the particular field of semantic segmentation using neural networks (MPEP 2106.05(f)) and/or mere data gathering (MPEP 2106.05(g)). Claim 3 recites additional elements “…extracting the vector hierarchy from a vector file of the vector image, the vector hierarchy comprising a plurality of nodes corresponding to vector objects in a tree structure…” and “…executing the search on the vector hierarchy utilizing a breadth first search algorithm to determine the user-tagged groups of objects based on tags of the plurality of nodes in the tree structure.”. The first limitation of Claim 3 is directed toward a mental process (MPEP 2106.04(a)(2)), specifically, a person, given a vector file of the vector image, can extract a vector hierarchy (e.g., vector file includes “animals”, “human”, “dog”, “cat”, “plant”, the hierarchy would be root node (null) with “animals” as left child node, “plant” as right child node, “animals” would have children “human” “dog” and “cat” nodes). The second limitation of claim 3 is directed toward a mental process and/or a mathematical algorithm (MPEP 2106.04(a)(2)), specifically, a person can search a tree in a breadth first manner (e.g., search order of previous tree would go animals->plants->human->dog->cat). Claim 4 recites additional limitations “determining a first user-tagged group of objects and a second user-tagged group of objects in response to executing the breadth first search algorithm; and generating a first group mask for the first user-tagged group of objects and a second group mask for the second user-tagged group of objects.”, which is directed toward a mental process (MPEP 2106.04(a)(2)), analogous to claim 1. Claim 5 recites additional limitations “determining an intersection-over-union metric for the group mask in relation to the one or more segmentation masks; and selecting the group mask in response to determining that the intersection-over-union metric meets a threshold value.”, which are directed toward mental process and/or mathematical calculations (MPEP 2106.04(a)(2)). Specifically, a person can approximate an IOU for two given masks (e.g., 50%), and determine if the IOU meets a threshold (e.g., threshold is >=40%, therefore, select the mask). Claim 6 recites “…wherein determining the group mask comprises filtering the group mask from the one or more masks by utilizing bipartite matching on the one or more group masks and the one or more segmentation masks to determine the intersection-over-union metric.”, which is directed toward a mental process and/or mathematical calculation (MPEP 2106.04(a)(2)). Specifically, a person can calculate a bipartite matching using a pen and paper to determine the IOU for masks. Claim 7 recites “…extracting the one or more masks comprises extracting, utilizing the group mask and the vector image, a partial mask, a full mask, or a color image corresponding to the semantically relevant set of objects.”, which is directed toward a mental process (MPEP 2106.04(a)(2)), analogous to limitation (4) of claim 1. Claim 9 recites “…filtering, from a vector image dataset, a plurality of vector images comprising the vector image by utilizing an image classifier model to determine that the vector image comprises a scene layout…” and “…determining distances between text embeddings representing elements in the plurality of vector images to image embeddings of the plurality of vector images…” and “…selecting the vector image from a subset of vector images having a similarity score above a threshold score based on the distances between the text embeddings and the image embeddings.”. Claim 10 recites additional limitations analogous to a combination of claims 1, and 5-6, and analogous mappings are applicable. Claim 11 likewise recites additional limitations analogous to a combination of claim 1, 2, and 5, and specifically notes it fails to integrate the judicial exception into a practical application because separating into a first and second segmentation network performed by different neural networks still constitutes mere instructions to apply the judicial exception of a mental process of segmentation to the field of multiple neural networks (MPEP 2106.05(f)), and may also constitute mere data acquiring (MPEP 2106.05(g)). Claim 12-16 recites additional limitations analogous to claims 1, 3, and 5-7, and analogous mappings are applicable. Likewise, claims 17-20 recites additional limitations analogous to those mapped in claims 1, 3, and 5-7. The examiner specifically notes that claims 8 and 9 both integrate the judicial exception into a practical application. Specifically, with regard to claim 8, the optimization of parameters of the neural networks based of the predicted masks integrates the judicial exception into a practical application because it constitutes an improvement to the neural network architecture and/or field of technology of neural networks (MPEP 2106.05(a)). Likewise, claim 9 integrates the judicial exception into a practical application because it allows for determination of a scene layout and directly improves the neural network masking by actively mapping text to image embeddings, which would improve the dataset for the neural network and constitutes an improvement to the neural network architecture and/or field of technology of neural networks (MPEP 2106.05(a))
At least for these reasons, claims 1-7, and 10-20 are ineligible under 35 U.S.C. 101.
Claim Rejections - 35 USC § 103
8. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
9. Claims 1, 7-8, and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over “Semantic Hierarchy-Aware Segmentation” to Li et al. (hereinafter Li), and further in view of U.S. Publication No. 2022/0375090 to Hwang et al. (hereinafter Hwang).
10. Regarding Claim 1, Li discloses a {computer-implemented} method comprising:
generating, {by at least one processor}, one or more group masks corresponding to {user-}tagged groups of objects in a vector image by executing a search on a vector hierarchy of the vector image ([pg. 2124, Fig. 1], [pg. 2127, Fig. 2], [pg. 2124, col. 2, par. 2, ln. 1-25] “To summarize, our main contributions are four-fold: We take the lead to tackle semantic segmentation in a class hierarchy-aware setting, which is long-neglected yet the desiderata for holistic scene understanding. We formulate our target task in an end-to-end deep learning framework which comprehensively exploits the rich dependencies between semantic concepts in order to produce predictions coherent with the semantic structures and improve seg mentation performance. We develop HSSN (and its variant HSSN+), a general HSS framework that cleverly converts such structured dense prediction task as a pixel-wise multi-label classification problem, with minimal architectural changes to existing standard segmentation models that do not take class hierarchy into account. We propose a pixel-wise hierarchical segmentation learning strategy, which explicitly enforces network predictions to respect the label structures by exploiting the meronymy and exclusion relations among visual concepts as training targets. We devise a pixel-wise hierarchical representation learning strategy which directly shapes the pixel embedding space by imposing two different kinds of class hierarchy-induced margin constraints to the relative arrangement of pixel samples.”, [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24] “Our goal is to leverage standard semantic segmentation networks for the HSS problem and then exploit structured class relations in order to generate hierarchy-coherent predictions and improve performance. Given our goal, we develop HIER ARCHICAL SEMANTIC SEGMENTATION NETWORKS (i.e., HSSN and HSSN+), a general framework for HSS network design (Section III-A) and training (Section III-B). A. Hierarchal Semantic Segmentation Networks… Rather than typical segmentation methods treating semantic classes as disjoint labels, in the HSS setting, the underlying dependencies between classes are considered and formalized in a form of a tree-structured hierarchy,
T
(
V
,
E
)
. Specifically, each node
v
∈
V
denotes a semantic class/concept, while each edge
(
u
,
v
)
∈
E
encodes the decomposition relationship between two classes,
u
,
v
∈
E
, i.e. parent node v is a more general, super class fo child node u, such that
u
,
v
=
(
b
i
c
y
c
l
e
,
v
e
h
i
c
l
e
)
. We assume
(
v
,
v
)
∈
E
, thus every class is both a subclass and superclass of itself. The root node of
T
, i.e.,
v
◻
, denotes the most general class. The leaf nodes, i.e.,
V
♢
, refer to the most fine-grained classes, such as
V
♢
=
t
r
e
e
,
b
i
c
y
c
l
i
s
t
,
…
in urban street scene parsing [2], [3], and
V
♢
=
h
e
a
d
,
l
e
g
,
…
in human parsing [33], [34]. For a typical hierarchy-agnostic segmentation network, an encoder
f
E
N
C
is first adopted to map an image I into a dense feature tensor
I
=
f
E
N
C
(
I
)
∈
R
H
×
W
×
C
, where
i
∈
I
is the embedding of pixel
i
∈
I
. Then a segmentation head
f
S
E
G
is used to get a score map
Y
=
s
o
f
t
m
a
x
f
S
E
G
I
∈
[
0,1
]
H
×
W
×
|
V
♢
|
, w.r.t. the leaf nose set
V
♢
. Given the score vector
y
=
[
y
V
♢
]
V
♢
∈
V
♢
∈
[
0,1
]
|
V
♢
|
and groundtruth leaf label
v
^
♢
∈
V
♢
for pixel i, the categorical cross-entropy loss is optimized:
L
C
C
E
y
=
-
l
o
g
y
v
^
♢
.
(
1
)
During inference, pixel i is associated to a single leaf node with the maximum probability:
v
♢
*
=
a
r
g
m
a
x
v
♢
(
y
v
♢
)
. To accommodate classical segmentation networks to the HSS setting without significant architectural change, our HSSN first formulates HSS as a pixel-wise multi-label classification task, i.e., map pixels with their corresponding classes in the hierarchy as a while. Specifically, only the segmentation head
f
S
E
G
is modified to predict an augmented score map
S
=
s
i
g
m
o
i
d
(
f
S
E
G
I
)
∈
[
0,1
]
H
×
W
×
|
V
|
w.r.t. the entire class hierarchy
V
. Given the score vector
s
=
[
s
v
]
v
∈
V
∈
{
0,1
}
|
V
|
for pixel i, the binary cross-entropy loss is optimized:
L
B
C
E
s
=
∑
v
∈
V
-
l
^
v
l
o
g
s
v
-
(
1
-
l
^
v
)
l
o
g
1
-
s
v
.
(
2
)
During inference, each pixel i is associated with the top-scoring root-to-leaf path in the class hierarchy
T
:
v
1
*
,
…
,
v
P
*
=
arg
max
P
⊆
T
Σ
v
P
∈
P
s
v
P
(3) where
P
=
{
v
1
…
v
P
*
}
⊆
T
denotes a feasible root-to-leaf path of
T
, i.e.,
v
1
∈
V
♢
,
v
|
P
|
=
v
◻
, and
∀
v
p
,
v
p
+
1
∈
P
⟹
(
v
p
,
v
p
1
+
1
)
∈
E
. Although (3) ensures the coherence between pixel-wise prediction and the class hierarchy during the inference stage, there is no any class relation information used for segmentation network training, as the binary cross-entropy loss in (2) is computed over each class independently. To alleviate this issue, we propose a hierarchy-aware segmentation learning scheme (Section III-B), which incorporates the semantic structures into the training of our HSSN.”);
determining, {by at least one processor} and from the one or more group masks, a group mask comprising a semantically relevant set of objects based on semantic information from one or more segmentation masks comprising a plurality of semantic segmentations generated utilizing one or more segmentation neural networks ([pg. 2124, Fig. 1], [pg. 2127, Fig. 2], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16] “Network Architecture: HSSN is a general HSS framework and readily applied to any hierarchy-agnostic segmentation models in principle. It consists of two major components: The segmentation encoder
f
E
N
C
(Section III-A) maps each input image I into a dense feature representation
I
∈
R
H
×
W
×
C
, and can be implemented by any backbone networks. In Section IV, we experimented with CNN-based (i.e., ResNet [50], HRNet [41]) and Transformer-based (i.e., Swing[42]) backbones. The segmentation head
f
S
E
G
(Section III-A) projects I to a structured score map
S
∈
R
H
×
W
×
|
V
|
for all the semantic classes in
V
. In our experiments, segmentation heads used in recent representative segmentation models (i.e., DeepLabV3+ [37], OCRNet [38], MaskFormer [39], DeepLabV3 [36]) are adopted and modified for our HSS setting.”, [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4] “Inference: For each pixel, the label assignment follows (3).As depicted in Algorithm 1, the assignment process starts from the root node
v
◻
and traverses down to the leaf nodes
v
♢
∈
V
♢
of the hierarchy 𝓣 and aggregates the prediction scores of the nodes along each feasible root-to-leaf path. There are |
V
♢
| feasible root-to-leaf paths in total, where |
V
♢
| denotes the number of the leaf nodes. The pixel is assigned the class labels of the root-to-leaf path with the largest aggregated prediction score. It is noteworthy that such hierarchy-aware pixel label assignment procedure can be executed in parallel, with only negligible computational overhead… Datasets: We conduct experiments on two popular urban street scene parsing datasets [2], [3], two human body parsing datasets [33], [34], and one object parsing dataset [35]. Please note that for Mapillary Vistas 2.0 [3], Cityscapes [2], and PASCAL-Part [35], the corresponding class hierarchies are employed as officially provided. For LIP [33] and PASCAL Person-Part [34], the class hierarchies are generated in accordance with the conventions [139], [140]. Mapillary Vistas 2.0 [3] is an urban egocentric street-view dataset with high-resolution images. It contains 18,000, 2,000 and 5,000 images for train, val and test, respectively. It provides annotations for 144 semantic concepts, which are organized in a three-level hierarchy, covering 4/16/124 concepts, respectively. Cityscapes [2] contains 5,000 elaborately annotated urban scene images, which are split into 2,975/500/1,524 for train/val/test. It is associated with 19 fine-grained concepts, which are grouped into 6 super-classes. PASCAL-Person-Part [34] has 1,716 and 1,817 images for train and test, with precise annotations for 6 human parts. Following [23], [24], we group 20 fine-grained parts (e.g., head, left-arm) into two superclasses upper body and lower-body, which are further combined to full-body. LIP [33] includes 50,462 single-person images gathered from real-world scenarios , with 30,462/10,000/10,000 for train/val/test splits. The hierarchy is similar to the one in PASCAL-Person-Part, but the leaf layer has 19 fine grained semantic parts. PASCAL-Part [35], augmented from PASCAL VOC 2010 dataset [141], provides dense object part annotations. It includes 4,998 images for training and 5,105 for testing. As in [139], [140], [142], we consider 58 and 108 part categories as the subclasses of 20 PASCAL VOC object categories (i.e., superclasses), leading to PASCAL-Part-58 and PASCAL-Part-108, respectively… Testing: The inference procedure adheres to (3) and is elaborated in Section III-C and Algorithm 1.Asin[23], [24], [38], [39], [66], [87], we report the segmentation scores at multiple scales {0.5,0.75,1.0,1.25,1.5,1.75} with horizontal flipping…”, [pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person],); and
extracting, {by at least one processor} and from the group mask, one or more masks corresponding to the semantically relevant set of objects ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]).
Li does not specifically disclose wherein the method is computer-implemented and performed by at least one processor or wherein the groups are user-tagged.
However, Hwang specifically teaches the method is performed by at least one processor and wherein the groups are user-tagged ([par. 0136, ln. 1-15] “As shown, the series of acts within the unknown object subclass labeler 536 includes an act 542 of the panoptic segmentation system 106 determining a label of an unknown object subclass… a user provides input labeling the unknown object subclasses… the panoptic segmentation system 106 provides one or more exemplar object instances in the unknown object subclass to another object detection network or image search model to predict a label for the unknown object subclass. To illustrate, an unknown object subclass includes an exemplar object instance of anteaters. Upon providing one or more of the exemplar object instances to a client device associated with a user, the panoptic segmentation system 106 receives input from the client device indicating a label of “anteaters.””, [par. 0199, ln. 1-13] “As shown in FIG. 10, the computing device 1000 may include one or more processor(s) 1002, memory 1004, a storage device 1006, input/output (“I/O”) interfaces 1008, and a communication interface 1010, which may be communicatively coupled by way of a communication infrastructure (e.g., bus 1012). While the computing device 1000 is shown in FIG. 10, the components illustrated in FIG. 10 are not intended to be limiting.”). One of ordinary skill in the art, before the effective filling date of the claimed invention, would specifically recognize Li and Hwang as within the same filed of semantic segmentation, and as analogous to the claimed invention. The motivation to combine would have been obvious to one of ordinary skill in the art, and is disclosed in Hwang ([par. 0136, ln. 1-15]), wherein users can define unknown classes of objects, and would also allow for a real-world implementation of Li. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor and user-tagged groups of Hwang, through known means, with no change to their respective functions, and the combination would have yielded nothing more than predictable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a computer and processor, and to operate using user-tagged groups as taught in Hwang.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the method performed by at least one processor and user-tagged groups of Hwang to obtain the invention as specified in claim 1.
11. Regarding Claim 7, a combination of Li and Hwang teaches the method of claim 1. Li further discloses wherein extracting the one or more masks comprises extracting, utilizing the group mask and the vector image, a partial mask, a full mask, or a color image corresponding to the semantically relevant set of objects ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]).
12. Regarding Claim 8, a combination of Li and Hwang teaches the method of claim 8. Li further discloses further comprising: determining a set of predicted masks and color images generated for the vector image utilizing an image processing neural network ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]); and optimizing parameters of the image processing neural network to reduce differences between the set of predicted masks and color images and a set of ground truth masks and color images comprising the partial mask, the full mask, and the color image corresponding to the semantically relevant set of objects ([pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24] see loss in equation (1) and (2), see also [pg. 2127, col. 1 and 2, equations (4) and (5)]).
13. Regarding Claim 17, Li discloses {a non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform} operations comprising: generating, utilizing one or more segmentation neural networks, one or more segmentation masks comprising a plurality of semantic segmentations ([pg. 2127, Fig 2], [pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]); generating one or more group masks corresponding to {user-}tagged groups of objects in a vector image by executing search algorithm on a vector hierarchy of the vector image ([pg. 2127, Fig 2], [pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]); determining, from the one or more group masks, a group mask comprising a semantically relevant set of objects based on semantic information from the one or more segmentation masks ([pg. 2127, Fig 2], [pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]); and extracting, from the group mask, a partial mask or a full mask corresponding to the semantically relevant set of objects ([pg. 2127, Fig 2], [pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]). Li does not specifically disclose a non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform the method or user-tagged groups.
However, Hwang specifically teaches a non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform a method and user-tagged groups ([par. 0136, ln. 1-15], [par. 0199, ln. 1-13]). The motivation to combine remains analogous to claim 1. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor and user-tagged groups of Hwang, through known means, with no change to their respective functions, and the combination would have yielded nothing more than predictable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a non-transitory medium and to operate using user-tagged groups as taught in Hwang.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the non-transitory computer readable medium and user-tagged groups of Hwang to obtain the invention as specified in claim 17.
14. Regarding Claim 18, a combination of Li and Hwang teaches the non-transitory medium of claim 17. Li further disclose wherein generating the one or more group masks comprises: executing the search algorithm on the vector hierarchy to determine a first user-tagged group of objects from a first node in a tree structure and a second user-tagged group of objects from a second node in the tree structure; and generating a first group mask for the first user-tagged group of objects and a second group mask for the second user-tagged group of objects ([pg. 2132, Fig. 2], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]).
15. Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over “Semantic Hierarchy-Aware Segmentation” to Li, and further in view of U.S. Publication No. 2022/0375090 to Hwang, and further in view of “Image semantic segmentation algorithm based on a multi-expert system” to Ma et al. (hereinafter Ma).
16. Regarding Claim 2, a combination of Li and Hwang teaches the method of claim 1. Li and Hwang do not specifically disclose generating the one or more segmentation masks by generating the plurality of semantic segmentations utilizing a plurality of separate segmentation neural networks.
However, Ma specifically teaches generating the one or more segmentation masks by generating the plurality of semantic segmentations utilizing a plurality of separate segmentation neural networks ([pg. 3, Fig. 1, see Expert1, Expert2, Expert3], see also [pg. 4, Fig. 2], [pg. 5, Fig. 3], [pg. 5, Fig. 4], [pg. 4, 3 Proposed Methods, 3.1 Overall Network Architecture, par. 1, ln. 1-9] “The network architecture of the proposed algorithm is shown in Fig. 1. The three different expert models are built on the basis of Deeplabv3plus, and the multi-expert results are fused by ensemble learning to produce the final segmentation results; ensemble learning refers to the combination of multiple expert models to achieve superior performance over a single expert model. Inspired by the idea of ensemble learning in Ref. 36, voting is used to decide the final class labels when adjudicating the multi-expert results. Specifically, one of the expert model results is used as a benchmark for voting, and for each pixel, the class with the highest count is used as the label for that pixel. When three expert models generate three different labels for a given pixel, the benchmark is used as the adjudication result, and in the experiment, expert model 2 is set as the benchmark.”, [pg. 4, 3 Proposed Methods, 3.1 Overall Network Architecture, par. 3, ln. 1 to pg. 5, par. 3, ln. 8] “The network architecture of expert model 1 is shown in Fig. 2. Here, the input image is passed through the backbone network to extract semantic features and then input to the C-ASPP module to capture the semantic information of different receptive fields. Then, the steps in the baseline network are followed using 1 × 1 convolution to reduce the number of channels to 256. In addition, in the decoder part of the expert model 1 network architecture, the proposed decoding structure in Deeplabv3plus22 is used, as shown in the black dashed box part of Fig. 2. Here, the output features of the encoder are channel concatenated with the feature maps with the same resolution in the backbone network after 4 times upsampling, and the features are refined using 3 × 3 convolution after the concatenation and finally upsampled to recover the image size. In the network architecture of expert model 2, the encoding part is the same as the baseline network, whereas in the decoding part, a new decoder is designed, as shown in Fig. 3. The decoder in the red dashed box relies on the feature fusion module, which reweights and fuses features at different levels to obtain more comprehensive and detailed information, and the detailed design of this module is shown in Sec. 3.3. Expert model 3 introduces a new loss function in the baseline network, and the network architecture of this expert model is Deeplabv3plus,22 which can be seen in the red dashed box in Fig. 4, using both detail loss on top of the Deeplabv3plus network architecture’s segmentation loss, in which the segmentation head uses 1 × 1 convolution and the detail head contains 3 × 3 convolution, Batch normalization, ReLU, and 1 × 1 convolution. In addition, when generating the detail loss, the corresponding true labels are generated by the detail guidance module, which is proposed by STDCSeg38 and implemented by 2D convolution kernel named Laplacian kernel and a trainable 1 × 1 convolution.”). One of ordinary skill in the art, before the effective filing date of the claimed invention, would specifically recognize Li, Hwang, and Ma as within the same filed of semantic segmentation, and as analogous to the claimed invention. The motivation to combine is disclosed in Ma, wherein multiple expert models achieve superior performance over a single expert model ([pg. 4, 3 Proposed Methods, 3.1 Overall Network Architecture, par. 1, ln. 1-9]). One of ordinary skill in the art, before the effective filing date of the claimed invention, would have combined the method of Li with the method performed by at least one processor and user-tagged groups of Hwang, and further combined the method of the combination of Li and Hwang with the plurality of separate segmentation networks as taught in Ma, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have used multiple DeepLabV3+ segmentations heads as taught in Ma in place of the single DeepLabV3+ segmentation head as taught in the method of the combination of Li and Hwang ([Li, pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16] see “…In Section IV, we experimented with CNN-based (i.e., ResNet [50], HRNet [41]) and Transformer-based (i.e., Swing[42]) backbones. The segmentation head
f
S
E
G
(Section III-A) projects I to a structured score map
S
∈
R
H
×
W
×
|
V
|
for all the semantic classes in
V
. In our experiments, segmentation heads used in recent representative segmentation models (i.e., DeepLabV3+ [37], OCRNet [38], MaskFormer [39], DeepLabV3 [36]) are adopted and modified for our HSS setting…”).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to combine the method of Li with the method performed by at least one processor and user-tagged groups of Hwang and the plurality of separate segmentation neural networks of Ma to obtain the invention as specified in claim 2.
17. Claims 3-4 are rejected under 35 U.S.C. 103 as being unpatentable over “Semantic Hierarchy-Aware Segmentation” to Li, and further in view of U.S. Publication No. 2022/0375090 to Hwang, and further in view of “Navigating the Forest: A Comprehensive Look at Tree Traversal Techniques” to Sirhith et al. (hereinafter Sirhith).
18. Regarding Claim 3, a combination of Li and Hwang teaches the method of claim 1. Li teaches wherein generating the one or more group masks comprises: extracting the vector hierarchy from a vector file of the vector image ([pg. 2127, Fig. 2], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24]), the vector hierarchy comprising a plurality of nodes corresponding to vector objects in a tree structure ([pg. 2127, Fig. 2], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24]); and executing the search on the vector hierarchy utilizing a {breadth first} search algorithm to determine the {user-}tagged groups of objects based on tags of the plurality of nodes in the tree structure ([pg. 2127, Fig. 2], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]). Li does not specifically disclose that the search is breadth first, or a user-tagged group.
However, Hwang teaches a user-tagged group ([par. 0136, ln. 1-15]). The motivation to combine remains analogous to claim 1. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor and user-tagged groups of Hwang, through known means, with no change to their respective functions, and the combination would have yielded nothing more than predictable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a computer and processor, and to operate using user-tagged groups as taught in Hwang.
Hwang does not specifically disclose a breadth first search. Therefore, the computer-implemented method of the combination of Li and Hwang does not specifically disclose a breadth first search.
However, Sirhith specifically teaches a breadth first search to traverse a tree structure ([pg. 335, Iterative Tree Traversal, par. 1, ln. 1-5] “Iterative tree traversal is another method of traversing a tree data structure that does not use recursion. Iterative traversal techniques are particularly useful for trees with large or deep structures, as they avoid the overhead associated with recursive function calls. In this section, we will discuss the principles of iterative tree traversal and provide some examples to illustrate how it works.”, [pg. 336, ln. 11-14] “Iterative tree traversal is a useful technique for traversing a tree data structure, particularly for large or deep trees. However, it can be more difficult to implement and understand than recursive traversal, and it can be less efficient for small or shallow”, [pg. 339, ln. 12-15] “The choice of traversal technique depends on the specific use case and the characteristics of the tree being traversed. If memory usage is a concern, depth-first traversal may be preferred, while breadth-first traversal may be preferred for finding the shortest path in an unweighted graph. In-order traversal is useful for traversing binary search trees in sorted order, and pre-order and post-order traversal can be useful for creating or deleting trees, respectively.”, [pg. 339, Choosing between recursion and iteration, par. 1, ln. 1-3] “As mentioned earlier, tree traversal can be implemented using either recursion or iteration. Recursion can be simpler to implement and understand, but it can lead to stack overflow errors for very large trees. Iteration can be more efficient for large trees, but it may require more complex code.”, [pg. 338, see Table 1 and Table 2]). One of ordinary skill in the art, before the effective filling date of the claimed invention, would specifically recognize the method of the combination of Li and Hwang as within the same field of programing using tree structures as Sirhith, and as analogous to the claimed invention. Specifically, one of ordinary skill in the art, before the effective filling date of the claimed invention, would recognize a breadth first search (BFS) to be a standard and often used method of full tree traversal along with depth first search (DFS). One of ordinary skill in the art, before the effective filling date of the claimed invention, given that Li teaches to traverse the entire tree and find the maximum leaf-node path ([pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24]), would specifically recognize BFS as an appropriate means of tree traversal to accomplish this function. Specifically, one benefit of BFS over DFS is it allows for computation of all root-node paths for a given level effectively simultaneously, because you search the breadth of the tree before continuing to the following level, and thus allows for immediate determination of the current root-node path that is highest scoring at a given level. In the case of DFS, this is not possible, since only the one, full depth root-node path is determined at a given time. Another motivation to combine is provided in Sirhith, wherein it allows for less overhead ([pg. 335, Iterative Tree Traversal, par. 1, ln. 1-5]). One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor and user-tagged groups of Hwang, and further combine the method of the combination of Li and Hwang with the breadth first tree traversal as taught in Sirhith, through known means, with no change to their respective unction, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have used the BFS traversal of Sirhith to traverse the tree of the method of the combination of Li and Hwang.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the method performed by at least one processor and user-tagged groups of Hwang and the breadth first tree traversal as taught in Sirhith to obtain the invention as specified in claim 3.
19. Regarding Claim 4, a combination of Li, Hwang, and Sirhith teaches the method of claim 3. Li further discloses wherein generating the one or more group masks comprises: determining a first {user-}tagged group of objects and a second {user-}tagged group of objects in response to executing the breadth first search algorithm ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]); and generating a first group mask for the first {user-}tagged group of objects and a second group mask for the second {user-}tagged group of objects ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]). Li does not specifically disclose that the groups are user-tagged.
However, Hwang teaches a user-tagged group ([par. 0136, ln. 1-15]). The motivation to combine remains analogous to claim 1. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor and user-tagged groups of Hwang, and further combine the method of the combination of Li and Hwang with the breadth first tree traversal as taught in Sirhith, through known means, with no change to their respective unction, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a computer and processor, and to operate using user-tagged groups as taught in Hwang, and further used the BFS tree traversal of Sirhith to traverse the tree of the method of the combination of Li and Hwang.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the method performed by at least one processor and user-tagged groups of Hwang and the breadth first tree traversal as taught in Sirhith to obtain the invention as specified in claim 4.
20. Claims 5-6, 10, 12-14, 16, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over “Semantic Hierarchy-Aware Segmentation” to Li, and further in view of U.S. Publication No. 2022/0375090 to Hwang, and further in view of “Combinatorial Optimization for Panoptic Segmentation: A Fully Differentiable Approach” to Abbas et al. (hereinafter Abbas).
21. Regarding Claim 5, a combination of Li and Hwang teaches the method of claim 1. Li and Hwang do not specifically disclose determining the group mask comprises: determining an intersection-over-union metric for the group mask in relation to the one or more segmentation masks; and selecting the group mask in response to determining that the intersection-over-union metric meets a threshold value.
However, Abbas specifically teaches wherein determining the group mask comprises: determining an intersection-over-union metric for the group mask in relation to the one or more segmentation masks ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23] “Panoptic quality (PQ) [37] is a size-invariant evaluation metric defined between a set of predicted masks and ground-truth masks for each semantic class
l
∈
[
K
]
. For each class, it requires to match predicted and object masks to each other w.r.t intersection-over-union (IoU) since instance labels are permutation invariant. A pair of predicted and ground truth binary masks p and g of the same class l is matched (i.e. true-positive) if
I
o
U
p
,
g
≥
0.5
.
We write
(
p
,
g
)
∈
T
P
l
. For the unmatched masks, each prediction (ground-truth) is marked as false positive
F
P
l
(false negative
F
N
l
). Since at most one match exists per ground truth mask, this matching process is well-defined [37]. The PQ metric is defined as the mean of class specific PQ scores
P
Q
l
=
∑
(
p
,
g
)
∈
T
P
l
I
o
U
(
p
,
g
)
T
P
l
+
0.5
(
|
F
P
l
|
+
|
F
N
l
|
)
(5). Note that the PQ score (5) can be arbitrarily low just by the presence of small sized false predictions [16, 58, 73]. A common practice to avoid such issue is to reject small predictions before computing the PQ score with some dataset specific size thresholds, before evaluation. However, this rejection mechanism is not incorporated during training. The PQ metric (5) cannot be straightforwardly used for training due to the discontinuity of the hard threshold based matching and the rejection mechanism. Therefore we replace the hard threshold matching process for each class l by computing correspondences via a maximum weighted bipartite matching with IoU as weights. The corresponding matches are
T
P
-
l
, the unmatched prediction masks
F
P
-
l
and the unmatched ground truth masks
F
N
-
l
. The hard thresholding is smoothed via soft thresholding function
h
u
=
u
4
u
4
+
(
1
-
u
)
4
centered around 0.5. The small prediction rejection mechanism for mask p is smoothed via
σ
l
p
=
[
1
+
e
x
p
-
0.1
1
T
p
-
t
l
]
-
1
centered at area threshold
t
l
for class l. The overall surrogate PQ for class l is
P
Q
-
l
=
∑
(
p
,
g
)
∈
T
P
-
l
h
(
I
o
U
p
,
g
)
σ
l
p
I
o
U
(
p
,
g
)
∑
(
p
,
g
)
∈
T
P
-
l
h
I
o
U
p
,
g
σ
l
p
+
0.5
{
∑
p
∈
F
P
-
l
σ
l
p
+
|
F
N
-
l
|
}
(6) where the term h(IoU(p,g)) models the probability of a predicted mask p being true positive.”); and selecting the group mask in response to determining that the intersection-over-union metric meets a threshold value ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23] see specifically thresholding function). One of ordinary skill in the art, before the effective filling date of the claimed invention, would specifically recognize Li, Hwang, and Abbas as within the same field of semantic image segmentation, and as analogous to the claimed invention. The motivation to combine would have been obvious to one of ordinary skill in the art, and is taught in Abbas, wherein the IoU and thresholding allows for size invariant accuracy determination of ground truth vs. predicted masks ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]). One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li and Hwang with the IoU metric and threshold value of Abbas through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have incorporated the IoU metric and threshold value of the loss of Abbas as a quality metric for the generated group masks vs. segmentation masks of the method of the combination of Li and Hwang.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the method performed by at least one processor and user-tagged groups of Hwang and the IoU metric and threshold value of Abbas to obtain the invention as specified in claim 5.
22. Regarding Claim 6, a combination of Li, Hwang, and Abbas teaches the method of claim 5. Li and Hwang do not specifically disclose wherein determining the group mask comprises filtering the group mask from the one or more masks by utilizing bipartite matching on the one or more group masks and the one or more segmentation masks to determine the intersection-over-union metric.
However, Abbas specifically teaches to filter the group mask from the one or more masks by utilizing bipartite matching on the one or more group masks and the one or more segmentation masks to determine the intersection-over-union metric ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]). The motivation to combine remains analogous to claim 5. Furthermore, one of ordinary skill in the art, before the effective filling date of the claimed invention, would recognize that formulating the matching problem as a bipartite matching would also result in a well-defined match, as disclosed in Abbas ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]). One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li and Hwang with the IoU metric, threshold value, and bipartite matching of Abbas through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have incorporated the bipartite graph with IoU metrics as weights and threshold value of the loss of Abbas as a quality metric for the generated group masks vs. segmentation masks of the method of the combination of Li and Hwang.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the method performed by at least one processor and user-tagged groups of Hwang and the bipartite matching, IoU metric, threshold value of Abbas to obtain the invention as specified in claim 6.
23. Regarding Claim 10, Li discloses {a system comprising: one or more memory devices; and one or more processors configured to cause the system to}: generate one or more group masks corresponding to {user-}tagged groups of objects in a vector image by executing a search on a vector hierarchy of the vector image ([pg. 2124, Fig. 1], [pg. 2127, Fig. 2], [pg. 2124, col. 2, par. 2, ln. 1-25], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24]); determine, {utilizing bipartite matching, intersection-over-union} metrics for the one or more group masks and one or more segmentation masks comprising a plurality of semantic segmentations generated utilizing one or more segmentation neural networks ([pg. 2124, Fig. 1], [pg. 2127, Fig. 2], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]); determine, from the one or more group masks, a group mask comprising a semantically relevant set of objects {according to the intersection-over-union metrics} ([pg. 2124, Fig. 1], [pg. 2127, Fig. 2], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]); and extract, from the group mask, one or more masks corresponding to the semantically relevant set of objects ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]).
Li does not specifically disclose a system comprising: one or more memory devices; and one or more processors configured to cause the system to perform the method, user-tagged groups, determining utilizing bipartite matching, and intersection-over-union metric, or determining one or more group masks according to the IoU metric.
However, Hwang specifically teaches a system comprising: one or more memory devices; and one or more processors configured to cause the system to perform the method ([par. 0199, ln. 1-13]), and user-tagged groups ([par. 0136, ln. 1-15]). The motivation to combine remains analogous to claim 1. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor and user-tagged groups of Hwang, through known means, with no change to their respective functions, and the combination would have yielded nothing more than predictable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a processor and memory of the system of Hwang, and to operate using user-tagged groups as taught in Hwang.
Hwang does not specifically teach determining utilizing bipartite matching, and intersection-over-union metric or determining one or more group masks according to the IoU metric.
However, Abbas specifically teaches determining, utilizing bipartite matching, intersect over union metrics for the one or more group masks and the one or more segmentation masks ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]), and wherein a group mask comprising a semantically relevant set of objects is determined according to the intersect-over-union metrics ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]). The motivation to combine remains analogous to claim 5 and 6. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li and Hwang with the IoU metric and bipartite matching of Abbas through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have incorporated the bipartite graph with IoU metrics as weights of Abbas as a quality metric and/or loss for the generated group masks vs. segmentation masks of the method of the combination of Li and Hwang.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the system and user-tagged groups of Hwang and the bipartite matching and IoU metric of Abbas to obtain the invention as specified in claim 10.
24. Regarding Claim 12, a combination of Li, Hwang, and Abbas teaches the method of claim 10. Rejections analogous to claim 5, 6, and 10 are further applicable to claim 12. Specifically, Li and Hwang do not specifically teach determining a first intersection-over-union metric for a first group mask relative to the plurality of semantic segmentations; and determining a second intersection-over-union metric for a second group mask relative to the plurality of semantic segmentations.
However, Abbas teaches determining a first intersection-over-union metric for a first group mask relative to the plurality of semantic segmentations; and determining a second intersection-over-union metric for a second group mask relative to the plurality of semantic segmentations ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]). The motivation to combine remains analogous to claim 5. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li and Hwang with the IoU metric and bipartite matching of Abbas, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, the examiner notes that Abbas determines the IoU metric for each class l. Given that the groups of the method of the combination of Li and Hwang are also based on classes ([Li, pg. 2124, Fig. 1], [Li, pg. 2127, Fig. 2], [Li, pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [Li, pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [Li, pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]), and Li discloses more than two classes ([Li, pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [Li, pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]), it would have been apparent to one of ordinary skill in the art, in applying the IoU metric of Abbas to each class in the method of the combination of Li and Hwang, that there would be a first intersection-over-union metric for a first group mask relative to the plurality of semantic segmentations (i.e., specifically for the first class of the method of Li and Hwang); and determining a second intersection-over-union metric for a second group mask relative to the plurality of semantic segmentations (i.e., specifically for the second class of the method of Li and Hwang).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the system and user-tagged groups of Hwang and the bipartite matching and IoU metric of Abbas to obtain the invention as specified in claim 12.
25. Regarding Claim 13, a combination of Li, Hwang, and Abbas teaches the method of claim 12. Rejections analogous to claim 12 are further applicable to claim 13. Li and Hwang do not specifically teach to determine the group mask comprising the semantically relevant set of objects by determining that the first group mask comprises semantically relevant objects in response to determining that the first intersection-over-union metric meets a threshold value.
However, Abbas teaches to determine the group mask comprising the semantically relevant set of objects by determining that the first group mask comprises semantically relevant objects in response to determining that the first intersection-over-union metric meets a threshold value ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23] see thresholding function h(u)). Specifically, each class is also compared to a threshold in Abbas, and therefore, it is determined to be a true positive or false positive based on said threshold, and thus it can be determined for each class if a first group mask comprises a semantically relevant set of objects via the threshold comparison. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li and Hwang with the IoU metric, bipartite matching, and IoU threshold of Abbas, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, in combining the thresholding of Abbas with the method of the combination of Li and Hwang, it would have been apparent to one of ordinary skill in the art to that for a first intersection-over-union metric for a first group mask relative to the plurality of semantic segmentations (i.e., specifically for the first class of the method of Li and Hwang), to determine the group mask comprising the semantically relevant set of objects by determining that the first group mask comprises semantically relevant objects in response to determining that the first intersection-over-union metric meets a threshold value (i.e., IoU is > h(u), therefore it’s a true-positive).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the system and user-tagged groups of Hwang and the bipartite matching and IoU metric of Abbas to obtain the invention as specified in claim 13.
26. Regarding Claim 14, rejections analogous to claim 12 are further applicable to claim 14. Li and Hwang do not specifically teach to determine that the second group mask does not comprise semantically relevant objects in response to determining that the second intersection-over-union metric does not meet a threshold value.
However, Abbas teaches to determine that the second group mask does not comprise semantically relevant objects in response to determining that the second intersection-over-union metric does not meet a threshold value ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23] see thresholding function h(u)). Specifically, each class is also compared to a threshold in Abbas, and therefore, it is determined to be a true positive or false positive based on said threshold, and thus it can be determined for each class if a second group mask comprises a semantically relevant set of objects via the threshold comparison. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li and Hwang with the IoU metric, bipartite matching, and IoU threshold of Abbas, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, in combining the thresholding of Abbas with the method of the combination of Li and Hwang, it would have been apparent to one of ordinary skill in the art to that for a second intersection-over-union metric for a second group mask relative to the plurality of semantic segmentations (i.e., specifically for the second class of the method of Li and Hwang), to determine the group mask does not comprise a semantically relevant set of objects in response to determining that the second intersection-over-union metric meets a threshold value (i.e., IoU is < h(u), therefore it’s not a true-positive).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the system and user-tagged groups of Hwang and the bipartite matching and IoU metric of Abbas to obtain the invention as specified in claim 14.
27. Regarding Claim 16, a combination of Li, Hwang, and Abbas teaches the method of claim 10. Rejections analogous to claim 7 are further applicable to claim 16 in view of the analogous claim language. Specifically, Li discloses to extract the one or more masks corresponding to the semantically relevant set of objects by extracting a partial mask, a full mask, and a color image of the semantically relevant set of objects based on the group mask and the vector image ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4] [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24] see loss in equation (1) and (2), see also [pg. 2127, col. 1 and 2, equations (4) and (5)]).
28. Regarding Claim 19, a combination of Li and Hwang teaches the non-transitory medium of claim 18. Rejections analogous to claim 12, 13, and 14 are further applicable to claim 19. Specifically, Li discloses comparing the first group mask to the one or more segmentation masks {by generating an intersection-over-union metric utilizing bipartite matching}; and determining that a set of objects in the first {user-}tagged group of objects are semantically relevant {in response to determining that the intersection-over-union metric meets a threshold value} ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]). Li does not specifically disclose generating an intersection-over-union metric utilizing bipartite matching, determining that the intersection-over-union metric meets a threshold value, or a user-tagged group.
However, Hwang teaches user tagged groups ([par. 0136, ln. 1-15]). The motivation to combine remains analogous to claim 1. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the non-transitory medium and user-tagged groups of Hwang, through known means, with no change to their respective function, and the combination would have yielded nothing more than predictable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a non-transitory medium and to operate using user-tagged groups as taught in Hwang.
However, Abbas teaches comparing the first group mask to the one or more segmentation masks by generating an intersection-over-union metric utilizing bipartite matching ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]); and determining that a set of objects in the first user-tagged group of objects are semantically relevant in response to determining that the intersection-over-union metric meets a threshold value ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23] see thresholding function h(u)). The motivation to combine remains analogous to claim 13. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the non-transitory medium of the combination of Li and Hwang with the IoU metric, bipartite matching, and IoU threshold of Abbas, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, in combining the thresholding of Abbas with the non-transitory medium of the combination of Li and Hwang, it would have been apparent to one of ordinary skill in the art to compare the first group mask to the one or more segmentation masks by generating an intersection-over-union metric utilizing bipartite matching (i.e., specifically for the first class of the method of Li and Hwang), to determine that a set of objects in the first user-tagged group of objects are semantically relevant in response to determining that the intersection-over-union metric meets a threshold value (i.e., IoU is > h(u), therefore it’s a true-positive, and therefore semantically relevant).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the non-transitory medium and user-tagged groups of Hwang and the bipartite matching, IoU metric, and IoU thresholding of Abbas to obtain the invention as specified in claim 19.
29. Regarding Claim 20, a combination of Li and Hwang teaches the non-transitory medium of claim 18. Rejections analogous to claim 12, 13, 14, and 19 are further applicable to claim 20. Specifically, Li discloses comparing the second group mask to the one or more segmentation masks {by generating an intersection-over-union metric utilizing bipartite matching}; and determining that a set of objects in the second user-tagged group of objects are not semantically relevant {in response to determining that the intersection-over-union metric does not meet a threshold value} ([pg. 2132, Fig. 6-8], [Table III, see Head, Torso, etc. columns corresponding to parts of person], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24], [pg. 2129, col. 2, C. Implementation Detail, par. 1, ln. 1-16], [pg. 2130, col. 1, par. 2, ln. 1 to col. 2, par. 3, ln. 4]). Li does not specifically disclose generating an intersection-over-union metric utilizing bipartite matching, determining that the intersection-over-union metric meets a threshold value, or a user-tagged group.
However, Hwang teaches user tagged groups ([par. 0136, ln. 1-15]). The motivation to combine remains analogous to claim 1. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the non-transitory medium and user-tagged groups of Hwang, through known means, with no change to their respective function, and the combination would have yielded nothing more than predictable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a non-transitory medium and to operate using user-tagged groups as taught in Hwang.
However, Abbas teaches comparing the second group mask to the one or more segmentation masks by generating an intersection-over-union metric utilizing bipartite matching ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23]); and determining that a set of objects in the second user-tagged group of objects are not semantically relevant in response to determining that the intersection-over-union metric does not meet a threshold value ([pg. 6, 3.3.2 Panoptic Quality Surrogate Loss, par. 1, ln. 1-23] see thresholding function h(u)). The motivation remains analogous to claim 14. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the non-transitory medium of the combination of Li and Hwang with the IoU metric, bipartite matching, and IoU threshold of Abbas, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, in combining the thresholding of Abbas with the method of the combination of Li and Hwang, it would have been apparent to one of ordinary skill in the art to comparing the second group mask to the one or more segmentation masks by generating an intersection-over-union metric utilizing bipartite matching (i.e., specifically for the second class of the method of Li and Hwang), to determine that a set of objects in the second user-tagged group of objects are not semantically relevant in response to determining that the intersection-over-union metric does not meet a threshold value (i.e., IoU is < h(u), therefore it’s not a true-positive, therefore not semantically relevant).
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the non-transitory medium and user-tagged groups of Hwang and the bipartite matching, IoU metric, and IoU thresholding of Abbas to obtain the invention as specified in claim 19.
30. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over “Semantic Hierarchy-Aware Segmentation” to Li, and further in view of U.S. Publication No. 2022/0375090 to Hwang, and further in view of “Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation” to Che et al. (hereinafter Che)
31. Regarding Claim 9, a combination of Li and Hwang teaches the method of claim 1. Li does not specifically discloses filtering, from a vector image dataset, a plurality of vector images comprising the vector image by utilizing an image classifier model to determine that the vector image comprises a scene layout; determining distances between text embeddings representing elements in the plurality of vector images to image embeddings of the plurality of vector images; and selecting the vector image from a subset of vector images having a similarity score above a threshold score based on the distances between the text embeddings and the image embeddings.
However, Hwang teaches filtering, from a vector image dataset, a plurality of vector images comprising the vector image by utilizing an image classifier model to determine that the vector image comprises a scene layout ([par. 0084, ln. 1-16] “… filtering out sparse and noisy clusters, in various implementations, the panoptic segmentation system 106 compares the combined-cluster object distance score for each cluster to a cluster object distance threshold. For example, if the combined-cluster object distance score for an object feature cluster is below a cluster object distance threshold of 0.025, the panoptic segmentation system 106 marks the object feature cluster as having high sparsity and/or removing the cluster from consideration as an unknown object subclass. Otherwise, if the combined-cluster object distance score for an object feature cluster is at or above the cluster object distance threshold, the panoptic segmentation system 106 maintains the object feature cluster… the panoptic segmentation system 106 removes the high sparsity object feature cluster.”, [par. 0090, ln. 1-15] “To illustrate, the storage manager shows Unknown Object Subclass A 312a having Exemplar Object Instance 1 314a, Exemplar Object Instance 2 314b, . . . , Exemplar Object Instance n 314n. Accordingly, for each unknown object instance remaining in a cluster after the filtering actions described above, the panoptic segmentation system 106 stores these unknown object instances as exemplars (e.g., Exemplar Object Instance 1-n 314a-n), Exemplar Object Instance 2 314b, . . . , Exemplar Object Instance n 314n located in a corresponding unknown object subclass (e.g., Unknown Object Subclass A 312a). If the unknown object subclass currently exists from previous iterations of training, the panoptic segmentation system 106 can add the newly discovered unknown object instances to the unknown object subclass as additional exemplars.”, [par. 0180, ln. 1-15] “…the sub-act 922 includes determining a plurality of unknown object subclasses based on the plurality of object feature clusters and a plurality of object feature vectors generated from the subset of unknown object instances by the panoptic segmentation neural network… the sub-act 922 includes filtering out object feature vectors from the first object feature cluster that are associated with unknown object instances originating from a same digital image within the set of digital images.”), determining distances between {text} embeddings representing elements in the plurality of vector images to image embeddings of the plurality of vector images ([par. 0082, ln. 1-15] “… unknown object subclass clustering manager 310 includes an act 326 of the panoptic segmentation system 106 determining cluster distances for each object feature cluster… 106 determines a combined-cluster object distance score for each cluster based on the distance between each unknown object instance in the cluster and a central point of the cluster… 106 determines a combined-cluster object distance score for a cluster based on the average cosine distance score between the centroid and the elements (e.g., unknown object instances) within the cluster… 106 determines the distance in object feature vector space.”, [par. 0084, ln. 1-16] “To illustrate further filtering out sparse and noisy clusters… 106 compares the combined-cluster object distance score for each cluster to a cluster object distance threshold. For example, if the combined-cluster object distance score for an object feature cluster is below a cluster object distance threshold of 0.025… 106 marks the object feature cluster as having high sparsity and/or removing the cluster from consideration as an unknown object subclass. Otherwise, if the combined-cluster object distance score for an object feature cluster is at or above the cluster object distance threshold, the … 106 maintains the object feature cluster. As shown in connection with the act 326… 106 removes the high sparsity object feature cluster.”, [par. 0102, ln. 1-19] “If an additional object feature vector 420 has a cosine similarity 422 with an exemplar object feature vector 418 that satisfies a similarity threshold, the panoptic segmentation system 106 can associate the unknown object instances of the additional object feature vector 420 with the unknown object subclass of the exemplar object feature vector 418.”, [par. 0103, ln. 1-13] “… the panoptic segmentation system 106 determines a correlation between the additional unknown object instances 414 and the unknown object subclasses 312 based on comparing individual exemplars (or their corresponding object feature vectors) to an object feature cluster or unknown object subclass… 106 determines a distance within the object feature space between the additional unknown object instances 414 and the centroid of the object feature cluster or unknown object subclass… 106 can then compare the distance to determine whether it satisfies a similarity threshold (e.g., is within a similarity threshold distance).”); and selecting the vector image from a subset of vector images having a similarity score above a threshold score based on the distances between the {text} embeddings and the image embeddings ([par. 0082, ln. 1-15], [par. 0084, ln. 1-16], [par. 0102, ln. 1-19], [par. 0103, ln. 1-13]). Specifically, one of ordinary skill in the art, before the effective filling date of the claimed invention, would recognize Hwang discloses filtering by utilizing an image classifier model to determine that the vector image comprises a scene layout by determining if a vector image contains known or unknown objects, and further determines distance between vectors to select if an object belongs to a cluster, and a selecting an vector image as belong to a cluster if a similarity score meets a threshold. The motivation to combine would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, and is disclosed in Hwang, in that by filtering known and unknown object, you can improve flexibility and ability of the model ([par. 0030, ln. 1-14] “Further, the panoptic segmentation system can also improve flexibility relative to conventional systems. As mentioned above, rather than being limited to a limited closed-set, the panoptic segmentation system can extend panoptic segmentation to the open-world by facilitating open-set panoptic segmentation. As a result of this flexible extension, the panoptic segmentation system can build a panoptic segmentation neural network that detects and classifies a wider range of object instances within digital images. Indeed, the panoptic segmentation system improves the ability of a panoptic segmentation neural network to detect any object instance seen during training, regardless of whether the object instance belongs to an unknown object class.”). One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor, user-tagged groups, and image classifier model, distance, and similarity determination of Hwang, through known means, with no change to their respective function, and the combination would have yielded nothing more than predictable results. Specifically, one of ordinary skill in the art would have incorporated an analogous object clustering/grouping mechanic implemented by an image classifier as taught in Hwang to filter out images containing unknown and known objects before performing segmentation as disclosed in the method of Li.
Hwang does not specifically disclose the distances are between text embeddings and image embeddings. Therefore, a combination of Li and Hwang does not specifically teach text embeddings.
However, Che specifically teaches determining distances between text embeddings representing elements in the plurality of vector images to image embeddings of the plurality of vector images ([pg. 4, col. 1, par. 1, ln. 1 to par. 5, ln. 6] “For each image embedding, Model B employs a K-nearest neighbors (KNN) model to retrieve a set of nearest neighbor text embeddings. The KNN-based approach introduces a distance-based weighting mechanism, which captures the relevance and contextual significance of each neighbor. This distance-based weighting is calculated using the following formula:
W
i
g
h
t
i
=
1
D
i
s
t
a
n
c
e
(
i
)
D
i
s
t
a
n
c
e
_
D
i
m
+
δ
×
C
o
e
f
(4) Where Distance(𝑖) is the Euclidean distance between the query image embedding and the 𝑖 th neighbor text embedding, Distance_Dim controls the influence of distance, and 𝛿 is a small positive constant to ensure numerical stability, and Coef is a coefficient to prevent overflow. The weighted contributions from all neighbors are then aggregated to form the final text embedding. Mathematically, given an image embedding Image_Emb, the KNN-based fusion mechanism generates the text embedding Text_Emb as
T
e
x
t
E
m
b
=
1
K
∑
i
=
1
K
W
e
i
g
h
t
(
i
)
×
K
N
N
_
T
e
x
t
E
m
b
1
(5), Here,𝐾 represents the number of nearest neighbors, KNN_Text_Emb𝑖 refers to the 𝑖th nearest neighbor text embedding. In summary, Model B extends the zero-shot learning capabilities of the CLIP model to image-to-text transformation. The KNN-based fusion approach intelligently combines image and text emb, with distance-based weighting to capture contextual relevance, resulting in an enriched text representation that reflects the underlying semantics”). Specifically, one of ordinary skill in the art, before the effective filling date of the claimed invention, would recognize Li, Che, and Hwang as within the same field of semantic image segmentation, and as analogous to the claimed invention. The motivation to combine is disclosed in Che, wherein it allows for combining both image and text embeddings to better capture underlying semantics ([pg. 4, col. 1, par. 1, ln. 1 to par. 5, ln. 6]). One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of Li with the method performed by at least one processor, user-tagged groups, and image classifier model, distance, and similarity determination of Hwang, and further combined the method of the combination of Li and Hwang with the text embedding distance of Che, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art, before the effective filling date of the claimed invention, would have used the text embeddings as another metric to determine clustering as disclosed in the method of the combination of Li and Hwang, such that a distance between the text embeddings and image embeddings was computed analogous to Che.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the method performed by at least one processor, user-tagged groups, and image classifier model, distance, and similarity determination of Hwang and the text embedding distance of Che to obtain the invention as specified in claim 9.
32. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over “Semantic Hierarchy-Aware Segmentation” to Li, and further in view of U.S. Publication No. 2022/0375090 to Hwang, and further in view of “Combinatorial Optimization for Panoptic Segmentation: A Fully Differentiable Approach” to Abbas, and further in view of “Navigating the Forest: A Comprehensive Look at Tree Traversal Techniques” to Sirhith.
33. Regarding Claim 15, a combination of Li, Hwang, and Abbas teaches the method of claim 10. Li discloses extracting, from a vector file of the vector image, the vector hierarchy comprising a plurality of nodes corresponding to vector objects ([pg. 2124, Fig. 1], [pg. 2127, Fig. 2], [pg. 2124, col. 2, par. 2, ln. 1-25], [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24]); and executing a {breadth first} search algorithm on the vector hierarchy to determine the {user-}tagged groups of objects ([pg. 2124, Fig. 1], [pg. 2127, Fig. 2], [pg. 2124, col. 2, par. 2, ln. 1-25] [pg. 2126, col. 1, III. Our Approach, par. 1, ln. 1 to col. 2, par. 3, ln. 24]). Li does not specifically disclose the search is breadth first, or a user-tagged group. Likewise, Abbas does not disclose a breadth first search of a user-tagged group.
However, Hwang specifically teaches a user-tagged group ([par. 0136, ln. 1-15]). The motivation to combine remains analogous to claim 1. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li and Hwang with the IoU metric, bipartite matching, and IoU threshold of Abbas, through known means, with no change to their respective function, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have modified the method of Li to operate on a computer and processor, and to operate using user-tagged groups as taught in Hwang, such that the nodes of the tree of the combination of Li, Hwang, and Abbas corresponded to vector objects and user-tagged groups could be determined via traversing the tree.
However, Hwang does not specifically disclose a breadth first search. Therefore, a combination of Li, Hwang, and Abbas does not specifically disclose a breadth first search.
However, Sirhith specifically teaches a breadth first search to traverse a tree ([pg. 335, Iterative Tree Traversal, par. 1, ln. 1-5], [pg. 336, ln. 11-14], [pg. 339, ln. 12-15], [pg. 339, Choosing between recursion and iteration, par. 1, ln. 1-3], [pg. 338, see Table 1 and Table 2]). The motivation to combine remains analogous to claim 3 and 4. One of ordinary skill in the art, before the effective filling date of the claimed invention, would have combined the method of the combination of Li, Hwang, and Abbas with the breadth first tree traversal as taught in Sirhith, through known means, with no change to their respective unction, and the combination would have yielded nothing more than predicable results. Specifically, one of ordinary skill in the art would have used the BFS traversal of Sirhith to traverse the tree of the method of the combination of Li, Hwang, and Abbas.
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filling date of the claimed invention, to combine the method of Li with the system and user-tagged groups of Hwang, the bipartite matching and IoU metric of Abbas, and the BFS tree traversal of Sirhith to obtain the invention as specified in claim 15.
Allowable Subject Matter
34. Claim 11 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
35. The following is a statement of reasons for the indication of allowable subject matter:
While individual aspects of the limitations presented in claim 11 are taught in recommendations of record (e.g., Ma teaches multiple segmentation networks, Abbas teaches intersection-over-union), the examiner specifically notes that references of record fail to specifically disclose combining the semantic segmentations of separate neural networks and comparing them to the group masks. Specifically, Ma teaches wherein a majority voting mechanism is used, which the examiner notes would not be analogous to a combination set of the semantic segmentations, since only the highest confidence segmentation of each segmentation network would be considered.
Conclusion
36. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. See PTO-892.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAULO ANDRES GARCIA whose telephone number is (703)756-5493. The examiner can normally be reached Mon-Fri, 8-4:30PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached on (571)272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAULO ANDRES GARCIA/Examiner, Art Unit 2669
/SUMATI LEFKOWITZ/Supervisory Patent Examiner, Art Unit 2672