Prosecution Insights
Last updated: August 06, 2026
Application No. 18/851,880

COARSE ANNOTATION BASED SEMANTIC SEGMENTATION MODEL TRAINING METHOD AND DEVICE

Non-Final OA §101§103
Filed
Sep 27, 2024
Priority
Apr 22, 2022 — nonprovisional of PCTCN2022088384
Examiner
LANTZ, KARSTEN FOSTER
Art Unit
Tech Center
Assignee
Hangzhou Innovation Institute Beihang University
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
4 granted / 4 resolved
+40.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
19 currently pending
Career history
28
Total Applications
across all art units

Statute-Specific Performance

§101
1.8%
-38.2% vs TC avg
§103
79.0%
+39.0% vs TC avg
§102
8.8%
-31.2% vs TC avg
§112
10.5%
-29.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged that application is a National Stage application of PCT PCT/CN2022/088384. Priority to PCT/CN2022/088384 with a priority date of 4/22/2022 is acknowledged under 35 USC 119(e) and 37 CFR 1.78. Information Disclosure Statement The IDS dated 9/27/2024 has been considered and placed in the application file. The NPL/Foreign reference(s) Wenxin Chen et al. “Multi-Branch Supervised Learning on Semantic Segmentation” on the IDS filed 9/27/2024 were not considered. See 37 CFR 1.98 (a) (2) for the requirements for a foreign patent or publication cited on an IDS. MPEP 609. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1- 10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The limitations, under their broadest reasonable interpretation, cover mental process using images/ drawings (concept performed in a human mind, including as observation, evaluation, judgment, opinion, prediction, etc.), and mathematical calculations for likelihood/ probability (e.g., - P(A) = f / N Where P(A) = Probability of an event (event A) occurring; f = Number of ways an event can occur (frequency); N = Total number of outcomes possible). This judicial exception is not integrated into a practical application because the steps do not add meaningful limitations to be considered specifically applied to a particular technological problem to be solved. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be done mentally and no additional features in the claims would preclude them from being performed as such. According to the USPTO guidelines, a claim is directed to non-statutory subject matter if: STEP 1: the claim does not fall within one of the four statutory categories of invention (process, machine, manufacture or composition of matter), or STEP 2: the claim recites a judicial exception, e.g. an abstract idea, without reciting additional elements that amount to significantly more than the judicial exception, as determined using the following analysis: STEP 2A (PRONG 1): Does the claim recite an abstract idea, law of nature, or natural phenomenon? STEP 2A (PRONG 2): Does the claim recite additional elements that integrate the judicial exception into a practical application? STEP 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? Using the two-step inquiry, it is clear that claims 1 and 10 are directed to an abstract idea as shown below: STEP 1: Do the claims fall within one of the statutory categories? YES. Claims 1-9 are directed to a method, i.e., process, and claim 10 is directed to a system i.e., a machine. STEP 2A (PRONG 1): Is the claim directed to a law of nature, a natural phenomenon or an abstract idea? YES, the claims are directed toward a mental process (i.e., abstract idea). With regard to STEP 2A (PRONG 1), the guidelines provide three groupings of subject matter that are considered abstract ideas: Mathematical concepts – mathematical relationships, mathematical formulas or equations, mathematical calculations; Certain methods of organizing human activity – fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions); and Mental processes – concepts that are practicably performed in the human mind (including an observation, evaluation, judgment, opinion). The method in claim 1, for example, comprises a mental process that can be practicably performed in the human mind therefore, an abstract idea. Claim 1 recites: 1. A coarse annotation based semantic segmentation model training method, comprising the following steps: sending an original image into a semantic segmentation model for processing; acquiring a semantic feature map output by a specified convolution layer in the semantic segmentation model; sending the semantic feature map to a first training branch for training to obtain a first cross entropy loss value, wherein the first training branch is a fully supervised branch depending on a ground-truth map; sending the semantic feature map to a second training branch for training to obtain a second cross entropy loss value, wherein the second training branch is an unsupervised branch without a ground-truth map; and determining an overall loss function according to the first cross entropy loss value and the second cross entropy loss value. These limitations, as drafted, under their broadest reasonable interpretation, cover performance of the limitations in the mind or by a human. The Examiner notes that under MPEP 2106.04(a)(2)(III), the courts consider a mental process (thinking) that “can be performed in the human mind, or by a human using a pen and paper" to be an abstract idea. CyberSource Corp. v. Retail Decisions, Inc., 654 F.3d 1366, 1372, 99 USPQ2d 1690, 1695 (Fed. Cir. 2011). As the Federal Circuit explained, "methods which can be performed mentally, or which are the equivalent of human mental work, are unpatentable abstract ideas the ‘basic tools of scientific and technological work’ that are open to all.’" 654 F.3d at 1371, 99 USPQ2d at 1694 (citing Gottschalk v. Benson, 409 U.S. 63, 175 USPQ 673 (1972)). See also Mayo Collaborative Servs. v. Prometheus Labs. Inc., 566 U.S. 66, 71, 101 USPQ2d 1961, 1965 ("‘[M]ental processes and abstract intellectual concepts are not patentable, as they are the basic tools of scientific and technological work’" (quoting Benson, 409 U.S. at 67, 175 USPQ at 675)); Parker v. Flook, 437 U.S. 584, 589, 198 USPQ 193, 197 (1978) (same). The mere nominal recitation that the various steps are being executed by a processor (e.g., processing unit) does not take the limitations out of the mental process grouping. Thus, the claims recite a mental process. If a claim limitation, under its broadest reasonable interpretation, covers performance of a mental step which could be performed with a simple tool such as a pen and paper, then it falls within the “mental steps” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. STEP 2A (PRONG 2): Does the claim recite additional elements that integrate the judicial exception into a practical application? NO, the claims do not recite additional elements that integrate the judicial exception into a practical application. With regard to STEP 2A (prong 2), whether the claim recites additional elements that integrate the judicial exception into a practical application, the guidelines provide the following exemplary considerations that are indicative that an additional element (or combination of elements) may have integrated the judicial exception into a practical application: an additional element reflects an improvement in the functioning of a computer, or an improvement to other technology or technical field; an additional element that applies or uses a judicial exception to affect a particular treatment or prophylaxis for a disease or medical condition; an additional element implements a judicial exception with, or uses a judicial exception in conjunction with, a particular machine or manufacture that is integral to the claim; an additional element effects a transformation or reduction of a particular article to a different state or thing; and an additional element applies or uses the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception. While the guidelines further state that the exemplary considerations are not an exhaustive list and that there may be other examples of integrating the exception into a practical application, the guidelines also list examples in which a judicial exception has not been integrated into a practical application: an additional element merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea; an additional element adds insignificant extra-solution activity to the judicial exception; and an additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use. Thus, Claims 1- 10 do not recite any of the exemplary considerations that are indicative of an abstract idea having been integrated into a practical application. Thus, since Claims 1 and 10 are/is: (a) directed toward an abstract idea, (b) do not recite additional elements that integrate the judicial exception into a practical application, and (c) do not recite additional elements that amount to significantly more than the judicial exception, claims 1 and 10 are not eligible subject matter under 35 U.S.C 101. Similar analysis is made for the dependent claims 2-9 and the dependent claims are similarly identified as: being directed towards an abstract idea, not reciting additional elements that integrate the judicial exception into a practical application, and not reciting additional elements that amount to significantly more than the judicial exception. 1st Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 2, and 10 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2020 0234447 A1, (Karmatha et al.) in view of US Patent Publication 2022 11450008 B1, (Tyagi et al.). Claim 1 Regarding claim 1, Karmatha et al. teach a coarse annotation based semantic segmentation model training method, comprising the following steps: sending an original image into a semantic segmentation model for processing; ("a method of training a model is provided, said model being for semantic segmenting an image," par. 49) acquiring a semantic feature map output by a specified convolution layer in the semantic segmentation model; ("the model comprising: a common processing stage to produce a first feature map," par. 49-50) sending the semantic feature map to a first training branch for training ("a parallel processing stage, said second processing stage comprising first and second parallel branches that receive the first feature map," par. 51) sending the semantic feature map to a second training branch for training ("a parallel processing stage, said second processing stage comprising first and second parallel branches that receive the first feature map," par. 51) and determining an overall loss function according to the first cross entropy loss value and the second cross entropy loss value ("the method further comprising training using the images as input and determining the loss by comparing with both the semantic segmented information at both the output and at the second output and updating the weights during training by using the determined losses from both outputs," par. 58). Karmatha et al. do not explicitly teach all of to obtain a first cross entropy loss value, wherein the first training branch is a fully supervised branch depending on a ground-truth map; and to obtain a second cross entropy loss value, wherein the second training branch is an unsupervised branch without a ground-truth map. However, Tyagi et al. teach to obtain a first cross entropy loss value, ("At action 440, the graph-cut segmentation mask and the first frame of image data including the bounding box may be used to determine a per-pixel cross-entropy loss," col. 8, line 48) wherein the first training branch is a fully supervised branch depending on a ground-truth map; ("The first example approach regularizes the attention mask using ground truth bounding boxes, by minimizing the L.sub.2 loss between the attention maps and the bounding boxes," col. 6, line 23) and to obtain a second cross entropy loss value, wherein the second training branch is an unsupervised branch without a ground-truth map ("back propagation of gradients for all pixels would be an issue for a weakly-supervised segmentation algorithm. To handle this, various embodiments described below predict a novel per-pixel class-specific attention map and pixel embeddings in addition to the per-pixel segmentation output. The attention map may be used to modulate the per-pixel cross-entropy loss," col. 3, line 35). Therefore, taking the teachings of Karmatha et al. and Tyagi et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the two branch semantic segmentation methods as taught by Karmatha et al. to use comparing multiple cross entropy losses as taught by Tyagi et al. The suggestion/motivation for doing so would have been that, “The attention map may be used to modulate the per-pixel cross-entropy loss to handle label noise and may reduce propagation of incorrect gradients” as noted by the Tyagi et al. disclosure in paragraph [14], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that the attention map would successfully focus the model on the most relevant features while suppressing incorrect gradients, ultimately yielding an improvement in pixel-wise classification accuracy; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. The rejection of method claim 1 above applies mutatis mutandis to the corresponding limitations of apparatus claim 10 while noting that the rejection above cites to both device disclosures. Claim 10 is mapped below for clarity of the record and to specify any new limitations not included in claim 1. Claim 2 PNG media_image1.png 264 686 media_image1.png Greyscale [AltContent: textbox (Figure 4 shows the network architecture.)][AltContent: textbox (Figure 3 shows the flow diagram of the neural network.)] PNG media_image2.png 532 225 media_image2.png Greyscale Regarding claim 2, Karmatha et al. teach wherein the specified convolution layer is the penultimate layer of the semantic segmentation model; ("In Step S303, the first layer of the convolutional neural network is a two-dimensional pooling convolution layer in order to achieve more efficient object classification," par. 79) and the first training branch is the last fully connected layer in the semantic segmentation model ("in FIG. 4. After the ‘Learning to downsample’ module the network branches out into two branches. The first network branch comprises 9 bottleneck residual blocks, 409, 411, 413, 415, 417, 419, 421, 423, 425, and 427 followed by a pyramid pooling module 427," par. 113). Karmatha et al. and Tyagi et al. are combined as per claim 1. Claim 10 Regarding claim 10, Karmatha et al. teach a coarse annotation based semantic segmentation model training device, comprising: an input module, which is configured to send an original image into a semantic segmentation model for processing; ("a method of training a model is provided, said model being for semantic segmenting an image," par. 49) an acquisition module, which is configured to acquire a semantic feature map output by a specified convolution layer in the semantic segmentation model; ("the model comprising: a common processing stage to produce a first feature map," par. 49-50) a first training branch, which is configured to train the semantic feature map ("a parallel processing stage, said second processing stage comprising first and second parallel branches that receive the first feature map," par. 51) a second training branch, which is configured to train the semantic feature map ("a parallel processing stage, said second processing stage comprising first and second parallel branches that receive the first feature map," par. 51) and a determining module, which is configured to determine an overall function according to the first cross entropy loss value and second cross entropy loss value ("the method further comprising training using the images as input and determining the loss by comparing with both the semantic segmented information at both the output and at the second output and updating the weights during training by using the determined losses from both outputs," par. 58). Karmatha et al. do not explicitly teach all of to obtain a first cross entropy loss value, wherein the first training branch is a fully supervised branch depending on a ground-truth map; and to obtain a second cross entropy loss value, wherein the second training branch is an unsupervised branch without a ground-truth map. However, Tyagi et al. teach to obtain a first cross entropy loss value, ("At action 440, the graph-cut segmentation mask and the first frame of image data including the bounding box may be used to determine a per-pixel cross-entropy loss," col. 8, line 48) wherein the first training branch is a fully supervised branch depending on a ground-truth map; ("The first example approach regularizes the attention mask using ground truth bounding boxes, by minimizing the L.sub.2 loss between the attention maps and the bounding boxes," col. 6, line 23) and to obtain a second cross entropy loss value, wherein the second training branch is an unsupervised branch without a ground-truth map ("back propagation of gradients for all pixels would be an issue for a weakly-supervised segmentation algorithm. To handle this, various embodiments described below predict a novel per-pixel class-specific attention map and pixel embeddings in addition to the per-pixel segmentation output. The attention map may be used to modulate the per-pixel cross-entropy loss," col. 3, line 35). Karmatha et al. and Tyagi et al. are combined as per claim 1. 2nd Claim Rejections - 35 USC § 103 Claim 3 is rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2020 0234447 A1, (Karmatha et al.) and US Patent Publication 2022 11450008 B1, (Tyagi et al.) in view of US Patent Publication 2021 0150281 A1, (Tsai et al.). Claim 3 Regarding claim 3, Tyagi et al. teach Yi is the true class of valid pixels in the ground-truth map Y ("The attention weighted loss is only defined for the L foreground classes. In addition, the loss is normalized with the size of the bounding box, to give similar weighting to each class," col. 5, line 63). Karmatha et al. and Tyagi et al. do not explicitly teach all of wherein the first cross entropy loss value is: Li=CE(Pi,Yi); where Pi= softmax(Ti), Ti =WfcOQi, Qi is the semantic feature map, Wfc is the model parameter of the last fully connected layer, the symbol O indicates matrix multiplication, Ti is the output result of the last fully connected layer. However, Tsai et al. teach wherein the first cross entropy loss value is: Li=CE(Pi,Yi); where Pi= softmax(Ti), Ti =WfcOQi, Qi is the semantic feature map, Wfc is the model parameter of the last fully connected layer, the symbol O indicates matrix multiplication, Ti is the output result of the last fully connected layer ("where L.sub.seg is the cross-entropy loss using ground truth annotations in the source domain, and L.sub.adv is the adversarial loss that adapts predicted segmentations of target images to the distribution of source predictions. λ.sub.adv is the weight used to balance the two losses," par. 56). Therefore, taking the teachings of Karmatha et al., Tyagi et al., and Tsai et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the two branch semantic segmentation methods as taught by Karmatha et al. and comparing multiple cross entropy losses as taught by Tyagi et al. to use the cross entropy loss calculation as taught by Tsai et al. The suggestion/motivation for doing so would have been that, “cross-entropy loss using ground truth annotations in the source domain, and L.sub.adv is the adversarial loss that adapts predicted segmentations of target images to the distribution of source predictions. λ.sub.adv is the weight used to balance the two losses. Although segmentation outputs are in the low-dimensional space, they contain rich information, e.g., scene layout and context” as noted by the Tsai et al. disclosure in paragraph [0056], which also motivates combination because the combination would predictably have a higher performance as there is a reasonable expectation that the calculated output distribution across the target domain would accurately reflect the scene layout and context learned from the source domain; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. 3rd Claim Rejections - 35 USC § 103 Claims 4, 5, and 6 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2020 0234447 A1, (Karmatha et al.) and US Patent Publication 2022 11450008 B1, (Tyagi et al.) in view of US Patent Publication 2021 0241109 A1, (Jie). Claim 4 Regarding claim 4, Karmatha et al. and Tyagi et al. do not explicitly teach all of wherein sending the semantic feature map to a second training branch for training comprises: performing unsupervised clustering on the semantic feature map, and classifying unannotated sample vectors; wherein the unannotated sample vectors are the feature vectors corresponding to the unannotated pixel in the semantic feature map. However, Jie teaches wherein sending the semantic feature map to a second training branch for training comprises: performing unsupervised clustering on the semantic feature map, and classifying unannotated sample vectors; ("a pixel clustering method may be used. K center points are first chosen," par. 164) wherein the unannotated sample vectors are the feature vectors corresponding to the unannotated pixel in the semantic feature map ("All points in an image are distributed to the K centers based on differences between each pixel point and the K pixels," par. 164). Therefore, taking the teachings of Karmatha et al., Tyagi et al., and Jie as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the two branch semantic segmentation methods as taught by Karmatha et al. and comparing multiple cross entropy losses as taught by Tyagi et al. to use the pixel clustering training methods as taught by Jie. The suggestion/motivation for doing so would have been that, “a pixel-level annotation does not need to be performed on massive images, to implement semantic image segmentation under a weakly-supervised condition. Only an image-level annotation is needed, and semantic segmentation precision comparable to that of an existing method may be achieved without expensive pixel-level information” as noted by the Jie disclosure in paragraph [0166], which also motivates combination because the combination would predictably have a higher efficiency as there is a reasonable expectation that the iterative pixel clustering method will accurately group similar features without manual pixel-level labels; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Claim 5 Regarding claim 5, Karmatha et al. and Tyagi et al. do not explicitly teach all of wherein the step of performing unsupervised clustering on the semantic feature map comprises: performing repeated iterations by a class-centered clustering algorithm; wherein each single iteration process comprises initialization, sample division and parameter updating. However, Jie teaches wherein the step of performing unsupervised clustering on the semantic feature map comprises: performing repeated iterations by a class-centered clustering algorithm; wherein each single iteration process comprises initialization, sample division and parameter updating ("each class center is recalculated, and iteration and optimization are performed based on the foregoing operations, so that all pixels in the image are classified into K classes," par. 164). Karmatha et al., Tyagi et al., and Jie are combined as per claim 4. Claim 6 Regarding claim 6, Karmatha et al. and Tyagi et al. do not explicitly teach all of wherein: in the initialization stage, an empty list of stored samples is established for the class corresponding to each class-centered vector; in the sample division stage, according to the similarity between the class-centered vector and the sample vector, each sample vector is divided into the sample list to which the class with the highest similarity belongs to; in the parameter updating stage, the sample vector and the class-centered vector are updated simultaneously by using a gradient back propagation method with reference to the fully supervised branch. However, Jie teaches wherein: in the initialization stage, an empty list of stored samples is established for the class corresponding to each class-centered vector; in the sample division stage, according to the similarity between the class-centered vector and the sample vector, each sample vector is divided into the sample list to which the class with the highest similarity belongs to; ("a pixel clustering method may be used. K center points are first chosen. All points in an image are distributed to the K centers based on differences between each pixel point and the K pixels," par. 164) in the parameter updating stage, the sample vector and the class-centered vector are updated simultaneously by using a gradient back propagation method with reference to the fully supervised branch ("each class center is recalculated, and iteration and optimization are performed based on the foregoing operations, so that all pixels in the image are classified into K classes," par. 164). Karmatha et al., Tyagi et al., and Jie are combined as per claim 4. 4th Claim Rejections - 35 USC § 103 Claim 8 is rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2020 0234447 A1, (Karmatha et al.) and US Patent Publication 2022 11450008 B1, (Tyagi et al.) in view of CN Patent Publication 2021 113469006 A, (Yang et al.) and US Patent Publication 2022 0164569 A1, (Kim et al.). Claim 8 Regarding claim 8, Karmatha et al. and Tyagi et al. do not explicitly teach all of wherein prior to sending the semantic feature map to a second training branch for training, the method further comprises: performing global average pooling operation on the whole semantic feature map to obtain the average vector of the whole semantic feature map Q, and taking the average vector as a common semantic vector B; uniformly subtracting the common semantic vector B from each vector Qi in the whole semantic feature map Q to obtain a unique semantic vector map R; enhancing each vector Ri in the unique semantic vector map R by using an enhancement function to obtain an enhanced vector R’i; adding the enhanced vector R’i to the initial feature vector Qi to obtain an enhanced feature map Q'. However, Yang et al. teach wherein prior to sending the semantic feature map to a second training branch for training, the method further comprises: performing global average pooling operation on the whole semantic feature map to obtain the average vector of the whole semantic feature map Q, and taking the average vector as a common semantic vector B; ("using the global comparison pool to convert the feature map into feature vector. As shown in FIG. 4, firstly according to the formula (1) and formula (2) to the characteristic map respectively global maximum pool and global average pool operation to obtain pmax and pavg," pg. 3, par. 6) uniformly subtracting the common semantic vector B from each vector Qi in the whole semantic feature map Q to obtain a unique semantic vector map R; ("the pmax and pavg are spliced together to obtain pcat; then according to the formula (4) the pcat through convolution and nonlinear activation function ReLU to obtain finally according to formula (5) subtracting pmax to obtain the final characteristic vector q," pg. 3, par. 6) enhancing each vector Ri in the unique semantic vector map R by using an enhancement function to obtain an enhanced vector R’i ("according to the formula (4) the pcat through convolution and nonlinear activation function ReLU," pg. 3, par. 6). Therefore, taking the teachings of Karmatha et al., Tyagi et al., and Yang et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the two branch semantic segmentation methods as taught by Karmatha et al. and comparing multiple cross entropy losses as taught by Tyagi et al. to use the pooling and vector enhancing methods as taught by Yang et al. The suggestion/motivation for doing so would have been because such a modification is the result of combining prior art elements according to known methods to yield predictable results. More specifically, this combination can yield a predictable result of enhanced feature representation and improved spatial context processing since REASON. Thus, a person of ordinary skill would have appreciated including in the unsupervised training branch the ability to do use average and max pooling since the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable. Additionally, Kim et al. teach adding the enhanced vector R’i to the initial feature vector Qi to obtain an enhanced feature map Q' ("feature vector by multiplying a scaling parameter to the spatial feature vector and adding the initial input video feature as shown in Equation 4 to output as the spatial feature map," par. 58). Therefore, taking the teachings of Karmatha et al., Tyagi et al., Yang et al. and Kim et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify the two branch semantic segmentation methods as taught by Karmatha et al., comparing multiple cross entropy losses as taught by Tyagi et al., and the pooling and vector enhancing methods as taught by Yang et al. to use adding multiple vectors to result in a feature map as taught by Kim et al. The suggestion/motivation for doing so would have been that, “The data transformation may be performed by a separate member other than the spatial attention module 200. Alternatively, the data transformation may just mean a selective use of only some portion of the video features stored in the memory rather than an actual data manipulation” as noted by the Kim et al. disclosure in paragraph [0053], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that the resulting enhanced feature map will capture more detailed image characteristics because performing a data transformation by modifying them through simple vector addition is a predictable operation in deep learning; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Allowable Subject Matter Claims 7 and 9 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 7, “in the initialization stage of each iteration, performing L2 norm-based normalization on the class-centered vector; in the sample division stage, defining the similarity between the class-centered vector and the sample vector as: Sm,n =AmOAn; where Am and An are the unit vectors corresponding to the class-centered vector and the sample vector, respectively, and the symbol O indicates a dot product operation between the vectors.” A person of ordinary skill in the art would not have been motivated to combine the prior art with the specific combination of L2 norm-based normalization of a class-centered vector and dot-product similarity, as the established understanding in the field would lead one away from discarding magnitude information. Regarding claim 9, “condensing the semantics of the unique semantic vector Ri through a fully connected layer of contraction parameters and a Swish activation function; converting the condensed semantic vectors into channel-by- channel enhancement coefficients by a fully connected layer of expansion parameters and a hyperbolic tangent TanH activation function; performing channel-by-channel multiplication of the obtained coefficients by the initial Ri to complete the enhancement of the unique semantic vector Ri.” One of ordinary skill in the art would not be motivated to sequentially combine a Swish and TanH activation function for channel-by-channel enhancement, as the prior art lacks any teaching of the resulting non-monotonic and zero-centered gradient flow. This specific layered architecture yields unexpected improvements in semantic vector scaling that are not suggested by the individual components. Reference Cited The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. CN Patent Publication 2020 110837836 A1 to Jin et al. discloses confidence-maximized semi-supervised semantic segmentation. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KARSTEN F LANTZ whose telephone number is (571) 272-4564. The examiner can normally be reached Monday-Friday 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached on 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Karsten F. Lantz/Examiner, Art Unit 2664 Date: 7/9/2026 /JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664
Read full office action

Prosecution Timeline

Sep 27, 2024
Application Filed
Jul 21, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 7m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month