DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The Declaration under 37 CFR 1.130(a) filed on 07/14/2026 is sufficient to overcome the rejection of claim 11 based on “SOHES: Self-supervised Open-world Hierarchical Entity Segmentation”. The declaration provides evidence [exception provisions of 35 U.S.C. 102(b)(1)A] that authors Liang-Yan Gui and Yu-Xiong Wang did not contribute to the claimed invention of the current application and that the joint inventors listed on this application are the sole inventors of the subject matter which is claimed in this application.
Applicant's request for reconsideration of the rejection of the last Office action is persuasive with respect to the previous rejections under 35 U.S.C. § 103 over Cao in view of SOHES and, therefore, the rejection of that action is withdrawn.
The Amendment filed on July 14th, 2026 has been entered. Claims 1, 6–11, 16, and 18 are currently pending with claims 2-5, 12-15, 17, and 19-20 withdrawn from consideration. Claims 1, 7, and 16 have been amended.
Response to Arguments
Applicant's arguments filed July 14th, 2026 have been fully considered. Applicant’s arguments are persuasive with respect to the previous rejections under 35 U.S.C. § 102, as explained below.
Applicant's arguments filed 05/15/2026 are also persuasive with respect to the previous rejections under 35 U.S.C. § 102. Applicant's arguments, see pages 13–15 of the Remarks, state that aamended independent claims 1, 7, and 16 now require "generating the pseudo-labels comprises determining mask hierarchies ... based on a reclustered pool of regions generated by cropping local images from candidate regions of a pool of regions and reclustering patches within the local images," and that Cao relies on global merging and does not teach this limitation. The Examiner agrees that the cited portions of Cao do not explicitly disclose cropping local images from candidate regions and performing an additional clustering operation using the cropped local images. Therefore, the previous rejections under 35 U.S.C. § 102 over Cao have been withdrawn.
However, a new ground of rejection under 35 U.S.C. § 103 is made in this Office Action applying Cao in view of newly cited Cheung [Cheung et al] to address the newly added limitation regarding cropping local images and reclustering patches. This modification to the rejection is directly necessitated by Applicant's amendment adding new limitations to independent claims 1, 7, and 16.
Applicant’s amendments to claims 1, 7, and 16 have been fully considered. The newly added limitations from the amended claims have been considered and addressed in the updated rejections utilizing newly cited prior art. Because the necessity to apply a new reference to independent claims 1, 7, and 16 was directly necessitated by Applicant’s substantive amendments adding new limitations, this action is properly made final in accordance with MPEP § 706.07(a).
Based on these facts, this action is made FINAL.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 6–10, 16 and 18 are rejected under 35 U.S.C. §103 as being unpatentable over Cao in view of Cheung (Cheung et al, “A Lightweight Clustering Framework for Unsupervised Semantic Segmentation.” arXiv.Org, 2023).
Regarding claim 7, Cao teaches a system comprising:
one or more memory devices; and
one or more processors coupled to the one or more memory devices that cause the system to perform operations comprising:
( [Appendix J, “Hyper-Parameters and Implementation Details”]: Cao discloses implementing HASSOD on computing hardware using PyTorch and Detectron2, with training performed on 4× NVIDIA A100 GPUs. )
extracting, utilizing an encoder neural network, features representing a digital image;
( [Sec. 3.1, “Hierarchical Adaptive Clustering”]: Cao discloses using a frozen ViT-B/8 model pre-trained by DINO to extract visual features for each image, where each spatial element in the feature map corresponds to an 8×8 patch in the original image. )
generating, for the digital image, pseudo-labels indicating segmentations of object entities of the digital image and a hierarchical segmentation indicating hierarchical relations among the object entities based on the features; and
( [Sec. 3, “Approach”], [Sec. 3.1, “Hierarchical Adaptive Clustering”], [Fig. 2], [Fig. 3]: Cao discloses generating initial pseudo-labels from unlabeled images using self-supervised visual representations; progressively merging adjacent regions with the highest feature similarities into object masks; recording masks at multiple merging thresholds; and identifying hierarchical levels by analyzing coverage relations between masks and constructing tree structures. )
wherein generating the pseudo-labels comprises determining mask hierarchies of the object entities based on a merging patches based on feature similarities at a plurality of merging thresholds; and
( [Sec. 3.1, “Hierarchical Adaptive Clustering”; Sec. 3.2, “Hierarchical Level Prediction”; Fig. 3; Appendix D, “Choosing Thresholds for Hierarchical Adaptive Clustering”]: Cao discloses initializing each image patch as an individual region, calculating cosine similarities between the features of adjacent regions, iteratively merging the adjacent regions having the highest feature similarity, and recording the resulting object masks at a plurality of preset merging thresholds. Cao combines the object masks generated at the different thresholds into a pool of regions and determines mask hierarchies from the pooled masks by analyzing coverage relations and constructing a forest of trees representing whole objects, parts, and subparts. )
generating, utilizing a segmentation model comprising parameters generated according to the pseudo-labels, a segmentation map comprising a predicted hierarchical segmentation of the object entities of the digital image
( [Sec. 3.2, “Hierarchical Level Prediction”], [Sec. 3.3, “Mean Teacher Training with Adaptive Targets”], [Fig. 2], [Fig. 4-5]; [Paragraph “Improvement over initial pseudo-labels”]: Cao discloses learning an object detection and instance segmentation model using the initial pseudo-labels generated in the first stage, and further discloses a hierarchical level prediction head that classifies each predicted object as a whole object, part object, or subpart object, such that the learned model outputs predicted hierarchical segmentation results based on parameters learned according to the pseudo-labels. )
Cao teaches generating a pooled set of segmentation regions through patch-feature clustering at multiple thresholds and determining mask hierarchies from the pooled regions, but does not clearly teach locally cropping candidate regions and performing a further clustering operation on features extracted from the cropped local images, where Cheung teaches:
wherein generating the pseudo-labels comprises determining mask hierarchies of the object entities based on a reclustered pool of regions generated by cropping local images from candidate regions of a pool of regions and reclustering patches within the local images; and
( [Abstract; Fig. 1; Fig. 2; Algorithm 1; Section “Mask Refinement and Labeling”]: Cheung discloses refining pseudo-masks utilizing a multi-level clustering framework. Specifically, Cheung discloses extracting binary patch-level pseudo-masks representing foreground objects (candidate regions of a pool of regions), and subsequently cropping the bounding boxes of these object regions (cropping local images from candidate regions). To finalize the masks and assign labels, Cheung discloses passing the cropped object regions back into a Vision Transformer, which processes the cropped regions as patches, and performing a subsequent, repeated clustering step on the extracted features/tokens of the cropped regions (which reads on reclustering patches within the local images to determine or refine the mask classifications/hierarchies). )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Cao’s hierarchical clustering process to crop candidate object regions and further cluster features extracted from the cropped regions, as taught by Cheung, in order to refine the segmentation and classification of locally identified objects. Such a modification would have predictably applied Cheung’s known crop-and-further-cluster technique to Cao’s known patch-clustering framework to improve local object discrimination while retaining Cao’s existing hierarchy determination based on the resulting region masks.
Regarding claim 8, Cao [as modified by Cheung] teaches the system of claim 7, wherein extracting the features representing the digital image comprises:
determining a plurality of patches of pixels within the digital image; and
extracting, for each patch of the plurality of patches, a feature vector representing visual features of the pixels within the patch.
( [Sec. 3.1, “Hierarchical Adaptive Clustering”], [Fig. 3]: Cao discloses extracting visual features with a frozen DINO ViT-B/8, where each spatial element of the resulting feature map corresponds to an 8×8 patch in the original image, and each patch is initialized as an individual region, thereby teaching patch-wise feature-vector extraction for a plurality of patches in the digital image. )
Regarding claim 9, Cao [as modified by Cheung] teaches the system of claim 8, wherein generating the pseudo-labels comprises:
merging at least some of the plurality of patches into a first group of clusters based on similarities of the features at a first merging threshold; and
( [Sec. 3.1, “Hierarchical Adaptive Clustering”] with [Fig. 3], [Sec. 4.5, “Quality of initial pseudo-labels”]: Cao discloses computing the pairwise cosine similarity between adjacent patch regions, iteratively identifying the pair of adjacent regions with the highest feature similarity, merging those regions, and stopping when the similarity becomes smaller than a pre-set threshold θmerge; further recording the derived object masks when one of the pre-set thresholds is reached. )
determining mask hierarchies of the object entities within the digital image from the first group of clusters.
( [Sec. 3.2, “Hierarchical Level Prediction”, Fig. 3]: Cao discloses leveraging coverage relations between masks to construct a forest of trees, where roots are whole objects and descendants are part or subpart objects [object entities], thereby determining mask hierarchies from the clustered masks. )
Regarding claim 10, Cao [as modified by Cheung] teaches the system of claim 9, wherein generating the pseudo-labels further comprises:
merging at least some of the plurality of patches into a second group of clusters based on similarities of the features at a second merging threshold;
( [Sec. 3.1, “Hierarchical Adaptive Clustering”, Fig. 3], [Sec. D, “Choosing Thresholds for Hierarchical Adaptive Clustering”]: Cao discloses that it is beneficial to use multiple pre-set thresholds {
θ
i
m
e
r
g
e
}, e.g., 0.1, 0.2, and 0.4, and to record the derived object masks when the currently highest feature similarity reaches one of those thresholds. )
combining the first group of clusters and the second group of clusters into a pool of regions;
( [Sec. 3.1, “Hierarchical Adaptive Clustering”, Fig. 3]: Cao discloses ensembling results from multiple pre-set thresholds {
θ
i
m
e
r
g
e
} and combining the resulting object masks to ensure better coverage of potential objects. )
determining a modified pool of regions by removing duplicate regions from the pool of regions; and
( [Sec. 3.1, “Hierarchical Adaptive Clustering”, Fig. 3], [Sec. 4.5, “Quality of initial pseudo-labels”]: Cao discloses performing post-processing to refine and select the object masks after combining the results from multiple thresholds, and further discloses that post-processing removes a substantial portion of the labels while improving quality. )
determining the mask hierarchies of the object entities within the digital image based on the modified pool of regions.
( [Sec. 3.2, “Hierarchical Level Prediction”], [Appendix J, “Hyper-Parameters and Implementation Details”]: Cao discloses that, after ensembling results from the multiple thresholds and after post-processing, HASSOD identifies the hierarchical levels of each object mask based on coverage relation analysis. )
Regarding claims 1, 6, 16 and 18. The rationale provided for claim 7-10 is incorporated herein. In addition, the method for hierarchical entity segmentation of claim 1 and 6 correspond to the system of claims 7-10, as well as the non-transitory computer-readable medium of claims 16 and 18, and performs the steps disclosed herein. Therefore, the claims are all rejected.
Allowable Subject Matter
Claim 11 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEN KUDO
Examiner
Art Unit 2671
/KEN KUDO/Examiner, Art Unit 2671
/VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671