DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-23 are pending in the application.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 7-13, 15 and 17-18 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Ren et al. (Ren, S., et al., “Segment anything, from space?”, arXiv:2304.13000v3 [cs.CV] 10 Sep 2023. Hereafter Ren).
As per claim 1, Ren teaches a computer-implemented method (Abstract) for processing image data, wherein the method comprises:
obtaining the image data (Fig. 1/3 “Input Image”),
processing the image data in a domain-specific machine-learned model in order to obtain a first prediction for image features in the image data, wherein the first prediction comprises a first localization of structures in the image data (Fig. 1 “Vision Model”; Fig. 3 “U-Net”; Ren trains a domain-specific machine-learning model U-Net specifically for overhead imagery. See Table 1 and the following attached para. from page 2 left col. and page 3 left col. The result from U-Net is class-level masks which represent localization of structures in the image data),
PNG
media_image1.png
339
589
media_image1.png
Greyscale
PNG
media_image2.png
142
584
media_image2.png
Greyscale
determining context information for a generic machine-learned foundation model based on the first prediction for the image features (Fig. 1 “Prompts”; Fig. 3 “Box, b”; See below para. from page 3 right col.),
PNG
media_image3.png
385
577
media_image3.png
Greyscale
based on the context information, processing the image data in the generic machine-learned foundation model in order to obtain a second prediction for the image features, wherein the second prediction comprises a second localization of the structures in the image data (Fig. 1/3 “SAM” being the generic machine-learned foundation model (see Abstract for description of SAM). The segmentation result from SAM is shown in Fig. 1 and Fig. 2).
As per claim 7, dependent upon claim 1, Ren teaches wherein the first localization has a first accuracy, wherein the second localization has a second accuracy, wherein the second accuracy is greater than the first accuracy (Fig. 4 shows the result. In DG-land-water and DG-land-agri cases, SAM Bounding Box result has better accuracy than the U-Net result).
As per claim 8, dependent upon claim 1, Ren teaches wherein the first localization has a first image space density, wherein the second localization has a second image space density, wherein the second image space density is greater than the first image space density (In Fig. 3 the input into SAM can be points, and result from SAM is predicted mask. Therefore the second image space density is greater than the first image space density).
As per claim 9, dependent upon claim 1, Ren teaches the first localization has a first localization degree of detail, wherein the second localization has a second localization degree of detail, wherein the second localization degree of detail is greater than the first localization degree of detail (In Fig. 3 the input into SAM can be points, and result from SAM is predicted mask. Therefore the second localization degree of detail is greater than the first localization degree of detail).
As per claim 10, dependent upon claim 1, Ren teaches wherein the first prediction and/or the second prediction furthermore comprises classifying the structures in the image data (See below attached para. from page 3 left col.).
PNG
media_image4.png
141
570
media_image4.png
Greyscale
As per claim 11, dependent upon claim 10, Ren teaches wherein the first prediction comprises a point localization of structures and associated class assignments to multiple classes, wherein the second prediction comprises multiple result masks of a semantic segmentation or an instance segmentation of the structures, wherein the multiple result masks correspond to the multiple classes (Fig. 3 “Point, p”, “ Fig. 2 shows the results; See Table 1 and below attached para. from page 3 right col.).
PNG
media_image5.png
250
576
media_image5.png
Greyscale
As per claim 12, dependent upon claim 1, Ren teaches, wherein determining the context information comprises: modifying the first prediction (See below attached para. from page 3 right col.).
PNG
media_image6.png
219
584
media_image6.png
Greyscale
As per claim 13, dependent upon claim 1, Ren teaches, wherein modifying the first prediction comprises subsampling the first prediction, optionally random subsampling (See below attached para. from page 3 right col.).
PNG
media_image7.png
223
571
media_image7.png
Greyscale
As per claim 15, dependent upon claim 12, Ren teaches, wherein the first prediction is modified based on a user input received from a user interface (Fig. 1 “Human”; See attached para. from page 5 right col.).
PNG
media_image8.png
421
591
media_image8.png
Greyscale
As per claim 17, dependent upon claim 1, Ren teaches wherein multiple instances of the second prediction for the image features are obtained by way of the generic machine-learned foundation model, wherein the multiple instances of the second prediction correspond to different instances of the context information that have been modified in relation to one another and/or different instances of the first prediction for the image features and/or different configurations of the machine-learned foundation model (Table 1; Fig. 2; Fig. 5).
As per claim 18, dependent upon claim 1, Ren teaches, wherein the method furthermore comprises: determining a confidence of the second prediction based on a variation between multiple instances of the second prediction (Fig. 2; Fig. 4; Fig. 5; See below attached para. from page 3 right col.).
PNG
media_image9.png
139
575
media_image9.png
Greyscale
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2-6, 16 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ren et al. (Ren, S., et al., “Segment anything, from space?”, arXiv:2304.13000v3 [cs.CV] 10 Sep 2023. Hereafter Ren), in view of Zhang et al. (Zhang, Y., et al., “SamDSK: Combining Segment Anything Model with Domain-Specific Knowledge for Semi-Supervised Learning in Medical Image Segmentation”, arXiv:2308.13759v1 [cs.CV] 26 Aug 2023. Hereafter Zhang).
As per claim 2, dependent upon claim 1, Ren does not teach performing retraining of the domain-specific machine-learned model based on training data that comprise a ground truth determined based on the second prediction.
Zhang in the same field of endeavor proposes a method that combines the segmentation foundation model (i.e., SAM) with domain-specific knowledge for reliable utilization of unlabeled images in building a medical image segmentation model (Abstract). The method is iterative and consists of two main stages: (1) segmentation model training; (2) expanding the labeled set by using the trained segmentation model, an unlabeled set, SAM, and domain-specific knowledge. These two stages are repeated until no more samples are added to the labeled set. A novel optimal-matching-based
method is developed for combining the SAM-generated segmentation proposals and pixel-level and image-level DSK for constructing annotations of unlabeled images in the iterative stage (2) (Abstract). Specifically, as shown in Fig. 2, a segmentation model (a domain-specific machine-learning model) is trained to generate task-specific segmentation prediction, i.e., annotations of unlabeled images. The SAM model takes the domain-specific knowledge and generates segmentation proposals for the unlabeled images. Matching segmentation proposals are generated by SAM with pixel-level DSK generated by the current trained segmentation model. The matching is constrained by additional image-level DSK, i.e., potential counts of RoIs. The generated label (i.e., ground truth) is added back to the labeled set to retrain the segmentation model. See descriptions in pages 4-7 section 3 Methodology.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teaching of Ren to incorporate the teaching of Zhang to retrain the domain-specific machine-learned model based on training data that comprise a ground truth determined based on the second prediction. The motivation would be “the segmentation foundation model can be harnessed as a valuable tool for label-efficient segmentation learning in medical image segmentation” (Zhang Abstract).
As per claim 3, dependent upon claim 1, Ren in view of Zhang teaches wherein the method furthermore comprises: based on the second prediction: setting a user-interactive annotation process that is used to generate ground truths for retraining the domain-specific machine-learned model (Ren Fig. 1 shows a user-interactive annotation process can be performed in an iterative manner based on the second prediction. Page 2 left col. lines 4-7 “The annotator can optionally provide additional prompts as feedback to SAM, to iteratively refine its segmentation, or mask, predictions”).
As per claim 4, dependent upon claim 3, Ren in view of Zhang teaches wherein the user-interactive annotation process comprises modifying the context information based on a user input and accordingly outputting the influence of the modification of the context information on the second prediction to the user (Ren Fig. 1 shows a user-interactive annotation process can be performed in an iterative manner based on the second prediction. Page 2 left col. lines 4-7 “The annotator can optionally provide additional prompts as feedback to SAM, to iteratively refine its segmentation, or mask, predictions”; Fig. 7).
As per claim 5, dependent upon claim 4, Ren in view of Zhang teaches, wherein, in order to determine the influence of the modification of the context information on the second prediction, only that part of the generic machine-learned foundation model that has a dependency on the context information is in each case inferred again (Fig. 2; See the below attached para. from page 6).
PNG
media_image10.png
581
591
media_image10.png
Greyscale
As per claim 6, dependent upon claim 4, Ren in view of Zhang teaches wherein setting the user-interactive annotation process comprises: based on an active learning process: selecting part of the first prediction in order to modify the context information based on the user input (Ren teaches selecting part of the first prediction in order to modify the context information based on the user input (Fig. 1 and 3). Zhang teaches iteratively training the domain-specific model (i.e., an active learning process) to generate a segmentation result used for providing domain-specific knowledge (first prediction) to SAM model.)
As per claim 16, dependent upon claim 15, Ren in view of Zhang teaches wherein the method furthermore comprises: outputting part of the first prediction to the user, said part being selected based on an active learning process, and receiving the user input, which concerns a modification of the part of the first prediction (Ren teaches outputting part of the first prediction to the user and receiving the user input, which concerns a modification of the part of the first prediction (Fig. 1 and 3). Zhang teaches iteratively training the domain-specific model (i.e., an active learning process) to generate a segmentation result used for providing domain-specific knowledge (first prediction) to SAM model.)
As per claim 19, dependent upon claim 18, Ren in view of Zhang teaches wherein the confidence is determined in the image space of the image data in a resolved manner (Ren Fig. 2; Fig. 4; Fig. 5; See below attached para. from page 3 right col.),
PNG
media_image9.png
139
575
media_image9.png
Greyscale
wherein the method furthermore comprises:
based on the confidence of the second prediction: setting a user-interactive annotation process that is used to generate ground truths for retraining the domain-specific machine-learned model (Zhang Abstract; Fig. 2; pages 4-7 section 3 Methodology).
Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ren et al. (Ren, S., et al., “Segment anything, from space?”, arXiv:2304.13000v3 [cs.CV] 10 Sep 2023. Hereafter Ren), in view of Zhang et al. (US 20090080729 A1, hereafter Zhang (2009)).
As per claim 14, dependent upon claim 12, Ren does not teach modifying the first prediction comprises applying noise to the first prediction.
Zhang (2009) teaches a method for evaluating image segmentation (Abstract). Specifically, as shown in FIG. 1, the method includes generating synthetic image data, segmenting generated synthetic image data and measuring visibility of the segmentation results. In the step of generating synthetic image data, ground truth images of an object are segmented, and the segmented ground truth images are combined with background images, resulting in synthetic images having a known foreground (ground truth object) and background. Additionally, varying levels of noise and contrast is added to the synthetic image data (para. [0028]).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teaching of Ren to incorporate the teaching of Zhang (2009) to modify the first prediction by applying noise to the first prediction. Doing so would allow robustness of an image segmentation technique to be tested under varying levels of noise and contrast” (Zhang (2009) para. [0028]).
Claim(s) 20-23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ren et al. (Ren, S., et al., “Segment anything, from space?”, arXiv:2304.13000v3 [cs.CV] 10 Sep 2023. Hereafter Ren), in view of Kanazawa et al. (US 20200218961 A1, hereafter Kanazawa).
As per claim 20, Ren does not teach preprocessing the image data before processing in the domain-specific machine-learned model and/or before processing in the generic machine-learned foundation model.
Kanazawa in an analogous field discloses a method for image segmentation by using a semantic segmentation neural network and an edge refinement neural network (Abstract). Specifically, as shown in FIG. 3, high resolution image data is downscaled before inputting into the semantic segmentation neural network (the first model). The edge refinement neural network takes the same high resolution image data and the semantic segmentation mask output from the semantic segmentation neural network as input to generate a refined semantic segmentation mask. Note before input into the edge refinement neural network, the high resolution image data is also pre-processed (para. [0035] “In some implementations, the high resolution image can be sampled, such as, for example, by randomly cropping a portion of the high resolution image and providing the cropped portion to the edge refinement neural network”).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teaching of Ren to incorporate the teaching of Kanazawa to preprocess the image data before processing in the domain-specific machine-learned model and/or before processing in the generic machine-learned foundation model. The motivation of such a treatment would be “computational intensity of determining the low resolution semantic segmentation mask 350 can be considerably less than inputting a high resolution version of the image 310 directly into the semantic segmentation neural network 340” (Kanazawa para. [0077]).
As per claim 21, dependent upon 20, Ren in view of Kanazawa teaches wherein the preprocessing comprises one or more of the following operations: rescaling (Kanazawa FIG. 3 “DOWN SCALING”; para. [0075]); intensity normalization; aberration correction; denoising; and deconvolution.
As per claim 22, dependent upon 20, Ren in view of Kanazawa teaches, wherein the image data are preprocessed differently before processing in the domain-specific machine-learned model than before processing in the generic machine-learned foundation model (FIG. 3 shows the preprocessing for the image data input into the semantic segmentation neural network is down-scaling. Para. [0035] describes the preprocessing for the image data input into the edge refinement neural network is cropping.)
As per claim 23, dependent upon claim 1, Ren teaches a method as recited in claim 1. Ren, however, does not further teach a data processing device having a processor and a memory, wherein the processor is configured to load program code from the memory and execute it, wherein the processor implements a method as claimed in claim 1, when it executes the program code.
Kanazawa in an analogous field discloses a method and system for image segmentation by using a semantic segmentation neural network and an edge refinement neural network (Abstract). The system comprises a processor and a memory, wherein the processor is configured to load program code from the memory and execute it, wherein the processor implements the method when it executes the program code (FIG. 1 process, memory and instructions. Para. [0059]).
It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teaching of Ren to incorporate the teaching of Kanazawa to include a system comprising processor, memory and program code in order to execute the method as recited in claim 1 (Kanazawa FIG. 1).
Conclusion
Prior art searched but not cited is recorded in PTO-892.
Additional reference Kirillov et al. (Kirillov et al., "Segment Anything," arXiv:2304.02643v1, 5 Apr 2023) discloses a foundation model for segmentation by introducing three interconnected components: a promptable segmentation task, a segmentation model (SAM) that powers data annotation and enables zero-shot transfer to a range of tasks via prompt engineering, and a data engine for collecting SA-1B, a dataset of over 1 billion masks (Abstract; Fig. 1).
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XUEMEI G CHEN whose telephone number is (571)270-3480. The examiner can normally be reached Monday-Friday 9am-6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XUEMEI G CHEN/Primary Examiner, Art Unit 2661