Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This action is in response to the applicant's communication filed on 10/10/2024. In virtue of this communication, claims 1-10 filed on 10/10/2024 are currently pending in the instant application.
Information Disclosure Statement
The information Disclosure statement (IDS) form PTO-1449, filed on 10/10/2024 are in compliance with the provisions of CFR 1.97. Accordingly, the information disclosed therein was considered by the examiner.
Drawings
The drawings are objected to under 37 CFR 1.83(a) because they fail to show details as described in the specification. Each empty boxes in the drawing needs to be named to help understanding the invention. Any structural detail that is essential for a proper understanding of the disclosed invention should be shown in the drawing. MPEP § 608.02(d). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Priority
Acknowledgment is made of applicant's claim for foreign priority under 35 U.S.C. 119(a)-(d).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 8 is rejected under 35 U.S.C. 101 as not falling within one of the four statutory categories of invention because the claimed invention is directed to computer program or software per se without reciting a physical or tangible medium containing the program, a machine incorporating the program, or the performance of process. See MPEP 2106(I). the recitation that the instructions cause a computer to carry out the method when the program is executed merely describes the function of the program and does not provide a present structural ;imitation placing the claimed program within the statutory category. A claim directed toward a non-transitory computer-readable medium having the program encoded thereon establishes a sufficient functional relationship between the program and a computer so as to remove it from the realm of “program per se”. MPEP 2111.05(III). Hence, adding the limitation of “stored on a non-transitory computer-readable medium” would resolve this issue.
Claim 8 (Amended) A non-transitory computer readable storage medium storing instructions that, when executed by computer to carry out the method of claim1.
Claims 10 is rejected under 35 U.S.C. § 101 because the claims are directed to non-statutory subject matter in the form of a “computer readable storage medium.” The claims fall outside the scope of patent-eligible subject matter at least because the claimed computer readable storage medium is broad enough to encompass non-transitory embodiments. (E.g., one of ordinary skill in the art could reasonably be expected to interpret the claimed computer readable medium as a carrier wave onto which instructions could be coded.)
See the precedential decision in Ex parte Mewherter and also the Mentor Graphics v. EVE-USA, Inc., 851 F.3d 1275, 112 USPQ2d 1120 (Fed. Cir. 2017) “Subject Matter Eligibility of Computer Readable Media” which states in relevant part “[i]n an effort to assist the patent community in overcoming a rejection or potential rejection under 35 U.S.C. § 101 in this situation, the USPTO suggests the following approach. A claim drawn to such a computer readable medium that covers both transitory and non-transitory embodiments may be amended to narrow the claim to cover only statutory embodiments to avoid a rejection under 35 U.S.C. § 101 by adding the limitation ‘non-transitory’ to the claim.”
Therefore, an amendment applicable to the claims, consistent with the recommendations in the above-noted Mentor Graphics v. EVE-USA, Inc., that would overcome the instant ‘101 rejection, follows:
Claim 10 (Amended) A non-transitory computer readable medium …
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 4-6, 8, 9, and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al., "Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images," in IEEE Transactions on Geoscience and Remote Sensing, 2022, in view of Shim et al. (US 2023/0093619), further in view of Kirillov et al. "Segment anything." 2023 IEEE/CVF international conference on computer vision (ICCV). IEEE, Oct 1, 2023.
As per claim 1, A method for providing a combined training data set for a machine learning model, comprising: “providing image data, wherein the image data comprises a non-labelled portion and a labelled portion;”(Li, Figure1, shows un-labeled images for self- supervised pre training, and few labeled data for supervised fine tuning. Page2, Col1, paragraph 2 discloses first learns knowledge from unlabeled image data and subsequently transfers the learned knowledge to a don stream task using a limited number of labeled samples.)
“training a base machine learning model based on the non-labelled portion of the image data to provide a generalized model;” (Li, page1, Col. 1 discloses pretraining a general model with a large number of unlabeled images. page 2, Col. 2, paragraph 2 discloses use contrastive learning to enhance the consistency of the sample on the label-free data to learn a General Remote Sensing vIsion model (G-RSIM). paragraph 3 discloses the self-supervised learning model(SSL) pretrained with unlabeled images. Further page 3, Section II.A discloses SSL learning potentially useful knowledge directly from a large amount of readily available unlabeled data and then transferring it to downstream tasks to achieve a better performance the downstream task is the semantic segmentation of RSIs, as such, we concentrate on designing a self-supervised model for the semantic segmentation of RSIs. In this article, we introduce contrastive learning to learn the general invariant representation
“training the generalized model based on the labelled portion of the image data to provide a semantic segmentation model;”(Li, page 1, Col.1, discloses fine-tuning the general model on a downstream task with labeled samples. Page 7, Col.2 section B.3) Implementation Details: discloses loading the encoder part from the self
supervised pretraining stage during the fine-tuning stage and using only a limited amount of annotated data for semantic segmentation training.)
“analyzing a training data set with the semantic segmentation model to provide a semantics for the training data set;”(Li, page 1, Col 2, discloses Semantic segmentation, as a pixel-level image analysis. Page 3, Col 1, Section II. A paragraph 1, discloses the downstream task is the semantic segmentation of remote sensing images(RSIs). Page 7, Col.1 Section III.B.2) “Evaluation Metrics” discloses applying downstream semantic segmentation model to image data and generating predicted pixels associated with respective classes, as a result providing class semantics for the analyzed images.)
However Li does not explicitly disclose the following which would have been obvious in view of Shim from similar filed of endeavor “analyzing the training data set with a segmentation model to provide segmentation for the training data set and providing the combined training data set based on a combination of the provided semantics and the provided segmentation of the training data set.”(Shim, ¶[0084] discloses generate a plurality of localization maps according to each classes¶[0116] discloses localization map can distinguish different objects or identities.(provided semantic) ¶[0086] The saliency map may provide object silhouettes that better represent object boundaries.¶[0092] discloses saliency map obtained from the existing saliency detector may be used as pseudo-masks¶[0116] discloses localization map can distinguish different objects or identities, The saliency map provides rich boundary information, but does not reveal the identity of the object.(foreground background segmentation). ¶[0102] discloses explicit pseudo-pixel supervision (EPS) that integrates the saliency map into the pseudo-pixel maps in the weakly supervised semantic segmentation and uses the integrated map as a clue for the boundary and the co-occurring pixel to train pixel-level feedback. ¶[0104] discloses matching the estimated foreground localization map with the foreground of the saliency map.¶[0116] disclose the localization map may distinguish different objects, but does not effectively distinguish boundaries. The saliency map provides rich boundary information, but does not reveal the identity of the object. In contrast, the present disclosure using both the localization map and the saliency map as illustrated in FIG. 6D may accurately classify people, trains, and cars.(combination). Further ¶[0096] discloses generate the pseudo-masks based on a plurality of second localization maps by the an updated classifier ¶[0097] discloses The pseudo-masks generator 270 may generate pseudo-masks for training a segmented network by joint training through the saliency loss and the classification loss. ¶[0139] discloses measuring the accuracy of the pseudo-masks of the train set is a common protocol in the WSSS because the pseudo-masks of the train set are used to guide the segmentation model.(combined training dataset).)
Before the effective filing date of the claimed invention it would have been obvious to a person of ordinary skill in the art to combine Shim technique of semantic segmentation into Li technique to provide the known and expected uses and benefits of Shim technique over learning model for semantic segmentation technique of Li. The proposed combination would have constituted a mere arrangement of old elements with each performing their known function, the combination yielding no more than one would expect from such an arrangement.
Therefore, it would have been obvious to a person of ordinary skill in the art to incorporate Shim to Li in order to improve performance of weakly supervised learning-based semantic segmentation by utilizing a localization map and a saliency map. (Refer to Shim paragraph [0023].)
However Li as modified by Shim does not explicitly disclose the following which would have been obvious in view of Kirillov from similar filed of endeavor “analyzing the training data set with a zero-shot segmentation model to provide segmentation for the training data set”(Kirillov, page 1, Col. 1 discloses the SAM model is designed and trained to be promotable, so it can transfer zero-shot to new image distributions and tasks. Page4, Col1. Section 2, “zero-shot transfer”, downstream segmentation tasks are performed through appropriate prompts at inference time, including automatic dataset labeling. Further discloses the promotable segmentation task returns a valid segmentation mask. Page 5, Col. 1, section 3, discloses SAM includes image encoder to process each image, prompt encoder, mask decoder produce mask from the image and prompt embedding, resolving ambiguity, wherein the mask decoder efficiently maps the image embedding, prompt embeddings, and an output token to a mask and computes the foreground probability at each image location. Further page 5, Col. 2, section 4 discloses fully automatic stage which SAM generates masks without annotator input. Page 6, Col. 1, “fully automatic stage”, discloses prompting SAM with a regular 32x32 grid of points, for each point predicted a set of masks that may correspond to valid objects and applying fully automatic mask to all images in the dataset. The resulting automatically generated masks are effective for training models and are included in the SA-1B segmentation dataset. )
Before the effective filing date of the claimed invention it would have been obvious to a person of ordinary skill in the art to combine Kirillov technique of zero-shot segmentation into Li as modified by Shim technique to provide the known and expected uses and benefits of Kirillov technique over learning model for semantic segmentation technique of Li as modified by Shim. The proposed combination would have constituted a mere arrangement of old elements with each performing their known function, the combination yielding no more than one would expect from such an arrangement.
Therefore, it would have been obvious to a person of ordinary skill in the art to incorporate Kirillov to Li as modified by Shim in order to improve model scale, dataset size and total training compute. (Refer to Kirillov Page1, Col.2, paragraph 1.)
Claim 8, 9, and 10 have been analyzed and are rejected for the reasons indicated in claim 1 above. Additionally, the rationale and motivation to combine the Li, Shim, and Kirillov references, presented in rejection of claim 1, apply to these claims.
As per claim 4, The method according to claim 1, Li as modified by Shim as modified by Kirillov further discloses “wherein analyzing the training data set with the semantic segmentation model comprises: “segmenting the training data set with the semantic segmentation model to provide first segmented areas in the training data set, and associating a respective characteristic with a respective first segmented area using the semantic segmentation model to provide the semantics for the training data set.”(Li, page1, Col. 2 discloses Semantic segmentation, is a pixel-level image analysis technology. page 7, Col. 2 discloses ac denotes the actual number of pixels of class c, and bc denotes the number of predicted pixels of class c . the pixels predicted as belonging to a respective class form the class associated regions produces by the semantic segmentation model. further Col. 2, section III. B. 3 discloses GLCNet train the complete encoder–decoder part of DeepLabV3+. Further discloses loading the self-supervised pretrained encoder during fine-tuning and using annotated data. Further last paragraph discloses the output of its semantic segmentation model includes class specific pixel predication. page 8, figure 7, discloses evaluation the semantic segmentation results for each class, including building, cars, water,… and other defined classes. Page 9, figure 8, Col. 2 discloses the model outputs spatially separated color-codes semantic regions corresponding to the class associated segmentation results. )
As per claim 5, The method according to claim 1, Li as modified by Shim as modified by Kirillov further discloses wherein analyzing the training data set with the zero-shot segmentation model comprises: “segmenting the training data set with the zero-shot segmentation model to provide the segmentation for the training data set based on second segmented areas in the training data set.”(Kirillov, page 1, Col. 1, discloses SAM is designed and trained to transfer Zero-shot to new image distributions and tasks. Page 4, Col1, Section 2. Zero shot task, discloses Zero-shot downstream segmentation tasks are performed by Supplying appropriate prompts to SAM at inference time. Page 5, Col. 1 section image encoder discloses the image encoder processes an image and produces an image embedding. Further in section Mask decoder, discloses the mask decoder maps the image embedding, prompt embedding and an output token to a segmentation mask. The mask decoder computes a foreground probability for each image location, resulting identifying the spatial region represented by the generated mask. Page 6, Col1, section fully automatic stage discloses prompting SAM with a regular 32x32 grid of points over the image, for each point predicted a set of masks that may correspond to valid objects selecting confident and stable masks from the predicted masks and filtering duplicate masks through non-maximum suppression. Applying fully automatic mask generation to all 11 million images in its dataset.)
As per claim 6, The method according to claim 4, wherein: Li as modified by Shim as modified by Kirillov further discloses “providing the combined training data set comprises comparing the first segmented areas of the training data set with the second segmented areas of the training data set to determine at least one matching area, and at least one characteristic of the first segmented areas is assigned to the second segmented areas based on the at least one determined matched range in order to combine the provided semantics and the provided segmentation of the training data set.”(Dai, ¶[0037-0038] discloses using a semantically labelled spatial region, namely a ground-truth bounding box annotation, to select a candidate segment mask that overlaps the labelled region to a desired degree. ¶[0039] discloses calculating intersection -over-union ration between the ground truth region and the spatial extent of the candidate segment. Also favoring candidate segments having both higher spatial overlap and a semantic label consistent with the labelled region. ¶[0040-0041] discloses using the resulting candidate segment and their estimated semantic labels as supervision for network training. ¶[0043] discloses selecting a candidate segment for the ground-truth, semantically labelled spatial region and assigning the semantic label associated with that region to the selected candidate segment.)
Claim(s) 2-3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al., "Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images," in IEEE Transactions on Geoscience and Remote Sensing, 2022, in view of Shim et al. (US 2023/0093619), further in view of Kirillov et al. "Segment anything." 2023 IEEE/CVF international conference on computer vision (ICCV). IEEE, Oct 1, 2023, further in view of Dai et al. (US 2017/0109625).
As per claim 2, The method according to claim 1, “wherein providing the image data further comprises: providing segments for a portion of the non-labelled image data using the zero-shot segmentation model,”(Kirillov, page 1, Col. 1 discloses the SAM model is designed and trained to be promotable, so it can transfer zero-shot to new image distributions and tasks. Page 4, Col1. Section 2, “zero-shot transfer”, downstream segmentation tasks are performed through appropriate prompts at inference time, including automatic dataset labeling. Further discloses the promotable segmentation task returns a valid segmentation mask in response to input prompt. Page 5, Col. 1, section 3, discloses SAM includes image encoder to process each image, prompt encoder, mask decoder produce mask from the image and prompt embedding, resolving ambiguity, wherein the mask decoder efficiently maps the image embedding, prompt embeddings, and an output token to a mask and computes the foreground probability at each image location. Mask decoder computes the foreground probability at each location of the image. Further page 5, Col. 2, section 4 discloses fully automatic stage which SAM generates masks without annotator input. Page 6, Col. 1, “fully automatic stage”, discloses prompting SAM with a regular 32x32 grid of points, for each point predicted a set of masks that may correspond to valid objects and applying fully automatic mask to all images in the dataset. )
However Li as modified by Shim as modified by Kirillov does not explicitly disclose the following which would have been obvious in view of Dai form similar filed of endeavor “and assigning labels to the provided segments to provide the labelled portion of the image data.” (Dai, ¶[0036] discloses A mask generator 224 generates several candidate segment masks 234 for the training image 232. The candidate segment masks 234 may be generated using various methods. ¶[0037] discloses objects in the candidate segment masks 234 can be assigned a label that can be, a semantic category label or background label. ¶[0040-0041] discloses using candidate segment masks and their estimated semantic labels, as pixel level supervision for a convolutional network. ¶[0043] discloses assigning ground truth semantic label associated with labeled bounding box to the selected candidate segment.)
Before the effective filing date of the claimed invention it would have been obvious to a person of ordinary skill in the art to combine Dai technique of network training for semantic segmentation into Li as modified by Shim as modified by Kirillov technique to provide the known and expected uses and benefits of Dai technique over learning model for semantic segmentation technique of Li as modified by Shim as modified by Kirillov. The proposed combination would have constituted a mere arrangement of old elements with each performing their known function, the combination yielding no more than one would expect from such an arrangement.
Therefore, it would have been obvious to a person of ordinary skill in the art to incorporate Dai to Li as modified by Shim as modified by Kirillov in order to increase the accuracy of object identification in an image. (Refer to Dai ¶[0002].)
As per claim 3, The method according to claim 2, wherein: “the labels are assigned using at least one prompt-based input of a user,”(Kirillov, page 4, Col1, section 2. discloses segmentation prompt maybe a foreground point, background points, a rough box or mask, free-form text, or, in general, any information indicating what to segment in an image. further discloses a rough bounding box may be supplied as the prompt for promotable segmentation task. Page 5, Col. 1, section 3, prompt encoder discloses its prompt encoder accepts sparse prompts that include points, boxes, and text. )
“and the labels are assigned to the provided segments by the at least one prompt-based input.” (Dai, ¶[0035] discloses the training images 230 are labeled with ground-truth bounding boxes of objects (e.g. “person,” “car,” “boat”). ¶[0039] discloses a ground-truth bounding box may be provided by a human. Using ground truth bounding box annotation to select a candidate segment mask that overlaps the bounding box. ¶[0043] discloses assigned the ground-truth semantic label associated with that bounding box to the selected candidate segment. )
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al., "Global and Local Contrastive Self-Supervised Learning for Semantic Segmentation of HR Remote Sensing Images," in IEEE Transactions on Geoscience and Remote Sensing, 2022, in view of Shim et al. (US 2023/0093619), further in view of Kirillov et al. "Segment anything." 2023 IEEE/CVF international conference on computer vision (ICCV). IEEE, Oct 1, 2023, further in view of Jeong et al. "Winclip: Zero-/few-shot anomaly classification and segmentation." 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, June 17, 2023.
As per claim 7, The method according to claim 1, wherein: “the machine learning model is trained based on the combined training data set for classification and/or detection based on image information, the image information is pixels of an image recording and/or represents at least one recorded object,.”(Dai, ¶[0040] discloses using candidate segment masks and their estimated semantic labels to supervise a deep convolutional network. P is a pixel index, l(p) is the ground-truth semantic label at a pixel, and Xθ(p) is the per-pixel labeling produced by the fully convolutional network and the network parameters θ can be updated by back-propagation and stochastic gradient descent (SGD).¶[0041] discloses the estimated semantic label used as supervision for the network training and that the estimated candidate segment is used as the regression target during training. )
However Li as modified by Shim as modified by Kirillov does not explicitly disclose the following which would have been obvious in view of Jeong from similar field of endeavor “the detection comprises a detection of a defective assembly in a production environment, and the detection is performed based on a semantic segmentation and/or a pixel-based classification”(Jeong, page 1, Col. 1 discloses localizing defects in industrial manufacturing and performs they operation by predicting whether a pixel is normal or anomalies identifies. industrial domains, including aerospace, automobile, pharmaceutical, and electronics. Page 1, Col. 2 discloses missing component on circuit board as a defect relative to normal circuit board having all component present. Page 3, Col. 2 section 3 discloses anomaly segmentation is the pixel level extension of anomaly detection and outputs the location of anomalies for an image of height h and Width w. page 4, Col. 1, section compositional prompt ensemble discloses missing transistor is anomalous in circuit board and identifies bad soldering on printed circuit board as a task specific defect. Further page 4, Col. 2, section 4.2, disclose Zero-shot anomaly segmentation that predicts pixel level anomalies and generates an anomaly segmentation map for an industrial inspection image. receiving an image x of resolution h × w and extracting dense visual features from the image to obtain an anomaly segmentation map. Page 5, Col 1, in Harmonic aggregation of windows discloses distributing an anomaly score to every pixel of local image window and aggregating the score at each pixel to improve segmentation.)
Before the effective filing date of the claimed invention it would have been obvious to a person of ordinary skill in the art to combine Jeong technique of anomaly classification and segmentation into Li as modified by Shim as modified by Kirillov technique to provide the known and expected uses and benefits of Jeong technique over learning model for semantic segmentation technique of Li as modified by Shim as modified by Kirillov. The proposed combination would have constituted a mere arrangement of old elements with each performing their known function, the combination yielding no more than one would expect from such an arrangement.
Therefore, it would have been obvious to a person of ordinary skill in the art to incorporate Jeong to Li as modified by Shim as modified by Kirillov in order to improve accuracy and performance of anomaly detection. (Refer to Jeong page1, Col. 2.)
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAGHAYEGH AZIMA whose telephone number is (571)272-1459. The examiner can normally be reached Monday-Friday, 9:30-6:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at (571)272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHAGHAYEGH AZIMA/Examiner, Art Unit 2671