Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This communication is in response to the Application Filed on 1/10/2025.
Claims 1-15 are pending in this application.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 1/10/2025, 8/4/2025, and 1/2/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Positive Statement regarding 35 U.S.C. 101: Claims 1-15 are determined to be eligible under 35 U.S.C. 101. Claim 1, for example, at lines 4-6 recites " generating a unified guidance map that indicates the at least one object to be segmented based on the one or more user inputs; generating a complex supervision image based on the unified guidance map." It is given the weight of the description in the specification paragraph [0057] that generating a unified guidance map is a result of encoding different types of input interactions into one map (See Fig. 3E) and the complex supervision image is produced through an analysis of edge, color, and geometry of the image object (See Fig. 4). Both the unified guidance map and complex supervision image are used to achieve better segmentation. Because the claims recite specific, claimed steps and structural elements that produce tangible technical results, they are found to be not directed to an abstract idea. Therefore, the claims are determined to be eligible under 35 U.S.C. 101.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claims 1-6, 9, 14, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Price et al. (US 2019/0236394 A1, hereinafter, “Price”) in view of Majumder et al. (“Content-Aware Multi-Level Guidance for Interactive Instance Segmentation”, 2019, hereinafter, “Majumder”).
Regarding claim 1, Price teaches a method for interactive image segmentation (See Price, Fig. 1; [0040] interactive deep learning approach; [0058) The multi-modal object selection neural network 108 can then generate an object segmentation output 110) by an electronic device (See Price, Fig. 8 computing device 801), the method comprising:
receiving one or more user inputs (See Price, Fig. 1, 102, 104, and 106; [0028] select an object within an image based on user input of various combinations of different input modalities) for segmenting at least one object from among a plurality of objects in an image (See Price, Fig. 1. Different inputs describing an object to be segmented are fed into 108 to then output an object segmentation; [0041] identify target objects in a digital image with only a handful of user interactions);
[generating a unified guidance map that indicates the at least one object to be segmented] based on the one or more user inputs (See Price, Fig. 1, 102, 104, and 106; [0028] select an object within an image based on user input of various combinations of different input modalities);
generating a complex supervision image based on the unified guidance map (See Price, [0036] the multi-modal selection system generates training image/user interaction pairs and trains a neural network utilizing the training image/user interaction pairs. Examiner considers the training image as a complex supervision image);
segmenting the at least one object from the image by inputting the image (See Price, [0029] digital image), the complex supervision image (See Price, [0036] the multi-modal selection system generates training image/user interaction pairs and trains a neural network utilizing the training image/user interaction pairs. Examiner considers the training image as a complex supervision image) [and the unified guidance map] into an adaptive Neural Network (NN) model (See Price, Fig. 5A; [0040] interactive deep learning approach);
and storing the at least one segmented object from the image (See Price, Fig. 8; [0173] the storage manager 814 may also include training image repository 816 and digital visual media 818. Examiner considers the object being segmented from an image to be stored just as the image itself is being stored. The method, 800, is being performed by the electronic device, 801, which also holds a storage location, 814).
However, Price does not teach generating a unified guidance map that indicates the at least one object to be segmented, and [segmenting the at least one object from the image by inputting the image, the complex supervision image and the unified guidance map into an adaptive Neural Network (NN) model.]
Majumder teaches generating a unified guidance map that indicates the at least one object to be segmented (See Majumder, Pg. 11597, caption of [Figure 3. Example of guidance maps], We transform the user-provided positive (shown as green dots) and negative (shown as red dots)
clicks into guidance maps. Examiner considers the guidance maps show in Fig. 3 to be generated using the user-inputs, or clicks, as a guide for segmenting the desired object), and [segmenting the at least one object from the image by inputting the image, the complex supervision image] and the unified guidance map (See Majumder, Pg. 11597, caption of [Figure 3. Example of guidance maps], We transform the user-provided positive (shown as green dots) and negative (shown as red dots)
clicks into guidance maps. Examiner considers the guidance maps show in Fig. 3 to be generated using the user-inputs, or clicks, as a guide for segmenting the desired object) [into an adaptive Neural Network (NN) model.]
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to generate a unified guidance map that indicates the at least one object to be segmented and segment the at least one object from the image by inputting the unified guidance map based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price to obtain the invention as specified in claim 1.
Regarding claim 2, in which claim 1 is incorporated, Price fails to explicitly disclose extracting input data based on the one or more user inputs;
obtaining one or more guidance maps corresponding to the one or more user inputs based on the input data; and
generating the unified guidance map by concatenating the one or more guidance maps..
However, Majumder teaches extracting input data based on the one or more user inputs (See Majumder, Pg. 11599, [Interaction Loop], a user provides positive and negative clicks sequentially to segment the object of interest);
obtaining one or more guidance maps corresponding to the one or more user inputs based on the input data (See Majumder, Pg. 11599, [Interaction Loop], After each click is added, the guidance maps are recomputed); and
generating the unified guidance map by concatenating the one or more guidance maps (See Majumder, Pg. 11599, [Interaction Loop], newly generated guidance map is concatenated with the image and given as input to the FCN 8s network which produces an updated segmentation map).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to extracting input data based on the one or more user inputs; obtaining one or more guidance maps corresponding to the one or more user inputs based on the input data; and generating the unified guidance map by concatenating the one or more guidance maps based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price to obtain the invention as specified in claim 2.
Regarding claim 3, in which claim 2 is incorporated, Price teaches based on the input data comprising one or more set of coordinates (See Price, [0033] a position corresponding to a user input (e.g., pixels indicated by a user input)),
obtaining traces of the one or more user inputs on the image using the input data (See Price, [0046] an indication as to how the one or more pixels correspond to a target object portrayed in the digital image), and
[encoding the traces into the one or more guidance maps, and wherein the traces represent user interaction locations on the image.]
However, Price does not teach encoding the traces into the one or more guidance maps and wherein the traces represent user interaction locations on the image.
Majumder teaches encoding the traces into the one or more guidance maps, and wherein the traces represent user interaction locations on the image (See Majumder, Pg. 11597, caption of [Figure 3. Example of guidance maps], We transform the user-provided positive (shown as green dots) and negative (shown as red dots) clicks into guidance maps for the instance segmentation network. Examiner considers the clicks as one type of format that is being encoded, or converted, into a guidance map which is a different type of format).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to encoding the traces into the one or more guidance maps, and wherein the traces represent user interaction locations on the image based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price to obtain the invention as specified in claim 3.
Regarding claim 4, in which claim 2 is incorporated, Price teaches based on the input data comprising text indicating the at least one object in the image (See Price, Fig. 1 - 106; [0057] a language user input 106a),
determining a segmentation mask based on a category of the text using an instance model (See Price, Fig. 1; [0058] The multi-modal object selection neural network 108 can then generate an object segmentation output 110. As illustrated the object segmentation output 110 can include a segmentation mask 110a and/or a segmentation boundary 110b. Examiner considers the neural network to be an instance model. Examiner considers the category of text to be “dog” as seen in Fig. 1, which is in line with Applicant Spec [0065], based on a category (e.g. dogs, cars, food, etc.) of the text); and [converting the segmentation mask into the one or more guidance maps.]
However, Price does not teach converting the segmentation mask into the one or more guidance maps.
Majumder teaches converting the segmentation mask into the one or more guidance maps (See Majumder, Pg. 11594, caption of [Figure 1.] our proposed technique exploits image structures such as superpixels and object proposals, allowing us to generate more informative guidance maps. Examiner considers the segmentation mask, which is at pixel-level, to be a type of image structure, such as superpixels, which are converted, or generated, into guidance maps).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to converting the segmentation mask into the one or more guidance maps based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price to obtain the invention as specified in claim 4.
Regarding claim 5, in which claim 4 is incorporated, Price teaches based on the input data comprising an audio (See Price, [0165] a specific voice command),
converting the audio into the text (See Price, [0057] language user input 106a comprises a word (i.e., “dog”); [0058] the language user input 106a corresponding to the verbal input modality 106. Examiner considers the verbal input to be converted into a word for use as a language input); and
determining the segmentation mask based on the category of the text using the instance model (See Price, Fig. 1; [0058] The multi-modal object selection neural network 108 can then generate an object segmentation output 110. As illustrated the object segmentation output 110 can include a segmentation mask 110a and/or a segmentation boundary 110b. Examiner considers the neural network to be an instance model. Examiner considers the category of text to be “dog” as seen in Fig. 1, which is in line with Applicant Spec [0065], based on a category (e.g. dogs, cars, food, etc.) of the text).
Regarding claim 6, in which claim 1 is incorporated, Price teaches determining a plurality of complexity parameters comprising at least one of a color complexity, an edge complexity or a geometry map of the at least one object to be segmented (See Price, Fig. 2A; [0048] a distance map can include a database or digital file that includes distances between pixels in a digital image and pixels indicated by user input. Examiner considers the distance map as a geometry map);
and generating the complex supervision image (See Price, [0036] training image/user interaction pairs. Examiner considers the training image as the complex supervision image) by concatenating [a weighted low frequency image] obtained using the color complexity (See Price, [0069] each color channel 218-222 comprises a two-dimensional matrix with entries for each pixel in the digital image 200. Examiner considers the color indication of the image in the color channels as color complexity) and [the unified guidance map], [a weighted high frequency image] obtained using the edge complexity (See Price, [0089] identify an object boundary for the target object (e.g., by identifying pixels along the edge of the target object). Examiner considers the identification of the pixels of the object's edge as the edge complexity) and [the unified guidance map], and the geometry map (See Price, Fig. 2A; [0048] a distance map can include a database or digital file that includes distances between pixels in a digital image and pixels indicated by user input).
However, Price does not teach a weighted low frequency image [obtained using the color complexity] and the unified guidance map, and a weighted high frequency image [obtained using the edge complexity] and the unified guidance map.
Majumder teaches a weighted low frequency image (See Majumder, Pg. 11596, [Figure 2 Outline]. Examiner considers, as seen in Figure 2, the input image, which can be a low frequency image, and the guidance maps concatenated, combined, or convolved together) [obtained using the color complexity] and the unified guidance map (See Majumder, Pg. 11597, caption of [Figure 3. Example of guidance maps], We transform the user-provided positive (shown as green dots) and negative (shown as red dots) clicks into guidance maps. Examiner considers the guidance maps show in Fig. 3 to be generated using the user-inputs, or clicks, as a guide for segmenting the desired object), and a weighted high frequency image (See Majumder, Pg. 11596, [Figure 2 Outline]. Examiner considers, as seen in Figure 2, the input image, which can be a high frequency image, and the guidance maps concatenated, combined, or convolved together) [obtained using the edge complexity] and the unified guidance map (See Majumder, Pg. 11597, caption of [Figure 3. Example of guidance maps], We transform the user-provided positive (shown as green dots) and negative (shown as red dots) clicks into guidance maps. Examiner considers the guidance maps show in Fig. 3 to be generated using the user-inputs, or clicks, as a guide for segmenting the desired object).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference wherein a weighted low frequency image [obtained using the color complexity] and the unified guidance map, and a weighted high frequency image [obtained using the edge complexity] and the unified guidance map based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price to obtain the invention as specified in claim 6.
Regarding claim 9, in which claim 6 is incorporated, Price teaches identifying a color at a location on the image (See Price, [0034] color channels (i.e., data sets reflecting the color of pixels in digital images). Color channels are used to identify colors);
tracing the color within a reference range of color at the location (See Price, [0065] the multi-modal selection system calculates a geodesic distance that utilizes curved paths to stay inside (e.g., traverse along) a particular color (or range of colors));
obtaining the geometry map (See Price, Fig. 2A; [0048] a distance map can include a database or digital file that includes distances between pixels in a digital image and pixels indicated by user input. Examiner considers the distance map as a geometry map) comprising a union of the traced color with an edge map of the at least one object (See Price, Fig. 2A. Examiner considers the combination of distance maps, color channels, and additional maps, like an edge map, to be the union as seen in Fig. 2A); and
estimating a span of the at least one object by determining a size of a bounding box of the at least one object in the geometry map (See Price, Fig. 2A; [0079] the multi-modal selection system transforms a bounding shape (e.g., a bounding box) to a signed Euclidean distance transform map as an additional channel (e.g., the additional map(s) 217)); [0081] the multi-modal selection system can limit the object segmentation output to within the bounding shape),
and wherein the span corresponds to a larger side of the bounding box in a rectangle shape (See Price, [0081] analyzes the digital image within the bounding shape for an object segmentation output. This approach zooms into a tight region surrounding the target object. The span is shown to cover a region enclosing the target object, which is seen as well by zooming into a tight region surrounding the target).
Regarding claim 14, Price teaches a method for encoding different types of user interactions (See Price, Fig. 1, [0054] the multi-modal object selection neural network 108 can consider a variety of different user inputs corresponding to different input modalities. Examiner considers the user interactions to be encoded, or converted, with the neural network) by an electronic device (See Price, Fig. 8 computing device 801), the method comprising:
detecting a plurality of user inputs performed on an image (See Price, [0028] user input of various combinations of different input modalities (e.g., regional clicks, boundary clicks, language input, bounding boxes, attention masks, and/or soft clicks));
[obtaining a plurality of guidance maps by converting each of the plurality of user inputs to one of the plurality of guidance maps] based on a type of the respective user input (See Price, [0028] based on user input of various combinations of different input modalities (e.g., regional clicks, boundary clicks, language input, bounding boxes, attention masks, and/or soft clicks));
[unifying the plurality of guidance maps to generate a unified guidance map representing a unified feature space;]
determining an object complexity (See Price, Fig. 2A. Examiner considers the distance map, color channels, and additional maps to cover the complexity of the object) [based on the unified guidance map] and the image (See Price, [0029] digital image);
and inputting the object complexity (See Price, Fig. 2A. Examiner considers the distance map, color channels, and additional maps to cover the complexity of the object) and the image ([0029] digital image) to an interactive segmentation engine (See Price, Fig. 5A; [0040] interactive deep learning approach; [0142] multi-modal selection system utilizes the neural network 506 to analyze the image/user interaction pairs 504a-504n and generate a predicted object segmentation output corresponding to the target objects).
However, Price does not teach obtaining a plurality of guidance maps by converting each of the plurality of user inputs to one of the plurality of guidance maps [based on a type of the respective user input]; unifying the plurality of guidance maps to generate a unified guidance map representing a unified feature space.
Majumder teaches obtaining a plurality of guidance maps (See Majumder, Pg. 11596, [Proposed Approach], par. 2 on the left col., generation of multiple guidance maps) by converting each of the plurality of user inputs to one of the plurality of guidance maps (See Majumder, Pg. 11594, [Introduction], par. 2 on the right col., The guidance map helps to select the specific instance to segment; Pg. 11595, [Introduction] first bullet point on the left col., user-provided clicks which generates guidance maps. Examiner considers the user input of clicks to be used to generate guidance maps. Any other type of user input can also generate a guidance map to aid in segmentation) [based on a type of the respective user input]; unifying the plurality of guidance maps to generate a unified guidance map representing a unified feature space (See Majumder, Pg. 11594, [Introduction], par. 2 on the right col., The guidance map helps to select the specific instance to segment; Pg. 11595, [Introduction] first bullet point on the left col., user-provided clicks which generates guidance maps. Examiner considers the user input of clicks to be used to generate guidance maps. Any other type of user input can also generate a collective guidance space to aid in segmentation).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to obtaining a plurality of guidance maps by converting each of the plurality of user inputs to one of the plurality of guidance maps [based on a type of the respective user input]; unifying the plurality of guidance maps to generate a unified guidance map representing a unified feature space based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price to obtain the invention as specified in claim 14.
Regarding claim 15, in which claim 14 is incorporated, Price teaches wherein the type of the user inputs is at least one of a touch, a contour, a scribble, a stroke, text, an audio, an eye gaze, or an air gesture (See Price, Fig. 6A; [0083] [0083] The multi-modal selection system can also accommodate soft clicks/scribbles as an input modality).
Claims 7 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Price et al. (US 2019/0236394 A1, hereinafter, “Price”) in view of Majumder et al. (“Content-Aware Multi-Level Guidance for Interactive Instance Segmentation”, 2019, hereinafter, “Majumder”) further in view of Roy et al. (“A Review on Automated Brain Tumor Detection and Segmentation from MRI of Brain”, 2022, hereinafter, “Roy”).
Regarding claim 7, in which claim 6 is incorporated, Price teaches wherein the determining the color complexity (See Price, [0069] each color channel 218-222 comprises a two-dimensional matrix with entries for each pixel in the digital image 200. Examiner considers the color indication of the image in the color channel as color complexity) of the at least one object comprises:
[obtaining a low frequency image by inputting the image into a low pass filter;]
[determining a weighted map by normalizing the unified guidance map;]
[determining a weighted low frequency image by convolving the low frequency image with the weighted map;]
determining a standard deviation (See Price, [0143] a loss function to determine a measure of loss (or error). Examiner considers the loss as the standard deviation) [of the weighted low frequency image];
determining whether the standard deviation (See Price, [0143] a loss function to determine a measure of loss (or error). Examiner considers the loss as the standard deviation) [of the weighted low frequency image] is greater than a first threshold (See Price, [0065] identifies a distance between two pixels by staying within a particular color (or range of colors) and avoiding (i.e., going around) colors that are outside the particular color (or range of colors). Examiner considers the range of colors as the first threshold and colors outside the range are greater than the threshold);
and performing one of: detecting that the color complexity is high, based on the standard deviation [of the weighted low frequency image] being greater than the first threshold, and detecting that the color complexity is low, based on the standard deviation of [the weighted low frequency image] being less than or equal to the first threshold (See Price; [0065] identifies a distance between two pixels by staying within a particular color (or range of colors) and avoiding (i.e., going around) colors that are outside the particular color (or range of colors); [0143] a loss function to determine a measure of loss (or error). Examiner considers the loss as the standard deviation. Examiner considers the range of colors as the first threshold and colors outside the range are greater than the threshold).
However, Price does not teach obtaining a low frequency image by inputting the image into a low pass filter;
determining a weighted map by normalizing the unified guidance map;
determining a weighted low frequency image by convolving the low frequency image with the weighted map; and [detecting that the color complexity is high, based on the standard deviation] of the weighted low frequency image [being greater than the first threshold.]
Majumder teaches [obtaining a low frequency image by inputting the image into a low pass filter;]
determining a weighted map by normalizing the unified guidance map (See Majumder, Pg. 11596, [3.1. Superpixel-base guidance map] the guidance maps values are scaled between [0,255]);
determining a weighted low frequency image by convolving the low frequency image with the weighted map (See Majumder, Pg. 11599, [Interaction Loop], generated guidance map is concatenated with the image. Examiner considers the input image, which can be a low frequency image, to be combined, or convolved, as shown in Figure 2);
[determining a standard deviation] of the weighted low frequency image (See Majumder, Pg. 11599, [Interaction Loop], generated guidance map is concatenated with the image. Examiner considers the input image, which can be a low frequency image, to be combined, or convolved, as show in Figure 2); and [detecting that the color complexity is high, based on the standard deviation] of the weighted low frequency image (See Majumder, Pg. 11599, [Interaction Loop], generated guidance map is concatenated with the image. Examiner considers the input image, which can be a low frequency image, to be combined, or convolved, as shown in Figure 2) [being greater than the first threshold.]
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference wherein determining a weighted map by normalizing the unified guidance map; determining a weighted low frequency image by convolving the low frequency image with the weighted map; and [detecting that the color complexity is high, based on the standard deviation] of the weighted low frequency image [being greater than the first threshold] based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
However, Price and Majumder do not teach obtaining a low frequency image by inputting the image into a low pass filter.
Roy teaches obtaining a low frequency image by inputting the image into a low pass filter (See Roy, Pg. 5, [5.2.3 Low pass Filter:], Low-pass filtering zeroes out the frequency components above intensity, if f(x,y) is a function, for example the brightness in an image, its Fourier transform is given by threshold. It suppresses all frequencies higher than the cut-off frequency and leaves smaller frequencies unchanged).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s or Majumder’s reference to obtaining a low frequency image by inputting the image into a low pass filter based on the method of Roy’s reference. The suggestion/motivation would have been to low-pass filter zeroes out the frequency components above intensity, if f(x,y) is a function, for example the brightness in an image and leaver smaller frequencies unchanged as suggested by Roy at Pg. 5. [5.2.3 Low pass Filter].
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Roy with Price and Majumder to obtain the invention as specified in claim 7.
Regarding claim 8, in which claim 6 is incorporated, Price teaches wherein the determining the edge complexity (See Price, [0089] identify an object boundary for the target object (e.g., by identifying pixels along the edge of the target object). Examiner considers the identification of the pixels of the object's edge as the edge complexity) of the at least one object comprises:
[obtaining a high frequency image by inputting the image into a high pass filter;]
[determining a weighted map by normalizing the unified guidance map;]
[determining a weighted high frequency image by convolving the high frequency image with the weighted map;]
determining a standard deviation (See Price, [0143] a loss function to determine a measure of loss (or error). Examiner considers the loss as the standard deviation) [of the weighted high frequency image] for analyzing the edge complexity (See Price, [0089] identify an object boundary for the target object (e.g., by identifying pixels along the edge of the target object). Examiner considers the identification of the pixels of the object's edge as the edge complexity);
determining whether the standard deviation (See Price, [0143] a loss function to determine a measure of loss (or error). Examiner considers the loss as the standard deviation) [of the weighted high frequency image] is greater than a second threshold (See Price, [0130] threshold distance);
and performing one of:
detecting that the edge complexity is high, based on the standard deviation [of the weighted high frequency image] being greater than the second threshold, and detecting that the edge complexity is low, based on the standard deviation (See Price, [0143] a loss function to determine a measure of loss (or error). Examiner considers the loss as the standard deviation) of [the weighted high frequency image] being less than or equal to the second threshold (See Price, [0130] threshold distance)
However, Price does not teach obtaining a high frequency image by inputting the image into a low pass filter;
determining a weighted map by normalizing the unified guidance map;
determining a weighted high frequency image by convolving the high frequency image with the weighted map; and [detecting that the edge complexity is high, based on the standard deviation] of the weighted high frequency image [being greater than the second threshold.]
Majumder teaches [obtaining a high frequency image by inputting the image into a high pass filter;]
determining a weighted map by normalizing the unified guidance map (See Majumder, Pg. 11596, [3.1. Superpixel-base guidance map] the guidance maps values are scaled between [0,255]);
determining a weighted high frequency image by convolving the high frequency image with the weighted map (See Majumder, Pg. 11599, [Interaction Loop], generated guidance map is concatenated with the image. Examiner considers the input image, which can be a high frequency image, to be combined, or convolved, as shown in Figure 2);
[determining a standard deviation] of the weighted high frequency image (See Majumder, Pg. 11599, [Interaction Loop], generated guidance map is concatenated with the image. Examiner considers the input image, which can be a high frequency image, to be combined, or convolved, as show in Figure 2); and [detecting that the edge complexity is high, based on the standard deviation] of the weighted high frequency image (See Majumder, Pg. 11599, [Interaction Loop], generated guidance map is concatenated with the image. Examiner considers the input image, which can be a high frequency image, to be combined, or convolved, as shown in Figure 2) [being greater than the second threshold.]
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference wherein determining a weighted map by normalizing the unified guidance map; determining a weighted high frequency image by convolving the high frequency image with the weighted map; and [detecting that the edge complexity is high, based on the standard deviation] of the weighted high frequency image [being greater than the second threshold] based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
However, Price and Majumder do not teach obtaining a high frequency image by inputting the image into a high pass filter.
Roy teaches obtaining a low frequency image by inputting the image into a low pass filter (See Roy, Pg. 5, [5.2.4 High-pass Filter [24, 25, 26]:, working activities of this filter is some other than it passes high frequency and remain unchanged and block low frequency signal, the elements of the mask contain both positive and negative weights).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s or Majumder’s reference to obtaining a high frequency image by inputting the image into a high pass filter based on the method of Roy’s reference. The suggestion/motivation would have been for a high-pass filter to take care of undesirable small noise intensity and useful for emphasizing transitions in intensity (e.g., edges) as suggested by Roy at Pg. 5. [5.2.34 High-pass Filter].
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Roy with Price and Majumder to obtain the invention as specified in claim 8.
Claims 10-13 are rejected under 35 U.S.C. 103 as being unpatentable over Price et al. (US 2019/0236394 A1, hereinafter, “Price”) in view of Majumder et al. (“Content-Aware Multi-Level Guidance for Interactive Instance Segmentation”, 2019, hereinafter, “Majumder”) further in view of Ghosh et al. (“Understanding deep learning techniques for image segmentation”, 2019, hereinafter, “Ghosh”).
Regarding claim 10, in which claim 1 is incorporated, Price teaches determining optimal scales for the adaptive NN model based on a relationship between a receptive field of the adaptive NN model and a span of the at least one object (See Price, Fig. 2A and 5A; [0080] the multi-modal selection system feeds the distance map corresponding to the bounding box to the neural network for consideration with the other user inputs corresponding to the other input modalities; [0080] The multi-modal selection system can incorporate a variety of bounding shapes (including shapes with curve-based input). Bounding shapes are fed to the neural network which works with other maps to train a neural network);
[determining an optimal number of layers for the adaptive NN model] based on a color complexity (See Price, [0069] each color channel 218-222 comprises a two-dimensional matrix with entries for each pixel in the digital image 200. Examiner considers the color indication of the image in the color channel as color complexity);
[determining an optimal number of channels for the adaptive NN model] based on an edge complexity (See Price, [0089] identify an object boundary for the target object (e.g., by identifying pixels along the edge of the target object). Examiner considers the identification of the pixels of the object's edge as the edge complexity)
configuring the adaptive NN model based on the optimal scales, [the optimal number of layers, and the optimal number of channels];
and segmenting the at least one object from the image by inputting the image (See Price, [0029] digital image), the complex supervision image (See Price, [0036] the multi-modal selection system generates training image/user interaction pairs and trains a neural network utilizing the training image/user interaction pairs), [and the unified guidance map] through the configured adaptive NN model (See Price, Fig. 5A; [0040] interactive deep learning approach).
However, Price does not teach determining an optimal number of layers for the adaptive NN model; determining an optimal number of channels for the adaptive NN model; and configuring the adaptive NN model based on [the optimal scales], the optimal number of layers, and the optimal number of channels; and [segmenting the at least one object from the image by inputting the image, the complex supervision image,] and the unified guidance map [through the configured adaptive NN model.]
Majumder teaches segmenting the at least one object from the image by inputting the image, the complex supervision image,] and the unified guidance map (See Majumder, Pg. 11597, caption of [Figure 3. Example of guidance maps], We transform the user-provided positive (shown as green dots) and negative (shown as red dots) clicks into guidance maps. Examiner considers the guidance maps show in Fig. 3 to be generated using the user-inputs, or clicks, as a guide for segmenting the desired object) [through the configured adaptive NN model.]
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to segmenting the at least one object from the image by inputting the image, the complex supervision image,] and the unified guidance map [through the configured adaptive NN model based on the method of Majumder’s reference. The suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
However, Price and Majumder do not teach determining an optimal number of layers for the adaptive NN model; determining an optimal number of channels for the adaptive NN model; and configuring the adaptive NN model based on [the optimal scales], the optimal number of layers, and the optimal number of channels.
Ghosh teaches determining an optimal number of layers for the adaptive NN model (See Ghosh, Pg. 4, Figure 2. Legends for subsequent diagrams of popular deep learning architectures. Examiner considers the diagram to show multiple layers that can be used in a neural network); determining an optimal number of channels for the adaptive NN model (See Ghosh, Pg. 15, [CRF as RNN] required number of channels); and configuring the adaptive NN model based on [the optimal scales], the optimal number of layers (See Ghosh, Pg. 4, Figure 2. Legends for subsequent diagrams of popular deep learning architectures. Examiner considers the diagram to show multiple layers that can be used in a neural network), and the optimal number of channels (See Ghosh, Pg. 15, [CRF as RNN] required number of channels).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s or Majumder’s reference wherein determining an optimal number of layers for the adaptive NN model; determining an optimal number of channels for the adaptive NN model; and configuring the adaptive NN model based on [the optimal scales], the optimal number of layers, and the optimal number of channels based on the method of Ghosh’s reference. The suggestion/motivation would have been for the development of deep learning algorithms which are also efficient in other related tasks like object detection, localization, tracking, or as in this case image segmentation as suggested by Ghosh at Pg. 5, [3 Impact of Deep Learning on Image Segmentation].
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Ghosh with Price and Majumder to obtain the invention as specified in claim 10.
Regarding claim 11, in which claim 10 is incorporated, Price teaches [downscaling the image by a factor of two until the span of matches to the receptive field]; and
determining the optimal scales for the adaptive NN model based on a number of times the image has been downscaled to match the span with the receptive field (See Price, Fig. 2A and 5A; [0080] the multi-modal selection system feeds the distance map corresponding to the bounding box to the neural network for consideration with the other user inputs corresponding to the other input modalities; [0080] The multi-modal selection system can incorporate a variety of bounding shapes (including shapes with curve-based input). Bounding shapes are fed to the neural network which works with other maps to train a neural network).
However, Price does not teach downscaling the image by a factor of two until the span of matches to the receptive field.
Majumder teaches downscaling the image by a factor of two until the span of matches to the receptive field (See Majumder, [3.3. Scale-aware guidance] truncating distances exceeding some factor f of our scale measures, i.e. Equation 4. Examiner consider the factor to vary and can be chosen by person having ordinary skill in the art).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to downscaling the image by a factor of two until the span of matches to the receptive field based on the method of Majumder’s reference. suggestion/motivation would have been to provide the network with necessary cues on the whereabouts of the object of interest as suggested by Majumder in the Abstract.
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price and Ghosh to obtain the invention as specified in claim 11.
Regarding claim 12, in which claim 10 is incorporated, Price teaches performing one of:
[selecting a default number of layers as the optimal number of layers] based on detecting a first color complexity in the image (See Price, [0069] each color channel 218-222 comprises a two-dimensional matrix with entries for each pixel in the digital image 200. Examiner considers the color indication of the image in the color channel as color complexity. The first color complexity is 0 as seen in Fig. 2A, 224), and
[adding a reference layer offset value with the default number of layers for obtaining the optimal number of layers] based on detecting a second color complexity (See Price, [0069] each color channel 218-222 comprises a two-dimensional matrix with entries for each pixel in the digital image 200. Examiner considers the color indication of the image in the color channel as color complexity. The second color complexity is 1 as seen in Fig. 2A, 224), and wherein the first color complexity is lower than the second color complexity (See Price, [0069] each color channel 218-222 comprises a two-dimensional matrix with entries for each pixel in the digital image 200. Examiner considers the color indication of the image in the color channel as color complexity. The first color complexity is 0 as seen in Fig. 2A, 224 and is lower than the second color complexity of 1).
However, Price does not teach selecting a default number of layers as the optimal number of layers and adding a reference layer offset value with the default number of layers for obtaining the optimal number of layers.
Majumder teaches performing one of selecting a default number of layers as the optimal number of layers and adding a reference layer offset value with the default number of layers for obtaining the optimal number of layers (See Majumder, g. 20, [U-NET] one hundred layers. Examiner considers this to be a default number of layers).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to performing one of selecting a default number of layers as the optimal number of layers and adding a reference layer offset value with the default number of layers for obtaining the optimal number of layers based on the method of Majumder’s reference. The suggestion/motivation would have been for the development of deep learning algorithms which are also efficient in other related tasks like object detection, localization, tracking, or as in this case image segmentation as suggested by Majumder in Pg. 5, [3 Impact of Deep Learning on Image Segmentation].
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price and Ghosh to obtain the invention as specified in claim 12.
Regarding claim 13, in which claim 10 is incorporated, Price teaches [performing one of:]
[selecting a default number of channels as the optimal number of channels] based on detecting a first edge complexity [See Price, [0056] boundary input modality 104 (e.g., a first edge click and a second edge click). The boundary user inputs 104a-104b. Examiner considers the first edge click as the second edge complexity and seen in Fig. 1), and
[adding a reference channel offset value with the default number of channels for obtaining the optimal number of channels] based on detecting a second edge complexity (See Price, [0056] boundary input modality 104 (e.g., a first edge click and a second edge click). The boundary user inputs 104a-104b. Examiner considers the second edge click as the first edge complexity as seen in Fig. 1), and wherein the first edge complexity is lower than the second edge complexity (See Price, [0056] boundary input modality 104 (e.g., a first edge click and a second edge click). The boundary user inputs 104a-104b. Examiner considers the first edge click to be lower than the second edge click as seen in Fig. 1).
However, Price does not teach performing one of: selecting a default number of channels as the optimal number of channels based on detecting a first edge complexity and adding a reference channel offset value with the default number of channels for obtaining the optimal number of channels.
Majumder teaches performing one of: selecting a default number of channels as the optimal number of channels based on detecting a first edge complexity and adding a reference channel offset value with the default number of channels for obtaining the optimal number of channels (See Majumder, Pg. 27, [4.6.2 Deep Extreme Cut], a 4 channel input. Examiner considers this to be a default number of channels).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Price’s reference to performing one of: selecting a default number of channels as the optimal number of channels based on detecting a first edge complexity and adding a reference channel offset value with the default number of channels for obtaining the optimal number of channels based on the method of Majumder’s reference. The suggestion/motivation would have been for the development of deep learning algorithms which are also efficient in other related tasks like object detection, localization, tracking, or as in this case image segmentation as suggested by Majumder in Pg. 5, [3 Impact of Deep Learning on Image Segmentation].
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Majumder with Price and Ghosh to obtain the invention as specified in claim 13.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Athanasiadis et al. (“Semantic image segmentation and object labeling”, 2007) discloses a framework for simultaneous image segmentation and object labeling. With the use of contextual knowledge, labeling results are re-adjusted based on detected concepts. Knowledge helps improve individual region recognition of objects.
Wang et al. (US 2022/0327711 A1) discloses an image segmentation method and an electronic device. A first segmentation result is obtained which represents the probability of each pixel in the target image belonging to a category before its correction. A second segmentation result is obtained by correcting the first segmentation output using the category of correction and at least one correctio point. A convolutional neural network is used to obtain probability images of the target object and its category.
Xu et al. (US 10,346,986 B2) discloses a method for segmenting 3D images. The 3D image is used to create multiple stacks of 2D images with corresponding multiple planes. The stacks are segmented using a neural network model. This approach aims to improve the accuracy of automatic image segmentation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jasmin Marcelino Hernandez whose telephone number is (571) 270-0211. The examiner can normally be reached 7am-3pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JASMIN MARCELINO HERNAND/Examiner, Art Unit 2676
/SHEFALI D GORADIA/Primary Patent Examiner, Art Unit 2676