DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4, 20-21, 24 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (US 20200380696 A1), hereinafter Feng, in view of Zhang et al.: "Brain Tumor Segmentation From Multi-Modal MR Images via Ensembling UNets", Frontiers.org, Published 20 October 2021 [Retrieved on 08/16/2026]. Retrieved from the internet <https://www.frontiersin.org/journals/radiology/articles/10.3389/fradi.2021.704888/full>, hereinafter Zhang.
Regarding claim 1, Feng teaches A method for detecting and contouring structures of interest in a 3D image, the method comprising: (Abstract see "The present disclosure relates to a method and apparatus for automatic muscle segmentation. The method includes: receiving a high-resolution three-dimensional (3D) image obtained by a magnetic resonance imaging (MRI) system; splitting the high-resolution 3D image into high-resolution 3D sub-images; acquiring low-resolution 3D sub-images of the high-resolution 3D sub-images; determining a boundary box within each corresponding low-resolution 3D sub-image; determining a crop box for each boundary box; cropping each high-resolution 3D sub-image based on its corresponding crop box; and segmenting at least one muscle from the cropped box high-resolution 3D sub-image." Para. 6 see "The apparatus may include one or more processors, a display, and a non-transitory computer-readable memory storing instructions executable by the one or more processors." Para. 55 see "cascaded three dimensional (3D) deep convolutional neural network (DCNN) segmentation framework for fully automatic segmentation" Para. 90 see "Segmentation is done by individual networks which were trained for each target muscle to obtain more accurate contours from images cropped based on the output of the first stage" Para. 94 see "A two-stage process is disclosed that captures location and detailed features respectively"). detecting structures of interest from an input comprising at least one 3D image (Abstract see "receiving a high-resolution three-dimensional (3D) image obtained by a magnetic resonance imaging (MRI) system" Claim 3 see "using a trained detection-based neural network, wherein the neural network is trained to detect muscles in the low-resolution 3D sub-image and generate the boundary box based on pixel-wise prediction maps."). and generating a corresponding segmentation map of the detected structures using a first neural network; (Para. 83 see "a modified 3D-UNet is used to generate the bounding box from the low-resolution images. ... Therefore, a modified 3D-UNet was built to segment the low-resolution images and generate the bounding boxes based on the pixel-wise prediction maps." Para. 84 see "The modified 3D-UNet network follows the structure of 3D-UNet which consisted of an encoder and a decoder, each with four resolution levels. ... The final block of the network contains a 1x1x1 convolution layer to reduce the dimension of the features to match the number of label maps, followed by a pixel-wise soft-max classifier." Para. 97 see "FIGS. 6A and 6B shows the results of bounding box extraction and Segmentation." Para. 104 see "FIG. 11A is an output image from the first stage. Due to the loss of resolution, the contours were jagged. However, from this contour, the location of the ROI can be accurately identified and the images can be cropped for a more accurate segmentation in stage two." Claim 3 see "using a trained detection-based neural network, wherein the neural network is trained to detect muscles in the low-resolution 3D sub-image and generate the boundary box based on pixel-wise prediction maps." Examiner Note: The segmentation map (pixel-wise prediction map) is used to determine the bounding box and cropping region with the first neural network in the first stage as seen in Fig. 6.). extracting a plurality of cropped images from the at least one 3D image, (Abstract see "determining a crop box for each boundary box; cropping each high-resolution 3D sub-image based on its corresponding crop box;" Para. 5 see "wherein the crop box is the boundary box increased by a ratio along each dimension, cropping each high-resolution 3D sub-image based on its corresponding crop box" Para. 87 see "In step 320, each high-resolution 3D sub-image based on its corresponding crop box is cropped ... The high-resolution 3D sub-image is localized and cropped to match the crop box of the low-resolution 3D sub-image." Para. 90 see "a series of images were cropped at varied bounding boxes based on the first stage output and fed into the network."). each cropped image corresponding to a subregion of the at least one 3D image containing at least one of the detected structures contained in the segmentation map; (Para. 5 see "determining a boundary box within each corresponding low-resolution 3D sub-image, wherein the boundary box encloses at least one target muscle in the low-resolution 3D image, determining a crop box for each boundary box, wherein the crop box is the boundary box increased by a ratio along each dimension, cropping each high-resolution 3D sub-image based on its corresponding crop box" Para. 87 see "The high-resolution 3D sub-image is localized and cropped to match the crop box of the low-resolution 3D sub-image." Para. 90 see "Segmentation is done by individual networks which were trained for each target muscle to obtain more accurate contours from images cropped based on the output of the first stage ... a series of images were cropped at varied bounding boxes based on the first stage output and fed into the network." Para. 104 see "FIG. 11A is an output image from the first stage. Due to the loss of resolution, the contours were jagged. However, from this contour, the location of the ROI can be accurately identified and the images can be cropped for a more accurate segmentation in stage two." Claim 8 see "segmenting at least one small muscle based on the segmentation of at least one neighboring large muscle, wherein the small and large muscles are target muscles enclosed in the boundary box."). and estimating contours of the detected structures in the plurality of cropped images (Para. 90 see "Segmentation is done by individual networks which were trained for each target muscle to obtain more accurate contours from images cropped based on the output of the first stage, whose resolution is close to the originally acquired images. During deployment, an averaging method that includes augmentation was also used to improve the segmentation accuracy and contour smoothness. Specifically, the method included steps where a series of images were cropped at varied bounding boxes based on the first stage output and fed into the network." Para. 114 see "The robustness against the first stage error is improved and the contours are smoothed due to multiple averages." Claim 7 see "an averaging method that improves the segmentation accuracy and contour smoothness comprising: cropping a series of high-resolution 3D sub-image based on their corresponding crop box; segmenting at least one muscle from the series of cropped box high-resolution 3D sub-images"). and generating corresponding shape representations of the estimated contours using a second neural network. (Para. 89 see "In step 322, one or more of the GPU-executable programs is executed on the GPU in order to segment at least one muscle from the cropped box high-resolution 3D sub-image" Para. 90 see "Segmentation is done by individual networks which were trained for each target muscle to obtain more accurate contours from images cropped based on the output of the first stage ... an averaging method that includes augmentation was also used to improve the segmentation accuracy and contour smoothness. Specifically, the method included steps where a series of images were cropped at varied bounding boxes based on the first stage output and fed into the network. ... The outputs from each input were then averaged after putting back to the uncropped original images as the final label map" Para. 92 see "In step 324, the segmented high-resolution 3D image is displayed." Para. 97 see "FIGS. 6A and 6B shows the results of bounding box extraction and Segmentation." Claim 6 see "segmenting at least one muscle from the cropped box high-resolution 3D sub-image comprises: using a trained convolutional neural network, wherein the convolutional neural network is trained to segment individual muscles in the high-resolution 3D sub-image." Claim 7 see "averaging the segmented series of cropped box high-resolution 3D sub-images after replacing the un-segmented high-resolution 3D sub-images as a final label map" Examiner Note: A second neural network in the second stage performs the final segmentation or contour as seen in Fig. 6B where the shape representations (segmentation/contour) are shown on the display.).
While Feng teaches detecting structures of interest from an input comprising at least one 3D image, Feng does not teach detecting structures of interest from a multi-channel input comprising at least one 3D image.
However, Zhang teaches detecting structures of interest from a multi-channel input comprising at least one 3D image (Pg. 1, Para. 1 see "Similar to most existing works, the first UNet uses 3D patches of multi-modal MR images as the input. The second UNet uses brain parcellation as an additional input." Pg. 2, Col. 2, Para. 3 see "we employ a multi-input UNet (MI-UNet) that jointly uses MR images and brain parcellation (BP) as the input." Pg. 4, Col. 2, Para. 1 see "We train a 3D MI-UNet that uses not only the multi-modal MR images but also the corresponding BP as the inputs.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng to incorporate the teachings of Zhang to detect structures of interest from a multi-channel input that includes at least one 3D image. Doing so would predictably improve detection accuracy by allowing the network to combine information from multiple image channels at once.
Regarding claim 4, Feng in view of Zhang teaches The method according to claim 1.
While Feng teaches acquiring at least one radiographic 3D image and segmenting muscle structures within the image, Feng does not teach wherein the at least one 3D image is a radiographic image of a brain, and the multi-channel input comprises a brain parenchyma boundary inferred from the at least one 3D image.
However, Zhang teaches wherein the at least one 3D image is a radiographic image of a brain, and the multi-channel input comprises a brain parenchyma boundary inferred from the at least one 3D image. (Pg. 1, Para. 1 see "The second UNet uses brain parcellation as an additional input." Pg. 2, Col. 2, Para. 3 see "we employ a multi-input UNet (MI-UNet) that jointly uses MR images and brain parcellation (BP) as the input." Pg. 4, Col. 2, Para. 1 see "We train a 3D MI-UNet that uses not only the multi-modal MR images but also the corresponding BP as the inputs. A BP model is trained using a dataset published in one of our previous works which outputs a five-class BP with a T1 MR image as the inputs." Pg. 5, Col. 1, Para. 1 see "All MR images of the BraTS dataset have already been skull-stripped. As such, the BP of T1 MR images from the BraTS dataset consists of only four subregions: lateral ventricles, white matter, gray matter, and cerebrospinal fluid, an example of which is shown in Figure 3.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Zhang to incorporate the teachings of Zhang to use a radiographic image of a brain and include a brain parenchyma boundary that is inferred from the image in the multi-channel input. Doing so would predictably reduce false detections by limiting the network’s attention to the actual brain tissue region.
Claim 20 is rejected under the same analysis as claim 1 above.
Claim 21 is rejected under the same analysis as claim 1 above.
Claim 24 is rejected under the same analysis as claim 4 above.
Claims 2-3, 22-23 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (US 20200380696 A1), hereinafter Feng, in view of Zhang et al.: "Brain Tumor Segmentation From Multi-Modal MR Images via Ensembling UNets", Frontiers.org, Published 20 October 2021 [Retrieved on 08/16/2026]. Retrieved from the internet <https://www.frontiersin.org/journals/radiology/articles/10.3389/fradi.2021.704888/full>, hereinafter Zhang, and Mlynarski et al.: "3D Convolutional Neural Networks for Tumor Segmentation using Long-range 2D Context", arxiv.org, Published 23 Jul 2018 [Retrieved on 08/16/2026]. Retrieved from the internet <https://arxiv.org/abs/1807.08599>, hereinafter Mlynarski.
Regarding claim 2, Feng in view of Zhang teaches The method according to claim 1.
While Feng teaches detecting structures of interest in a 3D image through segmentation, Feng does not teach wherein the multi-channel input comprises feature descriptors inferred from the at least one 3D image using a third neural network.
However, Mlynarski teaches wherein the multi-channel input comprises feature descriptors inferred from the at least one 3D image using a third neural network. (Pg. 4, Para. 3 see "we propose an efficient system based on a 2D-3D model in which features extracted by 2D CNNs (capturing a rich information from a long range 2D context in three orthogonal directions) are used as an additional input to a 3D CNN." Pg. 5, Fig. 2 see "Features extracted by 2D CNNs (processing the image by axial, coronal and sagittal slices) are used as additional channels of the patch processed by a 3D CNN." Pg. 9, Para. 1 see "input is a 3D patch of the image along with a set of feature maps produced by networks trained on axial, coronal and sagittal slices (three versions of one 2D network). The extracted feature maps are concatenated to the input patch as additional channels.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Zhang to incorporate the teachings of Mlynarski to include feature descriptors that a third neural network extracts from the 3D image as part of the multi-channel input. Doing so would predictably increase detection robustness by giving the main network richer context that a single network might miss.
Regarding claim 3, Feng in view of Zhang and Mlynarski teaches The method according to claim 2.
While Feng teaches detecting structures of interest in a 3D image through segmentation, Feng does not teach wherein the at least one 3D image is a radiographic image of an organ and the feature descriptors comprise anatomical priors.
However, Zhang teaches wherein the at least one 3D image is a radiographic image of an organ and the feature descriptors comprise anatomical priors. (Pg. 1, Para. 1 see "The second UNet uses brain parcellation as an additional input." Pg. 2, Col. 2, Para. 3 see "we employ a multi-input UNet (MI-UNet) that jointly uses MR images and brain parcellation (BP) as the input." Pg. 4, Col. 2, Para. 1 see "We train a 3D MI-UNet that uses not only the multi-modal MR images but also the corresponding BP as the inputs. A BP model is trained using a dataset published in one of our previous works which outputs a five-class BP with a T1 MR image as the inputs." Pg. 5, Col. 1, Para. 1 see "All MR images of the BraTS dataset have already been skull-stripped. As such, the BP of T1 MR images from the BraTS dataset consists of only four subregions: lateral ventricles, white matter, gray matter, and cerebrospinal fluid, an example of which is shown in Figure 3.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Zhang and Mlynarski to incorporate the teachings of Zhang to use a radiographic image of an organ and make the feature descriptors anatomical priors. Doing so would predictably increase segmentation accuracy by supplying the network with known organ structure information that guides it toward realistic results.
Claim 22 is rejected under the same analysis as claim 2 above.
Claim 23 is rejected under the same analysis as claim 3 above.
Claims 5, 25 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (US 20200380696 A1), hereinafter Feng, in view of Zhang et al.: "Brain Tumor Segmentation From Multi-Modal MR Images via Ensembling UNets", Frontiers.org, Published 20 October 2021 [Retrieved on 08/16/2026]. Retrieved from the internet <https://www.frontiersin.org/journals/radiology/articles/10.3389/fradi.2021.704888/full>, hereinafter Zhang, and Salehi et al.: "Tversky loss function for image segmentation using 3D fully convolutional deep networks", arxiv.org, Published 18 Jun 2017 [Retrieved on 08/16/2026]. Retrieved from the internet <https://arxiv.org/abs/1706.05721>, hereinafter Salehi.
Regarding claim 5, Feng in view of Zhang teaches The method according to claim 1.
While Feng teaches detecting structures of interest in a 3D image by generating a segmentation map using a probability map, Feng does not teach wherein detecting structures of interest comprises generating the segmentation map by thresholding a probability map corresponding to an inferred probability of each voxel of the at least one 3D image being part of one of the structures of interest.
However, Salehi teaches wherein detecting structures of interest comprises generating the segmentation map by thresholding a probability map corresponding to an inferred probability of each voxel of the at least one 3D image being part of one of the structures of interest. (Pg. 3, Para. 2 see "At the final layer a 1×1×1 convolution with softmax output is used to reach the feature map with a depth equal to the number of classes (lesion or non-lesion tissue), where the loss function is calculated." Pg. 4, Para. 4 see "The test fold MRI volumes were segmented using feedforward through the network. The output of the last convolutional layer with softmax non-linearity consisted of a probability map for lesion and non-lesion tissues. Voxels with computed probabilities of 0.5 or more were considered to belong to the lesion tissue and those with probabilities < 0.5 were considered non-lesion tissue.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Zhang to incorporate the teachings of Salehi to generate the segmentation map of the detected structures by thresholding a probability map that shows the likelihood of each voxel belonging to one of the structures of interest. Doing so would predictably improve the accuracy and reliability of the detected regions by converting the soft probability outputs into a definite binary map that clearly marks which voxels form the structures.
Claim 25 is rejected under the same analysis as claim 5 above.
Claims 6, 8, 26-27 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (US 20200380696 A1), hereinafter Feng, in view of Zhang et al.: "Brain Tumor Segmentation From Multi-Modal MR Images via Ensembling UNets", Frontiers.org, Published 20 October 2021 [Retrieved on 08/16/2026]. Retrieved from the internet <https://www.frontiersin.org/journals/radiology/articles/10.3389/fradi.2021.704888/full>, hereinafter Zhang, and Salehi et al.: "Tversky loss function for image segmentation using 3D fully convolutional deep networks", arxiv.org, Published 18 Jun 2017 [Retrieved on 08/16/2026]. Retrieved from the internet <https://arxiv.org/abs/1706.05721>, hereinafter Salehi, and Kamnitsas et al.: "Efficient Multi-Scale 3D CNN with Fully Connected CRF for Accurate Brain Lesion Segmentation", arxiv.org, Published 18 Mar 2016 [Retrieved on 08/16/2026]. Retrieved from the internet <https://arxiv.org/abs/1603.05959>, hereinafter Kamnitsas.
Regarding claim 6, Feng in view of Zhang and Salehi teaches The method according to claim 5.
While Feng teaches detecting structures of interest in a 3D image by generating a segmentation map using a probability map, Feng does not teach wherein detecting structures of interest comprises at least one of: performing the steps of applying the first neural network N times using the same multi-channel input with dropout and data augmentation to generate N intermediate probability maps; and aggregating the N intermediate probability maps to generate the probability map; and performing the steps of sampling the at least one 3D image into patches; applying the first neural network to each of the patches to generate a plurality of corresponding intermediate patch probability maps; and aggregating the intermediate patch probability maps to generate the probability map.
However, Kamnitsas teaches wherein detecting structures of interest comprises at least one of: performing the steps of applying the first neural network N times using the same multi-channel input with dropout and data augmentation to generate N intermediate probability maps; and aggregating the N intermediate probability maps to generate the probability map; and performing the steps of sampling the at least one 3D image into patches; applying the first neural network to each of the patches to generate a plurality of corresponding intermediate patch probability maps; and aggregating the intermediate patch probability maps to generate the probability map. (Pg. 6, Para. 4 see "A commonly adopted approach is training on image patches that are equally sampled from each class." Pg. 9, Fig. 2 see "When the network segments an input it predicts multiple voxels simultaneously, one for each shift of its receptive field over the input." Pg. 10, Para. 2 see "Networks that are implemented as fully-convolutionals are capable of dense-inference, which is performed when input of size greater than ϕ CNN is provided" Pg. 10, Para. 3 see "patches of size ϕ CNN are extracted from the training images." Pg. 11, Para. 3 see "we devise a training strategy that exploits the dense inference technique on image segments. ... if an image segment of size greater than ϕ CNN is given as input to our network, the output is a posterior probability for multiple voxels" Pg. 16, Para. 1 see "the soft segmentation maps produced by the CNN tend to be smooth" Pg. 24, Para. 4 see "we form an ensemble of three similar networks, aggregating their output by averaging." Examiner Note: Disclosed is sampling the at least one 3D image into patches; applying the first neural network to each of the patches to generate a plurality of corresponding intermediate patch probability maps; and aggregating the intermediate patch probability maps to generate the probability map.).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Zhang and Salehi to incorporate the teachings of Kamnitsas to sample the at least one 3D image into patches, apply the first neural network to each of the patches to generate intermediate patch probability maps, and aggregate those maps to produce the final probability map when detecting structures of interest. Doing so would predictably improve accuracy by combining predictions across multiple image regions into one smoother result.
Regarding claim 8, Feng in view of Salehi and Zhang and Kamnitsas teaches The method according to claim 6.
While Feng teaches detecting structures of interest in a 3D image by generating a segmentation map using a probability map, Feng does not teach wherein detecting structures of interest comprises sampling the at least one 3D image, and wherein sampling the at least one 3D image comprises at least one of: sampling the at least one 3D image in a grid-like fashion; and applying a sliding window algorithm, such that regions of the at least one 3D image represented in adjacent patches at least partially overlap.
However, Kamnitsas teaches wherein detecting structures of interest comprises sampling the at least one 3D image, and wherein sampling the at least one 3D image comprises at least one of: sampling the at least one 3D image in a grid-like fashion; and applying a sliding window algorithm, such that regions of the at least one 3D image represented in adjacent patches at least partially overlap. (Pg. 9, Fig. 2 see "When the network segments an input it predicts multiple voxels simultaneously, one for each shift of its receptive field over the input." Pg. 10, Para. 2 see "Networks that are implemented as fully-convolutionals are capable of dense-inference, which is performed when input of size greater than ϕ CNN is provided ... This strategy significantly reduces the computational costs and memory loads since the otherwise repeated computations of convolutions on the same voxels in overlapping patches are avoided. Optimal performance is achieved if the whole image is scanned in one forward pass.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Salehi and Zhang and Kamnitsas to incorporate the teachings of Kamnitsas to sample the at least one 3D image in a grid-like fashion with a sliding window so that regions in adjacent patches at least partially overlap when detecting structures of interest. Doing so would predictably increase processing efficiency by avoiding repeated calculations on the same image voxels across overlapping areas.
Claim 26 is rejected under the same analysis as claim 6 above.
Claim 27 is rejected under the same analysis as claim 8 above.
Claims 10, 28 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (US 20200380696 A1), hereinafter Feng, in view of Zhang et al.: "Brain Tumor Segmentation From Multi-Modal MR Images via Ensembling UNets", Frontiers.org, Published 20 October 2021 [Retrieved on 08/16/2026]. Retrieved from the internet <https://www.frontiersin.org/journals/radiology/articles/10.3389/fradi.2021.704888/full>, hereinafter Zhang, and Sakinis et al.: "Interactive segmentation of medical images through fully convolutional neural networks", arxiv.org, Published 19 Mar 2019 [Retrieved on 08/16/2026]. Retrieved from the internet <https://arxiv.org/abs/1903.08205>, hereinafter Sakinis.
Regarding claim 10, Feng in view of Zhang teaches The method according to claim 1.
In addition, Feng teaches further comprising displaying the detected structures and providing controls allowing a user to: (Para. 7 see "displaying the segmented high-resolution 3D image." Para. 67 see "FIG. 2 shows a computing environment 210 coupled with an MRI system 200 and user interface 260." Para. 102 see "The images can be displayed to the user on a display, such as operator workstation 102.").
While Feng teaches segmenting structures from 3D images and displaying the segmented structures on a display through a user interface, Feng does not teach approve or reject the detected structures and/or modify contours of the detected structures to generate a new shape representation.
However, Sakinis teaches approve or reject the detected structures and/or modify contours of the detected structures to generate a new shape representation. (Pg. 1, Fig. 1 see "The user, who possesses domain knowledge, observes the image and clicks on the region of interest (green Gaussians). Almost instantly the 2D segmentation is generated. If the result needs correction, the user can drag the green Gaussian around to see live as the segmentation results change. If that is not adequate, the user can place any number of additional clicks inside (green) or outside (red) of the region of interest to steer the DL-based segmentation according to his or her requirement and preferences. The user keeps interacting with the segmentation results until segmentation can be considered correct (usually 1 to 3 clicks)." Pg. 4, Col. 1, Para. 1 see "The minimum amount of interaction for a single image is 1-click, which should be placed within the object of interest." Pg. 8, Col. 1, Para. 2 see "users can interact in real-time with the segmentation results by simply clicking on any structure that they want to segment. Any false-positive and false-negative regions can be corrected by placing any number of additional clicks.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Zhang to incorporate the teachings of Sakinis to provide controls that let a user approve or reject the detected structures and modify their contours to generate a new shape representation. Doing so would predictably improve the accuracy and user experience by enabling quick interactive corrections of any mistakes in the automatic results.
Claim 28 is rejected under the same analysis as claim 10 above.
Claims 11, 29 are rejected under 35 U.S.C. 103 as being unpatentable over Feng et al. (US 20200380696 A1), hereinafter Feng, in view of Zhang et al.: "Brain Tumor Segmentation From Multi-Modal MR Images via Ensembling UNets", Frontiers.org, Published 20 October 2021 [Retrieved on 08/16/2026]. Retrieved from the internet <https://www.frontiersin.org/journals/radiology/articles/10.3389/fradi.2021.704888/full>, hereinafter Zhang, and Wang et al. (US 20200167930 A1), hereinafter Wang.
Regarding claim 11, Feng in view of Zhang teaches The method according to claim 1.
In addition, Feng teaches wherein the detecting, extracting and estimating are carried out to automatically detect and contour a plurality of structures of interest, (Para. 55 see "The present disclosure relates to a cascaded three dimensional (3D) deep convolutional neural network (DCNN) segmentation framework for fully automatic segmentation of all 35 lower limb muscles (70 on both left and right legs)." Para. 85 see "Due to the relatively small training size and the large number of muscles, instead of using single models to segment all muscles in the abdomen, upper leg and lower leg regions at the same time, using dedicated separate models for each individual muscle or muscle groups can greatly improve the accuracy. ... The total 70 ROIs were divided into 10 groups, each containing about 4 adjacent muscles on both legs." Para. 90 see "Segmentation is done by individual networks which were trained for each target muscle to obtain more accurate contours from images cropped based on the output of the first stage ... The total number of models to be trained was then 35.").
While Feng teaches segmenting structures from 3D images and displaying the segmented structures on a display through a user interface, Feng does not teach further wherein the method comprises receiving a user input corresponding to a manually defined bounding box, extracting an additional cropped image from the at least one 3D image corresponding to a subregion of the at least one 3D image within the bounding box, and estimating contours of an additional structure within the additional cropped image using the second neural network.
However, Wang teaches further wherein the method comprises receiving a user input corresponding to a manually defined bounding box, (Para. 106 see "input image may be a region of a larger image, with the user selecting the relevant region, such as by using a bounding box to denote the region of the larger image to use as the input image." Para. 189 see "During a test stage, the bounding box is provided by the user, and the segmentation and the CNN are refined through unsupervised (no further user interactions) or supervised (with user-provided scribbles) image-specific fine-tuning." Para. 192 see "In the testing stage, let X denote the sub-image inside a user-provided bounding box and Y be the target label of X."). extracting an additional cropped image from the at least one 3D image corresponding to a subregion of the at least one 3D image within the bounding box, (Para. 106 see "input image may be a region of a larger image, with the user selecting the relevant region, such as by using a bounding box to denote the region of the larger image to use as the input image." Para. 189 see "the present approach uses a CNN that takes as input the content of a bounding box (or the entire image) of one instance (object image) and gives a segmented output." Para. 191 see "Xp and Ypq are cropped based on Bpq, and so T is converted into a cropped set"). and estimating contours of an additional structure within the additional cropped image using the second neural network. (Para. 86 see "At step 150 a second (refinement) segmentation network (R-Net hereafter) uses the information of the original input image (as provided at step 100), the initial segmentation (as proposed at step 120) and the user interactions (as received at step 140) to provide a refined segmentation." Para. 188 see "A CNN-based method described herein allows segmentation of previously unseen objects and uses image-specific fine-tuning (i.e. adaptation) to further improve accuracy. The CNN is incorporated into an interactive, optionally bounding box-based, binary segmentation and employed to segment a range of different organs" Para. 189 see "During a test stage, the bounding box is provided by the user, and the segmentation and the CNN are refined through unsupervised (no further user interactions) or supervised (with user-provided scribbles) image-specific fine-tuning.").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Feng and Zhang to incorporate the teachings of Wang to include the ability for a user to supply a manually defined bounding box, extract an extra cropped region from inside that box, and estimate the contours of an additional structure inside the crop with the second neural network. Doing so would predictably improve flexibility by letting a user quickly include or correct a missed structure without restarting the entire automatic process.
Claim 29 is rejected under the same analysis as claim 11 above.
Allowable Subject Matter
Claim(s) 12 is/are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 12, none of the cited references above teach a neural network trained using the specified modified dice loss function.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Le et al.: "Deep Recurrent Level Set for Segmenting Brain Tumors", arxiv.org, Published 10 Oct 2018 [Retrieved on 08/16/2026]. Retrieved from the internet <https://arxiv.org/abs/1810.04752> discloses a method for segmenting brain tumors from 3D images and includes contour estimation and shape representation using level-set functions.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER VAUGHN whose telephone number is (571) 272-5253. The examiner can normally be reached M-F 11am-7pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JENNIFER MEHMOOD can be reached on (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEXANDER VAUGHN/Examiner, Art Unit 2675
/JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664