DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Claims 1-12 are pending for examination in the application filed 10/24/2024.
Priority
Acknowledgement is made of Applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been for parent application EP23306860.0, filing date: 10/24/2023 has been received.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/16/2026 has been considered by the examiner.
Claim Objections
Claims 1 and 6 are objected to because of the following informalities: “a first cohort of subject” and “a second cohort of subject” should read “a first cohort of subjects” and “a second cohort of subjects”. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 11 and 12 are rejected under 35 U.S.C. 101 because the claimed inventions are directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because claim 11 is drawn to “A computer program product comprising instructions” and claim 12 is drawn to “A computer-readable storage medium comprising instructions”.
Per the MPEP 2106.03 Eligibility Step 1: The Four Categories of Statutory Subject Matter [R-07.2022], non-limiting examples of claims that are not directed to any of the statutory categories include:
Products that do not have a physical or tangible form, such as information (often referred to as "data per se”) or a computer program per se (often referred to as "software per se") when claimed as a product without any structural recitations;
Transitory forms of signal transmission (often referred to as "signals per se"), such as a propagating electrical or electromagnetic signal or carrier wave; and
Subject matter that the statute expressly prohibits from being patented, such as humans per se, which are excluded under The Leahy-Smith America Invents Act (AIA ), Public Law 112-29, sec. 33, 125 Stat.284 (September 16, 2011).
Therefore, since claim 11 recites a computer program per se as a product without any structural recitations, it does not fall within a statutory category. Claim 11 is not eligible subject matter under 35 USC § 101.
Therefore, since claim 12 recites a computer-readable storage medium comprising instructions, which could include transmission type media, such as wireless or radio waves, etc. it does not fall within a statutory category. The broadest reasonable interpretation of the claim in light of the specification and Official Gazette Notice (1251 OG 212, made available February 23, 2010), concludes that the claim as a whole covers a transitory signal, which does not fall within the definition of a process, machine, manufacture, or composition of matter (In re Nuijten). In view of the Official Gazette Notice (1251 OG 212, made available February 23, 2010), the examiner suggests amending the claims to recite "a non-transitory computer readable
storage medium".
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-4 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Cersovsky (US20240312010A1) in view of Tran (US20250029254A1).
Regarding claim 1, Cersovsky teaches a device for training a machine learning model for detecting a presence of at least one region of interest in a 3D image of a subject ([Abstract] The systems, methods, and computer programs disclosed herein relate to detecting and/or recognizing a tissue type in an image of the tissue using machine learning techniques. [0069] FIG. 1 schematically shows an embodiment of the method for training the machine learning model of the present disclosure. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
said device comprising: at least one input configured to receive a first set of 3D images acquired from a first cohort of subject and a second set of 3D images acquired from a second cohort of subject, wherein the 3D images of said first set differs from the 3D images of said second set for the presence of said at least one region of interest ([0105] The machine learning model of the present disclosure is trained on training data. The training data comprise a multitude of tissue images. The term “multitude” means more than 10, preferably more than 100. [0106] Each of the tissue images may show the same tissue(s), such as lung tissue, breast tissue, liver tissue, thyroid tissue, skin tissue, bone tissue, and/or other tissue. It is possible that the tissue(s) shown in the images is(are) from different individuals. [0107] Each image is annotated (labelled), i.e., there is information which of the class of the at least two classes the image is assigned to (class information). [0145] The class may indicate whether the tissue shown in the image is a lesion, or an edema or a tumor);
at least one processor configured to: for each received 3D image of the first cohort and the second cohort, generate at least two stacks of 2D slices ([0105] The machine learning model of the present disclosure is trained on training data. The training data comprise a multitude of tissue images. The term “multitude” means more than 10, preferably more than 100. [0108] The images—or more precisely patches generated from the image—are used as input data when training the machine learning model; the class information is used as target. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice);
generate a training dataset comprising said stacks of 2D slices; and train said machine learning model using said generated training dataset so to obtain training parameters ([0318] From each image a number of patches is generated. This is shown in FIG. 1 only for the image Ip…The selected patches P*(I.sub.p).sub.1, . . . , P*(I.sub.p).sub.m are inputted into the machine learning model MLM. The machine learning model MLM is configured to assign the selected patches P*(I.sub.p).sub.1, . . . , P*(I.sub.p).sub.m (to be more precise: a bag-level representation of patch embeddings of the selected patches) to one of two classes based on the model parameters MP…The one or more deviations are reduced by modifying the model parameters MP, e.g., in an optimization procedure, such as a gradient descent optimization procedure. The process is repeated for further images and/or further selected patches),
and optimize a loss function during said training so to optimize discrimination between the 2D slices comprising said at least one region from the 2D slices comprising no region of interest ([0103] In general, a loss function can be used for training, where the loss function can quantify the deviations between the output and the target. The loss function may be chosen in such a way that it rewards a wanted relation between output and target and/or penalizes an unwanted relation between an output and a target. Such a relation can be, e.g., a similarity, or a dissimilarity, or another relation);
and at least one output configured to provide said training parameters for the trained machine learning model ([0102] In the training process, training data are inputted into the machine learning model and the machine learning model generates an output. The output is compared with the (known) target. Parameters of the machine learning model are modified in order to reduce the deviations between the output and the (known) target to a (defined) minimum).
Cersovsky does not explicitly teach that an orientation of the slices in each stack is different from an orientation of the slices in another stack of said at least two stacks; said machine learning model being configured to receive as input a 2D slice and provide as output a slice-level prediction score representing a probability that said input 2D slice comprises at least one region of interest, said slice-level prediction score being used to compute and optimize a loss function.
Tran, in the same field of endeavor of medical image analysis, teaches that an orientation of the slices in each stack is different from an orientation of the slices in another stack of said at least two stacks ([0043] In embodiments, the neural network takes as input a standardised 3D volume. CT brain studies consist of a range of 2D images, and the number of these images varies a lot depending on the slice thickness, for example. Before these images are passed for training the AI model they need to be standardised by converting them into a fixed shape and voxel spacing. Describing from a different perspective, there is provided a plurality of images of a CT scan/study (such as e.g. from about 100 to about 450) to the neural network as an input to the entire machine learning pipeline. The neural network produces as output an indication of a plurality of visual findings being present in any one of the plurality of images. [0060] According to a further aspect, there is provided a method comprising: receiving, by a processor, the results of a step of analysing a series of anatomical images from a computed tomography (CT) scan of a head of a subject using one or more deep learning models trained to detect and localise in 3D space at least a first visual finding in anatomical images, wherein the results comprise a plurality of segmentation maps obtained with for at least one anatomical plane, wherein a segmentation map indicates the areas of a respective anatomical image where the first visual finding has been detected. [0069] The anatomical plane may be any one from the group consisting of: sagittal, coronal or transverse. [0201] 3D segmentation masks 362: used to generate the axial, sagittal, and coronal viewpoints (each view is saved separately as a list of PNGs 384));
said machine learning model being configured to receive as input a 2D slice and provide as output a slice-level prediction score representing a probability that said input 2D slice comprises at least one region of interest, said slice-level prediction score being used to compute and optimize a loss function ([0147] In some embodiments, as illustrated in FIG. 3A, a CNN model 200 comprises a CNN encoder 302 that is configured to process input anatomical images 304. The CNN encoder 302 functions as an encoder, and is connected to a CNN decoder 306 which functions as a decoder. At least one MAP 308 comprises a two-dimensional array of values representing a probability that the corresponding pixel of an anatomical slice, e.g. the input anatomical images 204, which exhibits a visual finding, e.g. as identified by the CNN model 200. The MAP 308 can be in a particular anatomical plane, and can be overlaid over a respective anatomical image 204 in the same anatomical plane. The anatomical image 204 may also be a 2D anatomical slice in the anatomical plane generated from a 3D spatial model of the subject. [0320] FIG. 6A to 6D show exemplary interactive user interface screens of the viewer component 701 in accordance with an example embodiment. Clinical findings detected by the deep-learning model are listed and a segmentation MAP 502 (identified as one with most colour e.g. purple) is presented in relevant pathological slices. Finding likelihood scores and confidence intervals are also displayed under the NCCTB scan, illustrated as a sliding scale at the bottom of absent versus present. [0142] Training on the dataset related to the problem at hand by initialising with pre-trained weights allows for certain features to already be recognised and increases the likelihood of finding a global, or reduced local, minimum for the loss function than otherwise. [0144] In this example, the CNN model 200 comprises an ensemble of five CNN components trained using five-fold cross-validation. In an example, the CNN model comprises three heads (modules): one for classification, one for left-right localization in 3D space, and one for segmentation in 3D space. Models were based on the ResNet, Y-Net and Vision Transformer (ViT) architectures. An attention-per-token ViT head (vision transformer 394) is included to significantly improving the performance of classification of radiological findings for CTB studies (denoted visual anomaly findings or visual findings). Class imbalance is mitigated using class-balanced loss weighting and oversampling).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Cersovsky with the teachings of Tran to use stacks of slices in different orientations and output a slice level prediction that a 2D slice comprises a region of interest to optimize a loss function because "A challenge is that predictions generated by deep learning models can be difficult to interpret by a user (such as, e.g., a clinician). Such models produce a score, probability or combination of scores for each class that they are trained to distinguish, which are often meaningful within a particular context related to the sensitivity/specificity of the deep learning model in detecting the clinically relevant feature associated with the class" [0008] and "This enables providing an ideal default view to the user for the user to confirm the presence of a finding detected by the model by displaying the medical prediction in the easiest manner to the user rather than the user having to search through the plurality of slices through a plurality of anatomical planes through a plurality of windows/user interface. More specifically, the system is configured to select an optimal combination for an ideal default view, of: i) “slice-method” (e.g. which one of the 3 anatomical planes: sagittal, coronal, transverse)” is the best for the user to visually confirm the presence/absence of the radiological finding; ii) slice image 204 (which image 204 out of X number of the stack of images 204 of a CT scan) is the default (or best) for the user to visually confirm the presence/absence of the radiological finding; and iii) window to display the slice image (e.g. which window from the brain, bone, soft tissue, bone, stroke, subdural windowing presets) is the best for the user to visually confirm the presence/absence of the radiological finding" [0318].
Regarding claim 2, Cersovsky and Tran teach the device of claim 1. Cersovsky further teaches wherein said machine learning model is a 2D deep neural network ([0124] Usually, the machine learning model includes convolution and pooling operations to generate a patch embedding from a patch. For example, the machine learning model may be or include a convolutional neural network (CNN). [0117] Patches can be created from 2D images by dividing the 2D image into smaller areas. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice).
Regarding claim 3, Cersovsky and Tran teach the device of claim 1. Cersovsky further teaches wherein the region of interest is a lesion present in the subjects of the first cohort and absent in the subjects of the second cohort ([0106] Each of the tissue images may show the same tissue(s), such as lung tissue, breast tissue, liver tissue, thyroid tissue, skin tissue, bone tissue, and/or other tissue. It is possible that the tissue(s) shown in the images is(are) from different individuals. [0107] Each image is annotated (labelled), i.e., there is information which of the class of the at least two classes the image is assigned to (class information). [0145] The class may indicate whether the tissue shown in the image is a lesion, or an edema or a tumor. [0094] In one embodiment, each image used for training is assigned to one of exactly two classes, one class representing images that show a specific tissue type and the other class representing images that do not show the specific tissue type (binary classification)).
Regarding claim 4, Cersovsky and Tran teach the device of claim 1. Cersovsky further teaches wherein generating the training dataset comprises generating a plurality of training samples, each training sample comprising a stack of 2D slices and at least one label, said label representing an absence or a present of at least one region of interest in the 3D image from which the stack of 2D slices is obtained ([0105] The machine learning model of the present disclosure is trained on training data. The training data comprise a multitude of tissue images. The term “multitude” means more than 10, preferably more than 100. [0107] Each image is annotated (labelled), i.e., there is information which of the class of the at least two classes the image is assigned to (class information). [0108] The images—or more precisely patches generated from the image—are used as input data when training the machine learning model; the class information is used as target. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice. [0094] In one embodiment, each image used for training is assigned to one of exactly two classes, one class representing images that show a specific tissue type and the other class representing images that do not show the specific tissue type (binary classification)).
Regarding claim 6, Cersovsky teaches a computer-implemented method for training a machine learning model for detecting a presence of at least one region of interest in a 3D image of a subject ([Abstract] The systems, methods, and computer programs disclosed herein relate to detecting and/or recognizing a tissue type in an image of the tissue using machine learning techniques. [0069] FIG. 1 schematically shows an embodiment of the method for training the machine learning model of the present disclosure. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
said method comprising: receive a first set of 3D images acquired from a first cohort of subject (21) and a second set of 3D images acquired from a second cohort of subject (22), wherein the 3D images of said first set differs from the 3D images of said second set for the presence of said at least one region of interest ([0105] The machine learning model of the present disclosure is trained on training data. The training data comprise a multitude of tissue images. The term “multitude” means more than 10, preferably more than 100. [0106] Each of the tissue images may show the same tissue(s), such as lung tissue, breast tissue, liver tissue, thyroid tissue, skin tissue, bone tissue, and/or other tissue. It is possible that the tissue(s) shown in the images is(are) from different individuals. [0107] Each image is annotated (labelled), i.e., there is information which of the class of the at least two classes the image is assigned to (class information). [0145] The class may indicate whether the tissue shown in the image is a lesion, or an edema or a tumor);
for each received 3D image of the first cohort and the second cohort, generate at least two stacks of 2D slices ([0105] The machine learning model of the present disclosure is trained on training data. The training data comprise a multitude of tissue images. The term “multitude” means more than 10, preferably more than 100. [0108] The images—or more precisely patches generated from the image—are used as input data when training the machine learning model; the class information is used as target. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice);
generate a training dataset comprising said stacks of 2D slices; train said machine learning model using said generated training dataset so to obtain training parameters ([0318] From each image a number of patches is generated. This is shown in FIG. 1 only for the image Ip…The selected patches P*(I.sub.p).sub.1, . . . , P*(I.sub.p).sub.m are inputted into the machine learning model MLM. The machine learning model MLM is configured to assign the selected patches P*(I.sub.p).sub.1, . . . , P*(I.sub.p).sub.m (to be more precise: a bag-level representation of patch embeddings of the selected patches) to one of two classes based on the model parameters MP…The one or more deviations are reduced by modifying the model parameters MP, e.g., in an optimization procedure, such as a gradient descent optimization procedure. The process is repeated for further images and/or further selected patches),
and optimize a loss function during said training so to optimize discrimination between the 2D slices comprising said at least one region from the 2D slices comprising no region of interest ([0103] In general, a loss function can be used for training, where the loss function can quantify the deviations between the output and the target. The loss function may be chosen in such a way that it rewards a wanted relation between output and target and/or penalizes an unwanted relation between an output and a target. Such a relation can be, e.g., a similarity, or a dissimilarity, or another relation);
and at least one output configured to provide said training parameters for the trained machine learning model ([0102] In the training process, training data are inputted into the machine learning model and the machine learning model generates an output. The output is compared with the (known) target. Parameters of the machine learning model are modified in order to reduce the deviations between the output and the (known) target to a (defined) minimum).
Cersovsky does not explicitly teach that an orientation of the slices in each stack is different from an orientation of the slices in another stack of said at least two stacks; said machine learning model being configured to receive as input a 2D slice and provide as output a slice-level prediction score representing a probability that said input 2D slice comprises at least one region of interest, said slice-level prediction score being used to compute and optimize a loss function.
Tran, in the same field of endeavor of medical image analysis, teaches that an orientation of the slices in each stack is different from an orientation of the slices in another stack of said at least two stacks ([0043] In embodiments, the neural network takes as input a standardised 3D volume. CT brain studies consist of a range of 2D images, and the number of these images varies a lot depending on the slice thickness, for example. Before these images are passed for training the AI model they need to be standardised by converting them into a fixed shape and voxel spacing. Describing from a different perspective, there is provided a plurality of images of a CT scan/study (such as e.g. from about 100 to about 450) to the neural network as an input to the entire machine learning pipeline. The neural network produces as output an indication of a plurality of visual findings being present in any one of the plurality of images. [0060] According to a further aspect, there is provided a method comprising: receiving, by a processor, the results of a step of analysing a series of anatomical images from a computed tomography (CT) scan of a head of a subject using one or more deep learning models trained to detect and localise in 3D space at least a first visual finding in anatomical images, wherein the results comprise a plurality of segmentation maps obtained with for at least one anatomical plane, wherein a segmentation map indicates the areas of a respective anatomical image where the first visual finding has been detected. [0069] The anatomical plane may be any one from the group consisting of: sagittal, coronal or transverse. [0201] 3D segmentation masks 362: used to generate the axial, sagittal, and coronal viewpoints (each view is saved separately as a list of PNGs 384));
said machine learning model being configured to receive as input a 2D slice and provide as output a slice-level prediction score representing a probability that said input 2D slice comprises at least one region of interest, said slice-level prediction score being used to compute and optimize a loss function ([0147] In some embodiments, as illustrated in FIG. 3A, a CNN model 200 comprises a CNN encoder 302 that is configured to process input anatomical images 304. The CNN encoder 302 functions as an encoder, and is connected to a CNN decoder 306 which functions as a decoder. At least one MAP 308 comprises a two-dimensional array of values representing a probability that the corresponding pixel of an anatomical slice, e.g. the input anatomical images 204, which exhibits a visual finding, e.g. as identified by the CNN model 200. The MAP 308 can be in a particular anatomical plane, and can be overlaid over a respective anatomical image 204 in the same anatomical plane. The anatomical image 204 may also be a 2D anatomical slice in the anatomical plane generated from a 3D spatial model of the subject. [0320] FIG. 6A to 6D show exemplary interactive user interface screens of the viewer component 701 in accordance with an example embodiment. Clinical findings detected by the deep-learning model are listed and a segmentation MAP 502 (identified as one with most colour e.g. purple) is presented in relevant pathological slices. Finding likelihood scores and confidence intervals are also displayed under the NCCTB scan, illustrated as a sliding scale at the bottom of absent versus present. [0142] Training on the dataset related to the problem at hand by initialising with pre-trained weights allows for certain features to already be recognised and increases the likelihood of finding a global, or reduced local, minimum for the loss function than otherwise. [0144] In this example, the CNN model 200 comprises an ensemble of five CNN components trained using five-fold cross-validation. In an example, the CNN model comprises three heads (modules): one for classification, one for left-right localization in 3D space, and one for segmentation in 3D space. Models were based on the ResNet, Y-Net and Vision Transformer (ViT) architectures. An attention-per-token ViT head (vision transformer 394) is included to significantly improving the performance of classification of radiological findings for CTB studies (denoted visual anomaly findings or visual findings). Class imbalance is mitigated using class-balanced loss weighting and oversampling).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Cersovsky with the teachings of Tran to use stacks of slices in different orientations and output a slice level prediction that a 2D slice comprises a region of interest to optimize a loss function because "A challenge is that predictions generated by deep learning models can be difficult to interpret by a user (such as, e.g., a clinician). Such models produce a score, probability or combination of scores for each class that they are trained to distinguish, which are often meaningful within a particular context related to the sensitivity/specificity of the deep learning model in detecting the clinically relevant feature associated with the class" [0008] and "This enables providing an ideal default view to the user for the user to confirm the presence of a finding detected by the model by displaying the medical prediction in the easiest manner to the user rather than the user having to search through the plurality of slices through a plurality of anatomical planes through a plurality of windows/user interface. More specifically, the system is configured to select an optimal combination for an ideal default view, of: i) “slice-method” (e.g. which one of the 3 anatomical planes: sagittal, coronal, transverse)” is the best for the user to visually confirm the presence/absence of the radiological finding; ii) slice image 204 (which image 204 out of X number of the stack of images 204 of a CT scan) is the default (or best) for the user to visually confirm the presence/absence of the radiological finding; and iii) window to display the slice image (e.g. which window from the brain, bone, soft tissue, bone, stroke, subdural windowing presets) is the best for the user to visually confirm the presence/absence of the radiological finding" [0318].
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Cersovsky in view of Tran and Melapudi (US20240212158A1).
Regarding claim 5, Cersovsky and Tran teach the device of claim 4. Cersovsky does not explicitly teach wherein the training is a weakly supervised training.
Melapudi, in the same field of endeavor of medical image analysis, teaches wherein the training is a weakly supervised training ([Abstract] In one embodiment, a method includes receiving a positive point selection or negative point selection for an ROI in a first image slice of an image sequence, mapping the first image slice and the positive point selection or negative point selection to a first segmentation mask of the ROI using a weakly supervised segmentation model).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Cersovsky with the teachings of Melapudi to use weakly supervised training "to provide “weak supervision” (e.g., supervision less than complete manual delineation of the region of interest) regarding which region(s) of the image to segment. As an example, an annotator may make several positive point selections inside of a tumor (e.g., by selecting said points using a user input device such as a mouse, touch screen, stylus, etc.) in an image slice and the annotator may make several negative point selections outside of the tumor. A weakly supervised segmentation model may then map the positive and negative point selections, along with the image slice, to a segmentation mask of the tumor. The annotator may refine the segmentation mask of the tumor by adding more positive or negative point selections, to provide more “supervision” to the weakly supervised segmentation model" [0019].
Claims 7-8 and 10-12 are rejected under 35 U.S.C. 103 as being unpatentable over Cersovsky in view of Tran and Ursella (US20210116395A1).
Regarding claim 7, Cersovsky and Tran teach the method of claim 6. Cersovsky further teaches a device for detecting at least one region of interest in a 3D image of a subject using a trained machine learning model obtained with the method for training ([0002] The systems, methods, and computer programs disclosed herein relate to detecting and/or recognizing a tissue type in an image of the tissue using machine learning techniques. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
wherein said device comprises: at least one input configured to receive a 3D image of a subject ([0011] Once the machine learning model is trained, it can be used to classify a new image into one of the two trained classes. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
at least one processor configured to: obtain a plurality of stacks of 2D slices from said 3D image ([0033] In another aspect, the present disclosure provides a computer system comprising: a processor. [0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice));
for each obtained stack: feed each 2D slice into said trained machine learning model for detecting at least one region of interest in the 3D image of the subject ([0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue, [0111] generate a patch embedding for each received patch, [0112] aggregate patch embeddings into a bag-level representation, with each patch embedding assigned a learnable attention weight, and [0113] classify the bag-level-representation into one of at least two classes. [0143] The class may indicate whether the tissue shown in the image is a specific tissue or not. [0145] The class may indicate whether the tissue shown in the image is a lesion, or an edema or a tumor).
Cersovsky does not explicitly teach a device for detecting and reconstructing at least one region of interest in a 3D image; wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks; obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest; and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject.
Tran, in the same field of endeavor of medical image analysis, teaches a device for detecting and reconstructing at least one region of interest in a 3D image (Fig. 1. [0144] In this example, the CNN model 200 comprises an ensemble of five CNN components trained using five-fold cross-validation. In an example, the CNN model comprises three heads (modules): one for classification, one for left-right localization in 3D space, and one for segmentation in 3D space…A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model);
wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks ([0043] In embodiments, the neural network takes as input a standardised 3D volume. CT brain studies consist of a range of 2D images, and the number of these images varies a lot depending on the slice thickness, for example. Before these images are passed for training the AI model they need to be standardised by converting them into a fixed shape and voxel spacing. Describing from a different perspective, there is provided a plurality of images of a CT scan/study (such as e.g. from about 100 to about 450) to the neural network as an input to the entire machine learning pipeline. The neural network produces as output an indication of a plurality of visual findings being present in any one of the plurality of images. [0060] According to a further aspect, there is provided a method comprising: receiving, by a processor, the results of a step of analysing a series of anatomical images from a computed tomography (CT) scan of a head of a subject using one or more deep learning models trained to detect and localise in 3D space at least a first visual finding in anatomical images, wherein the results comprise a plurality of segmentation maps obtained with for at least one anatomical plane, wherein a segmentation map indicates the areas of a respective anatomical image where the first visual finding has been detected. [0069] The anatomical plane may be any one from the group consisting of: sagittal, coronal or transverse. [0201] 3D segmentation masks 362: used to generate the axial, sagittal, and coronal viewpoints (each view is saved separately as a list of PNGs 384));
obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest ([0147] In some embodiments, as illustrated in FIG. 3A, a CNN model 200 comprises a CNN encoder 302 that is configured to process input anatomical images 304. The CNN encoder 302 functions as an encoder, and is connected to a CNN decoder 306 which functions as a decoder. At least one MAP 308 comprises a two-dimensional array of values representing a probability that the corresponding pixel of an anatomical slice, e.g. the input anatomical images 204, which exhibits a visual finding, e.g. as identified by the CNN model 200. The MAP 308 can be in a particular anatomical plane, and can be overlaid over a respective anatomical image 204 in the same anatomical plane. The anatomical image 204 may also be a 2D anatomical slice in the anatomical plane generated from a 3D spatial model of the subject. [0320] FIG. 6A to 6D show exemplary interactive user interface screens of the viewer component 701 in accordance with an example embodiment. Clinical findings detected by the deep-learning model are listed and a segmentation MAP 502 (identified as one with most colour e.g. purple) is presented in relevant pathological slices. Finding likelihood scores and confidence intervals are also displayed under the NCCTB scan, illustrated as a sliding scale at the bottom of absent versus present);
and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject ([0144] The localization is a 3D tensor that will be sent to the viewer component 701 and overlaid on thumbnail images/slices of the CTB study. The 3D tensor may be a static size. A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model. Large coloured segments visually indicate that there is a larger mass detected, whereas faint lines (thin segments) visually indicate that one or two consecutive slices with localization. The absence of coloured segments denotes a lack of radiological findings not detected by the model for these slices. The rest of the slices are loaded on-demand, when/if a radiologist scrolls through the other slices/regions).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Cersovsky with the teachings of Tran to use stacks of slices in different orientations, obtain a slice level prediction that a 2D slice comprises a region of interest, and output a tomographic reconstruction of the region of interest because "A challenge is that predictions generated by deep learning models can be difficult to interpret by a user (such as, e.g., a clinician). Such models produce a score, probability or combination of scores for each class that they are trained to distinguish, which are often meaningful within a particular context related to the sensitivity/specificity of the deep learning model in detecting the clinically relevant feature associated with the class" [0008] and "This enables providing an ideal default view to the user for the user to confirm the presence of a finding detected by the model by displaying the medical prediction in the easiest manner to the user rather than the user having to search through the plurality of slices through a plurality of anatomical planes through a plurality of windows/user interface. More specifically, the system is configured to select an optimal combination for an ideal default view, of: i) “slice-method” (e.g. which one of the 3 anatomical planes: sagittal, coronal, transverse)” is the best for the user to visually confirm the presence/absence of the radiological finding; ii) slice image 204 (which image 204 out of X number of the stack of images 204 of a CT scan) is the default (or best) for the user to visually confirm the presence/absence of the radiological finding; and iii) window to display the slice image (e.g. which window from the brain, bone, soft tissue, bone, stroke, subdural windowing presets) is the best for the user to visually confirm the presence/absence of the radiological finding" [0318].
Cersovsky does not explicitly teach calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; and for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest.
Ursella, in the same field of endeavor of tomographic reconstruction, teaches calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; and for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest ([0058] To perform an appropriate tomographic reconstruction, the detection step should be repeated on each object as many times as it takes until a sufficient number of radiographs (electronic two-dimensional maps) have been obtained by observing the object from different angles (generally, at least one hundred different radiographs should be detected, although fewer or more radiographs may be necessary depending on the given case). [0100] The device 1 comprises at least one computer (not shown) associated with the one or more detectors 5 to receive from them electronic data corresponding to two-dimensional pixel maps representative of the density of the object 2 crossed by the x-rays. [0061] One possible approach is therefore to use the “local tomography” technique, which can be adapted to any type of trajectory. This is a relatively quick and straightforward algorithm that enables to obtain a three-dimensional tomographic image which represents the density derivative of the object, differently from more common reconstruction techniques which set out to calculate actual density. [0065] In addition to being extremely flexible in terms of the trajectories that may be used, this technique is also particularly efficient in computational terms. In a nutshell, it involves filtering each two-dimensional map (or radiographic image; i.e. each projection on the x-ray detector) using a filter which performs the second derivative, and then backprojecting the projections into the voxels of interest…In these cases, in order to check the characteristics searched for, the analysis of the three-dimensional images produced using the “local tomography” technique (which therefore represent density derivative) gives results which are analogous to those that would be obtained using a usual tomographic reconstruction).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Cersovsky with the teachings of Ursella to calculate the second derivative of the slice-level prediction score and perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest because "When using “local tomography” (i.e. a reconstruction of the density derivative of the object), high derivative values are obtained at the edge of the object by inputting the correct trajectory, whereas lower values are obtained by inputting an incorrect trajectory, due to a “blurring” of the edge itself. The mere fact that the object is surrounded by air and therefore has an outer surface can therefore be used as a priori information concerning the object. The measure of the value of the reconstructed gradient provides a good measure as to the plausibility that the reconstruction is correct, and therefore of the plausibility that the trajectory being processed is also correct" [0094].
Regarding claim 8, Cersovsky, Tran, and Ursella teach the device of claim 7. Cersovsky does not explicitly teach wherein the at least one processor is further configured to apply a threshold on the tomographic reconstruction to obtain a segmentation of the at least one region of interest.
Tran, in the same field of endeavor of medical image analysis, teaches wherein the at least one processor is further configured to apply a threshold on the tomographic reconstruction to obtain a segmentation of the at least one region of interest ([0144] When segmentation output is present, the segmentation may be displayed to participants through the graphical user interface. The localization is a 3D tensor that will be sent to the viewer component 701 and overlaid on thumbnail images/slices of the CTB study. The 3D tensor may be a static size. A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model. Large coloured segments visually indicate that there is a larger mass detected, whereas faint lines (thin segments) visually indicate that one or two consecutive slices with localization. The absence of coloured segments denotes a lack of radiological findings not detected by the model for these slices. [0292] In an example, the following postprocessing layers are included for segmentation: i) Mask generator (via threshold); ii) Default slice calculator and non-empty image calculator; and iii) 3D tensor to tensor of PNG bytes).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Cersovsky with the teachings of Tran to apply a threshold on the tomographic reconstruction to obtain a segmentation of the at least one region of interest so that "Through a user device, a user can select a desired viewing of at least one of the segmentation maps overlaid on a representative anatomical slice of the subject for at least one anatomical plane" [0013].
Regarding claim 10, Cersovsky and Tran teach the method of claim 6. Cersovsky further teaches a computer-implemented method for detecting at least one region of interest in a 3D image of a subject using a trained machine learning model obtained with the method ([0002] The systems, methods, and computer programs disclosed herein relate to detecting and/or recognizing a tissue type in an image of the tissue using machine learning techniques. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
wherein said method comprises: receiving a 3D image of a subject ([0011] Once the machine learning model is trained, it can be used to classify a new image into one of the two trained classes. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
obtain a plurality of stacks of 2D slices from said 3D image ([0033] In another aspect, the present disclosure provides a computer system comprising: a processor. [0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice));
for each obtained stack: feed each 2D slice into said trained machine learning model for detecting at least one region of interest in the 3D image of the subject ([0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue, [0111] generate a patch embedding for each received patch, [0112] aggregate patch embeddings into a bag-level representation, with each patch embedding assigned a learnable attention weight, and [0113] classify the bag-level-representation into one of at least two classes. [0143] The class may indicate whether the tissue shown in the image is a specific tissue or not. [0145] The class may indicate whether the tissue shown in the image is a lesion, or an edema or a tumor).
Cersovsky does not explicitly teach a computer-implemented method for detecting and reconstructing at least one region of interest in a 3D image; wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks; obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest; and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject.
Tran, in the same field of endeavor of medical image analysis, teaches a computer-implemented method for detecting and reconstructing at least one region of interest in a 3D image (Fig. 1. [0144] In this example, the CNN model 200 comprises an ensemble of five CNN components trained using five-fold cross-validation. In an example, the CNN model comprises three heads (modules): one for classification, one for left-right localization in 3D space, and one for segmentation in 3D space…A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model);
wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks ([0043] In embodiments, the neural network takes as input a standardised 3D volume. CT brain studies consist of a range of 2D images, and the number of these images varies a lot depending on the slice thickness, for example. Before these images are passed for training the AI model they need to be standardised by converting them into a fixed shape and voxel spacing. Describing from a different perspective, there is provided a plurality of images of a CT scan/study (such as e.g. from about 100 to about 450) to the neural network as an input to the entire machine learning pipeline. The neural network produces as output an indication of a plurality of visual findings being present in any one of the plurality of images. [0060] According to a further aspect, there is provided a method comprising: receiving, by a processor, the results of a step of analysing a series of anatomical images from a computed tomography (CT) scan of a head of a subject using one or more deep learning models trained to detect and localise in 3D space at least a first visual finding in anatomical images, wherein the results comprise a plurality of segmentation maps obtained with for at least one anatomical plane, wherein a segmentation map indicates the areas of a respective anatomical image where the first visual finding has been detected. [0069] The anatomical plane may be any one from the group consisting of: sagittal, coronal or transverse. [0201] 3D segmentation masks 362: used to generate the axial, sagittal, and coronal viewpoints (each view is saved separately as a list of PNGs 384));
obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest ([0147] In some embodiments, as illustrated in FIG. 3A, a CNN model 200 comprises a CNN encoder 302 that is configured to process input anatomical images 304. The CNN encoder 302 functions as an encoder, and is connected to a CNN decoder 306 which functions as a decoder. At least one MAP 308 comprises a two-dimensional array of values representing a probability that the corresponding pixel of an anatomical slice, e.g. the input anatomical images 204, which exhibits a visual finding, e.g. as identified by the CNN model 200. The MAP 308 can be in a particular anatomical plane, and can be overlaid over a respective anatomical image 204 in the same anatomical plane. The anatomical image 204 may also be a 2D anatomical slice in the anatomical plane generated from a 3D spatial model of the subject. [0320] FIG. 6A to 6D show exemplary interactive user interface screens of the viewer component 701 in accordance with an example embodiment. Clinical findings detected by the deep-learning model are listed and a segmentation MAP 502 (identified as one with most colour e.g. purple) is presented in relevant pathological slices. Finding likelihood scores and confidence intervals are also displayed under the NCCTB scan, illustrated as a sliding scale at the bottom of absent versus present);
and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject ([0144] The localization is a 3D tensor that will be sent to the viewer component 701 and overlaid on thumbnail images/slices of the CTB study. The 3D tensor may be a static size. A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model. Large coloured segments visually indicate that there is a larger mass detected, whereas faint lines (thin segments) visually indicate that one or two consecutive slices with localization. The absence of coloured segments denotes a lack of radiological findings not detected by the model for these slices. The rest of the slices are loaded on-demand, when/if a radiologist scrolls through the other slices/regions).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Cersovsky with the teachings of Tran to use stacks of slices in different orientations, obtain a slice level prediction that a 2D slice comprises a region of interest, and output a tomographic reconstruction of the region of interest because "A challenge is that predictions generated by deep learning models can be difficult to interpret by a user (such as, e.g., a clinician). Such models produce a score, probability or combination of scores for each class that they are trained to distinguish, which are often meaningful within a particular context related to the sensitivity/specificity of the deep learning model in detecting the clinically relevant feature associated with the class" [0008] and "This enables providing an ideal default view to the user for the user to confirm the presence of a finding detected by the model by displaying the medical prediction in the easiest manner to the user rather than the user having to search through the plurality of slices through a plurality of anatomical planes through a plurality of windows/user interface. More specifically, the system is configured to select an optimal combination for an ideal default view, of: i) “slice-method” (e.g. which one of the 3 anatomical planes: sagittal, coronal, transverse)” is the best for the user to visually confirm the presence/absence of the radiological finding; ii) slice image 204 (which image 204 out of X number of the stack of images 204 of a CT scan) is the default (or best) for the user to visually confirm the presence/absence of the radiological finding; and iii) window to display the slice image (e.g. which window from the brain, bone, soft tissue, bone, stroke, subdural windowing presets) is the best for the user to visually confirm the presence/absence of the radiological finding" [0318].
Cersovsky does not explicitly teach calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; and for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest.
Ursella, in the same field of endeavor of tomographic reconstruction teaches calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; and for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest ([0058] To perform an appropriate tomographic reconstruction, the detection step should be repeated on each object as many times as it takes until a sufficient number of radiographs (electronic two-dimensional maps) have been obtained by observing the object from different angles (generally, at least one hundred different radiographs should be detected, although fewer or more radiographs may be necessary depending on the given case). [0100] The device 1 comprises at least one computer (not shown) associated with the one or more detectors 5 to receive from them electronic data corresponding to two-dimensional pixel maps representative of the density of the object 2 crossed by the x-rays. [0061] One possible approach is therefore to use the “local tomography” technique, which can be adapted to any type of trajectory. This is a relatively quick and straightforward algorithm that enables to obtain a three-dimensional tomographic image which represents the density derivative of the object, differently from more common reconstruction techniques which set out to calculate actual density. [0065] In addition to being extremely flexible in terms of the trajectories that may be used, this technique is also particularly efficient in computational terms. In a nutshell, it involves filtering each two-dimensional map (or radiographic image; i.e. each projection on the x-ray detector) using a filter which performs the second derivative, and then backprojecting the projections into the voxels of interest…In these cases, in order to check the characteristics searched for, the analysis of the three-dimensional images produced using the “local tomography” technique (which therefore represent density derivative) gives results which are analogous to those that would be obtained using a usual tomographic reconstruction).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Cersovsky with the teachings of Ursella to calculate the second derivative of the slice-level prediction scores and perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest because "When using “local tomography” (i.e. a reconstruction of the density derivative of the object), high derivative values are obtained at the edge of the object by inputting the correct trajectory, whereas lower values are obtained by inputting an incorrect trajectory, due to a “blurring” of the edge itself. The mere fact that the object is surrounded by air and therefore has an outer surface can therefore be used as a priori information concerning the object. The measure of the value of the reconstructed gradient provides a good measure as to the plausibility that the reconstruction is correct, and therefore of the plausibility that the trajectory being processed is also correct" [0094].
Regarding claim 11, Cersovsky and Tran teach the method of claim 6. Cersovsky further teaches a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method for training and a method for detecting at least one region of interest in a 3D image of a subject using a trained machine learning model obtained with said method for training ([0002] The systems, methods, and computer programs disclosed herein relate to detecting and/or recognizing a tissue type in an image of the tissue using machine learning techniques. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
wherein said method for detecting comprises: receiving a 3D image of a subject ([0011] Once the machine learning model is trained, it can be used to classify a new image into one of the two trained classes);
obtain a plurality of stacks of 2D slices from said 3D image ([0033] In another aspect, the present disclosure provides a computer system comprising: a processor. [0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice));
for each obtained stack: feed each 2D slice into said trained machine learning model for detecting at least one region of interest in the 3D image of the subject ([0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue, [0111] generate a patch embedding for each received patch, [0112] aggregate patch embeddings into a bag-level representation, with each patch embedding assigned a learnable attention weight, and [0113] classify the bag-level-representation into one of at least two classes. [0143] The class may indicate whether the tissue shown in the image is a specific tissue or not. [0145] The class may indicate whether the tissue shown in the image is a lesion, or an edema or a tumor).
Cersovsky does not explicitly teach a method for detecting and reconstructing at least one region of interest in a 3D image; wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks; obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest; and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject.
Tran, in the same field of endeavor of medical image analysis, teaches a method for detecting and reconstructing at least one region of interest in a 3D image (Fig. 1. [0144] In this example, the CNN model 200 comprises an ensemble of five CNN components trained using five-fold cross-validation. In an example, the CNN model comprises three heads (modules): one for classification, one for left-right localization in 3D space, and one for segmentation in 3D space…A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model);
wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks ([0043] In embodiments, the neural network takes as input a standardised 3D volume. CT brain studies consist of a range of 2D images, and the number of these images varies a lot depending on the slice thickness, for example. Before these images are passed for training the AI model they need to be standardised by converting them into a fixed shape and voxel spacing. Describing from a different perspective, there is provided a plurality of images of a CT scan/study (such as e.g. from about 100 to about 450) to the neural network as an input to the entire machine learning pipeline. The neural network produces as output an indication of a plurality of visual findings being present in any one of the plurality of images. [0060] According to a further aspect, there is provided a method comprising: receiving, by a processor, the results of a step of analysing a series of anatomical images from a computed tomography (CT) scan of a head of a subject using one or more deep learning models trained to detect and localise in 3D space at least a first visual finding in anatomical images, wherein the results comprise a plurality of segmentation maps obtained with for at least one anatomical plane, wherein a segmentation map indicates the areas of a respective anatomical image where the first visual finding has been detected. [0069] The anatomical plane may be any one from the group consisting of: sagittal, coronal or transverse. [0201] 3D segmentation masks 362: used to generate the axial, sagittal, and coronal viewpoints (each view is saved separately as a list of PNGs 384));
obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest ([0147] In some embodiments, as illustrated in FIG. 3A, a CNN model 200 comprises a CNN encoder 302 that is configured to process input anatomical images 304. The CNN encoder 302 functions as an encoder, and is connected to a CNN decoder 306 which functions as a decoder. At least one MAP 308 comprises a two-dimensional array of values representing a probability that the corresponding pixel of an anatomical slice, e.g. the input anatomical images 204, which exhibits a visual finding, e.g. as identified by the CNN model 200. The MAP 308 can be in a particular anatomical plane, and can be overlaid over a respective anatomical image 204 in the same anatomical plane. The anatomical image 204 may also be a 2D anatomical slice in the anatomical plane generated from a 3D spatial model of the subject. [0320] FIG. 6A to 6D show exemplary interactive user interface screens of the viewer component 701 in accordance with an example embodiment. Clinical findings detected by the deep-learning model are listed and a segmentation MAP 502 (identified as one with most colour e.g. purple) is presented in relevant pathological slices. Finding likelihood scores and confidence intervals are also displayed under the NCCTB scan, illustrated as a sliding scale at the bottom of absent versus present);
and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject ([0144] The localization is a 3D tensor that will be sent to the viewer component 701 and overlaid on thumbnail images/slices of the CTB study. The 3D tensor may be a static size. A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model. Large coloured segments visually indicate that there is a larger mass detected, whereas faint lines (thin segments) visually indicate that one or two consecutive slices with localization. The absence of coloured segments denotes a lack of radiological findings not detected by the model for these slices. The rest of the slices are loaded on-demand, when/if a radiologist scrolls through the other slices/regions).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the product of Cersovsky with the teachings of Tran to use stacks of slices in different orientations, obtain a slice level prediction that a 2D slice comprises a region of interest, and output a tomographic reconstruction of the region of interest because "A challenge is that predictions generated by deep learning models can be difficult to interpret by a user (such as, e.g., a clinician). Such models produce a score, probability or combination of scores for each class that they are trained to distinguish, which are often meaningful within a particular context related to the sensitivity/specificity of the deep learning model in detecting the clinically relevant feature associated with the class" [0008] and "This enables providing an ideal default view to the user for the user to confirm the presence of a finding detected by the model by displaying the medical prediction in the easiest manner to the user rather than the user having to search through the plurality of slices through a plurality of anatomical planes through a plurality of windows/user interface. More specifically, the system is configured to select an optimal combination for an ideal default view, of: i) “slice-method” (e.g. which one of the 3 anatomical planes: sagittal, coronal, transverse)” is the best for the user to visually confirm the presence/absence of the radiological finding; ii) slice image 204 (which image 204 out of X number of the stack of images 204 of a CT scan) is the default (or best) for the user to visually confirm the presence/absence of the radiological finding; and iii) window to display the slice image (e.g. which window from the brain, bone, soft tissue, bone, stroke, subdural windowing presets) is the best for the user to visually confirm the presence/absence of the radiological finding" [0318].
Cersovsky does not explicitly teach calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest.
Ursella, in the same field of endeavor of tomographic reconstruction, teaches calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest ([0058] To perform an appropriate tomographic reconstruction, the detection step should be repeated on each object as many times as it takes until a sufficient number of radiographs (electronic two-dimensional maps) have been obtained by observing the object from different angles (generally, at least one hundred different radiographs should be detected, although fewer or more radiographs may be necessary depending on the given case). [0100] The device 1 comprises at least one computer (not shown) associated with the one or more detectors 5 to receive from them electronic data corresponding to two-dimensional pixel maps representative of the density of the object 2 crossed by the x-rays. [0061] One possible approach is therefore to use the “local tomography” technique, which can be adapted to any type of trajectory. This is a relatively quick and straightforward algorithm that enables to obtain a three-dimensional tomographic image which represents the density derivative of the object, differently from more common reconstruction techniques which set out to calculate actual density. [0065] In addition to being extremely flexible in terms of the trajectories that may be used, this technique is also particularly efficient in computational terms. In a nutshell, it involves filtering each two-dimensional map (or radiographic image; i.e. each projection on the x-ray detector) using a filter which performs the second derivative, and then backprojecting the projections into the voxels of interest…In these cases, in order to check the characteristics searched for, the analysis of the three-dimensional images produced using the “local tomography” technique (which therefore represent density derivative) gives results which are analogous to those that would be obtained using a usual tomographic reconstruction).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the product of Cersovsky with the teachings of Ursella to calculate the second derivative of the slice-level prediction scores and perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest because "When using “local tomography” (i.e. a reconstruction of the density derivative of the object), high derivative values are obtained at the edge of the object by inputting the correct trajectory, whereas lower values are obtained by inputting an incorrect trajectory, due to a “blurring” of the edge itself. The mere fact that the object is surrounded by air and therefore has an outer surface can therefore be used as a priori information concerning the object. The measure of the value of the reconstructed gradient provides a good measure as to the plausibility that the reconstruction is correct, and therefore of the plausibility that the trajectory being processed is also correct" [0094].
Regarding claim 12, Cersovsky and Tran teach the method of claim 6. Cersovsky further teaches a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method for training and a method for detecting at least one region of interest in a 3D image of a subject using a trained machine learning model obtained with said method for training ([0002] The systems, methods, and computer programs disclosed herein relate to detecting and/or recognizing a tissue type in an image of the tissue using machine learning techniques. [0082] The term “image” as used herein means a data structure that represents a spatial distribution of a physical signal. The spatial distribution may be of any dimension, for example 2D, 3D, 4D or any higher dimension);
wherein said method for detecting comprises: receiving a 3D image of a subject ([0011] Once the machine learning model is trained, it can be used to classify a new image into one of the two trained classes);
obtain a plurality of stacks of 2D slices from said 3D image ([0033] In another aspect, the present disclosure provides a computer system comprising: a processor. [0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue. [0118] Patches can be created from 3D images (or even higher-dimensional images) by cutting slices of a defined thickness (e.g., a voxel) from the 3D image. A patch can be a single CT or MRI slice, such as an axial slice, a sagittal slice, or a coronal slice));
for each obtained stack: feed each 2D slice into said trained machine learning model for detecting at least one region of interest in the 3D image of the subject ([0109] The machine learning model is configured and trained to: [0110] receive a number of patches generated from an image of a tissue, [0111] generate a patch embedding for each received patch, [0112] aggregate patch embeddings into a bag-level representation, with each patch embedding assigned a learnable attention weight, and [0113] classify the bag-level-representation into one of at least two classes. [0143] The class may indicate whether the tissue shown in the image is a specific tissue or not. [0145] The class may indicate whether the tissue shown in the image is a lesion, or an edema or a tumor).
Cersovsky does not explicitly teach a method for detecting and reconstructing at least one region of interest in a 3D image; wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks; obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest; and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject.
Tran, in the same field of endeavor of medical image analysis, teaches a method for detecting and reconstructing at least one region of interest in a 3D image (Fig. 1. [0144] In this example, the CNN model 200 comprises an ensemble of five CNN components trained using five-fold cross-validation. In an example, the CNN model comprises three heads (modules): one for classification, one for left-right localization in 3D space, and one for segmentation in 3D space…A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model);
wherein an orientation of the 2D slices in one stack is different from an orientation of the 2D slices in another stack of the plurality of stacks ([0043] In embodiments, the neural network takes as input a standardised 3D volume. CT brain studies consist of a range of 2D images, and the number of these images varies a lot depending on the slice thickness, for example. Before these images are passed for training the AI model they need to be standardised by converting them into a fixed shape and voxel spacing. Describing from a different perspective, there is provided a plurality of images of a CT scan/study (such as e.g. from about 100 to about 450) to the neural network as an input to the entire machine learning pipeline. The neural network produces as output an indication of a plurality of visual findings being present in any one of the plurality of images. [0060] According to a further aspect, there is provided a method comprising: receiving, by a processor, the results of a step of analysing a series of anatomical images from a computed tomography (CT) scan of a head of a subject using one or more deep learning models trained to detect and localise in 3D space at least a first visual finding in anatomical images, wherein the results comprise a plurality of segmentation maps obtained with for at least one anatomical plane, wherein a segmentation map indicates the areas of a respective anatomical image where the first visual finding has been detected. [0069] The anatomical plane may be any one from the group consisting of: sagittal, coronal or transverse. [0201] 3D segmentation masks 362: used to generate the axial, sagittal, and coronal viewpoints (each view is saved separately as a list of PNGs 384));
obtaining slice-level prediction scores for each 2D slice; wherein the slice-level prediction score represents a probability that said input 2D slice comprises at least one portion of said region of interest ([0147] In some embodiments, as illustrated in FIG. 3A, a CNN model 200 comprises a CNN encoder 302 that is configured to process input anatomical images 304. The CNN encoder 302 functions as an encoder, and is connected to a CNN decoder 306 which functions as a decoder. At least one MAP 308 comprises a two-dimensional array of values representing a probability that the corresponding pixel of an anatomical slice, e.g. the input anatomical images 204, which exhibits a visual finding, e.g. as identified by the CNN model 200. The MAP 308 can be in a particular anatomical plane, and can be overlaid over a respective anatomical image 204 in the same anatomical plane. The anatomical image 204 may also be a 2D anatomical slice in the anatomical plane generated from a 3D spatial model of the subject. [0320] FIG. 6A to 6D show exemplary interactive user interface screens of the viewer component 701 in accordance with an example embodiment. Clinical findings detected by the deep-learning model are listed and a segmentation MAP 502 (identified as one with most colour e.g. purple) is presented in relevant pathological slices. Finding likelihood scores and confidence intervals are also displayed under the NCCTB scan, illustrated as a sliding scale at the bottom of absent versus present);
and at least one output providing said tomographic reconstruction of at least one region of interest present in a 3D image of a subject ([0144] The localization is a 3D tensor that will be sent to the viewer component 701 and overlaid on thumbnail images/slices of the CTB study. The 3D tensor may be a static size. A slice from a 3D tensor can be rendered by the viewer component 701 and scrolling through slices of the 3D tensors in the viewer component 701 is provided. The 3D tensor is reconstructed and all the required slices of the CTB study are stored in all required axes. The viewer component 701 provides a slice scrollbar graphical user interface component. The slice scrollbar may be oriented vertically, and comprises a rectangular outline. Coloured segments in the slice scrollbar, for example, purple colour, indicates that these slices have radiological findings predicted by the model. Large coloured segments visually indicate that there is a larger mass detected, whereas faint lines (thin segments) visually indicate that one or two consecutive slices with localization. The absence of coloured segments denotes a lack of radiological findings not detected by the model for these slices. The rest of the slices are loaded on-demand, when/if a radiologist scrolls through the other slices/regions).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the medium of Cersovsky with the teachings of Tran to use stacks of slices in different orientations, obtain a slice level prediction that a 2D slice comprises a region of interest, and output a tomographic reconstruction of the region of interest because "A challenge is that predictions generated by deep learning models can be difficult to interpret by a user (such as, e.g., a clinician). Such models produce a score, probability or combination of scores for each class that they are trained to distinguish, which are often meaningful within a particular context related to the sensitivity/specificity of the deep learning model in detecting the clinically relevant feature associated with the class" [0008] and "This enables providing an ideal default view to the user for the user to confirm the presence of a finding detected by the model by displaying the medical prediction in the easiest manner to the user rather than the user having to search through the plurality of slices through a plurality of anatomical planes through a plurality of windows/user interface. More specifically, the system is configured to select an optimal combination for an ideal default view, of: i) “slice-method” (e.g. which one of the 3 anatomical planes: sagittal, coronal, transverse)” is the best for the user to visually confirm the presence/absence of the radiological finding; ii) slice image 204 (which image 204 out of X number of the stack of images 204 of a CT scan) is the default (or best) for the user to visually confirm the presence/absence of the radiological finding; and iii) window to display the slice image (e.g. which window from the brain, bone, soft tissue, bone, stroke, subdural windowing presets) is the best for the user to visually confirm the presence/absence of the radiological finding" [0318].
Cersovsky does not explicitly teach calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest.
Ursella, in the same field of endeavor of tomographic reconstruction teaches calculate the second derivative of the slice-level prediction scores along an axis of the orientation of the stack; for each orientation of the plurality of stacks and 2D slice, perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest ([0058] To perform an appropriate tomographic reconstruction, the detection step should be repeated on each object as many times as it takes until a sufficient number of radiographs (electronic two-dimensional maps) have been obtained by observing the object from different angles (generally, at least one hundred different radiographs should be detected, although fewer or more radiographs may be necessary depending on the given case). [0100] The device 1 comprises at least one computer (not shown) associated with the one or more detectors 5 to receive from them electronic data corresponding to two-dimensional pixel maps representative of the density of the object 2 crossed by the x-rays. [0061] One possible approach is therefore to use the “local tomography” technique, which can be adapted to any type of trajectory. This is a relatively quick and straightforward algorithm that enables to obtain a three-dimensional tomographic image which represents the density derivative of the object, differently from more common reconstruction techniques which set out to calculate actual density. [0065] In addition to being extremely flexible in terms of the trajectories that may be used, this technique is also particularly efficient in computational terms. In a nutshell, it involves filtering each two-dimensional map (or radiographic image; i.e. each projection on the x-ray detector) using a filter which performs the second derivative, and then backprojecting the projections into the voxels of interest…In these cases, in order to check the characteristics searched for, the analysis of the three-dimensional images produced using the “local tomography” technique (which therefore represent density derivative) gives results which are analogous to those that would be obtained using a usual tomographic reconstruction).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the medium of Cersovsky with the teachings of Ursella to calculate the second derivative of the slice-level prediction scores and perform backprojection using said calculated second derivative and generate a tomographic reconstruction of the at least one region of interest because "When using “local tomography” (i.e. a reconstruction of the density derivative of the object), high derivative values are obtained at the edge of the object by inputting the correct trajectory, whereas lower values are obtained by inputting an incorrect trajectory, due to a “blurring” of the edge itself. The mere fact that the object is surrounded by air and therefore has an outer surface can therefore be used as a priori information concerning the object. The measure of the value of the reconstructed gradient provides a good measure as to the plausibility that the reconstruction is correct, and therefore of the plausibility that the trajectory being processed is also correct" [0094].
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Cersovsky in view of Tran, Ursella, and Taubmann (US20210093272A1).
Regarding claim 9, Cersovsky, Tran, and Ursella teach the device of claim 7. Cersovsky does not explicitly teach wherein the at least one processor is further configured to calculate a volume-level score by taking the maximum over the slice-level scores calculated for each stack.
Taubmann, in the same field of endeavor of medical image analysis, teaches wherein the at least one processor is further configured to calculate a volume-level score by taking the maximum over the slice-level scores calculated for each stack ([0107] In concrete embodiments, if the input data comprises sectional or slice images, each sectional or slice image may be separately input to the trained function, yielding a corresponding probability image. [0155] In a third step of the evaluation algorithm, the probability images are interpreted to determine the presence of a finding and, if a finding is present, corresponding finding information. In this embodiment, the maximum of all probabilities in the probability images is compared to a first threshold. If the maximum probability exceeds this first threshold, the detection is determined as positive. Additionally, the location, where the maximum probability has been determined, is used as a location of the finding, and the sub-set of the pixels around this location whose predicted probabilities exceed a second threshold, is interpreted as defining the extensions of the finding, in particular its dimensions).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the device of Cersovsky with the teachings of Taubmann to calculate a volume level score by taking the maximum over the slice-level scores calculated for each stack because "the position of the data point of the maximum probability can be determined as the location of the finding and/or an area surrounding the data point of the maximum probability" [0108].
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wei (US12456197B2) teaches machine learning to determine the probability of an image slice containing a lesion.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jacqueline R Zak whose telephone number is (571)272-4077. The examiner can normally be reached M-F 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACQUELINE R ZAK/Examiner, Art Unit 2666
/EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666