Prosecution Insights
Last updated: October 04, 2026
Application No. 18/267,884

METHODS FOR TRAINING A CNN AND FOR PROCESSING AN INPUTTED PERFUSION SEQUENCE USING SAID CNN

Final Rejection §102§103
Filed
Jun 16, 2023
Priority
Dec 18, 2020 — EU 20306625.3 +2 more
Examiner
LAU, KAITLYN RENEE
Art Unit
2148
Tech Center
2100 — Computer Architecture & Software
Assignee
Guerbet
OA Round
2 (Final)
60%
Grant Probability
Moderate
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 60% of resolved cases
60%
Career Allowance Rate
6 granted / 10 resolved
+5.0% vs TC avg
Strong +67% interview lift
Without
With
+66.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
27 currently pending
Career history
40
Total Applications
across all art units

Statute-Specific Performance

§101
28.8%
-11.2% vs TC avg
§103
34.3%
-5.7% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
21.9%
-18.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 10 resolved cases

Office Action

§102 §103
DETAILED ACTION This action is in response to the application filed 07/15/2026. Claims 1-16 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 3 and 9 objected to because of the following informalities: Regarding claims 3 and 9, these claims use the term “n+1-dimensional feature maps” while claims 1, 7, 10, and 11 use the term “n+1-dimensional features maps.” For consistency, all claims should use either “n+1-dimensional feature maps” or “n+1 dimensional features maps” Appropriate correction is required. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-5, 8-11, and 15-16 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Golden et al. (US 2018/0218502 A1) (hereafter referred to as Golden). Regarding Claim 1, Golden teaches A method for processing an inputted perfusion sequence presenting n≥3 dimensions including at least two spatial dimensions and one temporal dimension, wherein the inputted perfusion sequence is a temporal sequence of medical images depicting a passage of fluid through a tissue, by means of a convolutional neural network, CNN (Golden, page 56, paragraph 0070, “The trained CNN model may have been trained on one or more of functional cardiac images, myocardial delayed enhancement images or myocardial perfusion images” where “Receiving learning data may include receiving image data which may include at least one of steady-state free precession (SSFP) magnetic resonance imaging (MRI) data or 4D flow MRI data” (Golden, page 53, paragraph 0042) and where “SSFP cine studies contain of 4 dimensions of data (3 space, 1 time), and 4D Flow studies contain 5 dimensions of data (3 space, 1 time, 4 channels of information)” (Golden, page 64, paragraph 0228) where “In 4D Flow, flow velocity information is available at every time point of the acquisition for every patient. In order to make full use of this information, the standard deviation along the time axis may be computed at every voxel of the 3D image. Standard deviation magnitude is associated with the amount of blood flow variation of that pixel over the course of one heartbeat. This standard deviation image is then normalized according the previously described normalization pipeline: resizing, clipping, scaling, histogram equalization, centering. Note that several other approaches can be considered to encode the temporal information of the flow data” (Golden, page 67, paragraph 0260) Examiner notes that the space dimensions are the spatial dimensions and the time dimension is the temporal dimension. Examiner further notes that the image is a perfusion image which is used to encode the temporal information of the blood flowing through a body.), comprising an encoder branch, a decoder branch and skip connections between the encoder branch and the decoder branch (Golden, page 70, paragraph 0285, “FIG. 36 shows a schematic representation of a fully convolutional encoder-decoder architecture with skip connections that utilizes a smaller expanding path than contracting path” and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the encoder branch is highlighted in FIG. 36 in the left box and the decoder branch is highlighted in the right circle.), the method comprising the implementation, by a data processor of a second server, of steps of (Golden, page 69, paragraph 0275, “The system memory 2508 may also include communications programs 3540, for example a server and/or a Web client or browser for permitting the processor-based device 3504 to access and exchange data with other systems such as user computing systems.”): extracting, using the encoder branch of the CNN, a plurality of initial n+1-dimensional features maps representative of the inputted perfusion sequence at different scales, and projecting, using the skip connections of the CNN, each one of the plurality of initial n+1-dimensional features maps into one of a plurality of initial n-dimensional feature maps (Golden, page 58, paragraph 0126, “The network 600 is configured such that, after every pooling layer 608, the number of feature maps doubles and the spatial resolution is halved. After every upsampling layer 610, the number of feature maps is halved and the spatial resolution is doubled. With the scheme, the number of feature maps for each layer across the network 600 can be fully described by the number (e.g., between 1 and 2000 feature maps) in the first layer” where “ENet utilizes an expanding path that is smaller than its contracting path. ENet also makes use of bottleneck modules, which are convolutions with a small receptive field that are applied in order to project the feature maps into a lower dimensional space” where “subsequent to each upsampling layer, the CNN model may include a concatenation of feature maps from a corresponding layer in the contracting path through a skip connection” (Golden, page 52, paragraph 0028) and “U-Net, originally developed for use in the biomedical community where there are often fewer training images and even finer resolution is required, added the use of skip connections between the contracting and expanding paths to preserve details” (Golden, page 70, paragraph 0286) and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the plurality of n+1-dimensional features maps are the feature maps within the convolutional layers of the encoder branch shown in the left box. Examiner further notes that the n-dimensional feature maps are the feature maps that are within the convolutional layer that the skip connection is connected or projected to. Examiner additionally notes that the feature maps are projected into a lower dimensional space and thus are n-dimensional feature maps. Examiner further notes that the different scales can be seen in Fig. 36 by the different sizes of layers.); generating, using said decoder branch of the CNN, a plurality of enriched n- dimensional feature maps also representative of the inputted perfusion sequence at different scales, an enriched n-dimensional feature map at a particular scale incorporating information from the initial n-dimensional feature maps at smaller or equal scale (Golden, page 59, paragraph 0130, “Upsampling the activation volumes back to the original resolution is necessary in a fully convolutional network for pixel-wise segmentation. To increase the resolution of the activation volumes in the network, some systems may use an upsampling operation, then a 2x2 convolution, then a concatenation of feature maps from the corresponding contracting layer through a skip connection, and finally two 3x3 convolutions” where “ENet utilizes an expanding path that is smaller than its contracting path. ENet also makes use of bottleneck modules, which are convolutions with a small receptive field that are applied in order to project the feature maps into a lower dimensional space” and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the enriched n-dimensional feature maps are the upsampled activation volumes); generating at least one quantitative map of the inputted perfusion sequence from the enriched n-dimensional feature maps at the largest scale among the different scales (Golden, page 60, paragraph 0175, “Unlike previous models which were only concerned with two classes for a cell discrimination task, foreground and background, the SSFP model disclosed herein attempts to distinguish four classes, namely, background, LV Endocardium, LV Epicardium and RV Endocardium. To accomplish this, the network output may include three probability maps, one for each non-background class” and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the quantitative map is the probability map as shown in Figure 36.). wherein the quantitative map is an image representing spatial values of a parameter characterizing said perfusion over the tissue (Golden, page 61, paragraph 0188, “Inference is performed at the slice locations and time points in the requested batch. At 912, a forward pass through the network is computed. For a given image, the model generates per-class probabilities for each pixel during the forward pass, which results in a set of probability maps, one for each class, with values ranging from 0 to 1. The probability maps are transformed into a single label mask by setting the class of each pixel to the class with the highest label map probability.” Examiner notes that the quantitative map is the probability map and the spatial values of a parameter charactering perfusion is the slice locations in the batch of images.) Regarding Claim 2, Golden teaches A method according to claim 1, wherein, for each enriched n- dimensional feature map, an initial n-dimensional feature map of the same scale is provided from the encoder branch to the decoder branch via a dedicated skip connection (Golden, page 58, paragraph 0126, “The network 600 is configured such that, after every pooling layer 608, the number of feature maps doubles and the spatial resolution is halved. After every upsampling layer 610, the number of feature maps is halved and the spatial resolution is doubled. With the scheme, the number of feature maps for each layer across the network 600 can be fully described by the number (e.g., between 1 and 2000 feature maps) in the first layer” where “ENet utilizes an expanding path that is smaller than its contracting path. ENet also makes use of bottleneck modules, which are convolutions with a small receptive field that are applied in order to project the feature maps into a lower dimensional space” and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the plurality of n+1-dimensional features maps are the feature maps within the convolutional layers of the encoder branch shown in the left box. Examiner further notes that the n-dimensional feature maps are the feature maps that are within the convolutional layer that the skip connection is connected or projected to. Examiner additionally notes that the feature maps are projected into a lower dimensional space and thus are n-dimensional feature maps.). Regarding Claim 3, Golden teaches, The method according to one of claim 1,wherein, at step (c), the enriched n-dimensional feature map generated at the smallest scale among the different scales is generated from the initial n+1 -dimensional feature map at the smallest scale among the different scales, and each enriched n-dimensional feature map generated at another scale than the smallest scale is generated from the initial n-dimensional feature map at the same another scale and a enriched n-dimensional feature map at a smaller scale than the another scale (Golden, page 37, Figure 36, PNG media_image2.png 480 805 media_image2.png Greyscale Examiner notes that the circled layer is the enriched n-dimensional feature map at the smallest scale and the boxed layer is the initial n+1-dimensional feature map at the smallest scale. Examiner further notes that the other upsample layers or enriched feature maps are generated from the convolution layer or n-dimensional feature map at the same scale as well as upsample layers of a smaller scale.). Regarding Claim 4, Golden teaches The method according to claim 1,further comprising obtaining the perfusion sequence by stacking a plurality of successive images of a perfusion (Golden, page 73, paragraph 0314, “Perfusion imaging using gadolinium-based contrast is used to identify biomarkers of coronary stenosis. Late gadolinium enhancement imaging, also using gadolinium based contrast, is used to assess myocardial infarction. In all of these imaging protocols, and in others, the anatomical orientations and the need for contouring tend to be similar. Images are typically acquired both in short axis orientations, in which the imaging plane is parallel to the short axis of the left ventricle, and in long axis orientations, in which the imaging plane is parallel to the long axis of the left ventricle” where “The short axis (SAX) view, which consists of a series of slices along the long axis of the left ventricle. Each slice is in the plane of the short axis of the left ventricle, which is orthogonal to the ventricle's long axis;” (Golden, page 49, paragraph 0003) and where “Example LV endocardium contours are shown as images 100a-100k in FIG. 1, which shows the contours at a single time point over a full SAX stack. From 100a to 100k, the slices proceed from the apex of the left ventricle to the base of the left ventricle” (Golden, page 49, paragraph 0008). Examiner notes that the perfusion imaging uses SAX views which are stacks of successive images. ). Regarding Claim 5, Golden teaches The method according to claim 4, wherein said successive images of a perfusion are acquired by a medical imaging device connected to the second server (Golden, page 61, paragraph 0185, “Inference is the process of utilizing a trained model for prediction on new data. In at least some implementations, a web application (or "web app") may be used for inference. FIG. 9 displays an example pipeline or process 900 by which predictions may be made on new SSFP studies. At 902, after a user has loaded a study in the web application, the user may invoke the inference service (e.g., by clicking a "generate missing contours" icon), which automatically generates any missing (not yet created) contours. Such contours may include LV Endo, LV Epi, or RV Endo, for example. In at least some implementations, inference may be invoked automatically when the study is either loaded by the user in the application or when the study is first uploaded by the user to a server. If inference is performed at upload time, the predictions may be stored in a nontransitory processor-readable storage medium at that time but not displayed until the user opens the study” where “Depending on the type of acquisition, these views may be captured directly in the scanner (e.g., steady-state free precession (SSFP) MRI) or may be created via multi planar reconstructions (MPRs) of a volume aligned in a different orientation (such as the axial, sagittal or coronal planes, e.g., 4D Flow MRI). The SAX view has multiple spatial slices, usually covering the entire volume of the heart, but the 2CH, 3CH and 4CH views often only have a single spatial slice. All series are cine, and have multiple time points encompassing a complete cardiac cycle” (Golden, page 49, paragraph 0004). Examiner notes that an MRI acquires the successive images and is connected to a server through inferencing.). Regarding Claim 8, Golden teaches The method according to claim 1,wherein said CNN is fully convolutional (Golden, page 51, paragraph 0028, “A machine learning system may be summarized as including at least one nontransitory processor-readable storage medium that stores at least one of processor-executable instructions or data; and at least one processor communicably coupled to the at least one nontransitory processor-readable storage medium, the at least one processor: … trains a fully convolutional neural network (CNN) model to segment at least one part of the anatomical structure utilizing tl1e received learning data”). Regarding Claim 9, Golden teaches The method according to claim 1, wherein the number of spatial dimensions is n-1, wherein the at least one quantitative map only presents the spatial dimensions of the inputted perfusion sequence, (Golden, page 64, paragraph 0228, “SSFP cine studies contain of 4 dimensions of data (3 space, 1 time), and 4D Flow studies contain 5 dimensions of data (3 space, 1 time, 4 channels of information). These 4 channels of information are the anatomy (i.e. signal intensity), x axis phase, y axis phase, and z axis phase. The simplest way to build a model uses only signal intensities at each 3D spatial point, and does not incorporate the temporal information or, for 4D Flow, the flow information” where “To incorporate time data, time may be added as an additional ‘channel’ to the intensity data. In such implementations, the model then takes as input 3D data blobs of shape (X, Y, NTIMES) or 4D data blobs of shape (X,Y,Z, NTIMES) where NTIMES is the number of time points to include. This may be all time points, or a few time points surrounding the time point of interest. If all time points are include, it may be desirable or necessary to pad the data with a few ‘wrapped around’ time points, since time represents a cardiac cycle and is intrinsically cyclical. The model then either involve 2D/3D convolutions with time points as additional ‘channels’ of the data, or 3D/4D convolutions” (Golden, page 65, paragraph 0230) and “generate per-class probabilities for each pixel of each image of the image data, each class corresponding to one of a plurality of parts of the anatomical structure represented by the image data; and for each image of the image data, generates a probability map for each of the plurality of classes using the generated per-class probabilities” (Golden, page 53, paragraph 0043). Examiner notes that the number of spatial dimensions is 3 and since the probability map presents a plurality of classes corresponding to parts of the structure in the image data, it only presents spatial dimensions of the perfusion sequence.), wherein the initial n+1-dimensional feature maps present n+1 dimensions consisting of the n-1 spatial dimensions, the temporal dimension and a semantic depth dimension (Golden, page 64, paragraph 0228, “SSFP cine studies contain of 4 dimensions of data (3 space, 1 time), and 4D Flow studies contain 5 dimensions of data (3 space, 1 time, 4 channels of information). These 4 channels of information are the anatomy (i.e. signal intensity), x axis phase, y axis phase, and z axis phase. The simplest way to build a model uses only signal intensities at each 3D spatial point, and does not incorporate the temporal information or, for 4D Flow, the flow information.” Examiner notes that the spatial dimensions are the 3 space, the temporal dimension is the 1 time and the semantic depth dimension is the number of channels of information as interpreted from page 21, 1st paragraph of the instant specification.), and wherein the initial n-dimensional feature maps and the enriched n-dimensional feature maps present n dimensions consisting of the n-1 spatial dimensions and the semantic depth dimensions (Golden, page 64, paragraph 0228, “SSFP cine studies contain of 4 dimensions of data (3 space, 1 time), and 4D Flow studies contain 5 dimensions of data (3 space, 1 time, 4 channels of information). These 4 channels of information are the anatomy (i.e. signal intensity), x axis phase, y axis phase, and z axis phase. The simplest way to build a model uses only signal intensities at each 3D spatial point, and does not incorporate the temporal information or, for 4D Flow, the flow information” where “To incorporate time data, time may be added as an additional ‘channel’ to the intensity data. In such implementations, the model then takes as input 3D data blobs of shape (X, Y, NTIMES) or 4D data blobs of shape (X,Y,Z, NTIMES) where NTIMES is the number of time points to include. This may be all time points, or a few time points surrounding the time point of interest. If all time points are include, it may be desirable or necessary to pad the data with a few ‘wrapped around’ time points, since time represents a cardiac cycle and is intrinsically cyclical. The model then either involve 2D/3D convolutions with time points as additional ‘channels’ of the data, or 3D/4D convolutions” (Golden, page 65, paragraph 0230) and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that by building a model without the time or temporal dimension, the feature maps in the model depicted in Fig. 36 present n dimensions consisting of spatial dimensions and a semantic depth.). Regarding Claim 10, Golden teaches, The method according to claim 9, wherein the projecting using said skip connections comprises performing a reduction in the temporal dimension using a pooling operation on the initial n+1 dimensional features maps to project each one of the plurality of initial n+1-dimensional features maps into the plurality of initial n-dimensional feature maps (Golden, page 64, paragraph 0228, “SSFP cine studies contain of 4 dimensions of data (3 space, 1 time), and 4D Flow studies contain 5 dimensions of data (3 space, 1 time, 4 channels of information). These 4 channels of information are the anatomy (i.e. signal intensity), x axis phase, y axis phase, and z axis phase. The simplest way to build a model uses only signal intensities at each 3D spatial point, and does not incorporate the temporal information or, for 4D Flow, the flow information” where “To incorporate time data, time may be added as an additional ‘channel’ to the intensity data. In such implementations, the model then takes as input 3D data blobs of shape (X, Y, NTIMES) or 4D data blobs of shape (X,Y,Z, NTIMES) where NTIMES is the number of time points to include. This may be all time points, or a few time points surrounding the time point of interest. If all time points are include, it may be desirable or necessary to pad the data with a few ‘wrapped around’ time points, since time represents a cardiac cycle and is intrinsically cyclical. The model then either involve 2D/3D convolutions with time points as additional ‘channels’ of the data, or 3D/4D convolutions” (Golden, page 65, paragraph 0230) where “FIG. 36 shows a schematic representation of a fully convolutional encoder-decoder architecture with skip connections that utilizes a smaller expanding path than a contracting path” (Golden, page 70, paragraph 0285) where “Furthermore, the ENet authors show that the primary purpose of the expanding path in FCNs is to upsample and fine-tune the details learned by the contracting path rather than to learn complicated upsampling features; hence, ENet utilizes an expanding path that is smaller than its contracting path….ENet also uses a path parallel to the bottleneck path that solely includes zero or more pooling layers to directly pass information from a higher resolution layer to the lower resolution layers” (Golden, page 70, paragraph 0287) and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that by building a model without the time or temporal dimension and passing information via a pooling layer from a higher resolution to a lower resolution, there is a reduction in the temporal dimension from the initial n+1 dimensional feature maps to the n-dimensional feature maps.) Regarding Claim 11, Golden teaches, A method for training a convolution neural network, CNN, for processing an inputted perfusion sequence presenting n≥3 dimensions including at least two spatial dimensions and one temporal dimension, (Golden, page 56, paragraph 0070, “The trained CNN model may have been trained on one or more of functional cardiac images, myocardial delayed enhancement images or myocardial perfusion images” where “Receiving learning data may include receiving image data which may include at least one of steady-state free precession (SSFP) magnetic resonance imaging (MRI) data or 4D flow MRI data” (Golden, page 53, paragraph 0042) and where “SSFP cine studies contain of 4 dimensions of data (3 space, 1 time), and 4D Flow studies contain 5 dimensions of data (3 space, 1 time, 4 channels of information)” (Golden, page 64, paragraph 0228). Examiner notes that the space dimensions are the spatial dimensions and the time dimension is the temporal dimension.), Wherein the CNN comprises an encoder branch, a decoder branch and skip connections between the encoder branch and the decoder branch (Golden, page 70, paragraph 0285, “FIG. 36 shows a schematic representation of a fully convolutional encoder-decoder architecture with skip connections that utilizes a smaller expanding path than contracting path” and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the encoder branch is highlighted in FIG. 36 in the left box and the decoder branch is highlighted in the right circle.), Wherein the method comprises the implementation, by a data processor of a first server, for each of a plurality of training perfusion sequence from a base of training perfusion sequences each associated to an expected quantitative map of the perfusion, of steps of (Golden, page 74, paragraph 0327 “In some implementations, the data on which the CNN model 4604 has been trained is data from functional cardiac magnetic resonance imaging (e.g., via a contrast-free SSFP imaging sequence) and the cardiac image data is data from a cardiac perfusion or myocardial delayed enhancement study” where “The system memory 2508 may also include communications programs 3540, for example a server and/or a Web client or browser for permitting the processor-based device 3504 to access and exchange data with other systems such as user computing systems” (Golden, page 69, paragraph 0275) where “the output of the network is eighteen scalars corresponding to three coordinates for each of the six landmarks in the input image. Such an architecture may be trained in a similar fashion to the previously described landmark detector. The only update needed is the re-formulation of the loss to account for the change in the network output format (a spatial point in this implementation, as opposed to the probability map used in the first implementation). One reasonable loss function may be the L2 (squared) distance between the output of the network and the real landmark coordinate, but other loss functions may be used as well, as long as the loss functions are related to the quantity of interest, namely the distance error” (Golden, page 68, paragraph 0263) and “Typically, around 20 3D volumetric images are acquired throughout a single cardiac cycle, each corresponding to one snapshot of the heartbeat. The initial database thus corresponds to the 3D images of different patients at different time steps. Each 3D MRI presents a number of landmark annotations, from zero landmark to six landmarks, placed by the user of the web application. The landmark annotations, if present, are stored as vectors of coordinates (x, y, z, t) indicating the position (x, y, z) of the landmark in the 3D MRI corresponding to the time point t” (Golden, page 65, paragraph 0237). Examiner notes that the real landmark coordinate is the expected quantitative map.) extracting, using the encoder branch of the CNN, a plurality of initial n+1-dimensional features maps representative of the training perfusion sequence at n≥3 different scales, and projecting, using the skip connections of the CNN, each one of the plurality of initial n+1-dimensional features maps into one of a plurality of initial n-dimensional feature maps (Golden, page 58, paragraph 0126, “The network 600 is configured such that, after every pooling layer 608, the number of feature maps doubles and the spatial resolution is halved. After every upsampling layer 610, the number of feature maps is halved and the spatial resolution is doubled. With the scheme, the number of feature maps for each layer across the network 600 can be fully described by the number (e.g., between 1 and 2000 feature maps) in the first layer” where “ENet utilizes an expanding path that is smaller than its contracting path. ENet also makes use of bottleneck modules, which are convolutions with a small receptive field that are applied in order to project the feature maps into a lower dimensional space” where “subsequent to each upsampling layer, the CNN model may include a concatenation of feature maps from a corresponding layer in the contracting path through a skip connection” (Golden, page 52, paragraph 0028) and “U-Net, originally developed for use in the biomedical community where there are often fewer training images and even finer resolution is required, added the use of skip connections between the contracting and expanding paths to preserve details” (Golden, page 70, paragraph 0286) and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the plurality of n+1-dimensional features maps are the feature maps within the convolutional layers of the encoder branch shown in the left box. Examiner further notes that the n-dimensional feature maps are the feature maps that are within the convolutional layer that the skip connection is connected or projected to. Examiner additionally notes that the feature maps are projected into a lower dimensional space and thus are n-dimensional feature maps. Examiner further notes that the different scales can be seen in Fig. 36 by the different sizes of layers.); generating, using said decoder branch of the CNN, a plurality of enriched n- dimensional feature maps also representative of the inputted perfusion sequence at different scales, an enriched n-dimensional feature map at a particular scale incorporating information from the initial n-dimensional feature maps at smaller or equal scale (Golden, page 59, paragraph 0130, “Upsampling the activation volumes back to the original resolution is necessary in a fully convolutional network for pixel-wise segmentation. To increase the resolution of the activation volumes in the network, some systems may use an upsampling operation, then a 2x2 convolution, then a concatenation of feature maps from the corresponding contracting layer through a skip connection, and finally two 3x3 convolutions” where “ENet utilizes an expanding path that is smaller than its contracting path. ENet also makes use of bottleneck modules, which are convolutions with a small receptive field that are applied in order to project the feature maps into a lower dimensional space” and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the enriched n-dimensional feature maps are the upsampled activation volumes); generating at least one quantitative map of the inputted perfusion sequence from the enriched n-dimensional feature maps at the largest scale among the different scales (Golden, page 60, paragraph 0175, “Unlike previous models which were only concerned with two classes for a cell discrimination task, foreground and background, the SSFP model disclosed herein attempts to distinguish four classes, namely, background, LV Endocardium, LV Epicardium and RV Endocardium. To accomplish this, the network output may include three probability maps, one for each non-background class” and Golden, page 37, Figure 36, PNG media_image1.png 480 805 media_image1.png Greyscale Examiner notes that the quantitative map is the probability map as shown in Figure 36.). And minimizing a distance with the expected quantitative map of the perfusion (Golden, page 68, paragraph 0263, “the output of the network is eighteen scalars corresponding to three coordinates for each of the six landmarks in the input image. Such an architecture may be trained in a similar fashion to the previously described landmark detector. The only update needed is the re-formulation of the loss to account for the change in the network output format (a spatial point in this implementation, as opposed to the probability map used in the first implementation). One reasonable loss function may be the L2 (squared) distance between the output of the network and the real landmark coordinate, but other loss functions may be used as well, as long as the loss functions are related to the quantity of interest, namely the distance error” where “Beginning with random noise as a model "input" and a real segmentation mask as the target, we perform backpropagation to update the pixel values in the input image such that the loss is minimized” (Golden, page 72, paragraph 0301). Examiner notes that the loss function computes the distance with the real landmark coordinate and since the loss is minimized, the distance that the loss function is computing is also minimized.). Regarding Claim 15, Golden teaches A non-transitory computer-readable medium comprising code instructions that, when executed by a computer, cause the computer to execute a method according to claim 1 for processing an inputted perfusion sequence (Golden, page 54, paragraph 0058, “A machine learning system may be summarized as including at least one nontransitory processor-readable storage medium that stores at least one of processor-executable instructions or data; and at least one processor communicably coupled to the at least one nontransitory processor-readable storage medium, the at least one processor: receives a plurality of sets of 3D MRJ images, the images in each of the plurality of sets represent an anatomical structure of a patient; receives a plurality of annotations for the plurality of sets of 3D MRI images, each annotation indicative of a landmark of an anatomical structure of a patient depicted in a corresponding image; trains a convolutional neural network (CNN) model to predict the locations of the plurality of landmarks utilizing the 3D MRI images; and stores the trained CNN model in the at least one nontransitory processor- readable storage medium of the machine learning system.”). Regarding Claim 16, Golden teaches A non-transitory computer-readable medium comprising code instructions that, when executed by a computer, cause the computer to execute a method according to claim 11 for training a convolutional neural network (Golden, page 54, paragraph 0058, “A machine learning system may be summarized as including at least one nontransitory processor-readable storage medium that stores at least one of processor-executable instructions or data; and at least one processor communicably coupled to the at least one nontransitory processor-readable storage medium, the at least one processor: receives a plurality of sets of 3D MRJ images, the images in each of the plurality of sets represent an anatomical structure of a patient; receives a plurality of annotations for the plurality of sets of 3D MRI images, each annotation indicative of a landmark of an anatomical structure of a patient depicted in a corresponding image; trains a convolutional neural network (CNN) model to predict the locations of the plurality of landmarks utilizing the 3D MRI images; and stores the trained CNN model in the at least one nontransitory processor- readable storage medium of the machine learning system.”). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 6-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Golden in view of Hess et al. (“Synthetic Perfusion Maps: Imaging Perfusion Deficits in DSC-MRI with Deep Learning”)(hereafter referred to as Hess). Regarding Claim 6, Golden teaches, The method according to claim 5, wherein said medical imaging device is a Magnetic Resonance Imaging, MRI, scanner (Golden, page 49, paragraph 0004, “Depending on the type of acquisition, these views may be captured directly in the scanner (e.g., steady-state free precession (SSFP) MRI) or may be created via multi planar reconstructions (MPRs) of a volume aligned in a different orientation (such as the axial, sagittal or coronal planes, e.g., 4D Flow MRI). The SAX view has multiple spatial slices, usually covering the entire volume of the heart, but the 2CH, 3CH and 4CH views often only have a single spatial slice. All series are cine, and have multiple time points encompassing a complete cardiac cycle”). Golden does not teach, but Hess does teach and the perfusion sequence is a Dynamic susceptibility Contrast, DSC, or a Dynamic Contrast Enhanced, DCE, perfusion sequence (Hess, page 1, abstract, “In this work, we present a novel convolutional neural network based method for perfusion map generation in dynamic susceptibility contrast-enhanced perfusion imaging”). Golden and Hess are considered analogous to the claimed invention because they both process perfusion sequences in CNNs. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Golden to use the DSC perfusion sequences from Golden. Doing so is a simple substitution of one known element (perfusion sequences from Golden) for another (DSC perfusion sequences from Hess) to obtain predictable results (processing perfusion sequences). Regarding Claim 7, Golden teaches the method of claim 4. Golden does not teach, but Hess does teach wherein obtaining the perfusion sequence comprises extracting patches of a predetermined size from the perfusion sequence and wherein the steps of extracting the plurality of initial n+1-dimensional features maps, generating the plurality of enriched n-dimensional feature maps and generating the at least one quantitative map are performed for each extracted patch (Hess, page 3, 1st paragraph, “Guided by the fact that those characteristics are captured best by large blood vessels entering the brain, we select the input to the BCS to be a patch sequence from the perfusion sequence, located at the transition between the basilar artery and the posterior cerebral artery. The location of this patch is globally fixed, i.e., it is not fine-tuned to the individual volume. Therefore, it may happen that this patch does not contain the desired blood vessels for specific instances in our data. The BCS processes the supplied patch sequence via two 3D convolutional layers, encoding each patch into a vector of size 16. The sequence of encoded patches is forwarded to the sequence encoder.” Examiner notes that the patches are from the perfusion sequence and thus is extracted. Examiner further notes that the patches are a predetermined size of 16 before being forwarded on for processing.). Golden and Hess are considered analogous to the claimed invention because they both process perfusion sequences in CNNs. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Golden to extract patches like in Hess prior to extracting feature maps, generating feature maps, and generating quantitative maps. Doing so is advantageous because “our method generates perfusion maps that are comparable to the target maps used for clinical routine, while being model-free, fast, and less noisy” (Hess, page 1, abstract). Claim(s) 12-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Golden in view of Ramon et al. (“Improving Diagnostic Accuracy in Low-Dose SPECT Myocardial Perfusion Imaging With Convolutional Denoising Networks”)(hereafter referred to as Ramon). Regarding claim 12, Golden teaches the method according to claim 11. Golden does not teach, but Ramon does teach further comprising generating at least one degraded version of at least one original training perfusion sequence of the base of training perfusion sequences (Ramon, page 2, 2nd column, 2nd paragraph, “Let vector x denote a reconstructed image volume from a low-dose SPECT-MPI acquisition, and vector y the corresponding image reconstructed from a standard full-dose acquisition of a given subject. Our goal is to determine a mapping from x to y such that y   ≈ f ( x ) . (1) In (1), the reconstructed image x, with lower data counts, represents a noisy version of image y (with full-dose data counts)” where “Specifically, assume a total of T such training image pairs are available. Then the training dataset is formed by input-output pairs as : {(x(i), y(i)), i=1,…,T} (2) where x(i) denotes the low-dose image from patient i, and y(i) the corresponding full-dose image” “Specifically, assume a total of T such training image pairs are available. Then the training dataset is formed by input-output pairs as : {(x(i), y(i)), i=1,…,T} (2) where x(i) denotes the low-dose image from patient i, and y(i) the corresponding full-dose image” (Ramon, page 2, 2nd column, 3rd paragraph). Examiner notes that the degraded version is x and the at least one original training perfusion sequence of the training base is y.), associating to said degraded version the expected quantitative map of the perfusion associated with the original training perfusion sequence (Ramon, page 2, 3rd paragraph, “Once optimized, f ^ is applied subsequently to a low-dose (unseen) image x to generate output y ~   =   f ^ ( x ) , which is desired to be similar to what would be obtained with a full-dose acquisition” and Ramon, page 3, Figure 1, PNG media_image3.png 425 1196 media_image3.png Greyscale Examiner notes that y ~ is the expected quantitative map of the perfusion and the full-dose acquisition is the original training perfusion sequence.), and enriching the base of training perfusion sequences by adding said degraded version (Ramon, page 2, 2nd column, 3rd paragraph, “Specifically, assume a total of T such training image pairs are available. Then the training dataset is formed by input-output pairs as : {(x(i), y(i)), i=1,…,T} (2) where x(i) denotes the low-dose image from patient i, and y(i) the corresponding full-dose image.” Examiner notes that the low-dose image is the degraded version. ). Golden and Ramon are considered analogous to the claimed invention because they both use encoders and decoders to process perfusion images. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Golden to create a degraded image and add it to the training base. Doing so is advantageous because it “can achieve substantial noise reduction and lead to improvement in the diagnostic accuracy of low-dose data” (Ramon, page 10, Conclusions, 2nd paragraph). Regarding claim 13, Golden in view of Ramon teaches the method according to claim 12. Golden in view of Ramon further teaches, wherein said original training perfusion sequence is associated to a contrast product dose, said degraded version of the original training perfusion sequence simulating a lower contrast product dose (Ramon, page 5, 1st column, 1st paragraph “In the first approach, we train the 3D CAE/CNN to learn the mapping from a specific reduced-dose level (i.e., 1/2, 1/4, 1/8, or 1/16 dose) to full dose. That is, the denoising network is obtained specifically for a given dose level” where “Specifically, assume a total of T such training image pairs are available. Then the training dataset is formed by input-output pairs as : {(x(i), y(i)), i=1,…,T} (2) where x(i) denotes the low-dose image from patient i, and y(i) the corresponding full-dose image” “Specifically, assume a total of T such training image pairs are available. Then the training dataset is formed by input-output pairs as : {(x(i), y(i)), i=1,…,T} (2) where x(i) denotes the low-dose image from patient i, and y(i) the corresponding full-dose image” (Ramon, page 2, 2nd column, 3rd paragraph). Examiner notes that the reduced dose level is simulated with a low-dose image, and the full-dose is associated to the full-dose image. ). Golden and Ramon are considered analogous to the claimed invention because they both use encoders and decoders to process perfusion images. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Golden to create a degraded image and add it to the training base. Doing so is advantageous because it “can achieve substantial noise reduction and lead to improvement in the diagnostic accuracy of low-dose data” (Ramon, page 10, Conclusions, 2nd paragraph). Allowable Subject Matter Claims 14 and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Specifically, regarding claim 14, “wherein the degraded version of the original training perfusion sequence simulating a lower contrast product dose is generated by calculating, for each voxel of the original training perfusion sequence, from a temporal signal S(t) of said voxel a degraded temporal signal Sd(t) using the formula Sd(t) = S(t) - (1 - d) -[ S ( t )   - - S(0)], wherein S ( t ) - is a local average of the temporal signal S(t), d is a dose reduction factor and S(0) is a baseline signal” in conjunction with the other limitations of the claims are not taught by the prior art of record. The closest prior art is Ramon and Wu et al. (“Physiological Modulations in Arterial Spin Labeling Perfusion Magnetic Resonance Imaging”)(hereafter referred to as Wu). Ramon discloses generating the degraded version of the original training perfusion sequence simulating a lower contrast product does using signals (Ramon, page 5, Section III A, 2nd paragraph). Ramon also discloses using voxels of the original training perfusion sequence (Ramon, page 2, 1st column, 3rd paragraph). Ramon does not disclose calculating for each voxel, from a temporal signal. Ramon also does not disclose the formula. Wu discloses an average of temporal signals (Wu, page 3, equation 3), but does not disclose a degraded version of the original training perfusion sequence. Therefore, the prior art of record, individually, or in combination, does not disclose claim 14 as a whole. Response to Arguments Examiner notes that the previous claim objections have been overcome in light of the instant amendments. Examiner further notes that new claim objections have been made in light of the instant amendments. Examiner notes that the previous 112(b) rejections have been overcome in light of the instant amendments On page 9, Applicant argues: Golden is entirely directed to automated anatomical segmentation (generating label masks or boundary contours of cardiac structures, such as ventricles, myocardium, or papillary muscles). See Golden, paragraphs [0043], [0306]. Golden's network outputs pixel-wise class probabilities to define anatomical boundaries. See Golden, paragraph [0043]. A segmentation mask that merely classifies pixels into anatomical classes (e.g., "myocardium" vs. "blood pool") does not depict the "passage of a fluid through a tissue," nor does it generate a "quantitative map representing spatial values of a parameter characterizing said perfusion" (such as CBV, CBF, MTT, or k-trans). Calculating physical/hemodynamic perfusion parameters is technically and fundamentally different from identifying anatomical boundaries. Applicant respectfully submits that Golden does NOT teach or suggest calculating parameters of fluid passage to generate a quantitative map. Regarding the Applicant’s argument that Golden does not teach the newly amended limitations, Examiner respectfully disagrees. Specifically Golden does teach these limitations by depicting blood flow at certain locations in the images (Golden, page 67, paragraph 0260; Golden, page 61, paragraph 0188). Examiner respectfully directs the Applicant to the above 102 rejection section. On pages 9-10, Applicant argues: Second, claims l and 11 require "projecting, using the skip connections of the CNN, each one of the plurality of initial n+ I-dimensional features maps into one of a plurality of initial n dimensional feature maps." Applicant respectfully submits that Golden does NOT teach or suggest performing a dimensional reduction (from n +I ton dimensions) using the skip connections. In Golden's neural network (DeepVentricle or FastVentricle), the skip connections are conventional lateral connections that simply copy or concatenate feature maps between the contracting path and the expanding path at identical spatial scales and dimensions. See Golden, claim 8, claim 18 and paragraph [0130]. Golden's skip connections do not modify the dimensionality of the feature maps to resolve dimensional inconsistency between an encoder and a decoder. By contrast, the stU-Net architecture of the present invention uses skip connections to project n + I-dimensional feature maps (which include a temporal dimension) into n-dimensional feature maps (where the temporal dimension is dropped by pooling) to make them dimensionally consistent with the n-dimensional decoder (see US publication of the present application, paragraphs [0160]-[0162]). This specific projection mechanism within the skip connections is completely absent from Golden. Regarding the Applicant’s argument that Golden does not disclose “projecting, using the skip connections of the CNN, -each one of the plurality of initial n+1 dimensional features maps into one of a plurality of initial n-dimensional feature maps,” Examiner respectfully disagrees. Specifically, Examiner notes that the plurality of n+1-dimensional features maps are the feature maps within the convolutional layers of the encoder branch. Examiner further notes that the n-dimensional feature maps are the feature maps that are within the convolutional layer that the skip connection is connected or projected to. Examiner additionally notes that the feature maps are projected into a lower dimensional space and thus are n-dimensional feature maps (Golden, page 70, paragraph 0287). Examiner respectfully directs the Applicant to the above 102 rejections. Examiner additionally notes that claim 1 does not recite dropping the temporal dimension. The amended claim 10 recites reducing the temporal dimension. However, claim 10 is also taught by Golden as discussed above in the 102 rejection section. On pages 10-11, Applicant argues: Each of claims 6, 7, 12 and 13 depends on corresponding independent claim 1 or independent claim 11, and therefore includes all of the limitations of claim 1 or claim 11, Applicant submits that the Examiner's reliance on additional secondary references (Hess and Ramon) as allegedly pertaining to certain dependent claims fails to make up for the deficiencies in Golden as discussed above. Specifically, neither Hess (which does not teach stU-Net dimension-projecting skip connections) nor Ramon (directed to denoising SPECT images) teaches or suggests the specific combination of an n+1-dimensional encoder and an n-dimensional decoder coupled with dimension-projecting skip connections to generate quantitative perfusion maps. Further, there is no motivation in the prior art to modify Golden's anatomical segmentation network to perform the specific spatio-temporal projection and quantitative parameter calculations of the present invention. Regarding the Applicant’s argument that the dependent claims are allowable at least due in part to their dependency on the independent claims, the Examiner respectfully disagrees and notes the instant rejections and response to arguments regarding the independent claims above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Kim et al. (“Improving resolution of MR images with an adversarial network incorporating images with different contrast”) uses encoders and decoders to generate higher resolution images from low resolution images. Kim et al. also uses a discriminator CNN to determine which images are generated and which are not. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KAITLYN R LAU whose telephone number is (571)272-1429. The examiner can normally be reached Monday - Thursday: 8:00 am - 6:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.R.L./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Jun 16, 2023
Application Filed
Apr 16, 2026
Non-Final Rejection mailed — §102, §103
Jul 15, 2026
Response Filed
Sep 04, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688298
FEATURE SELECTION FOR CYBERSECURITY THREAT DISPOSITION
4y 7m to grant Granted Jul 21, 2026
Patent 12602431
METHODS FOR PERFORMING INPUT-OUTPUT OPERATIONS IN A STORAGE SYSTEM USING ARTIFICIAL INTELLIGENCE AND DEVICES THEREOF
3y 10m to grant Granted Apr 14, 2026
Patent 12572828
METHOD FOR INDUSTRY TEXT INCREMENT AND ELECTRONIC DEVICE
4y 5m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
60%
Grant Probability
99%
With Interview (+66.7%)
3y 11m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 10 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month