DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted was filed on 01/23/2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Amendment
The amendment filed 06/26/2026 has been entered. Claims 1, 5, 15, 16, and 20 have been amended. Claims 1-20 remain pending in the application. Applicant’s amendments to the claims, specification, and drawings have overcome each and every objection and 112(b) rejections previously set forth in the Non-Final Office Action mailed 01/15/2026.
Response to Arguments
Applicant’s arguments with respect to the amended features of the pending claims have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument regarding the amended features.
Applicant’s arguments, see Remarks, filed 06/26/2026, with respect to the § 101 rejection of the pending claims have been fully considered and are persuasive. The § 101 rejections of claims 1-20 have been withdrawn as the claims are directed to an improvement in computer functionality and include a practical application of synthesizing images having a target contrast in magnetic resonance imaging.
Applicant argues, with regards to claim 11, that “Dalmaz’s ResVit is a fixed-channel architecture”, “The number and identity of input channels in Dalmaz is predetermined by the protocol; zeroing channels within a fixed-width input is not the same as accepting arbitrary number of input contrasts”, and “The Application describes a model ‘capable of taking any number and combination of input sequences’”.
Examiner respectfully disagrees. Claim 11 does not recite “any number of contrasts without limit” nor “a number not fixed at training time”. Under broadest reasonable interpretation (BRI), a model that accepts “any” number and combination of contrasts within its protocol, selected per subject at inference, reads on “capable of taking arbitrary number of contrasts as input”. Claim 11 recites “capable of taking arbitrary number of contrasts as input”. One is an arbitrary number as Dalmaz’s model is able one contrast as input.
Applicant argues, with regards to claim 12, that “Displaying synthesized outputs alongside references is not displaying an interpretation of how or why the model is generated the image”.
In response, examiner withdraws the 102(a)(1) rejection of claim 12, but a new ground of rejection is made in view of Rubin whom is previously cited for teaching the claims dependent on claim 12 (claims 13 and 15).
Applicant argues, with regards to claim 7, the modification of Dalmaz with Liu would not yield cross-contrast windows and is impermissible hindsight.
Examiner respectfully disagrees. Dalmaz teaches wherein the invention comprises an input of images with multiple contrasts, and under BRI, a shifted-window attention block operating on a representation derived from multiple contrasts is reasonably a “multi-contrast shifted window based attention block”. Dalmaz’s feature maps are already multi-contrast; Dalmaz discloses in § 3.1 Encoder: “its encoder receives as input the full set of modalities within the multi-modal imaging protocol, including both source and target modalities”. The output of the encoder is a feature map tensor in which information from all contrasts has already been encoded. Dalmaz’s transformer then patches the tensor (Eqs. 5-8). Dalmaz’s transformer is already attending across contrast because the contrasts were fused upstream into the tensor being partitioned/patched. Applying Liu’s windowing to that tensor gives windows that inherently span contrast information, thus no impermissible hindsight was used as argued. Moreover, the claims do not recite “partitioning a window across multiple contrasts”. In addition, the motivation to combine Dalmaz with Liu’s shifted windows is efficiency as stated in the last Office Action and disclosed by Liu (Abstract).
Applicant argues, with regards to claim 8, the modification of Dalmaz with Zhu “does not teach or suggest the claimed multi-contrast (cross-contrast) shifted window decoder block”, “neither reference teaches partitioning a window across multiple contrasts”, and “The Office’s conclusion that the combination ‘would predictably result wherein the encoder comprises a multi-contrast shifted window-based attention block’ supplies the ‘multi-contrast’ character solely from the claim itself, which is impermissible hindsight”.
Examiner respectfully disagrees. Applicant is arguing the references individually while the rejection is based on the combination of the two references. Zhu supplies the shifted-window decoder mechanism and Dalmaz supplies the multi-contrast content. Dalmaz discloses in § 3.1 Decoder, “its decoder synthesizes all contrasts within the multi-modal protocol” and outputs multi-modality images in separate channels. Dalmaz’s decoder thus operates on multi-contrast information. Therefore, applying Zhu’s shifted-window decoder blocks to it yields windows operating on multi-contrast derived features. Moreover, the claims do not recite “partitioning a window across multiple contrasts”. Zhu similarly provides motivation for the combination as it would provide greater computation flexibility and efficiency/performance ([0069], [0073]). Accordingly, impermissible hindsight has not been used.
Applicant argues, with regards to claims 13 and 15, that “Rubin’s model does not synthesize a target-contrast image and has no concept of input ‘contrasts’; its attention map identifies where an object is located for detection, not how a synthesis model attributed its output to input contrasts. The Office’s statement that Rubin’s “attention maps are synthesized images” (Office Action at p. 18) conflates a detection location map with the claimed synthesized image. The Application, by contrast, describes decoder attention scores that show ‘which region and/or which contrast contributes more or less to the prediction’”.
Examiner respectfully disagrees. As an initial matter, examiner withdraws the statement that Rubin’s attention maps are synthesized images. However, this characterization is not relied upon for the present rejection. Dalmaz teaches the synthesized image; Dalmaz’s decoder synthesizes all contrasts within the multi-modal protocol and produces multi-modality images (Abstract & Page 6 of Dalmaz). Applicant’s arguments that Rubin’s model does not synthesize a target-contrast image and “has no concept of input ‘contrasts” attack Rubin individually. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Dalmaz teaches synthesizing a target contrast images and input contrasts.
Claim 13 recites wherein the interpretation is “generated based at least in part on attention scores outputted by a decoder of the transformer model”. Rubin expressly discloses that “the attention map is generated by at least one decoder of the at least one layer” ([0015], [0138]). Rubin further teaches wherein its transformer layers are “based on an encoder-decoder transformer architecture” ([0008], [0125]).
Claim 15 recites “a visual representation of attention scores indicative of relevance of a region in the one or more images or a contrast from the one or more different contrasts to the synthesized image”. Rubin teaches attention scores indicative of region relevance ([0084], “The pixels depicted by the attention map depict the parts of the image to pay the highest ‘attention’ to (i.e., the ‘higher intensity’ pixels represent the pixels to pay the highest attention to)”) and that the attention maps “provide a high-content and visually explainable representation of object locations and/or appearances” ([0088]). Rubin therefore at least teaches wherein an interpretation may comprise a visual representation of attention scores indicative of relevance of a region.
Applicant argues, with regards to claim 14, that Lu does not synthesize images and does not quantify a contrast’s contribution to a synthesized image. Applicant argues the Application describes a synthesis-specific interpretation, and that there is no teaching or suggestion to apply Lu’s classification-importance ranking as an interpretation of a synthesis model.
Examiner respectfully disagrees. Dalmaz teaches the synthesized image. Lu is relied upon only for the recited quantitative analysis. Arguing that Lu does not synthesize images attacks Lu individually. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Moreover, the claim does not recite “contribution to a synthesized image”. Rather, the claim recites “quantitative analysis of a contribution or importance of each of the one or more different contrasts”. Lu’s teaches quantitative ranking of importance of each contrast via attention weights for each contrast. Therefore, Lu teaches the claimed feature.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claims 1 and 16 recite “a contrast encoding of a learnable contrast query”. The specification and drawings (Fig. 1B) describes contrast encodings (107) as being combined with partitioned small patches as input to an MMT Encoder (109), wherein the contrast encodings may include learnable parameters for each contrast in the input sequence and the target contrast, and may be an n-dimensional vector (¶ [0048] of specification filed 04/16/2024). Whereas, the specification and drawings (Fig. 1B) describe contrast queries (113) as a separate input into the MMT decoder (111), may similarly comprise vectors, and are learnable parameters that inform the decoder what contrast to synthesize (¶ [0049]). Nothing in the specification nor drawings show wherein the contrast encoding is “of” the learnable contrast query. For example, the specification does not describe a contrast encoding that is generated from, derived from, or otherwise comprised within a contrast query. Instead, the specification and drawings describe the contrast encoding and the contrast query as two distinct elements, situated at different locations within the disclosed transformer model and performing different functions. Moreover, the specification describes each element as independently learned. The specification describes the contrast encodings “may include learnable parameters for each contrast in the input sequence and the target contrast” which “may be learned during training process” (¶ [0048]). Separately, the specification discloses “The correspondence between a contrast query and a given contrast is learned during training” (¶ [0051]). Two independently trained parameter sets each described as acquiring its own correspondence to contrast, are not reasonably understood as being generated from or “of” the other. Accordingly, to the extent that claims 1 and 16 is interpreted as requiring a contrast encoding that is generated from, derived from, or comprised within a learnable contrast query, the specification does not reasonably convey possession of such subject matter and the claims are thus rejected under 35 U.S.C. § 112(a).
Claims 2-15 and 17-20 are rejected by virtue of dependency on rejected claims 1 and 16 respectively.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-6, 9, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Dalmaz (“ResViT: Residual vision transformers for multi-modal medical image synthesis”) in view of Strudel (“Segmenter: Transformer for Semantic Segmentation”) and Yang (“A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation”).
Regarding claim 1, Dalmaz teaches a computer-implemented method for synthesizing a contrast-weighted image (Abstract, wherein employing a residual vision transformer, i.e. ResVit, comprises a computer-implemented method; Pages 9-10, “Multi-Contrast MRI Synthesis”) comprising:
(a) receiving a multi-contrast image of a subject, wherein the multi-contrast image comprises one or more images of one or more different contrasts (Page 5 figure 2, wherein input images are shown to be of different contrasts);
(b) generating, by one or more processors, an input to a transformer model based at least in part on the multi- contrast image (Page 4 Figure 1, Page 5, “During training, ResViT takes as input the entire set of images within the multi-modal protocol, including both source and target modalities…Receiving as input the jth layer feature maps…where z0 ∈ ℝNP, ND denotes patch embeddings that the transformer encoder takes as input”; wherein the feature maps and patch embeddings are based on the multi-contrast images received earlier”, wherein training and executing a neural network, i.e. ResVit, is necessarily carried out by a processor); and
(c) generating, by the transformer model running on the one or more processors, a synthesized image having a target contrast (Page 10 Figure 3, “ResViT was demonstrated on the IXI dataset for two representative many-to-one synthesis tasks: a) T1, T2 → PD, b) T2, PD → T1”).
However, Dalmaz fails to teach wherein the input comprises a contrast encoding of a learnable contrast query and wherein the contrast encoding comprises learnable parameters that encode the one or more different contrasts and a target contrast that is different from the one or more different contrasts, wherein the target contrast is specified in the learnable contrast query.
In an analogous transformer model field of endeavor, Strudel teaches such a feature. Strudel teaches a transformer model for segmentation (Abstract). Strudel teaches inputting into the model learnable class embeddings (Page 2 left column middle paragraph, Page 3 right column first paragraph, Page 4 § 3.2 Decoder, Mask Transformer, “For the transformer-based decoder, we introduce a set of K learnable class embeddings cls = [cls1, …, clsk]..where K is the number of classes. Each class embedding is initialized randomly and assigned to a single semantic class. It will be used to generate the class mask”). Strudel therefore teaches wherein the class embeddings are trained/learned parameters. Strudel further teaches wherein the mask transformer (decoder) takes as input the output of the encoder and class embeddings to predict segmentation masks (Fig. 2 and its caption). Strudel shows in figure 8 in the appendix that a singular value decomposition of the class embeddings display an implicit clustering of semantically related categories. This corresponds to Applicant’s description that “vectors representing similar contrasts (e.g., T1 and T1Gd) lie closer and different contrasts (e.g., T1 and FLAIR) lie farther” (¶ [0048] spec filed 04/16/2024). Strudel therefore teaches wherein the class embeddings are a learned vector encoding the identity of a corresponding output, e.g. a target object or contrast, thus teaching a learnable query. As shown in Strudels’ figure 2, the class embeddings (i.e. “Tree, Sidewalk, Person”) encode the identity of an output and is input into the decoder (Mask Transformer) similar to Applicant’s figure 1B in which the MMT Decoder 111 accepts as input the contrast query 113.
Yang teaches cross-contrast image translation (Abstract). Yang teaches a one-hot code indicating MR contrast (Page 2, 2nd last paragraph). Yang teaches wherein the source and target contrasts are represented by a one-hot code, dubbed as a contrast code (Page 2, Method, last paragraph). Yang therefore teaches wherein designating both source contrasts and target contrasts by a code was conventional in multi-contrast MR synthesis. Yang further teaches wherein supplying a fixed contrast indicator may be insufficient to control the translation process (Page 2, 2nd paragraph).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to modify the multi-contrast synthesis model of Dalmaz to convey the source-target contrast configuration by learned parameter vectors supplied within the transformer input, as taught by Strudel (Abstract, Figures 2 & 8, Pages 2-3, Page 4, § Mask Transformer), thereby arriving at an input comprising a contrast encoding comprising learnable parameters that encode the one or more different contrasts and a target contrast that is different from the one or more different contrast, and a learnable contrast query in which the target contrast is specified. One of ordinary skill would have been motivated to do so because Yang above expressly identifies in the field of multi-contrast MRI synthesis that supplying a contrast indicator as a fixed input may be insufficient to control a cross-contrast image translation process (Page 2, 2nd paragraph). By processing using a learnable query and processing the image patches jointly with the class embeddings (i.e. learnable query), dynamical filters or outputs may be generated, changing with the input as recognized by Strudel (Page 4, left column last paragraph). The modification of Dalmaz with Strudel amounts to substitution of one known conditioning technique for another for performing the same function of informing the model which contrasts are available and which is to be synthesized. Both Dalmaz and Strudel teach transformer architecture operating on flattened image patch embeddings derived from a vision transformer. Dalmaz already sums a learned parameter vector with its patch tokens within the transformer input (Eq. 9, Page 5, “learnable positional embedding that carries information about patch location”). Dalmaz’s availability condition is defined per modality (Eq. 2); its unified model is trained across configurations (e.g. T1, T2 → PD; T2, PD → T1; T1, PD → T2, Page 4). Strudel assigns a separate trained embedding to each of the K categories among which its model distinguishes (Page 4, § 3.2 Decoder, “Each class embedding is initialized randomly and assigned to a single semantic class”). One of ordinary skill would accordingly have understood that a transformer required to distinguish among the several contrasts of Dalmaz’s protocol could be supplied with a corresponding trained embedding for each contrast, thereby arriving at learned parameters encoding the input contrasts and the target contrast within the transformer input. While no single reference discloses learned parameters encoding contrast identity, this element arises from the combination, Dalmaz supplying both the contrast-identity information its unified model requires and the practice of summing a learned parameter vector with its patch tokens within the transformer input, and Strudel supplying the assignment of an individual trained embedding to each category among which the model must distinguish.
Regarding claim 2, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 1.
Dalmaz further teaches wherein the multi-contrast image is acquired using a magnetic resonance (MR) device (Pages 7-8, “We demonstrated the proposed ResViT model on two multi-contrast brain MRI datasets”, wherein the IXI and BRATS datasets being MR images comprises the multi-contrast image being acquired by a MR device, Pages 9-10, “Multi-Contrast MRI Synthesis Experiments were conducted on the IXI and BRATS datasets to demonstrate synthesis performance in multi-modal MRI…many-to-one tasks of T1, T2 → PD; T1, PD → T2; T2, PD → T1…many-to-one tasks of T1, T2 → FLAIR; T1, FLAIR → T2; T2, FLAIR → T1 were considered”).
Regarding claim 3, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 1.
Dalmaz further teaches wherein the input to the transformer model comprises an image encoding generated by a convolutional neural network (CNN) model (Page 4 Figure 1, Page 6, “Finally, the feature maps are processed via a residual CNN (ResCNN) to distill learned structural and contextual representations…The decoder receives as input the feature maps distilled by the information bottleneck and produces multi-modality images in separate channels”).
Regarding claim 4, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 3.
Dalmaz further teaches wherein the image encoding is partitioned into image patches (Page 4 Figure 1, Page 5, “Accordingly, f’j is first split into non-overlapping patches of size (P, P), and the patches are then flattened”).
Regarding claim 5, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 3.
Dalmaz further teaches wherein the input to the transformer model comprises a combination of the image encoding and the contrast encoding (Page 4 Figure 1, wherein image patches are input and feature maps are output comprises image encoding, Pages 10-11, Figures 3 & Tables 1 & 2, wherein transforming an image of one contrast, i.e. T1, T2, PD, into a synthesized image of a different contrast necessarily comprises contrast encoding).
Regarding claim 6, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 1.
Dalmaz further teaches wherein the transformer model comprises: i) an encoder model receiving the input and outputting multiple representations of the input having multiple scales, ii) a decoder model receiving the query and the multiple representations of the input having the multiple scales and outputting the synthesized image (Page 4 Figure 1, wherein Figure 1 shows steps of down-sampling and up-sampling, resulting in outputting of multiple representations of input having multiple scales, and wherein figure 1 showing the Decoder stage coming after the down-sampling and up-sampling steps comprises the decoder model receiving the multiple representations of the input having multiple scales, Page 6, “The decoder receives as input the feature maps distilled by the information bottleneck”, Page 10, wherein task-specific synthesis implies given a task or query).
Regarding claim 9, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 1.
Dalmaz further teaches wherein the transformer model is trained utilizing a combination of synthesis loss, reconstruction loss, and adversarial loss (Page 7, “Loss Function”, wherein pixel-wise Lpix loss that is defined between acquired and synthesized target modalities comprises synthesis loss, Lrec is reconstruction loss, and Ladv is adversarial loss).
Regarding claim 11, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 1.
Dalmaz further teaches wherein the transformer model is capable of taking arbitrary number of contrasts as input (Page 10 last paragraph, wherein the ResViT model being capable of performing many-to-one and one-to-one tasks comprises it being capable of taking an arbitrary number of contrasts as input, i.e. one or many contrasts).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Dalmaz (“ResViT: Residual vision transformers for multi-modal medical image synthesis”) in view of Strudel (“Segmenter: Transformer for Semantic Segmentation”) and Yang (“A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation”) as applied to claim 6 above, and further in view Liu (“Swin Transformer: Hierarchical Vision Transformer using Shifted Windows”). Liu is cited in the IDS filed 08/14/2024.
Regarding claim 7, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 6.
However, Dalmaz fails to teach wherein the encoder model comprises a multi-contrast shifted window-based attention block.
In an analogous vision transformer field of endeavor, Liu teaches such a feature. Liu teaches a Swin transformer which uses shifted windows (Title, Abstract). Liu teaches transformer blocks with modified self-attention computation (Page 3, 3.1 Overall Architecture”). Liu teaches wherein the Swin transformer may comprise an encoder model (Page 4, Figure 3, wherein figure 3 depicts images being encoded) and wherein the Swin transformer blocks include a shifted window multi-head self attention module or block (SW-MSA) (Page 4, Figure 3, “Swin Transformer block”).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to have the encoder model include a shifted window-based attention module as taught by Liu (Page 4, Figure 3, “Swin Transformer block”). The shifted window scheme may have greater efficiency and flexibility as recognized by Liu (Abstract). Because Dalmaz teaches wherein inputs/outputs to their model are multi-contrast images, Dalmaz modified by the teachings of Liu to incorporate shifted window-based attention blocks into the encoder would predictably result wherein the encoder comprises a multi-contrast shifted window-based attention block.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Dalmaz (“ResViT: Residual vision transformers for multi-modal medical image synthesis”) in view of Strudel (“Segmenter: Transformer for Semantic Segmentation”) and Yang (“A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation”) as applied to claim 6 above, and further in view Zhu (US20230100413).
Regarding claim 8, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 6.
However, Dalmaz fails to teach wherein the decoder model comprises a multi-contrast shifted window-based attention block.
In an analogous transformer model architecture field of endeavor, Zhu teaches such a feature. Zhu teaches using transformer layers with shifted self-attention windows (Abstract, [0001]). Zhu teaches a transformer model including an encoder sub-network and decoder sub-network ([0070]). Zhu teaches the decoder is symmetric to the encoder ([0071]). Zhu teaches the encoder and decoder utilize modified self-attention computation with a shifted windows approach (Figs. 5A & 5B, [0072], [0116-0119]). Zhu teaches the shifted window transformer blocks of the decoder model can compute self-attention (Figs. 5B & 6A, [0119]) and thus comprise shifted window-based attention blocks.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to have the decoder model comprise shifted window-based attention blocks as taught by Zhu (Figs. 5B & 6A, [0119]). By having the decoder model include shifted window blocks which compute self attention, greater computation flexibility and efficiency/performance may be provided as recognized by Zhu ([0069], [0073]). Because Dalmaz teaches wherein inputs/outputs to their model are multi-contrast images, Dalmaz modified by the teachings of Zhu to incorporate shifted window-based attention blocks in the decoder would predictably result wherein the decoder comprises a multi-contrast shifted window-based attention block.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Dalmaz (“ResViT: Residual vision transformers for multi-modal medical image synthesis”) in view of Strudel (“Segmenter: Transformer for Semantic Segmentation”) and Yang (“A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation”) as applied to claim 1 above, and further in view of Hagi (US20220180527).
Regarding claim 10, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 1.
However, Dalmaz fails to explicitly teach wherein the transformer model is trained utilizing multi-scale discriminators.
In an analogous transformer model field of endeavor, Hagi teaches such a feature. Hagi teaches image synthesis by use of a transformer model (Abstract, [0061]). Hagi teaches wherein the trained model comprises a multi-scale discriminator comprising a plurality of single-scale discriminators which operate at different image scales comprising different resolutions ([0013], [0071]). Hagi teaches by using a multi-scale discriminator, samples/images that are indistinguishable from natural images may be generated ([0071]); multi-scale discriminators may help improve realism.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to train the transformer model by using a multi-scale discriminator as taught by Hagi ([0013], [0071]). Multi-scale discriminators may help improve realism of synthesized images as recognized by Hagi ([0071]).
Claims 12, 13, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Dalmaz (“ResViT: Residual vision transformers for multi-modal medical image synthesis”) in view of Strudel (“Segmenter: Transformer for Semantic Segmentation”) and Yang (“A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation”) as applied to claim 1 above, and further in view of Rubin (US20240371500).
Regarding claim 12, Dalmaz in view of Strudel and Yang teaches the invention as claimed above in claim 1.
However, Dalmaz fails to teach the invention further comprising displaying interpretation of the transformer model generating the synthesized image.
In an analogous vision transformer field of endeavor, Rubin teaches such a feature. Rubin teaches vision transformer models which may be based on an encoder-decoder architecture ([0008], [0047], [0053], [0057]). Rubin teaches the models receive input data comprising images ([0045-0046]). Rubin teaches the models output an attention map ([0075-0076]). Rubin teaches that the attention maps constitute an interpretation of the model’s operation, stating that the attention maps “provide a high-content and visually explainable representation of object locations and/or appearances” ([0088]), that they “increase the transparency of AI-based prediction by helping end-users understand the salient features used by the model for prediction” ([0088]), and that they “provide transparency into ‘black-box’ AI-based models and/or provide a mechanism for clinical review” ([0104]). Moreover, Rubin teaches displaying the generated attention map ([0016], [0104], [0140]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to generate and display an attention map as taught by Rubin (Fig. 4, [0007], [0015-0016], [0084]), thereby resulting in displaying an interpretation of the transformer model of Dalmaz generating the synthesized image. Attention maps may support human visual clinical interpretation, provide transparency to black-box ai models, and/or provide a mechanism for clinical review as recognized by Rubin ([0104]). Dalmaz teaches wherein the transformer model generates a synthesized image. Therefore, Dalmaz modified by Rubin to generate and display attention maps of the model’s operation would predictably result in the attention maps being an interpretation of the transformer model generating the synthesized image.
Regarding claim 13, Dalmaz in view of Strudel, Yang, and Rubin teaches the invention as claimed above in claim 12.
However, Dalmaz fails to teach wherein the interpretation is generated based at least in part on attention scores outputted by a decoder of the transformer model.
In an analogous vision transformer field of endeavor, Rubin teaches such a feature. Rubin teaches vision transformer models which may be based on an encoder-decoder architecture ([0008], [0047], [0053], [0057]). Rubin teaches the models receive input data comprising images ([0045-0046]). Rubin teaches the models output an attention map ([0075-0076]). Rubin teaches wherein the attention map (322) may be generated, i.e. outputted, by a decoder ([0015], [0075], [0100], [0138]). Rubin teaches wherein the attention maps are based on an input image (Fig. 4, [0084]). Rubin teaches the pixels depicted by the attention map depict parts of an image to pay the highest ‘attention’ to ([0084]) and are thus based on attention scores. Moreover, Rubin teaches displaying the generated attention map ([0016], [0104], [0140]). Since the attention map is a map of attention scores and generated by a decoder, Rubin therefore teaches generating an interpretation (attention map) based at least in part on attention scores outputted by a decoder of a transformer model.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to generate and display an attention map as taught by Rubin (Fig. 4, [0007], [0015-0016], [0084]). Attention maps may support human visual clinical interpretation, provide transparency to black-box ai models, and/or provide a mechanism for clinical review as recognized by Rubin ([0104]).
Regarding claim 15, Dalmaz in view of Strudel, Yang, and Rubin teaches the invention as claimed above in claim 12.
However, Dalmaz fails to teach wherein the interpretation comprises a visual representation of attention scores indicative of relevance of a region in the one or more images or a contrast from the one or more different contrasts to the synthesized image.
In an analogous vision transformer field of endeavor, Rubin teaches such a feature. Rubin teaches vision transformer models which may be based on an encoder-decoder architecture ([0008], [0047], [0053], [0057]). Rubin teaches the models receive input data comprising images ([0045-0046]). Rubin teaches the models output an attention map ([0075-0076]). Rubin teaches wherein the attention maps are based on an input image (Fig. 4, [0084]). Rubin teaches the pixels depicted by the attention map depict parts of an image to pay the highest ‘attention’ to, i.e. the ‘higher intensity’ pixels represent the pixels to pay the highest attention to, (Fig. 4, [0084]) and thus the higher intensity pixels comprise a visual representation of attention scores indicative of relevance of a region in the image.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to generate and display an attention map as taught by Rubin (Fig. 4, [0007], [0015-0016], [0084]). Attention maps may support human visual clinical interpretation, provide transparency to black-box ai models, and/or provide a mechanism for clinical review as recognized by Rubin ([0104]).
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Dalmaz (“ResViT: Residual vision transformers for multi-modal medical image synthesis”) in view of Strudel (“Segmenter: Transformer for Semantic Segmentation”), Yang (“A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation”), and Rubin (US20240371500) as applied to claim 12 above, and further in view of Lu (GAMER MRI: Gated-Attention Mechanism Ranking of Multi-Contrast MRI in Brain Pathology”).
Regarding claim 14, Dalmaz in view of Strudel, Yang, and Rubin teaches the invention as claimed above in claim 12.
However, Dalmaz fails to teach wherein the interpretation comprises quantitative analysis of a contribution or importance of each of the one or more different contrasts.
In an analogous multi-contrast imaging field of endeavor, Lu teaches such a feature. Lu teaches GAMER MRI: gated-attention mechanism ranking of multi-contrast MRI, which computes attention weights (AWs) as proxies of importance of features in image classification (Abstract). Lu teaches there is a need to address the selection of the most informative MR contrast since acquiring multiple MR contrasts requires significant time (Page 2, 1. Introduction first paragraph). Lu teaches wherein the model receives as input multiple contrasts consisting of combinations of ADC, FLAIR, and Trace (Page 8, Table 4, wherein rmAWs comprises reported mean attention weights). Moreover, Lu teaches computing and displaying attention weights for each contrast, thereby ranking each contrast (Page 8, Tables 4 & 5). Lu therefore teaches displaying an interpretation comprising a quantitative analysis of importance of each of the one or more different contrasts.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to calculate and display the attention weights of each inputted contrast as taught by Lu (Abstract, Page 8, Tables 4 & 5). By calculating the attention weight of each contrast, a clinician may know which contrast is the most important for classifying/diagnosing a certain pathology such as infarct strokes as recognized by Lu (Abstract, Page 10, 4.5 Conclusion).
Claims 16-20 are rejected under 35 U.S.C. 103 as being unpatentable over Dalmaz (“ResViT: Residual vision transformers for multi-modal medical image synthesis”) in view of Hagi (US20220180527), Strudel (“Segmenter: Transformer for Semantic Segmentation”) and Yang (“A Unified Hyper-GAN Model for Unpaired Multi-contrast MR Image Translation”).
Regarding claim 16, Dalmaz teaches:
(a) receiving a multi-contrast image of a subject, wherein the multi-contrast image comprises one or more images of one or more different contrasts (Page 5 figure 2, wherein input images are shown to be of different contrasts);
(b) generating an input to a transformer model based at least in part on the multi- contrast image (Page 4 Figure 1, Page 5, “During training, ResViT takes as input the entire set of images within the multi-modal protocol, including both source and target modalities…Receiving as input the jth layer feature maps…where z0 ∈ ℝNP, ND denotes patch embeddings that the transformer encoder takes as input”; wherein the feature maps and patch embeddings are based on the multi-contrast images received earlier”); and
(c) generating, by the transformer model, a synthesized image having a target contrast (Page 10 Figure 3, “ResViT was demonstrated on the IXI dataset for two representative many-to-one synthesis tasks: a) T1, T2 → PD, b) T2, PD → T1”).
However, Dalmaz fails to teach a non-transitory computer-readable storage medium including instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: (a), (b), and (c).
In an analogous transformer model field of endeavor, Hagi teaches such a feature. Hagi teaches image synthesis by use of a transformer model and wherein the input to the model comprise an image (Fig. 6, Abstract, [0061-0062], [0068]). Hagi teaches wherein the image synthesis or image transformation may be performed by a non-transitory computer-readable medium storing instructions which, when executed by a processor, causes the processor to perform the method/invention disclosed herein ([0026]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to have the method of image synthesis and/or transformation be stored as instructions in a non-transitory computer-readable storage medium that when executed, causes a processor to perform the invention as taught by Hagi (Fig. 6, [0026], [0061-0062], [0068]). Having the method be stored as instructions for a computer to perform may predictably allow for the method to be more automated, thereby minimizing manual work performed by a user.
However, the modified combination noted above fails to teach wherein the input comprises a contrast encoding of a learnable contrast query and wherein the contrast encoding comprises learnable parameters that encode the one or more different contrasts and a target contrast that is different from the one or more different contrasts, and wherein the target contrast is specified in the learnable contrast query.
In an analogous transformer model field of endeavor, Strudel teaches such a feature. Strudel teaches a transformer model for segmentation (Abstract). Strudel teaches inputting into the model learnable class embeddings (Page 2 left column middle paragraph, Page 3 right column first paragraph, Page 4 § 3.2 Decoder, Mask Transformer, “For the transformer-based decoder, we introduce a set of K learnable class embeddings cls = [cls1, …, clsk]..where K is the number of classes. Each class embedding is initialized randomly and assigned to a single semantic class. It will be used to generate the class mask”). Strudel therefore teaches wherein the class embeddings are trained/learned parameters. Strudel further teaches wherein the mask transformer (decoder) takes as input the output of the encoder and class embeddings to predict segmentation masks (Fig. 2 and its caption). Strudel shows in figure 8 in the appendix that the a singular value decomposition of the class embeddings display an implicit clustering of semantically related categories. This corresponds to Applicant’s description that “vectors representing similar contrasts (e.g., T1 and T1Gd) lie closer and different contrasts (e.g., T1 and FLAIR) lie farther” (¶ [0048] spec filed 04/16/2024). Strudel therefore teaches wherein the class embeddings are a learned vector encoding the identity of a corresponding output, e.g. a target object or contrast, thus teaching a learnable query. As shown in Strudels’ figure 2, the class embeddings (i.e. “Tree, Sidewalk, Person”) encode the identity of an output and is input into the decoder (Mask Transformer) similar to Applicant’s figure 1B in which the MMT Decoder 111 accepts as input the contrast query 113.
Yang teaches cross-contrast image translation (Abstract). Yang teaches a one-hot code indicating MR contrast (Page 2, 2nd last paragraph). Yang teaches wherein the source and target contrasts are represented by a one-hot code, dubbed as a contrast code (Page 2, Method, last paragraph). Yang therefore teaches wherein designating both source contrasts and target contrasts by a code was conventional in multi-contrast MR synthesis. Yang further teaches wherein supplying a fixed contrast indicator may be insufficient to control the translation process (Page 2, 2nd paragraph).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified the invention of Dalmaz to modify the multi-contrast synthesis model of Dalmaz to convey the source-target contrast configuration by learned parameter vectors supplied within the transformer input, as taught by Strudel (Abstract, Figures 2 & 8, Pages 2-3, Page 4, § Mask Transformer), thereby arriving at an input comprising a contrast encoding comprising learnable parameters that encode the one or more different contrasts and a target contrast that is different from the one or more different contrast, and a learnable contrast query in which the target contrast is specified. One of ordinary skill would have been motivated to do so because Yang above expressly identifies in the field of multi-contrast MRI synthesis that supplying a contrast indicator as a fixed input may be insufficient to control a cross-contrast image translation process (Page 2, 2nd paragraph). By processing using a learnable query and processing the image patches jointly with the class embeddings (i.e. learnable query), dynamical filters or outputs may be generated, changing with the input as recognized by Strudel (Page 4, left column last paragraph). The modification of Dalmaz with Strudel amounts to substitution of one known conditioning technique for another for performing the same function of informing the model which contrasts are available and which is to be synthesized. Both Dalmaz and Strudel teach transformer architecture operating on flattened image patch embeddings derived from a vision transformer. Dalmaz already sums a learned parameter vector with its patch tokens within the transformer input (Eq. 9, Page 5, “learnable positional embedding that carries information about patch location”). Dalmaz’s availability condition is defined per modality (Eq. 2); its unified model is trained across configurations (e.g. T1, T2 → PD; T2, PD → T1; T1, PD → T2, Page 4). Strudel assigns a separate trained embedding to each of the K categories among which its model distinguishes (Page 4, § 3.2 Decoder, “Each class embedding is initialized randomly and assigned to a single semantic class”). One of ordinary skill would accordingly have understood that a transformer required to distinguish among the several contrasts of Dalmaz’s protocol could be supplied with a corresponding trained embedding for each contrast, thereby arriving at learned parameters encoding the input contrasts and the target contrast within the transformer input. While no single reference discloses learned parameters encoding contrast identity, this element arises from the combination, Dalmaz supplying both the contrast-identity information its unified model requires and the practice of summing a learned parameter vector with its patch tokens within the transformer input, and Strudel supplying the assignment of an individual trained embedding to each category among which the model must distinguish.
Regarding claim 17, Dalmaz in view of Hagi, Strudel, and Yang teaches the invention as claimed above in claim 16.
Dalmaz further teaches wherein the multi-contrast image is acquired using a magnetic resonance (MR) device (Pages 7-8, “We demonstrated the proposed ResViT model on two multi-contrast brain MRI datasets”, wherein the IXI and BRATS datasets being MR images comprises the multi-contrast image being acquired by a MR device, Pages 9-10, “Multi-Contrast MRI Synthesis Experiments were conducted on the IXI and BRATS datasets to demonstrate synthesis performance in multi-modal MRI…many-to-one tasks of T1, T2 → PD; T1, PD → T2; T2, PD → T1…many-to-one tasks of T1, T2 → FLAIR; T1, FLAIR → T2; T2, FLAIR → T1 were considered”).
Regarding claim 18, Dalmaz in view of Hagi, Strudel, and Yang teaches the invention as claimed above in claim 16.
Dalmaz further teaches wherein the input to the transformer model comprises an image encoding generated by a convolutional neural network (CNN) model (Page 4 Figure 1, Page 6, “Finally, the feature maps are processed via a residual CNN (ResCNN) to distill learned structural and contextual representations…The decoder receives as input the feature maps distilled by the information bottleneck and produces multi-modality images in separate channels”).
Regarding claim 19, Dalmaz in view of Hagi, Strudel, and Yang teaches the invention as claimed above in claim 18.
Dalmaz further teaches wherein the image encoding is partitioned into image patches (Page 4 Figure 1, Page 5, “Accordingly, f’j is first split into non-overlapping patches of size (P, P), and the patches are then flattened”).
Regarding claim 20, Dalmaz in view of Hagi, Strudel, and Yang teaches the invention as claimed above in claim 18.
Dalmaz further teaches wherein the input to the transformer model comprises a combination of the image encoding and the contrast encoding (Page 4 Figure 1, wherein image patches are input and feature maps are output comprises image encoding, Pages 10-11, Figures 3 & Tables 1 & 2, wherein transforming an image of one contrast, i.e. T1, T2, PD, into a synthesized image of a different contrast necessarily comprises contrast encoding).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TOMMY T LY whose telephone number is (571) 272-6404. The examiner can normally be reached M-F 12:00pm-8:00pm eastern time.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Anhtuan Nguyen can be reached at 571-272-4963. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TOMMY T LY/ Examiner, Art Unit 3797
/SERKAN AKAR/ Primary Examiner, Art Unit 3797