DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-20 are pending regarding this application.
Priority
The present application claims foreign priority benefits from EP24188308.1 filed on 07/12/2024. The certified copies of the priority documents were electronically retrieved on 02/26/2025.
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/23/2025 is considered and attached.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
9. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
In claim 19: “a spatio-temporal encoder for encoding…”
After a careful analysis, as disclosed above, and a careful review of the specification, the above limitation in claim 19 is interpreted as computer-implemented 112(f). Below is the corresponding structure and algorithm which are being read into the above limitations:
“a spatio-temporal encoder” (Applicant’s specification states that “the spatio-temporal encoder 204 […] may be embodied (executed) by the processor 214” in para. [0143]. Therefore, the corresponding structure for the spatio-temporal encoder is a processor and the corresponding algorithm for the spatio-temporal encoder which can be found in para. [0135]-[138], [0159], [0160], [0165]-[0166], and [0178]).
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 1, claim 1 recites “a real-time time series of medical images” in line 1 and line 3. As such, it is unclear whether the real-time time series of medical images in line 1 is equivalent to or distinct from the real-time time series of medical images in line 3. Applicant discusses these images throughout the specification. However, none of these instances of a real-time time series of medical images clarify whether there are multiple distinct instances of real-time time series of medical images being claimed in claim 1. As such, claim 1 is rejected for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Additionally, regarding claim 1, claim 1 recites “a frame corresponds to a medical image” in line 7. However, claim 1 already recites “an encoded representation per frame” in line 6 and “medical images” in both line 6 and line 2. As such, it is unclear whether the frame in line 6 and the medical images in lines 6 and 2 are equivalent to or distinct from the frame/medical image claimed in line 7. Applicant discusses the frames and medical images throughout the specification. However, none of these instances of recitations of frames and medical images clarify whether there are multiple distinct instances of frames/medical images being claimed in claim 1.
Additionally regarding claim 1, applicant recites “at least one object” in line 14 and “an object” in line 1. As such, it is unclear whether the “at least one object” in line 14 is equivalent to or distinct from the “object” claimed in line 1. However, even if they are equivalent, it is unclear how there can be “at least one object” when only a single object is introduced in line 1. Applicant discusses the object throughout the specification. However, none of these instances clarify whether there are multiple distinct objects being claimed in claim 1, nor does the specification clarify how there can be “at least one object” when only one object is introduced. As such, claim 1 is rejected for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Similar analysis is applied to claim 15. Claim 15 is similarly rejected.
Claims 2-14 and 16-18 are rejected due to their dependence upon claims 1 and 15.
Regarding claim 13, claim 13 recites “wherein, using the multi-head cross-attention decoder in the decoding, a background is removed for spatial correlation, and a historical trajectory of the at least one object and/or the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features”. Here, it is unclear which clauses are separated by the “and/or”. Said differently, it is unclear whether the above subject matter is recited such that: “wherein,
(path 1) using the multi-head cross-attention decoder in the decoding, a background is removed for spatial correlation, and a historical trajectory of the at least one object and/or
(path 2) the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features”
OR
“wherein, using the multi-head cross-attention decoder in the decoding,
(path 1) a background is removed for spatial correlation, and a historical trajectory of the at least one object and/or
(path 2) the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features”.
Applicant discusses the above subject matter in para. [0087]-[0088], [0151], and [0177]. However, none of these sections clarify which of the two aforementioned possible pathways applicant intends to convey. As such, claim 13 is rejected for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Similar analysis is applied to corresponding claim 18. Claim 18 is similarly rejected.
Regarding claim 14, claim 14 recites “performing a further downstream task” in line 2. However, it is unclear what the “task” is downstream of, as there are no tasks introduced prior to the recitation of the task in claim 14 that the claimed task can be downstream of. As such, it is unclear whether the claimed task is downstream of the downstream NN, or a task performed within the downstream NN. Applicant discusses the downstream task in para. [0090]-[0096] and [0141]-[0143]. While applicant explains that the downstream NN system may include downstream task decoders for carrying out the downstream tasks, it remains unclear of what “task” the “further downstream task” is downstream (emphasis added). As such, claim 14 is rejected for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 3, 9, 10, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Demoustier et al. (“ConTrack: Contextual Transformer for Device Tracking in X-Ray’), hereinafter Demoustier, in view of Thawakar et al. (U.S. Publication No. 2024/0161334 A1), hereinafter Thawakar and Xie et al. (“VideoTrack: Learning to Track Objects via Video Transformer”), hereinafter Xie.
Regarding claim 1, Demoustier teaches a (Demoustier teaches “given a sequence of consecutive X-ray images and an initial location of the target catheter tip, our goal is to track the location of the target xt = (ut, vt), at any time t, t > 0.” in Section 2. Demoustier teaches a transformer network for tracking a catheter tip downstream of a ResNet-50 encoder as shown in Fig. 2), the method comprising:
receiving a real-time time series of medical images of a patient's anatomical region at an input layer of the downstream NN (Demoustier teaches “a sequence of consecutive X-ray images” in section 2, wherein “the test dataset is divided into two primary categories: fluoroscopic and angiographic sequences. Fluoroscopic sequences are real-time videos of internal movements captured by low-dose X-rays without radiopaque substances, while angiographic sequences display blood vessels in real-time after the introduction of radiopaque substances” in Section 3. Demoustier further teaches inputting a set containing historically selected frames for templates [selected from the sequence of consecutive X-ray images] and a search frame (current frame) into the encoding layer of the transformer in section 2.1. See Figure 2 wherein the transformer network is downstream from the ResNet-50 encoder and the images are of a patient’s anatomy);
encoding, using a spatio-(Demoustier teaches encoding a search frame (current frame) and templates of historically selected frames within in a sequence of real-time consecutive x-ray images wherein each frame has a corresponding time as shown in Sections 2 and 2.1, wherein the method involves utilizing medical images as shown in Section 3);
decoding, using a multi-head(Demoustier teaches decoding the encoded representation as shown in Fig. 2, which includes the search frame (most recent frame)), wherein the multi-head(Demoustier teaches “determin[ing] the target location by fusing information from multiple templates” in section 2.1, wherein the output of the encoder involves encoding both the search image (most recent frame) and a plurality of templates (number of preceding frames). As such, the decoder is utilized to utilize the output of the encoder in conjunction with the queries (catheter tip position and mask of catheter body) by comparing the encoder output and the decoder queries to promote regions with high similarities as shown in Section 2.1. Since Demoustier teaches a decoder which utilizes two object queries, wherein the tasks are completed simultaneously and use the same output from the encoder (see that “To guide the catheter tip tracking with spatial information, we incorporate additional contextual information by simultaneously segmenting the catheter body in the same frame”), it can be interpreted that the encoder as taught by Demoustier is a multi-head decoder); and
tracking at least one object comprised in the real-time time series of medical images, wherein the tracking comprises determining coordinates of the at least one object within an image plane based on the decoded most recent frame (Demoustier teaches “given a sequence of consecutive X-ray images and an initial location of the target catheter tip, our goal is to track the location of the target xt = (ut, vt), at any time t, t > 0.” in Section 2 using the decoder output as shown in Fig. 2).
Demoustier fails to teach that the method is implemented by a computer, wherein the encoder is a spatial-temporal encoder, and the decoder is a cross-attention decoder utilizing a predefined number of frames in the correlation process.
However, Thawakar teaches a method implemented by a computer (Thawakar teaches “the functions and processes of the in-vehicle computer system 114 may be implemented by one or more respective processing circuits 1126. A processing circuit includes a programmed processor as a processor includes circuitry” in para. [0067]), wherein the encoder is a spatial-temporal encoder (Thawakar teaches “the multi-scale spatio-temporal split (MS-STS) attention module FIG. 4 in the transformer encoder 310 effectively captures spatio-temporal feature relationships at multiple scales across frames in a video” in para. [0052]), and the decoder is a cross-attention decoder (Thawakar teaches “the transformer decoder in the base framework includes a series (layers) of alternating self- and cross-attention blocks, operating on the box queries BQ of individual frames” in para. [0060]).
Demoustier and Thawakar are both considered to be analogous to the claimed invention because they are in the same field of tracking objects of interest using a transformer architecture within videos. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier to incorporate the teachings of Thawakar and include that “the method is implemented by a computer, wherein the encoder is a spatial-temporal encoder, and the decoder is a cross-attention decoder”. The motivation for doing so would have been “improving the temporal consistency of the video mask predictions”, wherein “the encoder learns to better delineate foreground and background regions leading to improved video instance mask prediction” in order to “capture a continuous video sequence and simultaneously segment and track all object instances from a set of semantic categories”, as suggested by Thawakar in para. [0060], [0063], and [0040], respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier with Thawakar to obtain the invention specified in the above claim limitations.
Demoustier and Thawakar fail to explicitly teach that the decoder utilizes a predefined number of frames in the correlation process.
However, Xie teaches that the decoder utilizes a predefined number of frames in the correlation process (Xie teaches a process of, for each search frame, selecting a fixed number of frames from the set of historical frames as intermediate templates (see Section 4.1), wherein the method involves utilizing a ViTtrack which utilizes a decoder to decode a fused search/template image to locate and estimate the size of a target to be tracked (see Section 3.1), and wherein the method as taught by Xie adapts the fused image to include the predefined number of intermediate templates (frames) as taught in Section 3.4. Here, the process of decoding the fused image as taught by Xie is interpreted as equivalent to the process of determining a correlation between the frames).
Demoustier, Thawakar, and Xie are all considered to be analogous to the claimed invention because they are in the same field of tracking objects of interest using a transformer architecture within videos. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar) to incorporate the teachings of Xie and include that “the decoder utilizes a predefined number of frames in the correlation process”. The motivation for doing so would have been “ to integrate into the video backbone which decouples the redundant video information into the static & dynamic templates”, as suggested by Xie in Section 1. See also that Xie teaches “both the three-layer patterns IV (72.6%) and pattern V (71.5%), improve the performance by large margin comparing to the pattern II (62.6%). However, the performance degeneration when the frame number increases still exits. As our proposed disentangled dual-template mechanism (pattern VI) reduces the temporal redundancy in intermediate templates by cross attention, its performance has a rising tendency facing the longer video-clip (72.1% to 72.7%)” in Section 4.2 (Temporal Modelling). Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier and Thawakar with Xie to obtain the invention specified in claim 1.
Regarding claim 2, Demoustier, Thawakar, and Xie teach the method according to claim 1,
wherein the medical images are X-ray images (Demoustier teaches a “sequence of consecutive X-ray images” in Section 2) and/or wherein the medical images are chest images.
Regarding claim 3, Demoustier, Thawakar, and Xie teach the method according to claim 1, further comprising:
- providing the determined coordinates of the tracked at least one object (Demoustier teaches “given a sequence of consecutive X-ray images and an initial location of the target catheter tip, our goal is to track the location of the target xt = (ut, vt), at any time t, t > 0.” in Section 2 using the decoder output as shown in Fig. 2).
Regarding claim 9, Demoustier, Thawakar, and Xie teach the method according to claim 1, wherein the at least one object comprises:
- two or more objects, which are tracked separately; and/or
- two or more components and/or
parts of an extended object, which are tracked separately (Demoustier teaches tracking both the catheter tip and the catheter body separately (see the separate decoding tasks of the decoder as shown in Fig. 2), wherein the body and the tip are interpreted as parts of an extended object. See also that Thawakar teaches tracking multiple objects as shown in para. [0002], [0085], and FIG. 16).
Regarding claim 10, Demoustier, Thawakar, and Xie teach the method according to claim 1,
wherein the predefined number of preceding frames comprises between three and eight frames (Xie teaches that method involves selecting a fixed number (T=4) of frames from the set of historical frames to act as the intermediate templates in Section 4.1 (Online Inference)).
Similar motivations as applied to claim 1 can be applied here to claim 10.
Regarding claim 15, Demoustier teaches a system for tracking an object in a real-time time series of medical images (Demoustier teaches “given a sequence of consecutive X-ray images and an initial location of the target catheter tip, our goal is to track the location of the target xt = (ut, vt), at any time t, t > 0.” in Section 2. Demoustier teaches a transformer network for tracking a catheter tip downstream of a ResNet-50 encoder as shown in Fig. 2), the system comprising:
- an input layer configured for receiving a real-time time series of medical images of a patient's anatomical region (Demoustier teaches “a sequence of consecutive X-ray images” in section 2, wherein “the test dataset is divided into two primary categories: fluoroscopic and angiographic sequences. Fluoroscopic sequences are real-time videos of internal movements captured by low-dose X-rays without radiopaque substances, while angiographic sequences display blood vessels in real-time after the introduction of radiopaque substances” in Section 3. Demoustier further teaches inputting a set containing historically selected frames for templates [selected from the sequence of consecutive X-ray images] and a search frame (current frame) into the encoding layer of the transformer in section 2.1. See Figure 2 wherein the transformer network is downstream from the ResNet-50 encoder and the images are of a patient’s anatomy);
- a spatio-(Demoustier teaches encoding a search frame (current frame) and templates of historically selected frames within in a sequence of real-time consecutive x-ray images wherein each frame has a corresponding time as shown in Sections 2 and 2.1, wherein the method involves utilizing medical images as shown in Section 3);
- a multi-head (Demoustier teaches decoding the encoded representation as shown in Fig. 2, which includes the search frame (most recent frame)), wherein the multi-head (Demoustier teaches “determin[ing] the target location by fusing information from multiple templates” in section 2.1, wherein the output of the encoder involves encoding both the search image (most recent frame) and a plurality of templates (number of preceding frames). As such, the decoder is utilized to utilize the output of the encoder in conjunction with the queries (catheter tip position and mask of catheter body) by comparing the encoder output and the decoder queries to promote regions with high similarities as shown in Section 2.1. Since Demoustier teaches a decoder which utilizes two object queries, wherein the tasks are completed simultaneously and use the same output from the encoder (see that “To guide the catheter tip tracking with spatial information, we incorporate additional contextual information by simultaneously segmenting the catheter body in the same frame”), it can be interpreted that the encoder as taught by Demoustier is a multi-head decoder); and
- a tracking head configured for tracking at least one object comprised in the time series of medical images, wherein the tracking comprises determining coordinates of the at least one object within an image plane based on the decoded most recent frame (Demoustier teaches “given a sequence of consecutive X-ray images and an initial location of the target catheter tip, our goal is to track the location of the target xt = (ut, vt), at any time t, t > 0.” in Section 2 using the decoder output as shown in Fig. 2).
Demoustier fails to teach a processor configured to execute a downstream neural network (NN), wherein the encoder is a spatial-temporal encoder, and the decoder is a cross-attention decoder utilizing a predefined number of frames in the correlation process.
However, Thawakar teaches a processor configured to execute a downstream neural network (NN) (Thawakar teaches “the functions and processes of the in-vehicle computer system 114 may be implemented by one or more respective processing circuits 1126. A processing circuit includes a programmed processor as a processor includes circuitry” in para. [0067]), wherein the encoder is a spatial-temporal encoder (Thawakar teaches “the multi-scale spatio-temporal split (MS-STS) attention module FIG. 4 in the transformer encoder 310 effectively captures spatio-temporal feature relationships at multiple scales across frames in a video” in para. [0052]), and the decoder is a cross-attention decoder (Thawakar teaches “the transformer decoder in the base framework includes a series (layers) of alternating self- and cross-attention blocks, operating on the box queries BQ of individual frames” in para. [0060]).
Demoustier and Thawakar are both considered to be analogous to the claimed invention because they are in the same field of tracking objects of interest using a transformer architecture within videos. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier to incorporate the teachings of Thawakar and include that “the method is implemented by a computer, wherein the encoder is a spatial-temporal encoder, and the decoder is a cross-attention decoder”. The motivation for doing so would have been “improving the temporal consistency of the video mask predictions”, wherein “the encoder learns to better delineate foreground and background regions leading to improved video instance mask prediction” in order to “capture a continuous video sequence and simultaneously segment and track all object instances from a set of semantic categories”, as suggested by Thawakar in para. [0060], [0063], and [0040], respectively. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier with Thawakar to obtain the invention specified in the above claim limitations.
Demoustier and Thawakar fail to explicitly teach that the decoder utilizes a predefined number of frames in the correlation process.
However, Xie teaches that the decoder utilizes a predefined number of frames in the correlation process (Xie teaches a process of, for each search frame, selecting a fixed number of frames from the set of historical frames as intermediate templates (see Section 4.1), wherein the method involves utilizing a ViTtrack which utilizes a decoder to decode a fused search/template image to locate and estimate the size of a target to be tracked (see Section 3.1), and wherein the method as taught by Xie adapts the fused image to include the predefined number of intermediate templates (frames) as taught in Section 3.4. Here, the process of decoding the fused image as taught by Xie is interpreted as equivalent to the process of determining a correlation between the frames).
Demoustier, Thawakar, and Xie are all considered to be analogous to the claimed invention because they are in the same field of tracking objects of interest using a transformer architecture within videos. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar) to incorporate the teachings of Xie and include that “the decoder utilizes a predefined number of frames in the correlation process”. The motivation for doing so would have been “ to integrate into the video backbone which decouples the redundant video information into the static & dynamic templates”, as suggested by Xie in Section 1. See also that Xie teaches “both the three-layer patterns IV (72.6%) and pattern V (71.5%), improve the performance by large margin comparing to the pattern II (62.6%). However, the performance degeneration when the frame number increases still exits. As our proposed disentangled dual-template mechanism (pattern VI) reduces the temporal redundancy in intermediate templates by cross attention, its performance has a rising tendency facing the longer video-clip (72.1% to 72.7%)” in Section 4.2 (Temporal Modelling). Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier and Thawakar with Xie to obtain the invention specified in claim 15.
Claims 5, 7, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Demoustier et al. (“ConTrack: Contextual Transformer for Device Tracking in X-Ray”), hereinafter Demoustier, in view of Thawakar et al. (U.S. Publication No. 2024/0161334 A1), hereinafter Thawakar, Xie et al. (“VideoTrack: Learning to Track Objects via Video Transformer”), hereinafter Xie, and Cui et al. (CN 116385467 B, see attached English translation for citations), hereinafter Cui.
Regarding claim 5, Demoustier, Thawakar, and Xie teach the method according to claim 1,
wherein the spatio-temporal encoder is pretrained by performing a reconstruction task, wherein performing the reconstruction task comprises combining the spatio-temporal encoder with a reconstruction decoder.
However, Cui teaches wherein the spatio-temporal encoder is pretrained by performing a reconstruction task, wherein performing the reconstruction task comprises combining the spatio-temporal encoder with a reconstruction decoder (Cui teaches “the encoder is pre-trained based on the loss function corresponding to the vascular reconstruction task and the loss function of the vascular comparison learning task” in para. [0090], wherein the reconstruction loss is calculated by “inputting the blood vessel covering image into an encoder and a decoder, and reconstructing the covering blood vessel by the encoder and the decoder to obtain a cerebral blood vessel reconstruction image. And acquiring reconstruction loss according to the cerebral vessel image and the cerebral vessel reconstruction image” as shown in para. [0072]. See also that Thawakar teaches “an encoder that captures multi-scale spatio-temporal feature relationships” as shown in para.[0094]. As such, the teachings of Cui and Thawakar can be combined to teach the above subject matter, wherein the process of determining the reconstruction loss using both the decoder and encoder to pre-train the autoencoder is interpreted as equivalent to the claimed process of a task involving the combination of the encoder and decoder).
Demoustier, Thawakar, Xie, and Cui are all considered to be analogous to the claimed invention because they are in the same field of utilizing an encoder/decoder structure to analyze objects through image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Cui and include “wherein the spatio-temporal encoder is pretrained by performing a reconstruction task, wherein performing the reconstruction task comprises combining the spatio-temporal encoder with a reconstruction decoder”. The motivation for doing so would have been “so that the encoder with strong characteristic capturing capacity and expression capacity is obtained”, as suggested by Cui in para. [0090]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Cui to obtain the invention specified in claim 5.
Regarding claim 7, Demoustier, Thawakar, Xie, and Cui teach the method according to claim 5,
wherein an input to the spatio-temporal encoder is subject to tube masking and/or frame masking (Thawakar teaches “obtaining a video instance mask sequence from the sequence of image frames via […] a transformer encoder-decoder […], in which the encoder contains a multi-scale spatio-temporal split (MS-STS) attention module to capture spatio-temporal feature relationships at multiple scales across multiple frames of the sequence of image frames” in para. [0016]. See that the input to the encoder is the video sequence as shown in para. [0051]. See also para. [0052]-[0053]. Here, this masking process is interpreted as equivalent to the claimed frame masking) (Cui additionally teaches “after masking processing, the processed signals are input into an encoder of the multi-task self-supervision model for feature extraction” in para. [0114]).
Please note that it is ambiguous as to whether the masking occurs before or after being input into the encoder.
Similar motivations as applied to claims 1 and 5 can be applied here to claim 7.
Regarding claim 14, Demoustier, Thawakar, and Xie teach the method according to claim 1.
Demoustier, Thawakar, and Xie fail to teach further comprising:- performing a further downstream task comprising determining a stenosis, labelling a branch, determining a phase, and/or determining a vessel segmentation.
However, Cui teaches performing a further downstream task comprising determining
a stenosis,
labelling a branch,
determining a phase, and/or
determining a vessel segmentation (Cui teaches “segmenting a cerebral blood vessel based on self-supervised learning” as shown in para. [0091]. Since this process occurs after the self-supervised learning process, it is interpreted as equivalent to the claimed downstream task).
Demoustier, Thawakar, Xie, and Cui are all considered to be analogous to the claimed invention because they are in the same field of utilizing an encoder/decoder structure to analyze objects through image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Cui and include “performing a further downstream task comprising determining a stenosis, labelling a branch, determining a phase, and/or determining a vessel segmentation”. The motivation for doing so would have been so that “the cerebral blood vessel segmentation precision is improved”, as suggested by Cui in para. [0037]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Cui to obtain the invention specified in claim 14.
Claims 8, 11, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Demoustier et al. (“ConTrack: Contextual Transformer for Device Tracking in X-Ray’), hereinafter Demoustier, in view of Thawakar et al. (U.S. Publication No. 2024/0161334 A1), hereinafter Thawakar and Xie et al. (“VideoTrack: Learning to Track Objects via Video Transformer”), hereinafter Xie and Luengo Muntion et al. (U.S. Publication No. 2024/0206989 A1), hereinafter Luengo Muntion.
Regarding claim 8, Demoustier, Thawakar, and Xie teach the method according to claim 1.
Demoustier, Thawakar, and Xie fail to teach further comprising: initializing the tracking of the at least one object with application of a trained detection model for object detection on an initial frame of the received real-time time series of medical images.
However, Luengo Muntion teaches initializing the tracking of the at least one object with application of a trained detection model for object detection on an initial frame of the received real-time time series of medical images (While Demoustier teaches the received real-time time series of medical images (see claim 1), Luengo Muntion teaches “at block 206, the one or more machine-learning models 702 can detect one or more surgical instruments at least partially depicted in the video stream based on the input data, when such features are present. Detection of surgical instruments can include determining a presence or localization of one or more surgical instruments” wherein “upon surgical instrument detection through presence and/or localization, tracking can be performed to observe and predict positioning of the surgical instruments with respect to other structures” in para. [0068]. Luengo Muntion further teaches an input frame (interpreted as equivalent to the claimed initial frame) for which the object detection can take place as shown in para. [0131]. See also that Luengo Muntion teaches real-time video processing as shown in para. [0073]).
Demoustier, Thawakar, Xie, and Luengo Muntion are all considered to be analogous to the claimed invention because they are in the same field of utilizing an encoder/decoder structure to analyze objects through image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Luengo Muntion and include “initializing the tracking of the at least one object with application of a trained detection model for object detection on an initial frame of the received real-time time series of medical images”. The motivation for doing so would have been to “improve surgical procedures by improving the safety of the procedures”, as suggested by Luengo Muntion in para. [0099]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Luengo Muntion to obtain the invention specified in claim 8.
Regarding claim 11, Demoustier, Thawakar, and Xie teach the method according to claim 1.
Demoustier, Thawakar, and Xie fail to teach wherein the at least one object comprises a surgical instrument.
However Luengo Muntion teaches wherein the at least one object comprises a surgical instrument (Luengo Muntion teaches “the one or more machine-learning models 702 can include a plurality of feature encoders and task-specific decoders trained as an ensemble to detect the state and the one or more surgical instruments by sharing extracted features associated with the state and the one or more surgical instruments between the feature encoders and task-specific decoder” wherein “upon surgical instrument detection through presence and/or localization, tracking can be performed to observe and predict positioning of the surgical instruments with respect to other structures” as shown in para. [0068]).
Demoustier, Thawakar, Xie, and Luengo Muntion are all considered to be analogous to the claimed invention because they are in the same field of utilizing an encoder/decoder structure to analyze objects through image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Luengo Muntion and include “wherein the at least one object comprises a surgical instrument”. The motivation for doing so would have been to “improve surgical procedures by improving the safety of the procedures”, as suggested by Luengo Muntion in para. [0099]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Luengo Muntion to obtain the invention specified in claim 11.
Regarding claim 16, Demoustier, Thawakar, and Xie teach the system according to claim 15.
Demoustier, Thawakar, and Xie fail to teach wherein the downstream NN is further configured to initialize the tracking of the at least one object with application of a trained detection model for object detection on an initial frame of the received real-time time series of medical images.
However, Luengo Muntion teaches initializing the tracking of the at least one object with application of a trained detection model for object detection on an initial frame of the received real-time time series of medical images (While Demoustier teaches the received real-time time series of medical images (see claim 1), Luengo Muntion teaches “at block 206, the one or more machine-learning models 702 can detect one or more surgical instruments at least partially depicted in the video stream based on the input data, when such features are present. Detection of surgical instruments can include determining a presence or localization of one or more surgical instruments” wherein “upon surgical instrument detection through presence and/or localization, tracking can be performed to observe and predict positioning of the surgical instruments with respect to other structures” in para. [0068]. Luengo Muntion further teaches an input frame (interpreted as equivalent to the claimed initial frame) for which the object detection can take place as shown in para. [0131]. See also that Luengo Muntion teaches real-time video processing as shown in para. [0073]).
Demoustier, Thawakar, Xie, and Luengo Muntion are all considered to be analogous to the claimed invention because they are in the same field of utilizing an encoder/decoder structure to analyze objects through image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Luengo Muntion and include “wherein the downstream NN is further configured to initialize the tracking of the at least one object with application of a trained detection model for object detection on an initial frame of the received real-time time series of medical images”. The motivation for doing so would have been to “improve surgical procedures by improving the safety of the procedures”, as suggested by Luengo Muntion in para. [0099]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Luengo Muntion to obtain the invention specified in claim 16.
Claims 12 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Demoustier et al. (“ConTrack: Contextual Transformer for Device Tracking in X-Ray’), hereinafter Demoustier, in view of Thawakar et al. (U.S. Publication No. 2024/0161334 A1), hereinafter Thawakar, Xie et al. (“VideoTrack: Learning to Track Objects via Video Transformer”), hereinafter Xie, and Feng et al. (CN 105405152 B, see attached English translation for citations), hereinafter Feng.
Regarding claim 12, Demoustier, Thawakar, and Xie teach the method according to claim 1.
Demoustier, Thawakar, and Xie fail to teach further comprising:- symmetrically cropping any frame within the received real-time time series of medical images.
However, Feng teaches symmetrically cropping any frame within the received real-time time series of medical images (Feng teaches cropping each frame of a video sequence V into uniform size in para. [0044], which is interpreted as equivalent to the claimed “symmetrically cropping” process” as described in the above claim limitation and also Figures 15A and 15B of the applicant’s disclosure. See claim 1 wherein Demoustier teaches the received real-time time series of medical images. The teaching of Feng can be combined with the teaching of Demoustier to teach the above claim limitation).
Demoustier, Thawakar, Xie, and Feng are all considered to be analogous to the claimed invention because they are in the same field of analyzing objects in a video through machine learning and image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Feng and include “symmetrically cropping any frame within the received real-time time series of medical images”. The motivation for doing so would have been that “on the one hand[, uniform cropping] can reduce image [r]esolution ratio [] so as to reduce calculation amount, speed up processing, on the other hand different images frame can be unified for the same coordinate system from [a] unified reference standard”, as suggested by Feng in para. [0044]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Feng to obtain the invention specified in claim 12.
Regarding claim 17, Demoustier, Thawakar, and Xie teach the system according to claim 15.
Demoustier, Thawakar, and Xie fail to teach wherein the processor is configured to symmetrically crop any frame within the received real-time time series of medical images.
However, Feng teaches wherein the processor is configured to symmetrically crop any frame within the received real-time time series of medical images (Feng teaches cropping each frame of a video sequence V into uniform size in para. [0044], which is interpreted as equivalent to the claimed “symmetrically cropping” process” as described in the above claim limitation and also Figures 15A and 15B of the applicant’s disclosure. See claim 1 wherein Demoustier teaches the received real-time time series of medical images. The teaching of Feng can be combined with the teaching of Demoustier to teach the above claim limitation).
Demoustier, Thawakar, Xie, and Feng are all considered to be analogous to the claimed invention because they are in the same field of analyzing objects in a video through machine learning and image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Feng and include “wherein the processor is configured to symmetrically crop any frame within the received real-time time series of medical images”. The motivation for doing so would have been that “on the one hand[, uniform cropping] can reduce image [r]esolution ratio [] so as to reduce calculation amount, speed up processing, on the other hand different images frame can be unified for the same coordinate system from [a] unified reference standard”, as suggested by Feng in para. [0044]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Feng to obtain the invention specified in claim 17.
Claims 13 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Demoustier et al. (“ConTrack: Contextual Transformer for Device Tracking in X-Ray’), hereinafter Demoustier, in view of Thawakar et al. (U.S. Publication No. 2024/0161334 A1), hereinafter Thawakar, Xie et al. (“VideoTrack: Learning to Track Objects via Video Transformer”), hereinafter Xie, and Feng et al. (CN 105405152 B, see attached English translation for citations), hereinafter Feng.
Regarding claim 13, Demoustier, Thawakar, and Xie teach the method according to claim 1.
Demoustier and Thawakar teach using the multi-head cross-attention decoder in the decoding (see claim 1) which involves detecting spatial correlation and tracking at least one object (see claim 1).
Demoustier, Thawakar, and Xie fail to teach
wherein,
(path 1) using the multi-head cross-attention decoder in the decoding, a background is removed for spatial correlation, and a historical trajectory of the at least one object and/or
(path 2) the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features. Due to the “and/or” language above, only one of the pathways identified above need be found in the prior art.
However, Susnow teaches that the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features (Susnow teaches “temporal redundancies in video data between neighboring video frames or fields are identified so that an encoder only needs to pass a motion vector to a decoder” in para. [0004]).
Demoustier, Thawakar, Xie, and Susnow are all considered to be analogous to the claimed invention because they are in the same field of analyzing objects in a video through machine learning and image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Susnow and include that “the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features”. The motivation for doing so would have been “so that an encoder only needs to pass a motion vector to a decoder, instead of retransmitting redundant data”, as suggested by Susnow in para. [0004]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Susnow to obtain the invention specified in claim 13.
Regarding claim 18, Demoustier, Thawakar, and Xie teach the system according to claim 15.
Demoustier and Thawakar teach using the multi-head cross-attention decoder in the decoding (see claim 1) which involves detecting spatial correlation and tracking at least one object (see claim 1).
Demoustier, Thawakar, and Xie fail to teach
wherein
(path 1) the multi-head cross-attention decoder is configured to remove a background for spatial correlation, and a historical trajectory of the at least one object and/or
(path 2) the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features.
However, Susnow teaches that the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features (Susnow teaches “temporal redundancies in video data between neighboring video frames or fields are identified so that an encoder only needs to pass a motion vector to a decoder” in para. [0004]).
Demoustier, Thawakar, Xie, and Susnow are all considered to be analogous to the claimed invention because they are in the same field of analyzing objects in a video through machine learning and image analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Demoustier (as modified by Thawakar and Xie) to incorporate the teachings of Susnow and include that “the decoding based on the predefined number of preceding frames is applied solely on motion-preserved features”. The motivation for doing so would have been “so that an encoder only needs to pass a motion vector to a decoder, instead of retransmitting redundant data”, as suggested by Susnow in para. [0004]. Therefore, it would have been obvious to one of ordinary skill at the time the invention was filed to combine Demoustier, Thawakar, and Xie with Susnow to obtain the invention specified in claim 18.
Allowable Subject Matter
Claims 4 and 6 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
Claims 19-20 are allowed.
The following is a statement of reasons for the indication of allowable subject matter.
The best prior art of record is Demoustier, Thawakar, Xie, Shen et al. (CN 113870319 A, see attached English translation for citations), hereinafter Shen, Cui, Luengo Muntion, and Feng. Prior art applied alone or in combination with fails to anticipate or render obvious claims 4, 6, 19, and 20.
Claim 4
Regarding claim 4, Demoustier, Thawakar, and Xie teach the method according to claim 1.
Cui further teaches pretraining the spatio-temporal encoder using self-supervised learning (SSL) wherein the SSL comprises at least one of the tasks selected from the following group, consisting of: * determining a cardiac phase; * determining a stenosis; * determining a vessel segmentation.
However, neither Demoustier, nor Thawakar, nor Xie, nor Shen, nor Cui, nor Luengo Muntion, nor the combination, teaches wherein performing the at least one SSL task comprises combining the spatio-temporal encoder with a task-specific weak-label decoder for the at least one SSL task.
Claim 6 includes allowable subject matter by virtue of being dependent upon claim 4.
Claim 19
Regarding claim 19, Demoustier teaches a training system for training a downstream neural network (NN) for tracking an object in a real-time time series of medical images, the training system comprising:
- a spatio-temporal encoder for encoding the received real-time time series of medical images and obtaining an encoded representation per frame of the received real-time time series, wherein a frame corresponds to a medical image at a time instance within the real-time time series of medical images.
Thawakar further teaches wherein the encoder is a spatial-temporal encoder.
Shen further teaches a reconstruction decoder
However, neither Demoustier, nor Thawakar, nor Xie, nor Shen, nor Cui, nor Luengo Muntion, nor the combination, teaches the corresponding algorithm of the spatio-temporal encoder as found in para. [0135]-[138], [0159], [0160], [0165]-[0166], and [0178] of applicant’s specification.
Claim 20 includes allowable subject matter by virtue of being dependent upon claim 4.
***Please note that the spatio-temporal encoder is being interpreted under 112(f) as a computer-implemented means-plus-function limitation, wherein the corresponding algorithm of the spatio-temporal encoder is being read into the limitation: “a spatio-temporal encoder for encoding”. Applicant’s specification states that “the spatio-temporal encoder 204 […] may be embodied (executed) by the processor 214” in para. [0143]. Therefore, the corresponding structure for the spatio-temporal encoder is a processor. Claiming a means for performing a specific computer-implemented function and disclosing only a general-purpose computer as its structure amounts to pure functional claiming. Aristocrat, 521 F.3d 1328 at 1333, 86 USPQ2d at 1239. In this instance, the structure corresponding to a 35 U.S.C. 112(f) claim limitation for a computer-implemented function must include the algorithm needed to transform the general purpose computer or microprocessor disclosed in the specification. See MPEP 2181(II)(B). The specific information in the specification regarding the algorithm associated with the spatio-temporal encoder that makes this limitation allowable when analyzed in conjunction with the rest of the claim elements includes, but is not limited to para. [0135]-[138], [0159], [0160], [0165]-[0166], and [0178] of applicant’s specification.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner
should be directed to KYLA G ALLEN whose telephone number is (703)756-5315. The examiner can
normally be reached M-F 7:30am - 4:30pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a
USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use
the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor,
John Villecco can be reached on (571) 272-7319. The fax phone number for the organization where this
application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from
Patent Center. Unpublished application information in Patent Center is available to registered users. To
file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit
https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and
https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional
questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like
assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or
571-272-1000.
/Kyla Guan-Ping Tiao Allen/
Examiner, Art Unit 2661
/AARON W CARTER/Primary Examiner, Art Unit 2661