Prosecution Insights
Last updated: October 02, 2026
Application No. 18/438,640

TARGET TRACKING IN MEDICAL IMAGE DATA

Final Rejection §103
Filed
Feb 12, 2024
Priority
Mar 02, 2023 — provisional 63/487,961 +1 more
Examiner
HELCO, NICHOLAS JOHN
Art Unit
2667
Tech Center
2600 — Communications
Assignee
Siemens Healthineers AG
OA Round
2 (Final)
70%
Grant Probability
Favorable
3-4
OA Rounds
2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
33 granted / 47 resolved
+8.2% vs TC avg
Strong +43% interview lift
Without
With
+43.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
19 currently pending
Career history
71
Total Applications
across all art units

Statute-Specific Performance

§101
19.8%
-20.2% vs TC avg
§103
51.0%
+11.0% vs TC avg
§102
17.2%
-22.8% vs TC avg
§112
9.9%
-30.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 47 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Notice to Applicants This action is in response to the amendments and remarks filed on 06/24/2026. Claims 1 and 3-12 are pending. Corrective Actions by Applicant Claims 1, 3-4, and 11 have been amended. Claim 2 has been canceled. Response to Arguments The examiner has fully considered Applicant’s presented arguments. On page 6 of the remarks, Applicant argues that all claims are eligible under 35 U.S.C. 101 after the present amendments. This is persuasive, based on the inclusion of the eligible subject matter of claims 2 and/or 4-5 into the independent claims. All 101 rejections have thus been withdrawn. On page 7 of the remarks, Applicant argues that, regarding claim 1 and its dependents, Benseghir fails to disclose determining the optical flow based specifically on segmentations of the context, instead of raw image frames. This is persuasive, specifically because, as argued by Applicant, Benseghir does not appear to use the segmentation of the tool to determine the optical flow. Although Benseghir does separately segment the tool beforehand (see figure 3, step 325 and paragraph 0032), the subsequent optical flow calculation of the tool and its context does not appear to use said segmentation, but rather just the raw image data (see figure 3, step 330 and paragraph 0033). Thus, all previous 35 U.S.C. 103 rejections are withdrawn. However, the examiner argues that calculating optical flow of an object using successive segmentations/masks of said object is still well-known in the art, and that this would be a predictable modification to Benseghir and Zhang as both already disclose successive segmentations and involve optical flows, as shown in the updated 103 rejections below. On page 7 of the remarks, Applicant argues that, regarding claim 11 and its dependents, Benseghir fails to disclose the same elements as argued above, and also fails to disclose refining the segmentation of the context based on the optical flow and a vessel segmentation. This is persuasive, specifically because it includes the persuasive arguments from the first argument above, and thus the 103 rejections of claims 11-12 are withdrawn as well. The examiner notes that Benseghir does still disclose refining the context segmentation based on the optical flow and a vessel segmentation. The claim does not limit how the vessel segmentation is involved in correcting the context segmentation, only requiring the correction to be “based on” the vessel segmentation. Benseghir discloses, in cited paragraphs 0041-0048, a process involving registering/segmenting an image of patient anatomy, such as vessels, to an anatomical model, then registering the already segmented tool to the anatomy, such as by limiting the segmentation to be within boundaries of the vessels in paragraph 0045. Thus, the vessel segmentation is certainly related to refining the tool segmentation, as required by the claim. Claim Rejections – 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-8, and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. ("A Temporary Transformer Network for Guide-Wire Segmentation", CISP-BMEI paper, 7 December 2021) in view of Benseghir et al. (U.S. Publ. US-2023/0083936-A1) and Karmann et al. (U.S. Patent US-5034986-A). Regarding claim 1, Zhang discloses Zhang discloses a computer-implemented method (see figure 1) of tracking a target in medical imaging data (see Section I, where guide-wires are to be segmented from medical imaging data; see Section III.A, where a catheter image dataset is collected and subsequently analyzed), the method comprising: determining an encoded representation of a search image of the medical imaging data in a feature space using a feature encoding network (see figure 1, current frame and Section II.A, where the current frame/search image is input into a CNN to obtain feature representations/encoded representations), the search image depicting a target and a surrounding of the target (see figure 1, current frame and Section II.B, which specifies that the current frame depicts guide-wire information), determining encoded representations of one or more template images of the medical imaging data in the feature space using the feature encoding network (see figure 1, previous frame and Section II.A, where a previous frame/template image is input into another CNN to obtain feature representations/encoded representations), the one or more template images depicting the target (see figure 1, previous frame and Section II.B, which specifies that the previous frame also depicts guide-wire information), determining fused features by fusing the encoded representations of the one or more template images and the encoded representation of the search image using a fusion network (see figure 1, feature concatenation boxes and Section II.B, where the features of the current and previous frames are concatenated to generate a mixed feature representation), based on the fused features, determining a position prediction of the target in the search image, based on the fused features, determining a segmentation of a context of the target in the search image (see figure 1, segmentation result; figure 3, more segmentation results, and Section II.B, where the mixed features are used to generate a segmentation mask of the guide-wire; this reads on both a position prediction of the guide-wire, as well as a segmentation of the context/entire body of the wire), Zhang fails to disclose determining an optical flow based on the segmentation of the context and one or more previous segmentations of the context, and refining the position prediction of the target based on the optical flow. Pertaining to the same field of endeavor, Benseghir discloses determining a position prediction of the target in the search image (see figure 3, step 335 and paragraph 0034, where a position of an interventional tool/target is estimated), determining a segmentation of a context of the target in the search image (see figure 3, step 325 and paragraph 0032, where the interventional tool/target is additionally segmented), determining an optical flow (see figure 3, step 330 and paragraph 0033, where an optical flow is determined between the input and previous images, including movement of both tissue and the tool), and refining the position prediction of the target based on the optical flow (see figure 3, step 340 and paragraph 0034, where the estimated position & segmentation of the tool are corrected/refined based on the optical flow of the tissue and tool). Zhang and Benseghir are considered analogous art, as they are both directed to deep learning models for medical instrument segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Benseghir into Zhang because doing so increases accuracy of instrument segmentation by accounting for patient movement (see Benseghir paragraphs 0034, 0038, and 0052). Zhang in view of Benseghir fails to further disclose determining an optical flow based on the segmentation of the context and one or more previous segmentations of the context. More specifically, Benseghir does segment the context and separately determine an optical flow; they merely don’t use the segmentations as part of the optical flow generation. Pertaining to the same field of endeavor, Karmann discloses determining an optical flow based on the segmentation of the context and one or more previous segmentations of the context (see figure 1 and column 4, line 66 to column 5, line 21, where an object segmentation mask can be calculated for current and past images, then the gray values of said masks matched to determine motion vectors / optical flow). Zhang and Karmann are considered analogous art, as they are both directed to image-based object tracking. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Karmann into Zhang and Benseghir by using the object segmentations as part of the optical flow generation because using object segmentations for matching produces smooth motion fields (see Karmann column 7, lines 35-47). Regarding claim 3, Zhang in view of Karmann fails to disclose the limitations of claim 3. Pertaining to the same field of endeavor, Benseghir discloses wherein said refining of the position prediction comprises applying a refinement network to the optical flow and the position prediction of the target (see figure 3, step 340 and paragraph 0034, where the estimated position & segmentation of the tool are corrected/refined based on the optical flow of the tissue and tool; paragraph 0019 specifies that deep learning techniques can be used for the disclosed image processing techniques). Zhang and Benseghir are considered analogous art, as they are both directed to deep learning models for medical instrument segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Benseghir into Zhang and Karmann because doing so increases accuracy of instrument segmentation by accounting for patient movement (see Benseghir paragraphs 0034, 0038, and 0052). Regarding claim 4, Zhang in view of Karmann fails to disclose the limitations of claim 4. Pertaining to the same field of endeavor, Benseghir discloses refining the segmentation of the context of the target based on the optical flow (see figure 3, step 340 and paragraph 0034, where the estimated position & segmentation of the tool are corrected/refined based on the optical flow of the tissue and tool). Zhang and Benseghir are considered analogous art, as they are both directed to deep learning models for medical instrument segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Benseghir into Zhang and Karmann because doing so increases accuracy of instrument segmentation by accounting for patient movement (see Benseghir paragraphs 0034, 0038, and 0052). Regarding claim 5, Zhang in view of Karmann fails to disclose the limitations of claim 5. Pertaining to the same field of endeavor, Benseghir discloses wherein the segmentation of the context of the target is refined further based on a vessel segmentation (see figure 5, steps 505-560 and paragraphs 0041-0048, where the method can include identifying anatomical features such as vessels, registering the tool to a model of the anatomy, and correcting the tool segmentation based on the optical flow of the anatomy). Zhang and Benseghir are considered analogous art, as they are both directed to deep learning models for medical instrument segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Benseghir into Zhang and Karmann because doing so increases accuracy of instrument segmentation by accounting for patient movement (see Benseghir paragraph 0047). Regarding claim 6, Zhang in view of Karmann fails to disclose the limitations of claim 6. Pertaining to the same field of endeavor, Benseghir discloses wherein the segmentation of the context of the target is refined in a spatial-temporal mask refinement (the correction processes cited in regards to claim 4 above combine both spatial information from the segmentation with temporal information from the optical flow to correct/refine the segmentation mask). Zhang and Benseghir are considered analogous art, as they are both directed to deep learning models for medical instrument segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Benseghir into Zhang and Karmann because doing so increases accuracy of instrument segmentation by accounting for patient movement (see Benseghir paragraphs 0034, 0038, and 0052). Regarding claim 7, Zhang in view of Benseghir and Karmann discloses wherein the target is a tip of an interventional medical instrument, and wherein the context is a body of the interventional medical instrument extending away from the tip (see Zhang figure 1, segmentation result and figure 3, more segmentation results, where the segmentation includes both a position prediction of a tip of the instrument as well as of the entire body of the instrument). Regarding claim 8, Zhang in view of Karmann fails to disclose the limitations of claim 8. Pertaining to the same field of endeavor, Benseghir discloses wherein the context are predefined anatomical features in a surrounding of the target (see figure 5, steps 505-560 and paragraphs 0041-0048, where the method can include identifying anatomical features such as vessels, registering the tool to a model of the anatomy, and correcting the tool segmentation based on the optical flow of the anatomy). Zhang and Benseghir are considered analogous art, as they are both directed to deep learning models for medical instrument segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Benseghir into Zhang and Karmann because doing so increases accuracy of instrument segmentation by accounting for patient movement (see Benseghir paragraph 0047). Regarding claim 11, Zhang discloses a processing device (see figure 1, operating console 142) comprising: a processor (see figure 1, processor 181), and a memory storing program code (see figure 1, memory 182), wherein the processor is configured to load and execute the program code, upon executing the program code, the processor being configured to (see paragraph 0021): determine an encoded representation of a search image of medical imaging data in a feature space using a feature encoding network, the search image depicting a target and a surrounding of the target, determine encoded representations of one or more template images of the medical imaging data in the feature space using the feature encoding network, the one or more template images depicting the target, determine fused features by fusing the encoded representations of the one or more template images and the encoded representation of the search image using a fusion network, based on the fused features, determine a position prediction of the target in the search image, based on the fused features, determine a segmentation of a context of the target in the search image (see citations to identical limitations in claim 1 above), Zhang fails to disclose determine an optical flow based on the segmentation of the context and one or more previous segmentations of the context; refine the segmentation of the context of the target based on the optical flow and a vessel segmentation; and refine the position prediction of the target based on the refined segmentation of the context of the target. Pertaining to the same field of endeavor, Benseghir discloses determine an optical flow (see citations to identical limitation in claim 1 above); refine the segmentation of the context of the target based on the optical flow (see figure 3, step 340 and paragraph 0034, where the estimated position & segmentation of the tool are corrected/refined based on the optical flow of the tissue and tool) and a vessel segmentation (see figure 5, steps 505-560 and paragraphs 0041-0048, where the method can include identifying anatomical features such as vessels, registering the tool to a model of the anatomy, and correcting the tool segmentation based on the optical flow of the anatomy; paragraph 0045 in particular specifies that this can include imposing boundaries to the segmentation based on the vessels); and refine the position prediction of the target based on the refined segmentation of the context of the target (see figure 5, step 560 and paragraph 0048, where the registered position of the segmented tool to the model can be further updated). Zhang and Benseghir are considered analogous art, as they are both directed to deep learning models for medical instrument segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Benseghir into Zhang because doing so increases accuracy of instrument segmentation by accounting for patient movement (see Benseghir paragraphs 0034, 0038, and 0052). Zhang in view of Benseghir fails to further disclose determine an optical flow based on the segmentation of the context and one or more previous segmentations of the context. More specifically, Benseghir does segment the context and separately determine an optical flow; they merely don’t use the segmentations as part of the optical flow generation. Pertaining to the same field of endeavor, Karmann discloses determine an optical flow based on the segmentation of the context and one or more previous segmentations of the context (see citations to identical limitation in claim 1 above). Zhang and Karmann are considered analogous art, as they are both directed to image-based object tracking. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Karmann into Zhang and Benseghir by using the object segmentations as part of the optical flow generation because using object segmentations for matching produces smooth motion fields (see Karmann column 7, lines 35-47). Regarding claim 12, Zhang in view of Benseghir and Karmann discloses claim 12 as applied to claim 7 above. Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. ("A Temporary Transformer Network for Guide-Wire Segmentation", CISP-BMEI paper, 7 December 2021) in view of Benseghir et al. (U.S. Publ. US-2023/0083936-A1) and Karmann et al. (U.S. Patent US-5034986-A), and further in view of Yan et al. ("Learning Spatio-Temporal Transformer for Visual Tracking", ICCV paper, 28 February 2022). Regarding claim 9, Zhang in view of Benseghir and Karmann fails to disclose the limitations of claim 9. Pertaining to the same field of endeavor, Yan discloses wherein each of the one or more template images has at least one of a lower resolution or a smaller size than the search image (see figures 2 and 4, where the template images are smaller than the search regions/search images, while still including the object to be tracked; Section 4.1 specifies that the search images can be 320x320 pixels, whereas template images can be 128x128 pixels). Zhang and Yan are considered analogous art, as they are both directed to deep learning models for object segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Yan into Zhang, Benseghir, and Karmann because one of ordinary skill in the art would recognize that using lower resolutions/sizes for template images improves performance speed while still including the object to be tracked. Regarding claim 10, Zhang in view of Benseghir and Karmann fails to disclose the limitations of claim 10. Pertaining to the same field of endeavor, Yan discloses wherein the fused features are determined using a vision transformer network (see figures 2, 4, and Section 3.1, where a transformer encoder fuses the features of the search and template images using a multi-head self-attention module). Zhang and Yan are considered analogous art, as they are both directed to deep learning models for object segmentation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have integrated the teachings of Yan into Zhang, Benseghir, and Karmann because the transformer encoder captures dependencies among all elements in the input sequence, improving detection accuracy (see Yan Section 3.1). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS JOHN HELCO whose telephone number is (703)756-5539. The examiner can normally be reached on Monday-Friday from 9:00 AM to 5:00 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matthew Bella, can be reached at telephone number 571-272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from Patent Center. Status information for published applications may be obtained from Patent Center. Status information for unpublished applications is available through Patent Center for authorized users only. Should you have questions about access to Patent Center, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form. /NICHOLAS JOHN HELCO/Examiner, Art Unit 2667 /MATTHEW C BELLA/Supervisory Patent Examiner, Art Unit 2667
Read full office action

Prosecution Timeline

Feb 12, 2024
Application Filed
Apr 01, 2026
Non-Final Rejection mailed — §103
Jun 24, 2026
Response Filed
Aug 26, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718557
METHOD FOR DETECTING AND RESPONDING TO CONDITIONS WITHIN A FOREST
2y 8m to grant Granted Aug 25, 2026
Patent 12705856
GLOBAL CONTEXT VISION TRANSFORMER
3y 7m to grant Granted Aug 11, 2026
Patent 12705883
METHOD FOR DETERMINING AREAS OF LAND DISTINCT FROM PREDETERMINED OBSTACLES AND COMPATIBLE WITH THE INSTALLATION OF PHOTOVOLTAIC PANELS
2y 10m to grant Granted Aug 11, 2026
Patent 12670722
SELF-SUPERVISED COMPOSITIONAL FEATURE REPRESENTATION FOR VIDEO UNDERSTANDING
3y 6m to grant Granted Jun 30, 2026
Patent 12670713
INFORMATION PROVIDING SYSTEM, INFORMATION PROVIDING METHOD AND PROGRAM RECORDING MEDIUM
2y 9m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
70%
Grant Probability
99%
With Interview (+43.1%)
2y 10m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 47 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month