Prosecution Insights
Last updated: October 02, 2026
Application No. 17/918,146

AIDING A USER TO PERFORM A MEDICAL ULTRASOUND EXAMINATION

Non-Final OA §102§103
Filed
Oct 11, 2022
Priority
Apr 16, 2020 — provisional 63/010,818 +1 more
Examiner
MALDONADO, STEVEN
Art Unit
3797
Tech Center
3700 — Mechanical Engineering & Manufacturing
Assignee
Koninklijke Philips N.V.
OA Round
5 (Non-Final)
27%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
70%
With Interview

Examiner Intelligence

Grants only 27% of cases
27%
Career Allowance Rate
7 granted / 26 resolved
-43.1% vs TC avg
Strong +43% interview lift
Without
With
+42.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
34 currently pending
Career history
86
Total Applications
across all art units

Statute-Specific Performance

§101
6.6%
-33.4% vs TC avg
§103
56.9%
+16.9% vs TC avg
§102
13.2%
-26.8% vs TC avg
§112
22.0%
-18.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 26 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendments In view of the Appeal Brief filed on March 3rd, 2026, PROSECUTION IS HEREBY REOPENED. See the rejection set forth below. To avoid abandonment of the application, appellant must exercise one of the following two options: (1) file a reply under 37 CFR 1.111 (if this Office action is non-final) or a reply under 37 CFR 1.113 (if this Office action is final); or, (2) initiate a new appeal by filing a notice of appeal under 37 CFR 41.31 followed by an appeal brief under 37 CFR 41.37. The previously paid notice of appeal fee and appeal brief fee can be applied to the new appeal. If, however, the appeal fees set forth in 37 CFR 41.20 have been increased since they were previously paid, then appellant must pay the difference between the increased fees and the amount previously paid. A Supervisory Patent Examiner (SPE) has approved of reopening prosecution by signing below Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim 14 is rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Droste et al (R. Droste et al., “Ultrasound image representation learning by modeling sonographer visual attention,” Lecture Notes in Computer Science, pp. 592–604, 2019; hereinafter referred to as Droste) Regarding Claim 14, Droste teaches a method of training a model for use in aiding a user to perform a medical ultrasound examination (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract]), the method comprising: obtaining training data comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating a relevance of one or more image components in the respective example ultrasound image to the medical ultrasound examination (“This allocation of visual attention is typically quantified via the distribution of gaze points, which can be recorded with gaze tracking. There has been great interest in developing models of human visual attention that, given an image, predict the likelihood that each pixel is fixated upon, hereafter referred to as visual saliency map.” [Pg. 1], “Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]); training the model to predict a relevance to the medical ultrasound examination of one or more image components in an ultrasound image, based on the training data, wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when the radiologist would consider to inspect or further investigate the image frame, in a context of performing the medical ultrasound examination(“This allocation of visual attention is typically quantified via the distribution of gaze points, which can be recorded with gaze tracking. There has been great interest in developing models of human visual attention that, given an image, predict the likelihood that each pixel is fixated upon, hereafter referred to as visual saliency map.” [Pg. 1], “Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]); and outputting a relevance value or score for the one or more image components in the image frame based on the predicted relevance (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “Visual Saliency Prediction. Given an image and a gaze point set (X, G) ∈ , the idea is to generate a visual saliency map S ∈ ]0, 1]H×W, where Si,j is the probability that pixel Xi,j is fixated upon.” [Pg. 4]). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-6, 8, 13, & 16-21 are rejected under 35 U.S.C. 103 as being unpatentable over Baumgartner et al (C. F. Baumgartner et al., “SonoNet: Real-time detection and localisation of fetal standard scan planes in freehand ultrasound,” IEEE Transactions on Medical Imaging, vol. 36, no. 11, pp. 2204–2215, Nov. 2017; hereinafter referred to as Baumgartner) in view of Droste. Regarding Claim 1, Baumgartner discloses a system for aiding a user to perform a medical ultrasound examination (“a novel method based on convolutional neural networks, which can automatically detect 13 fetal standard views in freehand 2-D ultrasound data as well as provide a localization of the fetal structures via a bounding box.” [Abstract], “The proposed system can be used in a number of ways. It can be employed to provide real-time feedback about the content of a image frame to the operator.” [Pg. 2205]), the system comprising: a memory comprising a set of executable instructions; a processor; and a display; wherein the processor is configured to communicate with the memory and to execute the set of instructions, and wherein the set of instructions (“The optimisation typically converged after around 2 days of training on a Nvidia GeForce GTX 1080 GPU.” [Pg. 2209]), when executed by the processor, cause the processor to: i) receive a real-time sequence of ultrasound images captured by an ultrasound probe during the medical ultrasound examination (“we propose a novel system based on convolutional neural networks (CNNs) for real-time automated detection of 13 fetal standard scan planes, as well as localisa tion of the fetal structures associated with each scan plane in the images via bounding boxes.” [Pg. 2204]); iii) highlight to the user, in real-time on the display. image components that are predicted by the model to be relevant to the medical ultrasound examination, for further consideration by the user (“The network architecture is designed to operate in real-time while providing optimal output for the localization task.” [Abstract]), Baumgartner does not specifically disclose that ii) use a model comprising the executable instructions and trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed, wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination; wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame. However, in a similar field of endeavor, Droste teaches transferable representations of images can be learned without manual annotations by modeling human visual attention [Abstract]. Droste also teaches ii) use a model comprising the executable instructions and trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “This allocation of visual attention is typically quantified via the distribution of gaze points, which can be recorded with gaze tracking. There has been great interest in developing models of human visual attention that, given an image, predict the likelihood that each pixel is fixated upon, hereafter referred to as visual saliency map.” [Pg. 1], “Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]), wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “Visual Saliency Prediction. Given an image and a gaze point set (X, G) ∈ , the idea is to generate a visual saliency map S ∈ ]0, 1]H×W, where Si,j is the probability that pixel Xi,j is fixated upon.” [Pg. 4]); wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “Visual Saliency Prediction. Given an image and a gaze point set (X, G) ∈ , the idea is to generate a visual saliency map S ∈ ]0, 1]H×W, where Si,j is the probability that pixel Xi,j is fixated upon.” [Pg. 4]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above with ii) use a model comprising the executable instructions and trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed, wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination; wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 2, Baumgartner discloses that the processor is further caused to repeat blocks iii) for a plurality of image frames in the real-time sequence of ultrasound images (“The network architecture is designed to operate in real-time while providing optimal output for the localization task.” [Abstract]). Baumgartner does not specifically disclose that the processor is further caused to repeat blocks ii) for a plurality of image frames in the real-time sequence of ultrasound images. However, in a similar field of endeavor, Droste teaches that the processor is further caused to repeat blocks ii) for a plurality of image frames in the real-time sequence of ultrasound images (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “This allocation of visual attention is typically quantified via the distribution of gaze points, which can be recorded with gaze tracking. There has been great interest in developing models of human visual attention that, given an image, predict the likelihood that each pixel is fixated upon, hereafter referred to as visual saliency map.” [Pg. 1], “Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above the processor is further caused to repeat blocks ii) for a plurality of image frames in the real-time sequence of ultrasound images as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 3, Baumgartner discloses all limitations noted above except that the model is trained using the machine learning process on training data comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating a relevance of one or more image components in the respective example ultrasound image to the medical ultrasound examination. However, in a similar field of endeavor, Droste teaches that the model is trained using the machine learning process on training data comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating a relevance of one or more image components in the respective example ultrasound image to the medical ultrasound examination (“Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above with the model is trained using the machine learning process on training data comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating a relevance of one or more image components in the respective example ultrasound image to the medical ultrasound examination as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 4, Baumgartner discloses all limitations noted above except that the ground truth annotations are based on gaze tracking information obtained from observing a radiologist analysing the respective example ultrasound image for the purpose of the medical ultrasound examination. However, in a similar field of endeavor, Droste teaches that the ground truth annotations are based on gaze tracking information obtained from observing a radiologist analysing the respective example ultrasound image for the purpose of the medical ultrasound examination (“Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above with the ground truth annotations are based on gaze tracking information obtained from observing a radiologist analysing the respective example ultrasound image for the purpose of the medical ultrasound examination as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 5, Baumgartner discloses that the model is further trained to output an indication of a confidence associated with the predicted relevance for the one or more image components in the image frame ("After training we fed the network with cropped video frames with a size of 224×288. This resulted in K class score maps Fk with a size of 14×18. Those where averaged in the mean pooling layer to obtain a single class score ak for each category k. The softmax layer then produced the class confidence ck of each frame. The final prediction was given by the output with the highest confidence. For retrospective frame retrieval we calculated and recorded the confidence ck for each class over the entire duration of an input video. Subsequently, we retrieved the frame with the highest confidence for each class." [Pg. 2209]). Regarding Claim 6, Baumgartner discloses that the confidence reflects an estimated accuracy of the predicted relevance for the one or more image components as output by the model ("After training we fed the network with cropped video frames with a size of 224×288. This resulted in K class score maps Fk with a size of 14×18. Those where averaged in the mean pooling layer to obtain a single class score ak for each category k. The softmax layer then produced the class confidence ck of each frame. The final prediction was given by the output with the highest confidence. For retrospective frame retrieval we calculated and recorded the confidence ck for each class over the entire duration of an input video. Subsequently, we retrieved the frame with the highest confidence for each class." [Pg. 2209]). Regarding Claim 8, Baumgartner discloses block iii) comprises the processor being caused to: display the output of the model to the user in the form of a heatmap overlaid on the ultrasound image frame. and wherein the levels of the heatmap are based on the output confidences for image components in the image frame (“we use the activations hn k(X) in Fk to calculate the saliency maps as a weighted linear combination of the influence of each of the receptive fields of the neurons in Fk. In this manner, regions corresponding to highly activated neurons will have more importance than neurons with low activations in the resulting saliency map” [Pg. 2209], “we post-process saliency maps obtained using Eq. 4 to obtain confidence maps from which we then calculate bounding boxes. In a first step, we take the absolute value of the saliency map S and blur it using a 5×5Gaussian kernel. This produces confidence maps of the location of the structure in the image such as the onesshowninFig.4b.” [Pg. 2210]). Regarding Claim 13, Baumgartner discloses a method of aiding a user to perform a medical ultrasound examination (“a novel method based on convolutional neural networks, which can automatically detect 13 fetal standard views in freehand 2-D ultrasound data as well as provide a localization of the fetal structures via a bounding box.” [Abstract], “The proposed system can be used in a number of ways. It can be employed to provide real-time feedback about the content of a image frame to the operator.” [Pg. 2205]), the method comprising: receiving a real-time sequence of ultrasound images captured by an ultrasound probe during the medical ultrasound examination (“we propose a novel system based on convolutional neural networks (CNNs) for real-time automated detection of 13 fetal standard scan planes, as well as localisa tion of the fetal structures associated with each scan plane in the images via bounding boxes.” [Pg. 2204]); and highlighting to the user, in real-time on a display, image components that are predicted by the model to be relevant to the medical ultrasound examination, for further consideration by the user (“The network architecture is designed to operate in real-time while providing optimal output for the localization task.” [Abstract]), Baumgartner does not specifically disclose using a model trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed, wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination, wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame. However, in a similar field of endeavor, Droste teaches transferable representations of images can be learned without manual annotations by modeling human visual attention [Abstract]. Droste also teaches using a model trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “This allocation of visual attention is typically quantified via the distribution of gaze points, which can be recorded with gaze tracking. There has been great interest in developing models of human visual attention that, given an image, predict the likelihood that each pixel is fixated upon, hereafter referred to as visual saliency map.” [Pg. 1], “Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]), wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “Visual Saliency Prediction. Given an image and a gaze point set (X, G) ∈ , the idea is to generate a visual saliency map S ∈ ]0, 1]H×W, where Si,j is the probability that pixel Xi,j is fixated upon.” [Pg. 4]); wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “Visual Saliency Prediction. Given an image and a gaze point set (X, G) ∈ , the idea is to generate a visual saliency map S ∈ ]0, 1]H×W, where Si,j is the probability that pixel Xi,j is fixated upon.” [Pg. 4]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above with using a model trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed, wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination, wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 16, Baumgartner discloses A tangible, non-transitory computer readable medium that stores instructions, which when executed by a processor (“a novel method based on convolutional neural networks, which can automatically detect 13 fetal standard views in freehand 2-D ultrasound data as well as provide a localization of the fetal structures via a bounding box.” [Abstract], “The proposed system can be used in a number of ways. It can be employed to provide real-time feedback about the content of a image frame to the operator.” [Pg. 2205], “The optimisation typically converged after around 2 days of training on a Nvidia GeForce GTX 1080 GPU.” [Pg. 2209]), cause the processor to: i) receive a real-time sequence of ultrasound images captured by an ultrasound probe during the medical ultrasound examination (“we propose a novel system based on convolutional neural networks (CNNs) for real-time automated detection of 13 fetal standard scan planes, as well as localisa tion of the fetal structures associated with each scan plane in the images via bounding boxes.” [Pg. 2204]); iii) highlight to the user, in real-time on the display. image components that are predicted by the model to be relevant to the medical ultrasound examination, for further consideration by the user (“The network architecture is designed to operate in real-time while providing optimal output for the localization task.” [Abstract]), Baumgartner does not specifically disclose that ii) use a model comprising the executable instructions and trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed, wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination; wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame. However, in a similar field of endeavor, Droste teaches transferable representations of images can be learned without manual annotations by modeling human visual attention [Abstract]. Droste also teaches ii) use a model comprising the executable instructions and trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “This allocation of visual attention is typically quantified via the distribution of gaze points, which can be recorded with gaze tracking. There has been great interest in developing models of human visual attention that, given an image, predict the likelihood that each pixel is fixated upon, hereafter referred to as visual saliency map.” [Pg. 1], “Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]), wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “Visual Saliency Prediction. Given an image and a gaze point set (X, G) ∈ , the idea is to generate a visual saliency map S ∈ ]0, 1]H×W, where Si,j is the probability that pixel Xi,j is fixated upon.” [Pg. 4]); wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “Visual Saliency Prediction. Given an image and a gaze point set (X, G) ∈ , the idea is to generate a visual saliency map S ∈ ]0, 1]H×W, where Si,j is the probability that pixel Xi,j is fixated upon.” [Pg. 4]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above with ii) use a model comprising the executable instructions and trained using a machine learning process to take an image frame in the real-time sequence of ultrasound images as input, and output a predicted relevance of one or more image components in the image frame to the medical ultrasound examination being performed, wherein the predicted relevance comprises a level of importance that a radiologist would attribute to the image component or to different regions or groups of image components in the image frame when a clinician would consider, inspect or further investigate the image frame, in a context of performing the medical ultrasound examination; wherein the predicted relevance comprises a relevance value or score for the one or more image components in the image frame as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 17, Baumgartner discloses that the processor is further caused to repeat blocks iii) for a plurality of image frames in the real-time sequence of ultrasound images (“The network architecture is designed to operate in real-time while providing optimal output for the localization task.” [Abstract]). Baumgartner does not specifically disclose that the processor is further caused to repeat blocks ii) for a plurality of image frames in the real-time sequence of ultrasound images. However, in a similar field of endeavor, Droste teaches that the processor is further caused to repeat blocks ii) for a plurality of image frames in the real-time sequence of ultrasound images (“The basis of our analyses is a unique gaze tracking dataset of sonographers performing routine clinical fetal anomaly screenings. Models of sonographer visual attention are learned by training a convolutional neural network (CNN) to predict gaze on ultrasound video frames through visual saliency prediction or gaze-point regression.” [Abstract], “This allocation of visual attention is typically quantified via the distribution of gaze points, which can be recorded with gaze tracking. There has been great interest in developing models of human visual attention that, given an image, predict the likelihood that each pixel is fixated upon, hereafter referred to as visual saliency map.” [Pg. 1], “Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above the processor is further caused to repeat blocks ii) for a plurality of image frames in the real-time sequence of ultrasound images as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 18, Baumgartner discloses all limitations noted above except that the model is trained using the machine learning process on training data comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating a relevance of one or more image components in the respective example ultrasound image to the medical ultrasound examination. However, in a similar field of endeavor, Droste teaches that the model is trained using the machine learning process on training data comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating a relevance of one or more image components in the respective example ultrasound image to the medical ultrasound examination (“Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above with the model is trained using the machine learning process on training data comprising: example ultrasound images; and ground truth annotations for each example ultrasound image, the ground truth annotations indicating a relevance of one or more image components in the respective example ultrasound image to the medical ultrasound examination as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 19, Baumgartner discloses all limitations noted above except that the ground truth annotations are based on gaze tracking information obtained from observing a radiologist analysing the respective example ultrasound image for the purpose of the medical ultrasound examination. However, in a similar field of endeavor, Droste teaches that the ground truth annotations are based on gaze tracking information obtained from observing a radiologist analysing the respective example ultrasound image for the purpose of the medical ultrasound examination (“Sonographer visual attention is modeled by training a CNN to predict gaze on random video frames. We consider this to be self-supervised representation learning since it does not require any manual annotations and gaze data is acquired fully automatically. We extract high-resolution image features by introducing dilated convolutions [10,19] into a recently proposed image classification architecture [8]. Two methods for training the model for gaze prediction are evaluated: (i) Visual saliency prediction: Ground truth visual saliency maps are generated and used as training targets [2]. (ii) Gaze-point regression: The approach of gaze-point regression [14] is much less explored in the literature but is simpler since it does not require explicit modeling of foveal vision for ground truth saliency map generation. An existing mathematically differentiable method is based on a fully-connected layer [14] which does not scale well to high-resolution feature maps due to the exponentially increasing number of learnable parameters. Here, we propose a method based on the soft argmax algorithm by Levine et al. [11] with no additional learnable parameters compared to saliency prediction” [Pg. 2]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner as outlined above with the ground truth annotations are based on gaze tracking information obtained from observing a radiologist analysing the respective example ultrasound image for the purpose of the medical ultrasound examination as taught by Droste, because human gaze is inherently a strong prior for semantic information [Pg.3]. Regarding Claim 20, Baumgartner discloses that the model is further trained to output an indication of a confidence associated with the predicted relevance for the one or more image components in the image frame ("After training we fed the network with cropped video frames with a size of 224×288. This resulted in K class score maps Fk with a size of 14×18. Those where averaged in the mean pooling layer to obtain a single class score ak for each category k. The softmax layer then produced the class confidence ck of each frame. The final prediction was given by the output with the highest confidence. For retrospective frame retrieval we calculated and recorded the confidence ck for each class over the entire duration of an input video. Subsequently, we retrieved the frame with the highest confidence for each class." [Pg. 2209]). Regarding Claim 21, Baumgartner discloses that the confidence reflects an estimated accuracy of the predicted relevance for the one or more image components as output by the model ("After training we fed the network with cropped video frames with a size of 224×288. This resulted in K class score maps Fk with a size of 14×18. Those where averaged in the mean pooling layer to obtain a single class score ak for each category k. The softmax layer then produced the class confidence ck of each frame. The final prediction was given by the output with the highest confidence. For retrospective frame retrieval we calculated and recorded the confidence ck for each class over the entire duration of an input video. Subsequently, we retrieved the frame with the highest confidence for each class." [Pg. 2209]). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Baumgartner in view of Droste as applied to Claim 5 above, and further in view of Khrosravan et al (Gaze2Segment: A Pilot Study for Integrating Eye-Tracking Technology into Medical Image Segmentation. In: Müller, H., et al. Medical Computer Vision and Bayesian and Graphical Models for Biomedical Imaging. BAMBI MCV 2016.). Regarding Claim 7, Baumgartner in view of Droste discloses all limitations noted above except that the confidence for the one or more image components comprises a prediction of a priority with which a radiologist would investigate a region comprising that image component, compared to other regions when performing the medical ultrasound examination. However, in a similar field of endeavor, Khrosovan teaches integrating biological and computer vision techniques to support radiologists’ reading experience with an automatic image segmentation task [Abstract]. Khrosovan also teaches that the confidence for the one or more image components comprises a prediction of a priority with which a radiologist would investigate a region comprising that image component, compared to other regions when performing the medical ultrasound examination (“we extracted im age context information by predicting which point attracts the most attention. This step combines radiologist’s knowledge with image context. The context aware saliency explains the visual attention with feature-driven four principles, three of which were implemented in our study: (1) local low-level considerations, (2) global considerations, (3) visual organization rules, and (4) high-level factors. (1) For local low-level information, image was divided into local patches (pu) centered at pixel u, and for each pair of patches, their distance (dposition) and normalized intensity difference (dintensity) were used to assess saliency of a pixel u, as formulated below: d(pu,pv) = dintensity/(1 + λdposition), (2) where λ is a weight parameter. Pixel u was considered salient when it was highly dissimilar to all other image patches, d(pu,pv) is high ∀v. (2) For global considerations, a scale-space approach was utilized to suppress frequently occurring features such as background and maintain features that deviate from the norm. Saliency of any pixel in this configuration was defined as the average of its saliency in M scales {(r1,r2,...,rM),r ∈ R}: ¯ Su =(1/M) K Sr u = 1− exp{−(1/K) Sr u r∈R d(pr u, pr v)} for (r ∈ R). k=1 (3) (4) This scale-based global definition combined K most similar patches for the saliency definition and indicated a more salient pixel u when Sr u was large. (3) For visual organization rules, saliency was defined based on the Gestalt laws suggesting areas that were close to the foci of attention should be explored significantly more than far-away regions. Hence, assuming dfoci(u) is the Eu clidean distance between pixel u and the closest focus of attention pixel, then the saliency of the pixel was defined as ˆ Su = ¯Su(1 − dfoci(u)). A point was considered as a focus of attention if it was salient. (4) High-level factors such as recognized objects can be applied as a post processing step to refine saliency definition. In our current implementation, we did not apply this consideration.” [Pg. 5-6]) It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner in view of Droste as outlined above with the confidence for the one or more image components comprises a prediction of a priority with which a radiologist would investigate a region comprising that image component, compared to other regions when performing the medical ultrasound examination as taught by Khrsosovan, because eye-tracking can be used as an effective recognition strategy for the medical image segmentation problems [Pg. 2]. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Baumgartner in view of Droste as applied to Claim 1 above, and further in view of Payer et al (C. Payer, D. Štern, H. Bischof, and M. Urschler, “Integrating spatial configuration into heatmap regression based CNNS for landmark localization,” Medical Image Analysis, vol. 54, pp. 207–219, May 2019; hereinafter referred to as Payer). Regarding Claim 9, Baumgartner in view of Droste discloses all limitations noted above except that the processor being caused to take a relative spatial context of different anatomical features into account to predict the relevance of the one or more image components in the image frame. However, in a similar field of endeavor, Payer teaches a CNN architecture that learns to split the localization task into two simpler sub-problems, reducing the overall need for large training datasets [Abstract] Payer also teaches that the processor being caused to take a relative spatial context of different anatomical features into account to predict the relevance of the one or more image components in the image frame (“Aiming for low landmark localization error in the presence of limited training datasets, our proposed SCN architecture directly combines the corresponding outputs of two interacting compo- nents. As illustrated in Fig. 1 , the interaction of these components is made possible by multiplying the predictions from both com- ponents, when our fully convolutional network architecture based on heatmap regression ( Section 2.1 ) is trained in an end-to-end manner. Due to this interaction, our SCN learns to dedicate its local appearance component to deliver locally accurate but poten- tially ambiguous candidate predictions, and its spatial configura- tion component to focus on the improvement of robustness to- wards landmark misidentification by eliminating ambiguities (see Section 2.2 ).” [Pg. 209], “Schematic representation of our proposed SCN. In the local appearance component, the input image I is transformed into H LA , representing local appearance heatmaps for each of the N landmarks. The dashed black line indicates that H LA is used as an input for the spatial configuration component, where H LA is transformed into the spatial configuration heatmaps H SC . A multiplication of H LA and H SC results in the final heatmaps H . Empty boxes represent intermediate images;” [Fig. 3]. It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner in view of Droste as outlined above with the processor being caused to take a relative spatial context of different anatomical features into account to predict the relevance of the one or more image components in the image frame as taught by Payer, because it allows for efficient work with limited amounts of training data, and does not need any dataset specific postprocessing [Pg. 209]. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Baumgartner in view of Droste as applied to Claim 1 above, and further in view of Quarfordt et al (US20110206283A1; hereinafter referred to as Quarfordt). Regarding Claim 10, Baumgartner in view of Droste discloses all limitations noted above except that the set of instructions, when executed by the processor, further cause the processor to: determine gaze information of the user; and wherein block iii) further comprises the processor being caused to: highlight to the user, in real time on the display, one or more portions of the image frame that the gaze information indicates that the user has not yet looked at. However, in a similar field of endeavor, Quarfordt teaches an image analysis system that includes a processor and memory and displays an image to a first user [Abstract]. Quarfordt also teaches that the set of instructions, when executed by the processor, further cause the processor to: determine gaze information of the user; and wherein block iii) further comprises the processor being caused to: highlight to the user, in real time on the display, one or more portions of the image frame that the gaze information indicates that the user has not yet looked at (“The image analysis system further tracks gaze of the first user; and collects initial gaze data for the first user. The initial gaze data includes a plurality of gaze points. The image analysis system also identifies one or more ignored regions of the image based on a distribution of the gaze data within the image and displays at least a first subset of the image. In some embodiments the first subset is displayed to the first user. In some embodiments the first subset is displayed to a second user that is distinct from the first user. The first subset of the image is selected so as to include a respective ignored region of the one or more ignored regions and the first subset of the image is displayed in a manner that draws attention to the respective ignored region.” [0005]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner in view of Droste as outlined above with the set of instructions, when executed by the processor, further cause the processor to: determine gaze information of the user; and wherein block iii) further comprises the processor being caused to: highlight to the user, in real time on the display, one or more portions of the image frame that the gaze information indicates that the user has not yet looked at as taught by Quarfordt, because it ensures full examination of all relevant regions of an image so as to reduce error and improve target identification in visual search tasks [0003]. Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Baumgartner in view of Droste as applied to Claim 1 above, and further in view of Gafner et al (US20170360404A1; hereinafter referred to as Gafner). Regarding Claim 11, Baumgartner in view of Droste discloses that the processor being caused to: display markings highlighting the image components that are predicted by the model to be relevant to the medical ultrasound examination (“a novel method based on convolutional neural networks, which can automatically detect 13 fetal standard views in freehand 2-D ultrasound data as well as provide a localization of the fetal structures via a bounding box.” [Baumgartner Abstract], “The proposed system can be used in a number of ways. It can be employed to provide real-time feedback about the content of a image frame to the operator.” [Baumgartner Pg. 2205]). Baumgartner in view of Droste does not specifically teach that the markings are removed or faded after a predetermined time interval; display markings highlighting the image components that are predicted by the model to be relevant to the medical ultrasound examination, and wherein the markings are added or increased in prominence after a predetermined time interval; and highlight the image components that are predicted by the model to be relevant to the medical ultrasound examination using augmented reality. However, in a similar field of endeavor, Gafner teaches aspects of the technology described herein relate to techniques for guiding an operator to use an ultrasound device [Abstract]. Gafner also teaches that the markings are removed or faded after a predetermined time interval; display markings highlighting the image components that are predicted by the model to be relevant to the medical ultrasound examination, and wherein the markings are added or increased in prominence after a predetermined time interval; and highlight the image components that are predicted by the model to be relevant to the medical ultrasound examination using augmented reality (“the method further comprises displaying a plurality of composite images in real time. In some embodiments, the composite images are displayed on an augmented reality display. In some embodiments, the method further comprises providing instructions in real time based on the plurality of composite images, wherein the instructions guide a user of the ultrasound probe in acquisition of subsequent ultrasound images of the portion of the patient's body.” [0110]) It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner in view of Droste as outlined above with the markings are removed or faded after a predetermined time interval; display markings highlighting the image components that are predicted by the model to be relevant to the medical ultrasound examination, and wherein the markings are added or increased in prominence after a predetermined time interval; and highlight the image components that are predicted by the model to be relevant to the medical ultrasound examination using augmented reality at as taught by Gafner, because it may allow a technician to infer medical information about the patient [0003]. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Baumgartner in view of Droste as applied to Claim 1 above, and further in view of Schein et al (US20210100526A1; hereinafter referred to as Schein). Regarding Claim 12, Baumgartner in view of Droste discloses all limitations noted above except that the set of instructions, when executed by the processor, further cause the processor to: use a pixel-wise flow model to link the predicted relevance of image components in the image frame to a predicted relevance of image components in another image frame in the real-time sequence of ultrasound images. However, in a similar field of endeavor, Schein teaches methods and systems are provided for tracking anatomical features across multiple images [Abstract]. Schein also teaches that the set of instructions, when executed by the processor, further cause the processor to: use a pixel-wise flow model to link the predicted relevance of image components in the image frame to a predicted relevance of image components in another image frame in the real-time sequence of ultrasound images ("For example, first image 710 may be input into the segmentation and tracking model (and optionally along with a prior output of the model, if the prior output is available) to generate the annotations overlaid on first image 810, and second image 720 may be input into the segmentation and tracking model along with the output of the model used to generate the annotations overlaid on first image 810, in order to generate the annotations overlaid on second image 820" [0105], “The motion tracker may determine motion (whether for an entire image or for each separate identified feature of an image) using a suitable technique, such as changes in pixel brightness, movement of an associated tracking boundary for each identified feature, changes in edges of identified features, or the like.” [0077], “For example, the model may output a plurality of pixels each having a value that reflects a calculation/prediction of whether or not that pixel is part of an anatomical feature, and all pixels having a value below a threshold may be “thresholded” out (e.g., given a pixel value of zero)” [0086]). It would have been obvious to an ordinary skilled person in the art before the effective filing date of the claimed invention to modify the system of Baumgartner in view of Droste as outlined above with the set of instructions, when executed by the processor, further cause the processor to: use a pixel-wise flow model to link the predicted relevance of image components in the image frame to a predicted relevance of image components in another image frame in the real-time sequence of ultrasound images at as taught by Schein, because it may assist a clinician performing a medical procedure on the patient [0002]. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEVEN MALDONADO whose telephone number is 703-756-1421. The examiner can normally be reached 8:00 am-4:00 pm PST M-Th Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Christopher Koharski can be reached on (571) 272-7230. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Steven Maldonado/ Patent Examiner, Art Unit 3797 /CHRISTOPHER KOHARSKI/Supervisory Patent Examiner, Art Unit 3797
Read full office action

Prosecution Timeline

Show 16 earlier events
Apr 08, 2026
Response after Non-Final Action
Apr 08, 2026
Response after Non-Final Action
Apr 08, 2026
Response after Non-Final Action
Apr 10, 2026
Response after Non-Final Action
May 06, 2026
Response after Non-Final Action
May 11, 2026
Response after Non-Final Action
Jul 02, 2026
Non-Final Rejection (signed) — §102, §103
Aug 27, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12685446
SEMI-COMPACT PHOTOACOUSTIC DEVICES AND SYSTEMS
3y 7m to grant Granted Jul 21, 2026
Patent 12653416
WIRELESS MEDICAL LOCATION TRACKING
3y 0m to grant Granted Jun 16, 2026
Patent 12635910
METHOD AND SYSTEM FOR TRACKING OF ACOUSTIC VIBRATIONS USING OPTICAL COHERENCE TOMOGRAPHY
3y 4m to grant Granted May 26, 2026
Patent 12551289
Tracker-Based Surgical Navigation
4y 1m to grant Granted Feb 17, 2026
Patent 12496034
SYSTEMS AND METHODS FOR PATIENT MONITORING
3y 0m to grant Granted Dec 16, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
27%
Grant Probability
70%
With Interview (+42.9%)
3y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 26 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month