Prosecution Insights
Last updated: October 02, 2026
Application No. 18/869,608

Computer Vision Based Two-Stage Surgical Phase Recognition Module

Non-Final OA §102§103
Filed
Nov 26, 2024
Priority
May 26, 2022 — provisional 63/345,990 +1 more
Examiner
PATEL, JAYESH A
Art Unit
2677
Tech Center
2600 — Communications
Assignee
Verily Life Sciences LLC
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
764 granted / 913 resolved
+21.7% vs TC avg
Minimal +5% lift
Without
With
+4.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
34 currently pending
Career history
937
Total Applications
across all art units

Statute-Specific Performance

§101
9.1%
-30.9% vs TC avg
§103
46.7%
+6.7% vs TC avg
§102
15.6%
-24.4% vs TC avg
§112
22.1%
-17.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 913 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification The substitute specification along with the clean copy filed on 11/26/2024 has been entered and made of record. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 6-8, 10-12, 14, and 18-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Wolf et al. (US20200237452) hereinafter Wolf 1. Regarding Claim 1, Wolf discloses a system configured to automatically identify a surgical phase of a surgery captured in a video of the surgery (paragraph 115 for computer analysis may be used to identify surgical phases [identifying surgical phases], intraoperative events, event characteristics, and/or other features appearing in the video footage), the system comprising a non-transitory computer readable medium having stored thereon a plurality of instructions (paragraphs 0187-0188 discloses a non-transitory computer readable medium having stored thereon a plurality of instructions): wherein the surgery comprises a plurality of sequential phases (paragraphs 115 for features that may be detected In video footage for placing markers may include [using computer readable medium], motions of a surgeon or other medical professional, patient characteristics, surgeon characteristics or characteristics of other medical professionals, sequences of operations being performed [using sequential phases in the surgery], timings of operations or events, characteristics of anatomical structures, paragraph 287 discloses “the elapsed time between two surgical phases and/or intraoperative surgical events”); wherein the video comprises a sequence of video frames (paragraph 114 for Computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN); and wherein the instructions are executed with one or more processors (paragraphs 0187-0188 discloses a non-transitory computer readable medium having stored thereon a plurality of instructions executed by the one or more processors) so that the following steps are executed (paragraph 114 for markers may be automatically generated and included in the timeline based on information in the video at a given location. In some embodiments, computer analysis may be used to analyze frames of the video footage and identify markers to include at various locations in the timeline. Computer analysis may include any form of electronic analysis using a computing device): receiving, by a two-stage surgical phase recognition ("SPR") module, the video of the surgery (paragraph 115 for computer analysis may be used to identify surgical phases, intraoperative events, event characteristics [receiving a video for computer analysis of surgical phases], and/or other features appearing in the video footage with paragraph 118 for the marker may have a different color depending on what type of intraoperative surgical event the marker represents. In some example embodiments, markers associated with an incision, an excision, a resection, a ligation, a graft, or various other events may each be displayed with a different color [having multiple stages that are color coded with markers for surgery phases of incision & graft, and stages that are critical or an adverse event]); wherein the two-stage SPR module comprises a first stage and a second stage (paragraph 118-119 for the severity of an adverse event may be represented by on a color scale ranging from yellow to red, or other suitable color scales. In some embodiments, the location and/or size of the marker may be associated with a criticality level [multiple stages of criticality and stages of adverse events indicated by colors and markers in video]. The criticality level may represent the relative importance of an event, action, technique, phase or other occurrence identified by the marker with criticality level may include finite number of discrete levels such as "Level 0", "Level 1", "Level 2", "High Criticality", "Low Criticality", "Non-Critical [different types of stages of surgical phases]); extracting, using a neural network that forms the first stage, visual information content of a single frame based on the single frame (paragraph 116 for marker locations may be identified using a trained machine learning model. For example, a machine learning model may be trained using training examples, each training example may include video footage known to be associated with surgical procedures [using individual video frames as visual information of surgical content for training the neural network], surgical phases, intraoperative events, and/or event characteristics [extracting visual information from video], together with labels indicating locations within the video footage with models such as a gradient boosting algorithm, artificial neural networks such as deep neural networks, convolutional neural networks [using a neural network to identify a stage such as a procedure or an event]); and identifying, using a multi-stage temporal convolution network that forms the second stage (paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN [using a convolutional neural network] with paragraph 122 for decision making junction may be detected using the computer analysis described above. In some embodiments, video footage may be analyzed to identify particular actions or sequences of actions performed by a surgeon that may indicate a decision has been made. For example, if the surgeon pauses during a procedure [using pause to end the first stage and begin with the second stage of surgery], begins to use a different medical device, or changes to a different course of action), surgical phases captured in the frames of the video based on the visual information content from the first stage (paragraph 122 for the decision making junction may be identified based on a surgical phase or intraoperative event identified in the video footage at that location. For example, an adverse event [surgical phase], such as a bleed, may be detected which may indicate a decision must be made on how to address the adverse event); and wherein the two-stage SPR module has been trained using videos of surgeries involving complex anatomies and videos of surgeries involving adverse events (paragraph 120 for icons or other visual properties may be used to distinguish between unplanned events and planned events, types of errors e.g., miscommunication errors, judgment errors, or other forms of errors, specific adverse events [videos of adverse events] that occurred, types of techniques being performed, the surgical phase being performed, locations of intraoperative surgical events and paragraph 121 for the decision making junction marker may indicate a location of a video depicting a surgical procedure where multiple courses of action are possible, and a surgeon opts to follow one course over another. For example, the surgeon may decide whether to depart from a planned surgical procedure, to take a preventative action [complex surgery], to remove an organ or tissue, to use a particular instrument, to use a particular surgical technique, or any other intraoperative decisions a surgeon may encounter, see also paragraphs 79-82 for the machine learning algorithms in the present disclosure trained using training examples). 2. Regarding Claim 6, Wolf discloses the system further comprises a camera (paragraph 85 for the cameras may capture the video/image data at a location 127 of a body of patient 143 on which a surgical procedure is performed); and wherein the instructions are executed with the one or more processors so that the following step is also executed (paragraph 84 for analyzing image data (for example, by the methods, steps and modules described herein) may comprise analyzing pixels, voxels, point cloud, range data, etc. included in the image data): creating, using the camera, the video of the surgery (paragraph 85 for cameras may capture video/image data associated with surgical team personnel, such as an anesthesiologist, nurses, surgical tech and the like located in operating room 101. Additionally, operating room cameras may capture video/image data associated with medical equipment located in the room). 3. Regarding Claim 7, Wolf discloses identifying the surgical phases occurs in real-time or near-real-time (paragraph 536 for a method for providing decision support for surgical procedures may be performed in real time during a surgical procedure. Real-time recommendations may include providing recommendations via an interface in an operating room [identifying surgical stages in real time] e.g., an operating room depicted in FIG. 1. Real-time recommendations may be updated during a surgical procedure). 4. Regarding claim 8, Wolf discloses a system configured to identify an adverse event during a surgery based on a video of the surgery (paragraph 115 for computer analysis may be used to identify surgical phases [identifying surgical phases], intraoperative events, event characteristics, and/or other features appearing in the video footage), the system comprising a non-transitory computer readable medium having stored thereon a plurality of instructions (paragraphs 0187-0188 discloses a non-transitory computer readable medium having stored thereon a plurality of instructions): wherein the surgery is on a patient and comprises a plurality of sequential surgical phases (paragraph 115 for features that may be detected in video footage for placing markers may include [using computer readable medium], motions of a surgeon or other medical professional, patient characteristics, surgeon characteristics or characteristics of other medical professionals, sequences of operations being performed [using sequential phases in the surgery], timings of operations or events, characteristics of anatomical structures); wherein the video comprises a sequence of video frames paragraph 287 discloses “the elapsed time between two surgical phases and/or intraoperative surgical events”, (paragraph 114 for Computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN); wherein the instructions are executed with one or more processors (paragraphs 0187-0188 discloses a non-transitory computer readable medium having stored thereon a plurality of instructions executed by the one or more processors) so that the following steps are executed (paragraph 114 for markers may be automatically generated and included in the timeline based on information in the video at a given location. In some embodiments, computer analysis may be used to analyze frames of the video footage and identify markers to include at various locations in the timeline. Computer analysis may include any form of electronic analysis using a computing device): receiving, by a two-stage surgical phase recognition ("SPR") module, the video of the surgery (paragraph 115 for computer analysis may be used to identify surgical phases, intraoperative events, event characteristics [receiving a video for computer analysis of surgical phases], and/or other features appearing in the video footage with paragraph 118 for the marker may have a different color depending on what type of intraoperative surgical event the marker represents. In some example embodiments, markers associated with an incision, an excision, a resection, a ligation, a graft, or various other events may each be displayed with a different color [having multiple stages that are color coded with markers for surgery phases of incision & graft, and stages that are critical or an adverse event]); wherein the two-stage SPR module comprises a first stage and a second stage (paragraph 118-119 for the severity of an adverse event may be represented by on a color scale ranging from yellow to red, or other suitable color scales. In some embodiments, the location and/or size of the marker may be associated with a criticality level [multiple stages of criticality and stages of adverse events indicated by colors and markers in video]. The criticality level may represent the relative importance of an event, action, technique, phase or other occurrence identified by the marker with criticality level may include finite number of discrete levels such as "Level 0", "Level 1", "Level 2", "High Criticality", "Low Criticality", "Non Critical [different types of stages of surgical phases]); extracting, using a neural network that forms the first stage, visual information content of a single frame based on the single frame (paragraph 116 for marker locations may be identified using a trained machine learning model. For example, a machine learning model may be trained using training examples, each training example may include video footage known to be associated with surgical procedures [using individual video frames as visual information of surgical content for training the neural network], surgical phases, intraoperative events, and/or event characteristics [extracting visual information from video], together with labels indicating locations within the video footage with models such as a gradient boosting algorithm, artificial neural networks such as deep neural networks, convolutional neural networks [using a neural network to identify a stage such as a procedure or an event]); identifying, using a multi-stage temporal convolution network that forms the second stage (paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN [using a convolutional neural network] with paragraph 122 for decision making junction may be detected using the computer analysis described above. In some embodiments, video footage may be analyzed to identify particular actions or sequences of actions performed by a surgeon that may indicate a decision has been made. For example, if the surgeon pauses during a procedure [using pause to end the first stage and begin with the second stage of surgery], begins to use a different medical device, or changes to a different course of action), surgical phases captured in the frames of the video based on the visual information content from the first stage (paragraph 122 for the decision making junction may be identified based on a surgical phase or intraoperative event identified in the video footage at that location. For example, an adverse event [surgical phase], such as a bleed, may be detected which may indicate a decision must be made on how to address the adverse event); and identifying, using the two-stage SPR module and the identified surgical phases, an adverse event during the surgery (paragraph 120 for icons or other visual properties may be used to distinguish between unplanned events and planned events, types of errors e.g., miscommunication errors, judgment errors, or other forms of errors, specific adverse events [videos of adverse events] that occurred, types of techniques being performed, the surgical phase being performed, locations of intraoperative surgical events and paragraph 121 for the decision making junction marker may indicate a location of a video depicting a surgical procedure where multiple courses of action are possible, and a surgeon opts to follow one course over another. For example, the surgeon may decide whether to depart from a planned surgical procedure, to take a preventative action [complex surgery], to remove an organ or tissue, to use a particular instrument, to use a particular surgical technique, or any other intraoperative decisions a surgeon may encounter); wherein the adverse event comprises at least one of: the absence of a surgical phase in the plurality of sequential surgical phases (paragraph 122 for training example may include a video clip, together with a label indicating locations of decision making junctions within the video clip, or together with a label indicating an absent of decision making junctions in the video clip with paragraph 162 for identify the event location of the particular intraoperative surgical event within the surgical phase. An example of such training example may include a video clip together with a label indicating a location of a particular event within the video clip, or an absence of such event [using machine learning to analyze the video for the absence of an event]); and an injury to the patient (paragraph 129 for the first prior event may include an adverse event or complication, such as bleeding, mesenteric emphysema, injury, conversion to unplanned open, incision significantly larger than planned, hypertension, hypotension, bradycardia, hypoxemia, adhesions, hernias, atypical anatomy, dural tears, periorator injury, arterial occlusions). 5. Regarding Claim 10, Wolf discloses the instructions are executed with the one or more processors so that the following step (paragraphs 0187-0188 discloses a non-transitory computer readable medium having stored thereon a plurality of instructions executed by the one or more processors) is also executed: in response to the identification of the adverse event (paragraph 124 for the decision marker may open a menu or otherwise display options for viewing the alternative video clips. For example, selecting the decision naming marker may pop up an alternative video menu containing depictions of the conduct in the associated alternative video clips), annotating the video to indicate the video includes the adverse event (paragraph 119 for measure of an immediate need for an action to prevent hazardous result within a surgical procedure. For example, criticality level may include a numerical measure such as "1.12", "3.84", "7", "-4.01", etc., for example within a particular range of values [making numerical notes]. In another example, criticality level may include finite number of discrete levels such as "Level 0", "Level 1", "Level 2", "High Criticality", "Low Criticality", "Non Critical" [adding notes on levels or criticality to events in the video]). 6. Regarding Claim 11, Wolf discloses the instructions are executed with the one or more processors so that the following step (paragraphs 0187-0188 discloses a non-transitory computer readable medium having stored thereon a plurality of instructions executed by the one or more processors) is also executed: in response to the identification of the adverse event (paragraph 120 for icons or other visual properties may be used to distinguish between unplanned events and planned events, types of errors e.g., miscommunication errors, judgment errors, or other forms of errors, specific adverse events [videos of adverse events] that occurred, types of techniques being performed, the surgical phase being performed, locations of intraoperative surgical events), automatically providing an indication to a user of the system (paragraph 114 for paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames [using markers as indications]), wherein the indication indicates the existence of the adverse event (paragraph 124 for the decision marker may open a menu or otherwise display options for viewing the alternative video clips. For example, selecting the decision naming marker may pop up an alternative video menu containing depictions of the conduct in the associated alternative video clips). 7. Regarding Claim 12, Wolf discloses identifying the surgical phases occurs in real-time or near-real-time (paragraph 536 for a method for providing decision support for surgical procedures may be performed in real time during a surgical procedure. Real-time recommendations may include providing recommendations via an interface in an operating room [identifying surgical stages in real time] e.g., an operating room depicted in FIG. 1. Real-time recommendations may be updated during a surgical procedure). 8. Regarding Claim 14, Wolf discloses a method of identifying an adverse event using a video of a surgery (paragraph 115 for computer analysis may be used to identify surgical phases [identifying surgical phases], intraoperative events, event characteristics, and/or other features appearing in the video footage) and a two-stage surgical phase recognition ("SPR") module, the method comprising (paragraph 115 for computer analysis may be used to identify surgical phases, intraoperative events, event characteristics [receiving a video for computer analysis of surgical phases], and/or other features appearing in the video footage with paragraph 118 for the marker may have a different color depending on what type of intraoperative surgical event the marker represents. In some example embodiments, markers associated with an incision, an excision, a resection, a ligation, a graft, or various other events may each be displayed with a different color [having multiple stages that are color coded with markers for surgery phases of incision & graft, and stages that are critical or an adverse event]): receiving, by the two-stage SPR module, a video of the surgery (paragraph 115 for computer analysis may be used to identify surgical phases, intraoperative events, event characteristics [receiving a video for computer analysis of surgical phases], and/or other features appearing in the video footage with paragraph 118 for the marker may have a different color depending on what type of intraoperative surgical event the marker represents. In some example embodiments, markers associated with an incision, an excision, a resection, a ligation, a graft, or various other events may each be displayed with a different color [having multiple stages that are color coded with markers for surgery phases of incision & graft, and stages that are critical or an adverse event]); wherein the video comprises a sequence of video frames (paragraph 114 for Computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN); wherein the surgery is on a patient and comprises a plurality of sequential surgical phases (paragraph 114 for markers may be automatically generated and included in the timeline based on information in the video at a given location. In some embodiments, computer analysis may be used to analyze frames of the video footage and identify markers to include at various locations in the timeline. Computer analysis may include any form of electronic analysis using a computing device); and wherein the two-stage SPR module comprises a first stage and a second stage (paragraph 118-119 for the severity of an adverse event may be represented by on a color scale ranging from yellow to red, or other suitable color scales. In some embodiments, the location and/or size of the marker may be associated with a criticality level [multiple stages of criticality and stages of adverse events indicated by colors and markers in video]. The criticality level may represent the relative importance of an event, action, technique, phase or other occurrence identified by the marker with criticality level may include finite number of discrete levels such as "Level 0", "Level 1", "Level 2", "High Criticality", "Low Criticality", "Non Critical [different types of stages of surgical phases]); extracting, using a neural network that forms the first stage, visual information content of a single frame based on the single frame (paragraph 116 for marker locations may be identified using a trained machine learning model. For example, a machine learning model may be trained using training examples, each training example may include video footage known to be associated with surgical procedures [using individual video frames as visual information of surgical content for training the neural network], surgical phases, intraoperative events, and/or event characteristics [extracting visual information from video], together with labels indicating locations within the video footage with models such as a gradient boosting algorithm, artificial neural networks such as deep neural networks, convolutional neural networks [using a neural network to identify a stage such as a procedure or an event]); identifying, using a multi-stage temporal convolution network that forms the second stage, surgical phases captured in the frames of the video based on the visual information content from the first stage (paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN [using a convolutional neural network] with paragraph 122 for decision making junction may be detected using the computer analysis described above. In some embodiments, video footage may be analyzed to identify particular actions or sequences of actions performed by a surgeon that may indicate a decision has been made. For example, if the surgeon pauses during a procedure [using pause to end the first stage and begin with the second stage of surgery], begins to use a different medical device, or changes to a different course of action); and identifying, using the two-stage SPR module and the identified surgical phases, an adverse event during the surgery (paragraph 122 for the decision-making junction may be identified based on a surgical phase or intraoperative event identified in the video footage at that location. For example, an adverse event [surgical phase], such as a bleed, may be detected which may indicate a decision must be made on how to address the adverse event); and wherein the adverse event comprises at least one of: the absence of a surgical phase in the plurality of sequential surgical phases (paragraph 122 for training example may include a video clip, together with a label indicating locations of decision making junctions within the video clip, or together with a label indicating an absent of decision making junctions in the video clip with paragraph 162 for identify the event location of the particular intraoperative surgical event within the surgical phase. An example of such training example may include a video clip together with a label indicating a location of a particular event within the video clip, or an absence of such event [using machine learning to analyze the video for the absence of an event]); and an injury to the patient (paragraph 129 for the first prior event may include an adverse event or complication, such as bleeding, mesenteric emphysema, injury, conversion to unplanned open, incision significantly larger than planned, hypertension, hypotension, bradycardia, hypoxemia, adhesions, hernias, atypical anatomy, dural tears, periorator injury, arterial occlusions). 9. Regarding Claim 18, Wolf discloses in response to the identification of the adverse event (paragraph 124 for the decision marker may open a menu or otherwise display options for viewing the alternative video clips. For example, selecting the decision naming marker may pop up an alternative video menu containing depictions of the conduct in the associated alternative video clips), automatically providing a notification regarding the existence of the adverse event (paragraph 118 for the marker may have a different color depending on what type of intraoperative surgical event the marker represents. In some example embodiments, markers associated with an incision, an excision, a resection, a ligation, a graft, or various other events may each be displayed with a different color [using a notification of color] and paragraph 343 for further include determining an extent of variance from a scheduled time associated with completion, in response to a first determined extent, outputting a notification, and in response to a second determined extent, forgoing outputting the notification). 10. Regarding Claim 19, Wolf discloses identifying, using the two-stage SPR module and the video, the surgical phases occurs in real-time or near-real-time (paragraph 536 for a method for providing decision support for surgical procedures may be performed in real time during a surgical procedure. Real-time recommendations may include providing recommendations via an interface in an operating room [identifying surgical stages in real time] e.g., an operating room depicted in FIG. 1. Real-time recommendations may be updated during a surgical procedure); and wherein automatically providing a notification regarding the existence of the adverse event occurs in real-time or near-real-time (paragraph 536 for Real-time recommendations may include providing recommendations via an interface in an operating room e.g., an operating room depicted in FIG. 1. Real-time recommendations may be updated during a surgical procedure). 11. Regarding Claim 20, Wolf discloses before receiving the video of the surgery, training the two-stage SPR module (paragraph 120 for icons or other visual properties may be used to distinguish between unplanned events and planned events, types of errors e.g., miscommunication errors, judgment errors, or-other forms of errors, specific adverse events [videos of adverse events] that occurred, types of techniques being performed, the surgical phase being performed, locations of intraoperative surgical events, see also paragraphs 79-82 for the machine learning algorithms in the present disclosure trained using training examples); wherein training the two-stage SPR module comprises: collecting a dataset of surgery videos (paragraph 114 for paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames [using markers as indications]); annotating the frames of the surgery videos with one surgical phase from the plurality of sequential surgical phases (paragraph 119 for measure of an immediate need for an action to prevent hazardous result within a surgical procedure. For example, criticality level may include a numerical measure such as "1.12", "3.84", "7", "-4.01", etc., for example within a particular range of values [making numerical notes]. In another example, criticality level may include finite number of discrete levels such as "Level 0", "Level 1", "Level 2", "High Criticality", "Low Criticality", "Non Critical" [adding notes on levels or criticality to events in the video]); annotating at least a portion of the frames of the surgery videos with one adverse event from a plurality of adverse events (paragraph 124 for the decision marker may open a menu or otherwise display options for viewing the alternative video clips. For example, selecting the decision naming marker may pop up an alternative video menu containing depictions of the conduct in the associated alternative video clips); creating a training set comprising a portion of the annotated surgery videos (paragraph 122 for a trained machine learning model may be used to identify the decision-making junction. For example, a machine learning model may be trained using training examples to detect decision making junctions in videos, and the trained machine learning model may be used to analyze the video and detect the decision-making junction. An example of such training example may include a video clip, together with a label indicating locations of decision-making junctions within the video clip with paragraph 125 for alternative videos, the alternative possible decisions may be overlaid on the timeline and/or video, or may be displayed In a separate region, such as above, below and/or to the side of the video, in a separate window, on a separate screen, or in any other suitable manner. The alternative possible decisions may be a list of alternative decisions the surgeon could have made at the decision-making junction. The list may also include images e.g., depicting alternative actions, flow diagrams, statistics e.g., success rates, failure rates, usage rates, or other statistical information); and training the neural network using the training set (paragraph 114 for machine learning model may be trained using training examples to generate markers for videos, and the trained machine learning model may be used to analyze the video and generate markers for that video. Such generated markers may include locations within the video for the marker, type of the marker, properties of the marker, and so forth. An example of such training example may include a video clip). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2-5, 9, 13, and 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf in view of Nauta et al. (Casual discovery with Attention-Based Convolutional Neural Networks, Meike Nauta et al., MDPI, 2019, Pages 1-28,) hereinafter, "Nauta". 12. Regarding Claim 2, Wolf the system of claim 1, wherein the output of the neural network comprises a sequence of feature vectors that represent the video (paragraph 114 for a machine learning model may be trained using training examples to generate markers for videos, and the trained machine learning model may be used to analyze the video and generate markers for that video. Such generated markers may include locations within the video for the marker, type of the marker, properties of the marker [identifying features in the video frame], and so forth. An example of such training example may include a video clip depicting at least part of a surgical procedure), with each feature vector expressing visual information content of one single frame from the sequence of video frames (paragraphs 115 for computer analysis may be used to identify surgical phases, intraoperative events, event characteristics, and/or other features appearing in the video footage. For example, in some embodiments, computer analysis may be used to identify one or more medical instruments [analyzing frames for instrument feature] used in a surgical procedure, paragraphs 0079, 205, 587-589 discloses the visual information content expressed as structure (feature) vector); wherein the sequence of feature vectors is an input to the multi-stage temporal convolution network (paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames [detecting feature of motion]. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN). Wolf fails to explicitly disclose wherein the multi-stage temporal convolution network comprises temporal convolution layers with a dilation rate that increases across layers to capture temporal connections between the sequence of feature vectors. Nauta has insight into the causal associations for decision making in medical treatments (abstract) and teaches wherein the multi-stage temporal convolution network comprises temporal convolution layers with a dilation rate that increases across layers to capture temporal connections between the sequence of feature vectors (page 9, second and third paragraph for With an exponentially increasing dilation factor f, a network with stacked dilated convolutions can operate on a coarser scale without loss of resolution or coverage with this shows that dilated convolutions support an exponential increase [increasing dilated convolutions] of the receptive field while the number of parameters grows only linearly, which is especially useful when there is a large delay between cause and effect). Before the effective filing date of the invention was made, Wolf and Nauta are combinable because they are from the same field of endeavor and are analogous art of neural network data processing. The suggestion/motivation would be to modify Wolf with the insight into the causal associations for decision making in medical treatments teaching of Nauta for the purpose of allowing dilated convolutions support an exponential increase of the receptive field (Nauta, page 9, third paragraph). Therefore, it would be obvious and within one of ordinary skill in the art to have recognized the advantages of Nauta in the system of Wolf to obtain the invention as specified in claim 2. 13. Regarding Claim 3, Wolf and Nauta discloses the system of claim 2. Wolf fails to explicitly disclose the neural network extracts visual information content of the single frame based on the single frame and is temporal-agnostic; and wherein the multi-stage temporal convolution network is non-casual. Nauta has insight into the causal associations for decision making in medical treatments (abstract) and teaches the neural network extracts visual information content of the single frame based on the single frame and is temporal-agnostic (page 6, first paragraph for feature importance proposed by an interpretable LSTM already showed to be highly in line with results from the Granger causality test. Multiple deep learning models exist for non-temporal causal discovery: Variational Autoencoders to estimate causal effects, Causal Generative Neural Networks to learn functional causal models and the Structural Agnostic Model SAM [using agnostic models for feature importance] for causal graph reconstruction); and wherein the multi-stage temporal convolution network is non-casual (page-22, second paragraph for a non-causal correlation that arises because of the hidden confounder with unequal delays was too weak to be selected as potential cause by the attention mechanism, which indicates that our attention interpretation method to select potential causes is effective and strict enough). Before the effective filing date of the invention was made, Wolf and Nauta are combinable because they are from the same field of endeavor and are analogous art of neural network data processing. The suggestion/motivation would be to modify Wolf with the insight into the causal associations for decision making in medical treatments teaching of Nauta for the purpose of ensuring attention interpretation method to select potential causes is effective and strict enough (Nauta, page 22, second paragraph). Therefore, it would be obvious and within one of ordinary skill in the art to have recognized the advantages of Nauta in the system of Wolf to obtain the invention as specified in claim 3. 14. Regarding Claim 4, Wolf and Nauta discloses the system of claim 2. Wolf fails to explicitly disclose the multi-stage temporal convolution network comprises a plurality of stages; and wherein each stage in the plurality of stages comprises a plurality of temporal convolution layers with a dilation rate that increases across the layers of that stage. Nauta has insight into the causal associations for decision making in medical treatments (abstract) and teaches the multi-stage temporal convolution network comprises a plurality of stages (page 9, and figure 5 for having a convolutional neural network with three hidden layers and each layer having multiple stages shown by the square boxes and second paragraph for an exponentially increasing dilation factor f, a network with stacked dilated convolutions can operate on a coarser scale without loss of resolution or coverage); and wherein each stage in the plurality of stages comprises a plurality of temporal convolution layers with a dilation rate that increases across the layers of that stages (page 9, second and third paragraph for With an exponentially increasing dilation factor f, a network with stacked dilated convolutions can operate on a coarser scale without loss of resolution or coverage with this shows that dilated convolutions support an exponential increase [increasing dilated convolutions] of the receptive field while the number of parameters grows only linearly, which is especially useful when there is a large delay between cause and effect). Before the effective filing date of the invention was made, Wolf and Nauta are combinable because they are from the same field of endeavor and are analogous art of neural network data processing. The suggestion/motivation would be to modify Wolf with the insight into the causal associations for decision making in medical treatments teaching of Nauta for the purpose of using coarser scale without loss of resolution or coverage (Nauta, page 9, second paragraph). Therefore, it would be obvious and within one of ordinary skill in the art to have recognized the advantages of Nauta in the system of Wolf to obtain the invention as specified in claim 4. 15. Regarding Claim 5, Wolf and Nauta discloses the system of claim 2. Wolf discloses further the video of the surgery is created at a first location (paragraph 122 for decision making junction may be detected using the computer analysis described above. In some embodiments, video footage may be analyzed to identify particular actions or sequences of actions performed by a surgeon that may indicate a decision has been made with paragraph 123 for if the current video footage includes a compilation of differing procedures, the alternative footage may be drawn from a differing location [having a first and second location] of the current video footage being displayed); wherein the instructions are executed with the one or more processors so that the following step is also executed: before receiving the video of the surgery (paragraph 122 for surgeon pauses during a procedure, begins to use a different medical device, or changes to a different course of action, this may indicate a decision has been made. In some embodiments, the decision making junction may be identified based on a surgical phase or intraoperative event identified in the video footage at that location [having a video of the surgeon making a surgery decision at a first location]), retraining a plurality of last prediction layers of the neural network on a dataset associated with the first location (paragraph 122 for a trained machine learning model may be used to identify the decision making junction [using neural networks with datasets]. For example, a machine learning model may be trained using training examples to detect decision making junctions in videos, and the trained machine learning model may be used to analyze the video and detect the decision-making junction [analyzing the decision at the first location]). 16. Regarding Claim 9, Wolf discloses the system of claim 8. Wolf discloses further wherein output of the neural network comprises a sequence of feature vectors that represent the video (paragraph 114 for a machine learning model may be trained using training examples to generate markers for videos, and the trained machine learning model may be used to analyze the video and generate markers for that video. Such generated markers may include locations within the video for the marker, type of the marker, properties of the marker [identifying features in the video frame], and so forth. An example of such training example may include a video clip depicting at least part of a surgical procedure), with each feature vector expressing visual information content of one single frame from the sequence of video frames (paragraph 115 for computer analysis may be used to identify surgical phases, intraoperative events, event characteristics, and/or other features appearing in the video footage. For example, in some embodiments, computer analysis may be used to identify one or more medical instruments [analyzing frames for instrument feature] used in a surgical procedure); wherein the sequence of feature vectors is an input to the multi-stage temporal convolution network (paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames [detécting feature of motion]. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN). Wolf fails to explicitly disclose wherein the multi-stage temporal convolution network comprises temporal convolution layers with a dilation rate that increases across layers to capture temporal connections between the sequence of feature vectors. Nauta has insight into the causal associations for decision making in medical treatments (abstract) and teaches wherein the multi-stage temporal convolution network comprises temporal convolution layers with a dilation rate that increases across layers to capture temporal connections between the sequence of feature vectors (page 9, second and third paragraph for With an exponentially increasing dilation factor f, a network with stacked dilated convolutions can operate on a coarser scale without loss of resolution or coverage with this shows that dilated convolutions support an exponential increase [increasing dilated convolutions] of the receptive field while the number of parameters grows only linearly, which is especially useful when there is a large delay between cause and effect). Before the effective filing date of the invention was made, Wolf and Nauta are combinable because they are from the same field of endeavor and are analogous art of neural network data processing. The suggestion/motivation would be to modify Wolf with the insight into the causal associations for decision making in medical treatments teaching of Nauta for the purpose of allowing dilated convolutions support an exponential increase of the receptive field (Nauta, page 9, third paragraph). Therefore, it would be obvious and within one of ordinary skill in the art to have recognized the advantages of Nauta in the system of Wolf to obtain the invention as specified in claim 9. 17. Regarding Claim 13, Wolf discloses the system of claim 8. Wolf fails to explicitly disclose the neural network extracts visual information content of the single frame based on the single frame and is temporal-agnostic; and wherein the multi-stage temporal convolution network is non-casual. Nauta has insight into the causal associations for decision making in medical treatments (abstract) and teaches the neural network extracts visual information content of the single frame based on the single frame and is temporal-agnostic (page 6, first paragraph for feature importance proposed by an interpretable LSTM already showed to be highly in line with results from the Granger causality test. Multiple deep learning models exist for non-temporal causal discovery: Variational Autoencoders to estimate causal effects, Causal Generative Neural Networks to learn functional causal models and the Structural Agnostic Model SAM [using agnostic models for feature importance] for causal graph reconstruction); and wherein the multi-stage temporal convolution network is non-casual (page 22, second paragraph for a non-causal correlation that arises because of the hidden confounder with unequal delays was too weak to be selected as potential cause by the attention mechanism, which indicates that our attention interpretation method to select potential causes is effective and strict enough). Before the effective filing date of the invention was made, Wolf and Nauta are combinable because they are from the same field of endeavor and are analogous art of neural network data processing. The suggestion/motivation would be to modify Wolf with the insight into the causal associations for decision making in medical treatments teaching of Nauta for the purpose of ensuring attention interpretation method to select potential causes is effective and strict enough (Nauta, page 22, second paragraph). Therefore, it would be obvious and within one of ordinary skill in the art to have recognized the advantages of Nauta in the system of Wolf to obtain the invention as specified in claim 13. 18. Regarding Claim 15, Wolf discloses the method of claim 14. Wolf discloses further output of the neural network comprises a sequence of feature vectors that represent the video (paragraph 114 for a machine learning model may be trained using training examples to generate markers for videos, and the trained machine learning model may be used to analyze the video and generate markers for that video. Such generated markers may include locations within the video for the marker, type of the marker, properties of the marker [identifying features in the video frame], and so forth. An example of such training example may include a video clip depicting at least part of a surgical procedure), with each feature vector expressing visual information content of one single frame from the sequence of video frames (paragraph 115 for computer analysis may be used to identify surgical phases, intraoperative events, event characteristics, and/or other features appearing in the video footage. For example, in some embodiments, computer analysis may be used to identify one or more medical instruments [analyzing frames for instrument feature] used in a surgical procedure); wherein the sequence of feature vectors is an input to the multi-stage temporal convolution network (paragraph 114 for analyze frames of the video footage and identify markers to include at various locations in the timeline [using temporal analysis with a timeline] also computer analysis may be performed on individual frames [having a video with a sequence of frames], or may be performed across multiple frames, for example, to detect motion or other changes between frames [detecting feature of motion]. In some embodiments computer analysis may include object detection algorithms, such as Viola-Jones object detection, scale-invariant feature transform SIFT, histogram of oriented gradients HOG features, convolutional neural networks CNN). Wolf fails to explicitly disclose wherein the multi-stage temporal convolution network comprises temporal convolution layers with a dilation rate that increases across layers to capture temporal connections between the sequence of feature vectors. Nauta has insight into the causal associations for decision making in medical treatments (abstract) and teaches wherein the multi-stage temporal convolution network comprises temporal convolution layers with a dilation rate that increases across layers to capture temporal connections between the sequence of feature vectors (page 9, second and third paragraph for With an exponentially increasing dilation factor f, a network with stacked dilated convolutions can operate on a coarser scale without loss of resolution or coverage with this shows that dilated convolutions support an exponential increase [increasing dilated convolutions] of the receptive field while the number of parameters grows only linearly, which is especially useful when there is a large delay between cause and effect). Before the effective filing date of the invention was made, Wolf and Nauta are combinable because they are from the same field of endeavor and are analogous art of neural network data processing. The suggestion/motivation would be to modify Wolf with the insight into the causal associations for decision making in medical treatments teaching of Nauta for the purpose of allowing dilated convolutions support an exponential increase of the receptive field (Nauta, page 9, third paragraph). Therefore, it would be obvious and within one of ordinary skill in the art to have recognized the advantages of Nauta in the system of Wolf to obtain the invention as specified in claim 15. 19. Regarding Claim 16, Wolf and Nauta discloses the method of claim 15. Wolf discloses further wherein the video of the surgery is created at a first location (paragraph 122 for decision making junction may be detected using the computer analysis described above. In some embodiments, video footage may be analyzed to identify particular actions or sequences of actions performed by a surgeon that may indicate a decision has been made with paragraph 123 for if the current video footage includes a compilation of differing procedures, the alternative footage may be drawn from a differing location [having a first and second location] of the current video footage being displayed); and wherein the method further comprises before receiving the video of the surgery (paragraph 122 for surgeon pauses during a procedure, begins to use a different medical device, or changes to a different course of action, this may indicate a decision has been made. In some embodiments, the decision making junction may be identified based on a surgical phase or intraoperative event identified in the video footage at that location [having a video of the surgeon making a surgery decision at a first location]), retraining a plurality of last prediction layers of the neural network on a dataset associated with the first location (paragraph 122 for a trained machine learning model may be used to identify the decision making junction [using neural networks with datasets]. For example, a machine learning model may be trained using training examples to detect decision making junctions in videos, and the trained machine learning model may be used to analyze the video and detect the decision-making junction [analyzing the decision at the first location]). 20. Regarding Claim 17, Wolf and Nauta disclose the method of claim 16. Wolf discloses further comprising, in response to the identification of the adverse event, annotating the video to indicate the video includes the adverse event (paragraph 124 for the decision marker may open a menu or otherwise display options for viewing the alternative video clips. For example, selecting the decision naming marker may pop up an alternative video menu containing depictions of the conduct in the associated alternative video clips). Examiner's Note: Examiner has cited figures, and paragraphs in the references as applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested for the applicant, in preparing the responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner. Examiner has also cited references in PTO892 but not relied on, which are relevant and pertinent to the applicant’s disclosure, and may also be reading (anticipatory/obvious) on the claims and claimed limitations. Applicant is advised to consider the references in preparing the response/amendments in-order to expedite the prosecution. NOTE: Examiner’s preliminary review of WO2022047043A1 is also an anticipatory/obvious reference. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAYESH PATEL whose telephone number is (571)270-1227. The examiner can normally be reached IFW Mon-FRI. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at 571-270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAYESH PATEL/ Primary Examiner Art Unit 2677 /JAYESH A PATEL/Primary Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Nov 26, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750231
METHODS AND SYSTEMS FOR ENROLLMENT AND AUTHENTICATION
3y 0m to grant Granted Sep 29, 2026
Patent 12737900
ATTENTION-BASED REFINEMENT FOR DEPTH COMPLETION
3y 1m to grant Granted Sep 15, 2026
Patent 12731441
PERSON AUTHENTICATION SUPPORT SYSTEM, PERSON AUTHENTICATION SUPPORT METHOD, AND NON-TRANSITORY STORAGE MEDIUM
3y 1m to grant Granted Sep 08, 2026
Patent 12723964
PROCESS FOR IDENTIFYING A SUB-SAMPLE AND A METHOD FOR DETERMINING THE PETROPHYSICAL PROPERTIES OF A ROCK SAMPLE
3y 2m to grant Granted Sep 01, 2026
Patent 12718349
EVALUATION APPARATUS, INFORMATION PROCESSING APPARATUS, COMPUTER-READABLE STORAGE MEDIUM, FILM FORMING SYSTEM, AND ARTICLE MANUFACTURING METHOD
3y 2m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
89%
With Interview (+4.9%)
2y 11m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 913 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month