DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant's election of Invention II - Species VIII (claims 12–13, 15 and 20) with 2 new added claims 21-22 and claims 5 and 8 have been canceled in the reply filed on 04/20/2026 is acknowledged. Applicant's election is noted with traverse.
Applicant argues that restriction between Inventions I and II is improper because newly presented claim 22 combines the model-training subject matter of Invention I with the runtime surgical-stage-association subject matter of Invention II. Applicant further argues that election among Species VII-XII is improper because newly presented claim 21 combines the limitations corresponding to the respective species.
The Examiner respectfully disagrees as to the restriction between Invention I and Invention II. Invention I and Invention II remain directed to independent and distinct technical features requiring materially different searches. Invention I is directed to constructing and training a computational model using training images and corresponding labels, including labels identifying surgical-tool characteristics and surgical stages, negative training images, and particular CNN, HMM, supervised-training, or unsupervised-training techniques. Invention II is directed to executing a computational model on runtime surgical images to determine surgical-tool characteristics, associate the runtime images with stages of a surgical procedure, and generate output indicating the associated stages. Further, Examiner issued restriction between Inventions based on reading on independent claims’ scope.
Although the trained model produced according to Invention I may subsequently be used to perform the runtime operations of Invention II, the runtime method does not require that the model be trained using the particular training operations recited by Invention I, and the training method may be performed independently to produce a model for subsequent deployment. Accordingly, a search directed to labeled training datasets, model construction, CNN/HMM training, and supervised or unsupervised training techniques would not necessarily locate the most relevant prior art concerning runtime surgical-video analysis, tool-based procedural-stage association, and stage-output generation. The restriction between Invention I and Invention II is therefore maintained, and claims 1-4, 6-7, 9-11 remain withdrawn from further consideration as directed to the non-elected invention. Claims 5 and 8 remain canceled.
However, upon further consideration, the election-of-species requirement as between Species VII-XII is hereby withdrawn. Applicant's newly presented claims 21 and 22 constitute generic or "linking" claims under MPEP § 809 that combine limitations of Species VII-X (CNN/HMM-based tool type and pose determination) with Species XI-XII (instrument-specific operation-stage and reconstruction-stage association). In accordance with MPEP § 809.03, because these linking claims are being examined on the merits rather than held unallowable, the previously restricted species are hereby rejoined for examination. Accordingly, previously withdrawn claims 14 and 16-19, directed to Species VII, IX, X, XI, and XII, respectively, are rejoined and will be examined together with elected claims 12-13, 15, and 20-22.
Accordingly, examination proceeds on claims 12-22. Claims 1-4, 6-7, 9-11 remain withdrawn from further consideration as directed to the non-elected invention (Invention I). Claims 5 and 8 have been canceled.
Claim Rejections - 35 USC § 112(b)
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 21-22 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 21 recites the limitation “using a CNN to determine a type … wherein … using an HMM … based on a type” and “using the CNN to determine a pose … wherein … using the HMM … based on a pose” in claim. There is insufficient antecedent basis for this limitation in the claim. Because the second occurrences use “a” rather than “the”, the claim does not expressly say that the HMM uses the type and pose determined by the CNN. A clean amendment would replace the second occurrences with: “-- based on the determined type”; and “-- based on the determined pose”.
Claim 21 further recites the limitation “associating a first runtime image...” and “associating a second runtime image...” in claim. There is insufficient antecedent basis for this limitation in the claim. Parent claim 12 only introduces the plural “runtime images”. The applicant should amend this to read "a first runtime image of the runtime images" and "a second runtime image of the runtime images" to properly ground the limitations.
Claim 22 recites the limitation “accessing labels that indicate, for each of the training images, characteristics of the one or more surgical tools depicted.” in claim. There is insufficient antecedent basis for this limitation in the claim. It is unclear whether “the one or more surgical tools” refers to the surgical tools previously recited in claim 12 as (i) being depicted by the runtime images, (ii) to surgical tools depicted by each respective training image, or (iii) to other surgical tools. Additionally, the phrase does not identify what depicts the recited surgical tools. Consequently, the relationship between the training images, the labels, and the surgical tools whose characteristics are represented by the labels is unclear, and the metes and bounds of claim 22 cannot be determined with reasonable certainty. The Examiner interpret this limitation, considering surrounding claims, in light of the specification as: “accessing labels that indicate, for each respective training image, characteristics of one or more surgical tools depicted by the respective training image and a stage of the multiple stages of the surgical procedure depicted by the respective training image”.
Claim 22 is a newly added claim, and depends on Claim 12, but it adds steps (accessing training images, accessing labels, and training the computational model) that logically and chronologically must occur before the inferencing steps recited in the parent claim 12 (associating runtime images using the model). Appending prerequisite training steps via “further comprising” to a claim that already executes the trained model makes the chronological sequence of the claimed method confusing and indefinite.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 12–22 are rejected under 35 U.S.C. § 101 because the claimed invention is directed to a judicial exception without significantly more.
This rejection has been made in accordance with the current USPTO subject matter eligibility framework, including MPEP §§ 2103–2106.07, the 2019 Revised Patent Subject Matter Eligibility Guidance, the October 2019 Patent Eligibility Guidance Update, the 2024 Guidance Update on Patent Subject Matter Eligibility, Including on Artificial Intelligence, the July 2024 AI Subject Matter Eligibility Examples, the August 4, 2025 USPTO memorandum titled “Reminders on evaluating subject matter eligibility of claims under 35 U.S.C. § 101,” and the USPTO’s updated guidance concerning Ex parte Desjardins, Appeal No. 2024-000567. The claims have been evaluated under the broadest reasonable interpretation, and the claims have been considered as a whole.
Step 1: Statutory Category
Independent claim 12 is directed to a method and therefore falls within the statutory category of a process. Independent claim 20 is directed to a non-transitory computer-readable medium and therefore falls within the statutory category of a manufacture. Accordingly, the analysis proceeds to Step 2A.
Step 2A, Prong One (Judicial Exception)
Independent claim 12 recites associating, using a computational model, runtime images with a stage of a surgical procedure based on characteristics of one or more surgical tools depicted by the runtime images; and generating output that indicates the stage associated with each of the runtime images. These limitations recite an abstract idea, namely collecting or observing image information, evaluating characteristics of items depicted in the images, classifying the images according to a corresponding procedural stage, and reporting the classification result.
The claim recites observation, data translation, classification, and mathematical calculation. The "runtime images", "computational model", "stage", "characteristics", and "output" are used as mathematical constructs and information items in a data-analysis and classification process. The claim does NOT recite an improvement to the way runtime images are captured, digitized, or encoded by the surgical camera. Rather, the claim uses generic computer models to obtain images, mathematically determine tool characteristics, and classify those images into defined stages.
The claim is also similar in character to claims courts have found abstract where the focus is collecting information, analyzing the collected information, and presenting or acting on the result. In Electric Power Group, LLC v. Alstom S.A., the Federal Circuit recognized claims directed to collecting and analyzing information and displaying the results as abstract. The present claims similarly obtain or review image information, analyze characteristics represented by the image information, classify the images, and generate output containing the classification result.
Claim 13 further recites identifying the runtime images from among a set of images such that the runtime images depict a surgical tool, and associating the identified runtime images. These limitations recite an additional observation and classification process: reviewing images, determining which images depict a surgical tool, and selecting those images for further evaluation. Selecting relevant information from a larger collection of information does not alter the abstract character of the claimed process.
Claims 14–19, 21, and 22 further limit the method to using a "convolutional neural network (CNN)" to determine a type and a pose of the surgical tools, using a "hidden Markov model (HMM)" to associate the runtime images, and "training the computational model" using "training images" and "labels". These limitations recite the use of mathematical models and statistical algorithms as tools for performing the abstract image classification task. The claims do not recite any specific, unconventional hardware architecture or technical improvement to machine-learning hardware technology itself.
Independent claim 20 recites substantially the same abstract idea in computer-readable medium form using generic computer processors. Merely implementing the same mathematical classification process using a generic manufacture does not avoid the judicial exception.
The claims are also consistent with the reasoning of AI Visualize, Inc. v. Nuance Communications, Inc., where the Federal Circuit looked to the character of the claims as a whole and affirmed ineligibility where the claims were directed to obtaining, manipulating, and using information at a high level of generality rather than to a specific improvement in computer functionality. Here, the character of the elected claims as a whole is image data collection, tool observation, mathematical model classification, and generating a stage output, not an improvement to endoscopic camera technology, computer memory, or machine-learning computational architecture itself.
Accordingly, claims 12–22 recite an abstract idea under Step 2A, Prong One.
Step 2A, Prong Two (Practical Application)
The additional elements, considered individually and in combination, do not integrate the abstract idea into a practical application.
The recited "computational model", "convolutional neural network", "hidden Markov model", "runtime images", and "output" amount to data objects, mathematical models, and generic computer implementation of the abstract classification concept.
The claims do not recite a particular improvement to computer or imaging technology. They do not improve how a digital image is physically captured by an endoscope or stored. The claims merely require obtaining surgical images that have already been captured. The claims also do not recite a particular improvement to graphical processing hardware. Rather, the claims use standard digital images as input data for mathematical tool detection and algorithmic stage classification.
The claims further do not recite a particular improvement to artificial-intelligence or machine-learning technology. The claims do not train a model in a novel way that improves the computer's operation, update model parameters to reduce processing load, modify model architecture to save hardware resources, reduce model storage, preserve prior model knowledge, or improve inference speed by a specific claimed technical mechanism. The CNNs and HMMs are recited functionally as mathematical tools for identifying relationships between images, tools, and stages.
This analysis is consistent with the USPTO's 2024 AI subject matter eligibility guidance and AI examples, which emphasize that AI-related claims may be eligible when they recite a specific technological improvement or otherwise integrate a judicial exception into a practical application. The present claims do NOT recite such a specific technological improvement. Instead, the claims use generic computer operations to collect images, mathematically evaluate tool characteristics, and output a stage assignment.
This case is distinguishable from Ex parte Desjardins. In Desjardins, the claims were found to reflect an improvement in machine-learning technology itself, including training a machine-learning model on a series of tasks while preserving prior knowledge and reducing complexity/storage burdens. Here, the claims do NOT recite a particular parameter-update mechanism, memory-saving arrangement, or data structure that improves the physical or computational operation of a machine-learning model. The claimed neural networks and Markov models merely automate the mathematical task of mapping tool characteristics to surgical stages. Nor does limiting the abstract idea to the environment of surgical procedure classification make the claims eligible. In Recentive Analytics, Inc. v. Fox Corp., the Federal Circuit rejected the argument that applying machine learning to a new field of use was sufficient for eligibility where the claims did not recite a technical improvement to the machine-learning process itself. Similarly here, applying neural networks to surgical tool detection is a field-of-use limitation, not an integration of the abstract idea into a practical application.
Accordingly, the claims do not integrate the judicial exception into a practical application under Step 2A, Prong Two.
Step 2B (Inventive Concept)
The additional elements, considered both individually and as an ordered combination, do not amount to significantly more than the abstract idea.
The claims use generic computer components to perform ordinary computer functions, including obtaining images, identifying tools, determining type and pose, calculating mathematical probabilities using an HMM, and generating an output. These are conventional data-processing and mathematical operations performed using generic computer technology.
The ordered combination also does not provide an inventive concept. The ordered combination follows the abstract idea itself: obtain an image, use a mathematical model to identify a tool and its characteristics, use another mathematical model to associate it with a stage, and output the result. This is no more than the abstract mental and mathematical idea implemented on generic computer components.
Dependent claims 13–19, and 21-22 recite additional steps of identifying images containing tools, using a CNN for type and pose, using an HMM for stage association, detecting specific tools like a ring curette or cauterizer, and accessing training images with labels. These limitations merely specify the generic machine-learning tools, known statistical probability formulas, and mathematical gathering mechanisms used in the abstract evaluation and do not add significantly more.
The independent method and computer-readable medium claims recite generic counterparts using basic computing logic to perform substantially the same operations. The recitation of generic neural network frameworks and statistical classification calculations does not transform the abstract idea into patent-eligible subject matter.
Accordingly, claims 12–22 are directed to a judicial exception without significantly more and are therefore rejected under 35 U.S.C. § 101.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 12–15, 20 and 22 are rejected under 35 U.S.C. §102(a)(1) as being anticipated by Zisimopoulos (Zisimopoulos et al, US 2019/0164012 A1, 2019).
Regarding claim 12, Zisimopoulos teaches a method comprising:
associating, using a computational model, runtime images with a stage of a surgical procedure based on characteristics of one or more surgical tools depicted by the runtime images; and
( [0030], [0039–0042], [0053–0054]: Zisimopoulos teaches a machine-learning processing system that receives real-time image [runtime images] or video-frame data collected during performance of a surgical procedure and processes each image, frame, or block of frames using a trained machine-learning model. The model detects and characterizes surgical tools depicted in the image data, including tool presence, location, position, use, manipulation, movement, and object state. )
generating output that indicates the stage associated with each of the runtime images.
( [0043–0046], [0054], [0060]: Zisimopoulos further teaches a state detector uses the detected presence and characteristics of the surgical tools to identify a procedural state, expressly described as a “stage”, corresponding to the processed runtime image data. Zisimopoulos also teaches that, after identifying the procedural statecorresponding to the processed image, frame, or block, an output generator generates output based on the identified state. The output may include information corresponding to the current procedural state and may be visually presented, overlaid on the real-time surgical capture, or transmitted to a user device. )
Regarding claim 13, Zisimopoulos teaches the method of claim 12, further comprising, prior to the associating, identifying the runtime images from among a set of images such that the runtime images depict a surgical tool,
wherein associating the runtime images comprises associating the runtime images identified as depicting a surgical tool.
( [0040], [0042], [0053–0054]: Zisimopoulos teaches receiving stream data comprising a video stream or image time series and iteratively analyzing individual images, frames, or blocks of frames using a trained machine-learning model. For each processed image or frame, the model generates segmentation data identifying which, if any, surgical tools are depicted and characteristics of the detected tools, including their position or state. Zisimopoulos further teaches using the segmented data indicating the presence and characteristics of the detected surgical tools to identify the procedural state corresponding to the processed image or frame. Therefore, Zisimopoulos identifies, from among the set of runtime images or frames, those images or frames depicting a surgical tool before associating the identified images or frames with a procedural stage. )
Regarding claim 14, Zisimopoulos teaches the method of claim 12, further comprising using a convolutional neural network (CNN) to determine a type of the one or more surgical tools depicted by the runtime images, wherein associating the runtime images comprises associating the runtime images based on the type of the one or more surgical tools depicted by the runtime images.
( [0028–0030], [0040–0042], [0074–0075]: Zisimopoulos teaches using a fully convolutional network (a type of CNN) to perform semantic segmentation of runtime surgical images into classes corresponding to respective types of surgical tools, assigning the detected pixels a corresponding tool-class label. Zisimopoulos further teaches that a state detector uses this segmented data indicating the presence and characteristics of the detected tool to associate the runtime image with an estimated procedural-state node. )
Regarding claim 15, Zisimopoulos teaches the method of claim 12, further comprising using a convolutional neural network (CNN) to determine a pose of the one or more surgical tools depicted by the runtime images,
wherein associating the runtime images comprises associating the runtime images based on the pose of the one or more surgical tools depicted by the runtime images.
( [0028–0029], [0041–0042], [0053–0054], [0060], [0075–0077]: Zisimopoulos teaches using a fully convolutional neural network to process runtime surgical images or video frames and detect and characterize depicted surgical tools. The network generates segmentation data indicating the tool’s class, location, and position, and may further infer the tool’s use, manipulation, movement, or object state from the position data. Zisimopoulos also trains the network using simulated surgical images rendered with variations in camera pose, viewing angle, and instrument motion, thereby enabling detection and spatial characterization of tools appearing in different positions and angular configurations. The state detector then uses the segmented tool data, including the tool’s presence, position, movement, manipulation, and state, to identify the procedural stage corresponding to the runtime image. Accordingly, Zisimopoulos teaches determining spatial characteristics corresponding to the pose of the depicted surgical tool and associating the runtime image with a surgical stage based on those characteristics. )
Regarding claim 20, the rationale provided in the rejection of claim 12 is incorporated herein. In addition, Zisimopoulos teaches a computer-program product tangibly embodied in a computer system, including instructions configured to cause one or more data processors to perform the disclosed procedural-state detection method [0144–0146]. Accordingly, the method of claim 12 corresponds to the non-transitory computer-readable medium of claim 20, and performs the steps disclosed herein. Therefore, the claims are all rejected.
Regarding claim 22, Zisimopoulos teaches the method of claim 12, further comprising:
accessing training images that collectively depict multiple stages of the surgical procedure;
( [0048], [0050], [0055–0056], [0077]: Zisimopoulos teaches identifying a set of states represented in a surgical procedural workflow and, for each state, accessing one or more base images corresponding to that state. Zisimopoulos further generates a set of virtual training images including at least one image corresponding to each state. In the disclosed cataract-surgery example, the training images collectively correspond to multiple surgical phases, including patient preparation, phacoemulsification, and insertion of an intraocular lens. )
accessing labels that indicate, for each of the training images, characteristics of the one or more surgical tools depicted and a stage of the multiple stages of the surgical procedure depicted; and
( [0026], [0048], [0055–0057]: Zisimopoulos teaches that the training images are accompanied by corresponding metadata and image-segmentation data. For each training image, the segmentation data identifies surgical toolsdepicted in the image and characteristics of the tools, including tool presence and position, while corresponding state data indicates the procedural state or stage associated with that image. Accordingly, each training image is associated with labels indicating both characteristics of depicted surgical tools and the surgical stage represented by the image. )
training the computational model, using the training images and the labels, to associate the runtime images with a stage of the multiple stages based on characteristics of the one or more surgical tools that are depicted by the runtime images.
( [0040–0042], [0058–0060]: Zisimopoulos teaches training the machine-learning model using the training images and corresponding data that includes surgical-tool segmentation data and the indicated procedural state. The trained model is subsequently executed on real-time image data to identify surgical-tool presence, position, use, manipulation, or object state. A state detector uses the detected presence and characteristics of the surgical tools to identify the procedural state or stage corresponding to the runtime image data. )
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 16-19 and 21 are rejected under 35 U.S.C. §103 as being unpatentable over Zisimopoulos, in view of Twinanda (Twinanda et al, EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos. arXiv:1602.03012 [Cs], 2016), further in view of Marcus (Marcus et al, Pituitary society expert Delphi consensus: operative workflow in endoscopic transsphenoidal pituitary adenoma resection. Pituitary, 24(6), 839–853, 2021).
Regarding claim 21, Zisimopoulos teaches the method of claim 12, further comprising:
using a convolutional neural network (CNN) to determine a type of the one or more surgical tools depicted by the runtime images, wherein associating the runtime images comprises using a graph-based procedural state tracker to associate the runtime images based on a type of the one or more surgical tools depicted by the runtime images; and
( [0028–0030], [0040–0042], [0073–0075], [0078], [0082]: Zisimopoulos teaches using a fully convolutional network, including an FCN-VGG, to perform semantic segmentation of runtime surgical images into classes corresponding to respective types of surgical tools. The network detects the tool and assigns the detected pixels a corresponding tool-class label. Zisimopoulos further teaches a procedural tracking data structure comprising nodes corresponding to respective procedural states and directional connections representing the expected order of the states. Each node may identify one or more tools typically used during the corresponding state, and a state detector uses segmented data indicating the presence and characteristics of a detected tool to associate the runtime image with an estimated procedural-state node. )
using the convolutional neural network (CNN) to determine a pose of the one or more surgical tools depicted by the runtime images, wherein associating the runtime images comprises using the graph-based procedural state tracker to associate the runtime images based on a pose of the one or more surgical tools depicted by the runtime images,
( [0029], [0040–0042], [0053–0054], [0075–0077]: Zisimopoulos teaches using the convolutional model to process runtime surgical images and generate segmentation data indicating the location and position of a detected surgical tool, as well as an inferred use, manipulation, or object state based on the position data. The state detector uses the detected tool characteristics, including a typical type of tool movement, to associate the runtime image with a procedural-state node. Zisimopoulos further teaches training the tool-detection model using images generated with variations in camera pose, viewing angle, and instrument motion, thereby enabling detection and spatial characterization of tools appearing in different positions and configurations. )
wherein associating the runtime images comprises associating a first runtime image with an operation stage of the surgical procedure based on detecting phacoemulsifier handpiece within the first runtime image, and
( [0041–0042], [0074], [0077–0078]: Zisimopoulos teaches that a procedural-state node may identify tools typically used during that state and that the state detector associates runtime images with respective state nodes based on the detected tools. In the cataract-surgery embodiment, Zisimopoulos identifies phacoemulsification as one of the surgical phases and expressly trains its convolutional model to detect a phacoemulsifier handpiece as one of the surgical-tool classes. Accordingly, associating a runtime image depicting the phacoemulsifier handpiece with first runtime image in the phacoemulsification stage [operation stage]. )
wherein associating the runtime images comprises associating a second runtime image with a reconstruction stage of the surgical procedure based on detecting implant injector within the second runtime image.
( [0041–0042], [0077–0078]: Zisimopoulos teaches associating runtime images with procedural-state nodes based on the presence of tools typically used during the corresponding state. Zisimopoulos identifies insertion of the intraocular lens as one of the cataract-surgery phases and identifies an implant injector as one of the surgical-tool classes detected by the convolutional model. Accordingly, associating a runtime image depicting the implant injector with second runtime image in the intraocular-lens insertion stage [reconstruction stage]. )
Zisimopoulos teaches CNN-based determination of surgical-tool characteristics and association of runtime images with procedural stages using a graph-based procedural state tracker, but fails to expressly disclose implementing the procedural state tracker as a hidden Markov model, where Twinanda teaches:
using a convolutional neural network (CNN) to determine a type of the one or more surgical tools depicted by the runtime images, wherein associating the runtime images comprises using a hidden Markov model (HMM) to associate the runtime images based one or more surgical tools depicted by the runtime images; and
( [Secs. III-A and III-C], [Figs. 1–2 and 7]: Twinanda teaches an EndoNet convolutional neural network configured to jointly perform surgical-tool presence detection and surgical-phase recognition on images from laparoscopic video. The network includes an fc_tool layer having a respective node for each defined surgical-tool type, where each node outputs a confidence that the corresponding tool type is depicted in the image. The tool-presence confidence values are concatenated with CNN visual features in the fc8 layer to construct a final feature for phase recognition. The fc8 feature is processed to obtain phase-confidence values that are supplied as observations to a two-level hierarchical hidden Markov model, which determines the most likely surgical phase. Therefore, Twinanda uses a CNN to determine the types of surgical tools depicted in runtime images and uses an HMM to associate the images with surgical phases based at least partly on information representing the surgical tools depicted in the images. )
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Zisimopoulos’s graph-based procedural state tracker to incorporate Twinanda’s hierarchical hidden Markov model for associating runtime surgical images with procedural stages. Both references use CNN-derived surgical-tool information to determine stages in temporally ordered surgical video, and Twinanda’s HMM predictably applies inter-phase and intra-phase temporal dependencies to identify the most likely surgical stage. Such a modification would have provided temporally consistent stage associations while retaining Zisimopoulos’s existing tool-based stage-detection framework.
Zisimopoulos [as modified by Twinanda] teaches detecting surgical tools in runtime images and using an HMM-based procedural tracker to associate the images with surgical stages, but still fails to disclose the specific claimed relationships between a ring curette or rongeur and an operation stage, and between a cauterizer and a reconstruction stage, where Marcus teaches:
wherein associating the runtime images comprises associating a first runtime image with an operation stage of the surgical procedure based on detecting a ring curette or a rongeur within the first runtime image, and
( [Table 4]: Marcus teaches that the sellar phase of an endoscopic transsphenoidal surgical procedure includes operative resection of a pituitary tumor. For microadenoma resection, Marcus identifies a ring curette as a surgical instrument used during the resection step; and for macroadenoma piecemeal resection, Marcus identifies a ring curette and pituitary rongeurs as instruments used during the resection step. Therefore, the presence of a ring curette or pituitary rongeur provides a known indication that the surgical procedure is in the operative tumor-resection stage. )
wherein associating the runtime images comprises associating a second runtime image with a reconstruction stage of the surgical procedure based on detecting a cauterizer within the second runtime ima
( Table 5: Marcus teaches a closure phase comprising hemostasis and skull-base repair, including dural repair or reconstruction. Marcus identifies a bipolar electrosurgical instrument [a cauterizer] among the instruments used during the closure phase, including during hemostasis and graft harvesting. Therefore, detection of the bipolar cauterizing instrument provides an indication that the procedure is in the closure/ reconstruction phase.)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to configure Zisimopoulos [as modified by Twinanda]’s tool-based stage-recognition system to recognize Marcus’s known stage-specific surgical instruments. Marcus teaches that ring curettes and pituitary rongeurs are used during operative tumor resection, while a bipolar cauterizing instrument is used during the closure/ reconstruction phase. Using these known instrument-stage relationships in the existing CNN/ HMM recognition framework would have predictably enabled identification of the corresponding surgical stage from the detected instrument.
Regarding claims 16-19, the rationale provided in the rejection of claim 21 is incorporated herein. In addition, the method of claim 21 corresponds to the method of claims 16-19, and performs the steps disclosed herein. Therefore, the claims are all rejected.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEN KUDO whose telephone number is (571)272-4498. The examiner can normally be reached M-F 8am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
KEN KUDO
Examiner
Art Unit 2671
/KEN KUDO/Examiner, Art Unit 2671
/VINCENT RUDOLPH/Supervisory Patent Examiner, Art Unit 2671