Prosecution Insights
Last updated: July 27, 2026
Application No. 18/286,969

COMPUTER VISION-BASED SURGICAL WORKFLOW RECOGNITION SYSTEM USING NATURAL LANGUAGE PROCESSING TECHNIQUES

Final Rejection §103
Filed
Oct 13, 2023
Priority
Apr 14, 2021 — provisional 63/174,820 +1 more
Examiner
ABDI, AMARA
Art Unit
2668
Tech Center
2600 — Communications
Assignee
Csats Inc.
OA Round
2 (Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
688 granted / 831 resolved
+20.8% vs TC avg
Minimal -7% lift
Without
With
+-7.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
19 currently pending
Career history
857
Total Applications
across all art units

Statute-Specific Performance

§101
2.8%
-37.2% vs TC avg
§103
89.9%
+49.9% vs TC avg
§102
2.1%
-37.9% vs TC avg
§112
2.4%
-37.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 831 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Applicant's response to the last office action, filed February 17, 2026 has been entered and made of record. Claims 8, 12, 15, and 20 are amended. Claims 1-20 are pending for examination. Response to Arguments Applicant's arguments filed February 17, 2026 have been fully considered but they are not persuasive. -- Applicant asserted, (on Page 7, last paragraph), with respect to claims 1 and 13, that Wolf describing identifying video footage location associated with a surgical event is merely identifying a video footage location of a surgical event, but is not the same as associating "the plurality of images with a plurality of surgical activities." The Examiner respectfully disagrees because Wolf clearly discloses in Par. 0118, training machine learning model using video footage known to be associated with surgical procedures, surgical phases, intraoperative events, and/or event characteristics, together with labels indicating locations within the video footage, to identify similar phases and events in other video footage, [i.e., perform processing, “training machine learning model”, on the surgical video data, “video footage”, to associate the plurality of images with a plurality of surgical activities, “implicit associating the video footage with surgical procedures, surgical phases, intraoperative events, and/or event characteristics”]). -- Applicant further asserted, (Page 8, first paragraph), that Patel is silent as to associating "the plurality of images with a plurality of surgical activities," as recited in independent claims 1 and 13, because using natural language processing for speech recognition and/or text recognition is not the same as using natural language processing for associating a plurality of images with a plurality of surgical activities. The Examiner respectfully disagrees, because Wolf discloses already the associating the plurality of images with a plurality of surgical activities and the other limitations of claim 1; and Patel is used as secondary reference to teach the natural language processing, (Par. 0061), that Wolf failed to disclose. Wolf and Patel are combinable because they are both concerned with video data processing. Using Patel’s natural language processing in Wolf’s surgical video data represent a simple substitution of one well known element, (Wolf’s machine learning model) with another well-known element, (Patel’s natural language processing model), to produce predictable results generate the start times at which video data chunks start, and stop times at which video data chunks end, (Patel, Par. 0093). Further, [KSR type finding; e.g., "the claim would have been obvious because the substitution of one known element for another would have yielded predictable results to one of ordinary skill in the art at the time of the invention”. For the reasons stated above, the rejection of claims 1 and 13 was proper, and it is maintained. -- Applicant further asserted, with respect to claim 20, (on pages 8-9) that Médioni describes segmenting video and extracting spatiotemporal feature vectors using a 3D CNN (See Médioni col 20 line 55 to col 21 line 11), but there is no relation to any plurality of workflow activities. The Examiner respectfully disagrees, because Médioni clearly discloses at operation 205, determining a presence of a highlight moment within the video content, which the highlight moment technically includes plurality of temporal events, [i.e., determining, based on the extracted representation, “based on set of spatiotemporal feature vectors relative to video summary data”, a predicted grouping of video segments associated with a plurality of workflow activities, “a set of predicted spatiotemporal feature vectors corresponding to different set of predicted video segments or video highlights associated with the highlight moments, implicitly including temporal events workflow”]). For the reasons stated above, the rejection of claim 20 was proper, and it is maintained. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 6, 10-13, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf et al, (US-PGPUB 20200272660) in view of Patel et al, (US-PGPUB 20220129501) In regards to claim 1, Wolf et al discloses a computing system comprising: a processor, (see at least: Abstarct, and Par. 0021, the system may include at least one processor), configured to: obtain surgical video data comprising a plurality of images, (see at least: Par. 0107, the video of the surgical procedure may be recorded by an image capture device, such as a camera, in an operating room or in a cavity of a patient, where the video of a surgical procedure may include any series of still images that were captured during and are associated with the surgical procedure, [i.e., obtain surgical video data, “video of a surgical”, comprising a plurality of images, “series of still images”]); perform events in other video footage, [i.e., perform processing, “training machine learning model”, on the surgical video data, “video footage”, to associate the plurality of images with a plurality of surgical activities, “implicit associating the video footage with surgical procedures, surgical phases, intraoperative events, and/or event characteristics”]); generate, based at least in part on the performed Par. 0160, analyzing the video footage to identify the video footage location associated with at least one of the surgical events or the surgical phase, [i.e., generate, based at least in part on the performed processing, “implicit by analyzing the video footage of region of interest of patient”, a prediction result, “identify the video footage location”]’ and from Par. 0155, the video footage location may refer to a time index or timestamp, a time range, a particular starting time and/or ending time, [i.e., wherein the prediction result, “the identify the video footage location”, is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, “the video footage location may refer to …starting time and ending time”]); Although discloses the machine learning model or neural network, such as deep neural network, convolutional neural networks, etc., (see at least: Par. 0118, 0160); Wolf et al does not expressly disclose using the natural language processing. Patel et al discloses using the natural language processing, (see at least: Par. 0204, the DPU uses the context generator to generate contextual attributes or a text transcript associated with the video data chunks; and from Fig. 1b, and Par. 0061, the context generator (108) may include one or more context generation models, which may include natural language processing models and/or body language detection models, for generating contextual attributes associated with video data, [i.e., using the natural language processing, “context generator (108) including context generation models”, to associate the plurality of images, “video data chunks”, with a plurality of activities, “contextual attributes”]; and generate, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, (see at least: Par. 0066, the virtual blob generator (110) may include the functionality to generate … indexing metadata, and contextual metadata associated with video data chunks using video data chunks; and from Par. 0093, the indexing metadata (210) may include indexing information associated with each video data chunk, including … start times, end times, …, where start times may represent timestamps in the video stream at which video data chunks start, and stop times may represent timestamps in the video stream at which video data chunks end, [i.e., generate, based at least in part on the performed natural language processing, a prediction result, “generate … indexing metadata by the virtual blob generator (110), implicitly after performing, the natural language processing, by the context generator 108”, wherein the prediction result, is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, “the indexing metadata (210) may include indexing information associated with each video data chunk, including … start times at which video data chunks start, and stop times at which video data chunks end”]). Wolf and Patel are combinable because they are both concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify Wolf, to apply the context generator (108) including natural language processing, as though by Patel, to the Wolf’s surgical video data, in order to generate the indexing metadata, associated with each video data chunk, including the start times at which video data chunks start, and stop times at which video data chunks end, (Par. 0093). In regards to claim 6, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Wolf further discloses wherein the prediction result comprises at least one of an annotated surgical video or metadata associated with the surgical video, (Wolf, see at least: Par. Fig. 6, table, under “Footage location”, where the video footage comprises, metadata indicating time stamp) In regards to claim 10, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Wolf further discloses wherein the plurality of surgical activities indicates one or more of a surgical event, a surgical phase, a surgical task, a surgical step, an idle period, or usage of a surgical tool, (Wolf, see at least: Par. 0109, the surgical timeline may be a list of descriptions of intraoperative surgical events or surgical phases within a surgical procedure). In regards to claim 11, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Wolf further discloses wherein the video data is received from a surgical device, wherein the surgical device is a surgical computing system, a surgical hub, a surgical-site camera, or a surgical surveillance system, (see at least: Fig. 1, Par. 0086-0088, room 101 may include one or more microphones (e.g., audio sensor 111, as shown in FIG. 1), several cameras (e.g., overhead cameras 115, 121, and 123, and a tableside camera 125) for capturing video/image data during surgery, “surgical-site camera”). In regards to claim 12, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Wolf further discloses wherein the natural language processing is associated with detecting a surgical tool in the video data, (Wolf, see at least: Par. 0079-0081, trained machine learning algorithm may include a classification algorithm, the input may include an image, and the inferred output may include a classification of an item depicted in the image; and from Par. 0087, camera 115 may be configured to track a surgical instrument (also referred to as a surgical tool) within location 127, an anatomical structure, a hand of surgeon 131, an incision, a movement of anatomical structure, and the like, [i.e., the image acquired by the camera while tracking the surgical tool, is implicitly input to the trained machine learning algorithm. That is, the trained machine learning algorithm processing is associated with detecting a surgical tool in the video data]); wherein the prediction result is configured to indicate a start time associated with a use of the surgical tool in the surgical procedure and an end time associated with the use of the surgical tool in the surgical procedure, (see at least: Par. 0169, event characteristics may include a time associated with the event (such as start time, end time, etc.), type of the event, information related to medical instruments involved in the event, [i.e., start time and end time associated with a use of medical instruments or the surgical tool]). In the other hand, Patel discloses the natural language processing, (Patel, see at least: Par. 0061-0063) Regarding claim 13, claim 13 recites substantially similar limitations as set forth in claim 1. As such, claim 13 is rejected for at least similar rational. The Examiner further acknowledged the following additional limitation(s): “a method”. However, Wolf discloses the “method”, (see at least: Par. 0005, “a methods for analysis of surgical videos). Regarding claim 17, claim 17 recites substantially similar limitations as set forth in claim 6. As such, claim 17 is rejected for at least similar rational. Claims 2 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf and Patel et al, as applied to claim 1 above; and further in view of Chalana et al, (US-PGPUB 20220215052) In regards to claim 2, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Patel further discloses wherein the performed natural language processing comprises: extracting a representation summary of the surgical video data summary of the video data chunks, [i.e., extracting a representation summary of the video data, “textual summary of the video data chunks”, using a the context generator (108)]). The combine teaching Wolf and Patel as whole does not expressly disclose extracting a representation summary of the surgical video data using a transformer network. However, Chalana discloses extracting a representation summary of the surgical video data using a transformer network, (see at least: Par. 0039, audio-visual media synopsis module 400 may prepare a transcript summary of the transcript, and may prepare a video summary of audio-visual media using neural networks, wherein the neural networks may be arranged in a transformer-based machine learning process, [i.e., extracting a representation summary of the video data, “video summary of audio-visual media”, using a transformer network, “transformer-based machine learning process’]). Wolf, Patel, and Chalana are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify combine teaching Wolf and Patel, to use transformer-based machine learning process, as though by Chalana, in order to generate a video summary of audio-visual media efficiently, and faster, (Chalana, Par. 0039) Regarding claim 14, claim 14 recites substantially similar limitations as set forth in claim 2. As such, claim 14 is rejected for at least similar rational. Claims 3-4, and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf and Patel et al, as applied to claim 1 above; and further in view of Chalana et al, (US-PGPUB 20220215052); and further in view of in view of Médioni, (US Patent 9,836,853) Regarding claim 3, claim 3 recites substantially similar limitations as set forth in the above claim 2, (see claim 2: wherein the performed natural language processing comprises: extracting a representation summary of the surgical video data a transformer network). As such, claim 3 is rejected for at least similar rational. However, the combine teaching Wolf, Patel, and Chalana as whole does not expressly disclose wherein the performed natural language processing comprises: extracting a representation summary of the surgical video data using a three-dimensional convolutional neural network (3D CNN). Médioni discloses extracting a representation summary of the surgical video data using a three-dimensional convolutional neural network (3D CNN), (see at least: col. 6, lines 54-56, Processor 11 may be configured to execute one or more machine readable instructions 100 to facilitate uses of three-dimensional convolutional neural network for video highlight detection, [i.e., extracting a representation summary of the video data, “video highlight detection”, using a three-dimensional convolutional neural network (3D CNN)]). Wolf, Patel, Chalana, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf, Patel, and Chalana, to use the three-dimensional convolutional neural network, as though by Médioni, in order to detect the video highlight, (Médioni, col. 6, lines 54-56) In regards to claim 4, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Patel further discloses wherein the performed natural language processing comprises: extracting a representation summary of the surgical video data using natural language processing, (Patel, see at least: Par. 0061, context generator (108), “natural language processing”, may also be used to generate text transcripts associated with video data chunks; and from Par. 0204, the text transcript may be a text file that includes a textual summary of the video data chunks, [i.e., extracting a representation summary of the video data, “textual summary of the video data chunks”, using the natural language processing, “context generator (108)”]). The combine teaching Wolf and Patel as whole does not expressly disclose wherein extracting using natural language processing is associated with a transformer; generating a vector representation based on the extracted representation summary; and determining, based on the generated vector representation, a predicted grouping of video segments using natural language processing. However, Chalana discloses wherein extracting using natural language processing is associated with a transformer, (see at least: Par. 0039, audio-visual media synopsis module 400 may prepare a transcript summary of the transcript, and may prepare a video summary of audio-visual media using neural networks, wherein the neural networks may be arranged in a transformer-based machine learning process, [i.e., extracting a representation summary of the video data, “video summary of audio-visual media”, using a transformer network, “transformer-based machine learning process’]). Wolf, Patel, and Chalana are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify combine teaching Wolf and Patel, to use transformer-based machine learning process, as though by Chalana, in order to generate a video summary of audio-visual media efficiently, and faster, (Chalana, Par. 0039) The combine teaching Wolf, Patel, and Chalana as whole does not expressly disclose generating a vector representation based on the extracted representation summary; and determining, based on the generated vector representation, a predicted grouping of video segments using natural language processing. However, Médioni discloses generating a vector representation based on the extracted representation summary; and determining, based on the generated vector representation, a predicted grouping of video segments using natural language processing, (see at least: steps 201-203 of Fig. 2, and col. 20, line 55 through col. 21, line 11, At operation 202, the video content may be segmented into a set of video segment, and at operation 203, the set of video segments may be inputted into a three-dimensional convolutional neural network, which the three-dimensional convolutional neural network may output a set of spatiotemporal feature vectors corresponding to the set of video segments, [i.e., generating a vector representation, “set of spatiotemporal feature vectors”, based on the extracted representation summary, “set of video segments”]. Further, at operation 204, the set of spatiotemporal feature vectors may be inputted into a long short-term memory network, which the long short-term memory network may determine a set of predicted spatiotemporal feature vectors based on the set of spatiotemporal feature vectors, and at operation 205, a presence of a highlight moment within the video content may be determined, which the highlight moment implicitly includes plurality of temporal events, [i.e., determining, based on the extracted representation, “based on set of spatiotemporal feature vectors relative to video summary data”, a predicted grouping of video segments associated with a plurality of workflow activities, “a set of predicted spatiotemporal feature vectors corresponding to different set of predicted video segments associated with the highlight moment, implicitly including temporal event steps or events workflow”], using natural language processing, “3D CNN”). Wolf, Patel, Chalana, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf, Patel, and Chalana, to use the 3D CNN and LSTM, as though by Médioni, in order to generate a set of predicted spatiotemporal feature vectors corresponding to predicted set video segments, based on the set of spatiotemporal feature vectors, (Médioni, col. 21, lines 12-19) Regarding claim 15, claim 15 recites substantially similar limitations as set forth in claim 3. As such, claim 15 is rejected for at least similar rational. Regarding claim 16, claim 16 recites substantially similar limitations as set forth in claim 4. As such, claim 16 is rejected for at least similar rational. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Wolf and Patel et al, as applied to claim 1 above; and further in view of Médioni, (US Patent 9,836,853) The combine teaching Wolf and Patel as whole discloses the limitations of the claim 1. Patel further discloses wherein the performed natural language processing comprises: extracting a representation summary of the surgical video data, (Patel, see at least: Par. 0061, context generator (108), “natural language processing”, may also be used to generate text transcripts associated with video data chunks; and from Par. 0204, the text transcript may be a text file that includes a textual summary of the video data chunks, [i.e., extracting a representation summary of the video data, “textual summary of the video data chunks”, using the natural language processing, “context generator (108)”]). The combine teaching Wolf and Patel as whole does not expressly disclose generating a vector representation based on the extracted representation summary; determining, based on the generated vector representation, a predicted grouping of video segments; and filtering the predicted grouping of video segments using natural language processing. However, Médioni discloses generating a vector representation based on the extracted representation summary; and determining, based on the generated vector representation, a predicted grouping of video segments using natural language processing, (see at least: steps 201-203 of Fig. 2, and col. 20, line 55 through col. 21, line 11, At operation 202, the video content may be segmented into a set of video segment, and at operation 203, the set of video segments may be inputted into a three-dimensional convolutional neural network, which the three-dimensional convolutional neural network may output a set of spatiotemporal feature vectors corresponding to the set of video segments, [i.e., generating a vector representation, “set of spatiotemporal feature vectors”, based on the extracted representation summary, “set of video segments”]. Further, at operation 204, the set of spatiotemporal feature vectors may be inputted into a long short-term memory network, which the long short-term memory network may determine a set of predicted spatiotemporal feature vectors based on the set of spatiotemporal feature vectors, and at operation 205, a presence of a highlight moment within the video content may be determined, [i.e., determining, based on the generated vector representation, “set of spatiotemporal feature vectors”, a predicted grouping of video segments, “a set of predicted spatiotemporal feature vectors corresponding to different set of predicted video segments”, using natural language processing, “3D CNN”). Médioni further discloses filtering the predicted grouping of video segments using natural language processing, (col. 8, lines 13-26, a three-dimensional convolutional neural network may include filters that are self-optimized through learning for classification of faces within images, where first three-dimensional convolutional neural network may be trained for video highlight detection using video segments of sixteen video frames. A second three-dimensional convolutional neural network may be trained for video highlight detection using video segments of twenty-four video frames, [i.e., filtering the predicted grouping of video segments, “detecting first video highlight using video segments of sixteen video frames, and second video highlight detection using video segments of twenty-four video frames”, using natural language processing, “3D CNN”]). Wolf, Patel, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf and Patel, to use the 3D CNN and LSTM, as though by Médioni, in order to generate a set of predicted spatiotemporal feature vectors corresponding to predicted set video segments, based on the set of spatiotemporal feature vectors, (Médioni, col. 21, lines 12-19) Claims 7 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf and Patel et al, as applied to claims 1 and 13 above; and further in view Venkataraman et al, (US-Patent 11,205,508) In regards to claim 7, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Patel further discloses the natural language processing, (Patel, see at least: Par. 0061-0063) The combine teaching Wolf and Patel as whole does not expressly disclose wherein the natural language processing is associated with: determining, using natural language processing, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase; and generating an output, wherein the output indicates a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time. However, Venkataraman discloses wherein the natural language processing is associated with: determining, using natural language processing, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase, (see at least: col. 2, lines 35-46, and col. 11, lines 48-52, two consecutive phases of the set of predefined phases can be separated by an identifiable “phase boundary” in the surgical videos, which indicates the end of a current phase and the beginning of the next phase in the surgical procedure; and from col. 16, lines 48-63, generates trained machine learning classifiers 522 … to generate iteratively more accurate phase boundaries for video segments 518, [i.e., determining, using “trained machine learning classifiers associated with machine learning descriptors 534”, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase, “two consecutive phases of the set of predefined phases can be separated by an identifiable “phase boundary”]); and generating an output, wherein the output indicates a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time, (see at least: col. 2, lines 39-41, detects the phase boundary by detecting an initial appearance of a surgical tool as an indicator of the beginning of a given phase; and from col. 11, lines 48-52, identifiable “phase boundary” in the surgical videos, which indicates the end of a current phase and the beginning of the next phase in the surgical procedure, [i.e., indicating a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time, “first start time and second time are implicitly the initial appearance of a surgical tool in the first and second surgical phases”]). Wolf, Patel, and Venkataraman are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf and Patel, to detects the phase boundary between two phases, as though by Venkataraman, in order to detect an event in a set of events as an indicator of the beginning of a given phase, including a cautery event; a bleeding event; and an adhesion event, (Venkataraman, col. 2, lines 44-46), in order to facilitate improving outcomes of surgeries and skills of surgeons, (Venkataraman, col. 1, lines 11-12) Regarding claim 18, claim 18 recites substantially similar limitations as set forth in claim 7. As such, claim 18 is rejected for at least similar rational. Claims 8-9, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf and Patel et al, as applied to claim 1 above; and further in view Donhowe et al, (US-Patent 11,974,813) In regards to claim 8, the combine teaching Wolf and Patel as whole discloses the limitations of claim 1. Patel further discloses the natural language processing, (Patel, see at least: Par. 0061-0063) The combine teaching Wolf and Patel as whole does not expressly disclose wherein the natural language processing is associated with: identifying an idle period, wherein the idle period is associated with inactivity during the surgical procedure; generating an output, wherein the output indicates an idle start time and an idle end time; and refining the prediction result based on the identified idle period. However, Donhowe discloses identifying an idle period, wherein the idle period is associated with inactivity during the surgical procedure, (see at least: col. 3, lines 63-66, a surgical procedure analysis system is able to identify delays that occurred during the surgical procedure by determining idle periods in the surgical procedure; and from col. 7, lines 15-18, idle period can also include a time period when the robotic surgical system 102 is active but the tools on the robotic surgical system 102 are inactive. Further, col. 12, lines 25-29, the surgical procedure analysis system 104 may employ a machine learning model to identify the sequence of events that are associated with the delay based on the records in the surgical procedure data 114, [i.e., identifying an idle period, “determining idle periods in the surgical procedure implicitly using machine learning model”, wherein the idle period is associated with inactivity during the surgical procedure, “time period where the robotic surgical system 102 are inactive”]); generating an output, wherein the output indicates an idle start time and an idle end time, (see at least: col. 12, lines 25-29, the surgical procedure analysis system 104 may employ a machine learning model to identify the sequence of events that are associated with the delay based on the records in the surgical procedure data 114, [i.e., the sequence of event associated with idle implicitly include an idle start time and an idle end time]); and refining the prediction result based on the identified idle period, (see at least: col. 5, lines 11-15, the computing environment 100 further includes a server device 110 configured to present the recommendations to a surgeon 122 or other medical personal and to update the procedure setup plan 112 for future related surgical procedures, [i.e., refining the prediction result based on the identified idle period, “update the procedure setup plan 112 for future related surgical procedures, implicitly based on the identified idle period”]). Wolf, Patel, and Donhowe are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf and Patel, to surgical procedure analysis system 104, as though by Donhowe, in order to identify the sequence of events that are associated with the delay based on the records in the surgical procedure data 114, (Donhowe, col. 12, lines 25-29); and further determine recommendations for improving the surgical procedure, (Donhowe, col. 5, lines 8-11). The following prior art of record, Thurimella, (US-Patent 11,340,887), discloses also the following limitations of claim 8, as follow: -- Thurimella discloses identifying an idle period, wherein the idle period is associated with inactivity during an event, (see at least: col. 2, lines 33-65, during operation of the control unit, at least one idle time interval is detected, in which at least one software module of the control unit is currently not required, “i.e., idle time interval is associated with the control unit inactivity”); and generating an output, wherein the output indicates an idle start time and an idle end time, (col. 3, lines 61-64, analysis device determines the idle time interval on the basis of a machine learning method and/or a predictive analytics method by using historical operating pattern, which the machine learning method can be implemented on the basis of an artificial neural network; and from col. 10, lines 19-22, wherein predicting the idle time interval comprises predicting a start time and an end time of the idle time interval based at least in part on the at least one historical operating pattern, [i.e., generating an output, wherein the output indicates an idle start time and an idle end time, “predicting a start time and an end time of the idle time interval, based implicitly on training the machine learning with historical operating pattern]); and refining the prediction result based on the identified idle period, (see at least: col. 2, lines 33-37, after the idle time interval is predicted by the analysis device (i.e., ML), the software update is then started at the beginning of the idle time interval) In regards to claim 9, the combine teaching Wolf, Patel, and Donhowe as whole discloses the limitations of claim 8. Donhowe further discloses generate a surgical procedure improvement recommendation based on the identified idle period, (see at least: see at least: Abstract, and col. 10, lines 12-15, generating a recommendation for modifying the procedure setup plan for the robotic surgical procedure based on the determined cause of the idle period). Regarding claim 19, claim 19 recites substantially similar limitations as set forth in claim 8. As such, claim 19 is rejected for at least similar rational. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Wolf et al, (US-PGPUB 20200272660) in view of Patel et al, (US-PGPUB 20220129501); and further in view of Médioni, (US Patent 9,836,853) Wolf discloses a computing system comprising: a processor, (see at least: Abstarct, and Par. 0021, the system may include at least one processor), configured to: obtain surgical video data comprising a plurality of images, (see at least: Par. 0107, the video of the surgical procedure may be recorded by an image capture device, such as a camera, in an operating room or in a cavity of a patient, where the video of a surgical procedure may include any series of still images that were captured during and are associated with the surgical procedure, [i.e., obtain surgical video data, “video of a surgical”, comprising a plurality of images, “series of still images”]); extracting a representation summary of the video data generate, based at least in part on the performed Par. 0160, analyzing the video footage to identify the video footage location associated with at least one of the surgical events or the surgical phase, [i.e., generate, based at least in part on the performed processing, “implicit by analyzing the video footage of region of interest of patient”, a prediction result, “identify the video footage location”]’ and from Par. 0155, the video footage location may refer to a time index or timestamp, a time range, a particular starting time and/or ending time, [i.e., wherein the prediction result, “the identify the video footage location”, is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, “the video footage location may refer to …starting time and ending time”]). Wolf does not expressly disclose that the representation summary of the video data being generated using a natural language processing network; determining, based on the extracted representation, a predicted grouping of video segments associated with a plurality of workflow activities; and generate, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of workflow activities in the surgical video data. However, Patel discloses extracting a representation summary of the video data at least in part using a natural language processing network, (see at least: Par. 0061, context generator (108), “natural language processing”, may also be used to generate text transcripts associated with video data chunks; and from Par. 0204, the text transcript may be a text file that includes a textual summary of the video data chunks, [i.e., extracting a representation summary of the video data, “textual summary of the video data chunks”, using the natural language processing, “context generator (108)”]); and generate, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of workflow activities in the surgical video data, (see at least: Par. 0066, the virtual blob generator (110) may include the functionality to generate … indexing metadata, and contextual metadata associated with video data chunks using video data chunks; and from Par. 0093, the indexing metadata (210) may include indexing information associated with each video data chunk, including … start times, end times, …, where start times may represent timestamps in the video stream at which video data chunks start, and stop times may represent timestamps in the video stream at which video data chunks end, [i.e., generate, based at least in part on the performed natural language processing, a prediction result, “generate … indexing metadata by the virtual blob generator (110), implicitly after performing, the natural language processing, by the context generator 108”, wherein the prediction result, is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, “the indexing metadata (210) may include indexing information associated with each video data chunk, including … start times at which video data chunks start, and stop times at which video data chunks end”]). Wolf and Patel are combinable because they are both concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify Wolf, to apply the context generator (108), as though by Patel, in order to generate the indexing metadata, associated with each video data chunk, including the start times at which video data chunks start, and stop times at which video data chunks end, (Patel, Par. 0093). The combine teaching Wolf and Patel as whole does not expressly disclose determining, based on the extracted representation, a predicted grouping of video segments associated with a plurality of workflow activities. However, Médioni discloses determining, based on the extracted representation, a predicted grouping of video segments associated with a plurality of workflow activities, (see at least: steps 201-203 of Fig. 2, and col. 20, line 55 through col. 21, line 11, At operation 202, the video content may be segmented into a set of video segment, and at operation 203, the set of video segments may be inputted into a three-dimensional convolutional neural network, which the three-dimensional convolutional neural network may output a set of spatiotemporal feature vectors corresponding to the set of video segment; and at operation 204, the set of spatiotemporal feature vectors may be inputted into a long short-term memory network, for determining a set of predicted spatiotemporal feature vectors based on the set of spatiotemporal feature vectors, and at operation 205, a presence of a highlight moment within the video content may be determined, [i.e., determining, based on the extracted representation, “based on set of spatiotemporal feature vectors relative to video summary data”, a predicted grouping of video segments associated with a plurality of workflow activities, “a set of predicted spatiotemporal feature vectors corresponding to different set of predicted video segments associated with the highlight moment, implicitly including temporal event steps or events workflow”]). Wolf, Patel, Chalana, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf, Patel, and Chalana, to use the 3D CNN and LSTM, as though by Médioni, in order to generate a set of predicted spatiotemporal feature vectors corresponding to predicted set video segments, based on the set of spatiotemporal feature vectors, (Médioni, col. 21, lines 12-19) Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMARA ABDI whose telephone number is (571)272-0273. The examiner can normally be reached 9:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMARA ABDI/Primary Examiner, Art Unit 2668 04/22/2026
Read full office action

Prosecution Timeline

Oct 13, 2023
Application Filed
Nov 14, 2025
Non-Final Rejection mailed — §103
Feb 17, 2026
Response Filed
Apr 24, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682484
TARGET TRACKING METHOD AND APPARATUS, DEVICE, AND MEDIUM
2y 10m to grant Granted Jul 14, 2026
Patent 12651364
METHOD AND SYSTEM FOR ESTIMATING THE LENGTH OF A VESSEL
3y 2m to grant Granted Jun 09, 2026
Patent 12646325
VIDEO ANALYSIS SYSTEM USING EDGE COMPUTING
3y 0m to grant Granted Jun 02, 2026
Patent 12646356
AUTOMATED EYE TRACKING ASSESSMENT SOLUTION FOR SKILL DEVELOPMENT
2y 11m to grant Granted Jun 02, 2026
Patent 12646328
PARKING FACILITY SYSTEM FOR VEHICLE DETECTION AND IDENTIFICATION
1y 5m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
76%
With Interview (-7.2%)
2y 6m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 831 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month