Prosecution Insights
Last updated: October 04, 2026
Application No. 18/286,969

COMPUTER VISION-BASED SURGICAL WORKFLOW RECOGNITION SYSTEM USING NATURAL LANGUAGE PROCESSING TECHNIQUES

Non-Final OA §103
Filed
Oct 13, 2023
Priority
Apr 14, 2021 — provisional 63/174,820 +1 more
Examiner
ABDI, AMARA
Art Unit
2668
Tech Center
2600 — Communications
Assignee
Csats Inc.
OA Round
3 (Non-Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
697 granted / 840 resolved
+21.0% vs TC avg
Minimal -7% lift
Without
With
+-7.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
23 currently pending
Career history
859
Total Applications
across all art units

Statute-Specific Performance

§101
11.0%
-29.0% vs TC avg
§103
64.5%
+24.5% vs TC avg
§102
9.7%
-30.3% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 840 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on July 24, 2026 has been entered. Response to Amendment Applicant's response to the last office action, filed July 24, 2026 has been entered and made of record. Claims 1, 13, and 20 are amended. Claims 1-20 are pending in this application for examination. Response to Arguments Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 6, 10-14, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf et al, (US-PGPUB 20200272660) in view of Ravine, (US-PGPUB 20210099505) Regarding claim 1, Wolf et al discloses a computing system comprising: a processor, (see at least: Abstarct, and Par. 0021, the system may include at least one processor), configured to: obtain surgical video data comprising a plurality of images, (see at least: Par. 0107, the video of the surgical procedure may be recorded by an image capture device, such as a camera, in an operating room or in a cavity of a patient, where the video of a surgical procedure may include any series of still images that were captured during and are associated with the surgical procedure, [i.e., obtain surgical video data, “video of a surgical”, comprising a plurality of images, “series of still images”]); perform video footage, “i.e., the video footage implicitly relates to surgical video data”, [i.e., perform processing, “training machine learning model”, on an input, “video footage”, to associate the plurality of images with a plurality of surgical activities, “implicitly associating the video footage with surgical procedures, surgical phases, intraoperative events, and/or event characteristics”, wherein the input is at least one image in the surgical video data, “the video footage at input implicitly comprises to at least one image in the surgical video data”]); generate, based at least in part on the performed Par. 0160, analyzing the video footage to identify the video footage location associated with at least one of the surgical events or the surgical phase, [i.e., generate, based at least in part on the performed processing, “implicit by analyzing the video footage of region of interest of patient”, a prediction result, “identify the video footage location”]’ and from Par. 0155, the video footage location may refer to a time index or timestamp, a time range, a particular starting time and/or ending time, [i.e., wherein the prediction result, “the identify the video footage location”, is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, “the video footage location may refer to …starting time and ending time”]); Although Wolf discloses the machine learning model or neural network, such as deep neural network, convolutional neural networks, etc., (see at least: Par. 0118, 0160); Wolf et al does not expressly disclose performing the natural language processing on an input, wherein the input is at least one image in the surgical video data; and generating, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data. However, Ravine discloses performing the natural language processing on at least one image in the video data, (see at least: Par. 0118-0119, where a transformer machine-learning language model could be trained on videos and video summaries in a multi-modal fashion (e.g., being trained on text and videos at the same time) to output video summaries in an end-to-end single process; and on Par. 0117, summarization analysis techniques may be trained to detect and consider additional factors such as scene transitions and salient visual activity to determine portions of a target video that should be included in a summarized video, [i.e., performing the natural language processing, “implicit by using transformer machine-learning language model”, on at least one image in the video data, “implicitly by using one or more videos or video images”, to associate the plurality of images with a plurality of activities, “salient visual activity to determine portions of a target video that should be included in a summarized video”]); generating, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, (see at least: Par. 0119-0120, users can create video summaries with a manual summarization tool to allow the model to learn what users would like summarized, and then (2) the model can be used to generate multiple versions of video summaries that are then presented to users, with an editing tool that allows the users to choose which summarization is more accurate and/or provide direct feedback to the model, where the user my specify the time interval, and the model attempts to determine which information is more relevant and creates an even more compact summarized version that meets the time specification, [i.e., generating, based at least in part on the performed natural language processing, “implicit by using transformer machine-learning language model”, a prediction result, “generating summarized video”, wherein the prediction result, “the summarized video”, is configured to indicate a start time and an end time of the plurality of activities in the video data, “user’s specifying the time interval, which implicitly includes start time and end time”]). Wolf and Ravine are combinable because they are both concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify Wolf, to train the transformer machine-learning language model, on videos and video summaries in a multi-modal fashion, as though by Ravine, in order to generate summarized video that meets the user’s time specification, (Ravine, Par. 0120). Regarding claim 2, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. Ravine further discloses wherein the performed natural language processing comprises extracting a representation summary of the surgical video data using a transformer network, (see at least: Par. 0119-0120, implicitly creating summarized video, “i.e., representation summary of the video data”, using transformer machine-learning language model) Regarding claim 6, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. Wolf further discloses wherein the prediction result comprises at least one of an annotated surgical video or metadata associated with the surgical video, (Wolf, see at least: Par. Fig. 6, table, under “Footage location”, where the video footage comprises, metadata indicating time stamp) Regarding claim 10, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. Wolf further discloses wherein the plurality of surgical activities indicates one or more of a surgical event, a surgical phase, a surgical task, a surgical step, an idle period, or usage of a surgical tool, (Wolf, see at least: Par. 0109, the surgical timeline may be a list of descriptions of intraoperative surgical events or surgical phases within a surgical procedure). Regarding claim 11, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. Wolf further discloses wherein the video data is received from a surgical device, wherein the surgical device is a surgical computing system, a surgical hub, a surgical-site camera, or a surgical surveillance system, (see at least: Fig. 1, Par. 0086-0088, room 101 may include one or more microphones (e.g., audio sensor 111, as shown in FIG. 1), several cameras (e.g., overhead cameras 115, 121, and 123, and a tableside camera 125) for capturing video/image data during surgery, “surgical-site camera”). Regarding claim 12, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. Wolf further discloses wherein the natural language processing is associated with detecting a surgical tool in the video data, (Wolf, see at least: Par. 0079-0081, trained machine learning algorithm may include a classification algorithm, the input may include an image, and the inferred output may include a classification of an item depicted in the image; and from Par. 0087, camera 115 may be configured to track a surgical instrument (also referred to as a surgical tool) within location 127, an anatomical structure, a hand of surgeon 131, an incision, a movement of anatomical structure, and the like, [i.e., the image acquired by the camera while tracking the surgical tool, is implicitly input to the trained machine learning algorithm. That is, the trained machine learning algorithm processing is associated with detecting a surgical tool in the video data]); wherein the prediction result is configured to indicate a start time associated with a use of the surgical tool in the surgical procedure and an end time associated with the use of the surgical tool in the surgical procedure, (see at least: Par. 0169, event characteristics may include a time associated with the event (such as start time, end time, etc.), type of the event, information related to medical instruments involved in the event, [i.e., start time and end time associated with a use of medical instruments or the surgical tool]). In the other hand, Patel discloses the natural language processing, (Patel, see at least: Par. 0061-0063) Regarding claim 13, claim 13 recites substantially similar limitations as set forth in claim 1. As such, claim 13 is rejected for at least similar rational. The Examiner further acknowledged the following additional limitation(s): “a method”. However, Wolf discloses the “method”, (Wolf, see at least: Par. 0005, “a methods for analysis of surgical videos). Regarding claim 14, claim 14 recites substantially similar limitations as set forth in claim 2. As such, claim 14 is rejected for at least similar rational. Regarding claim 17, claim 17 recites substantially similar limitations as set forth in claim 6. As such, claim 17 is rejected for at least similar rational. Claims 3-5, 15-16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Wolf and Ravine, as applied to claim 1 above; and further in view of Médioni, (US Patent 9,836,853) Regarding claim 3, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. wherein the performed natural language processing comprises extracting a representation summary of the surgical video data using a transformer network, (see at least: Par. 0119-0120, implicitly creating summarized video, “i.e., representation summary of the video data”, using transformer machine-learning language model) The combination of teaching Wolf and Ravine as whole does not expressly disclose extracting a representation summary of the surgical video data using a three-dimensional convolutional neural network (3D CNN). Médioni discloses extracting a representation summary of the surgical video data using a three-dimensional convolutional neural network (3D CNN), (see at least: col. 6, lines 54-56, Processor 11 may be configured to execute one or more machine readable instructions 100 to facilitate uses of three-dimensional convolutional neural network for video highlight detection, [i.e., extracting a representation summary of the video data, “video highlight detection”, using a three-dimensional convolutional neural network (3D CNN)]). Wolf, Ravine, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf and Ravine, to use the three-dimensional convolutional neural network, as though by Médioni, in order to detect the video highlight, (Médioni, col. 6, lines 54-56). Regarding claim 4, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. Ravine further discloses wherein the performed natural language processing comprises extracting a representation summary of the surgical video data using natural language processing, wherein extracting using natural language processing is associated with a transformer, (see at least: Par. 0119-0120, implicit by creating summarized video, “i.e., representation summary of the video data”, using transformer machine-learning language model). The combination of Wolf and Ravine as whole does not expressly disclose generating a vector representation based on the extracted representation summary; and determining, based on the generated vector representation, a predicted grouping of video segments using natural language processing. However, Médioni discloses generating a vector representation based on the extracted representation summary; and determining, based on the generated vector representation, a predicted grouping of video segments using natural language processing, (see at least: steps 201-203 of Fig. 2, and col. 20, line 55 through col. 21, line 11, At operation 202, the video content may be segmented into a set of video segment, and at operation 203, the set of video segments may be inputted into a three-dimensional convolutional neural network, which the three-dimensional convolutional neural network may output a set of spatiotemporal feature vectors corresponding to the set of video segments, [i.e., generating a vector representation, “set of spatiotemporal feature vectors”, based on the extracted representation summary, “set of video segments”]. Further, at operation 204, the set of spatiotemporal feature vectors may be inputted into a long short-term memory network, which the long short-term memory network may determine a set of predicted spatiotemporal feature vectors based on the set of spatiotemporal feature vectors, and at operation 205, a presence of a highlight moment within the video content may be determined, which the highlight moment implicitly includes plurality of temporal events, [i.e., determining, based on the extracted representation, “based on set of spatiotemporal feature vectors relative to video summary data”, a predicted grouping of video segments associated with a plurality of workflow activities, “a set of predicted spatiotemporal feature vectors corresponding to different set of predicted video segments associated with the highlight moment, implicitly including temporal event steps or events workflow”], using natural language processing, “3D CNN”). Wolf, Ravine, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Wolf and Ravine, to use the 3D CNN and LSTM, as though by Médioni, in order to generate a set of predicted spatiotemporal feature vectors corresponding to predicted set video segments, based on the set of spatiotemporal feature vectors, (Médioni, col. 21, lines 12-19) Regarding claim 5, the combination of teaching Wolf and Ravine as whole discloses the limitations of claim 1. Ravine further discloses wherein the performed natural language processing comprises extracting a representation summary of the surgical video data, (see at least: Par. 0119-0120, implicitly creating summarized video, “i.e., representation summary of the video data”, using transformer machine-learning language model) The combination of Wolf and Ravine as whole does not expressly disclose generating a vector representation based on the extracted representation summary; determining, based on the generated vector representation, a predicted grouping of video segments; and filtering the predicted grouping of video segments using natural language processing. However, Médioni discloses generating a vector representation based on the extracted representation summary; and determining, based on the generated vector representation, a predicted grouping of video segments using natural language processing, (see at least: steps 201-203 of Fig. 2, and col. 20, line 55 through col. 21, line 11, At operation 202, the video content may be segmented into a set of video segment, and at operation 203, the set of video segments may be inputted into a three-dimensional convolutional neural network, which the three-dimensional convolutional neural network may output a set of spatiotemporal feature vectors corresponding to the set of video segments, [i.e., generating a vector representation, “set of spatiotemporal feature vectors”, based on the extracted representation summary, “set of video segments”]. Further, at operation 204, the set of spatiotemporal feature vectors may be inputted into a long short-term memory network, which the long short-term memory network may determine a set of predicted spatiotemporal feature vectors based on the set of spatiotemporal feature vectors, and at operation 205, a presence of a highlight moment within the video content may be determined, [i.e., determining, based on the generated vector representation, “set of spatiotemporal feature vectors”, a predicted grouping of video segments, “a set of predicted spatiotemporal feature vectors corresponding to different set of predicted video segments”, using natural language processing, “3D CNN”). Médioni further discloses filtering the predicted grouping of video segments using natural language processing, (col. 8, lines 13-26, a three-dimensional convolutional neural network may include filters that are self-optimized through learning for classification of faces within images, where first three-dimensional convolutional neural network may be trained for video highlight detection using video segments of sixteen video frames. A second three-dimensional convolutional neural network may be trained for video highlight detection using video segments of twenty-four video frames, [i.e., filtering the predicted grouping of video segments, “detecting first video highlight using video segments of sixteen video frames, and second video highlight detection using video segments of twenty-four video frames”, using natural language processing, “3D CNN”]). Wolf, Ravine, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Wolf and Ravine, to use the 3D CNN and LSTM, as though by Médioni, in order to generate a set of predicted spatiotemporal feature vectors corresponding to predicted set video segments, based on the set of spatiotemporal feature vectors, (Médioni, col. 21, lines 12-19) Regarding claim 15, claim 15 recites substantially similar limitations as set forth in claim 3. As such, claim 15 is rejected for at least similar rational. Regarding claim 16, claim 16 recites substantially similar limitations as set forth in claim 4. As such, claim 16 is rejected for at least similar rational. Regarding claim 20, Wolf discloses a computing system comprising: a processor, (see at least: Abstarct, and Par. 0021, the system may include at least one processor), configured to: obtain surgical video data comprising a plurality of images, (see at least: Par. 0107, the video of the surgical procedure may be recorded by an image capture device, such as a camera, in an operating room or in a cavity of a patient, where the video of a surgical procedure may include any series of still images that were captured during and are associated with the surgical procedure, [i.e., obtain surgical video data, “video of a surgical”, comprising a plurality of images, “series of still images”]); extracting a representation summary of the video data generate, based at least in part on the performed Par. 0160, analyzing the video footage to identify the video footage location associated with at least one of the surgical events or the surgical phase, [i.e., generate, based at least in part on the performed processing, “implicit by analyzing the video footage of region of interest of patient”, a prediction result, “identify the video footage location”]’ and from Par. 0155, the video footage location may refer to a time index or timestamp, a time range, a particular starting time and/or ending time, [i.e., wherein the prediction result, “the identify the video footage location”, is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, “the video footage location may refer to …starting time and ending time”]). Wolf does not expressly disclose that the representation summary of the video data being generated using a natural language processing network on an input comprising the plurality of images; determining, based on the extracted representation, a predicted grouping of video segments associated with a plurality of workflow activities; and generate, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of workflow activities in the surgical video data. However, Ravine discloses extracting a representation summary of the video data at least in part using a natural language processing network on an input comprising the plurality of images, (see at least: Par. 0118-0120, where a transformer machine-learning language model could be trained on videos, “i.e., an input comprising the plurality of images”, and video summaries in a multi-modal fashion (e.g., being trained on text and videos at the same time) to output video summaries in an end-to-end single process, [i.e., extracting a representation summary of the video data, “generating summarized video”, at least in part using a natural language processing network on an input comprising the plurality of images, “implicit by training transformer machine-learning language model on videos and video summaries”]); and generate, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of workflow activities in the surgical video data, (see at least: Par. 0119-0120, users can create video summaries with a manual summarization tool to allow the model to learn what users would like summarized, and then (2) the model can be used to generate multiple versions of video summaries that are then presented to users, with an editing tool that allows the users to choose which summarization is more accurate and/or provide direct feedback to the model, where the user my specify the time interval, and the model attempts to determine which information is more relevant and creates an even more compact summarized version that meets the time specification, [i.e., generating, based at least in part on the performed natural language processing, “implicit by using transformer machine-learning language model”, a prediction result, “generating summarized video”, wherein the prediction result, “the summarized video”, is configured to indicate a start time and an end time of the plurality of activities in the video data, “user’s specifying the time interval, which implicitly includes start time and end time”]). Wolf and Ravine are combinable because they are both concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify Wolf, to train the transformer machine-learning language model, on videos and video summaries in a multi-modal fashion, as though by Ravine, in order to generate summarized video that meets the user’s time specification, (Ravine, Par. 0120). The combination of Wolf and Ravine as whole does not expressly disclose determining, based on the extracted representation, a predicted grouping of video segments associated with a plurality of workflow activities. However, Médioni discloses determining, based on the extracted representation, a predicted grouping of video segments associated with a plurality of workflow activities, (see at least: steps 201-203 of Fig. 2, and col. 20, line 55 through col. 21, line 11, At operation 202, the video content may be segmented into a set of video segment, and at operation 203, the set of video segments may be inputted into a three-dimensional convolutional neural network, which the three-dimensional convolutional neural network may output a set of spatiotemporal feature vectors corresponding to the set of video segment; and at operation 204, the set of spatiotemporal feature vectors may be inputted into a long short-term memory network, for determining a set of predicted spatiotemporal feature vectors based on the set of spatiotemporal feature vectors, and at operation 205, a presence of a highlight moment within the video content may be determined, [i.e., determining, based on the extracted representation, “based on set of spatiotemporal feature vectors relative to video summary data”, a predicted grouping of video segments associated with a plurality of workflow activities, “a set of predicted spatiotemporal feature vectors corresponding to different set of predicted video segments associated with the highlight moment, implicitly including temporal event steps or events workflow”]). Wolf, Ravine, and Médioni are combinable because they are all concerned with video data processing. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Wolf, and Ravine, to use the 3D CNN and LSTM, as though by Médioni, in order to generate a set of predicted spatiotemporal feature vectors corresponding to predicted set video segments, based on the set of spatiotemporal feature vectors, (Médioni, col. 21, lines 12-19). Allowable Subject Matter Claims 7-9, and 18-19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. With respect to claim 7, the prior art of record, alone or in reasonable combination, does not teach or suggest, the following underlined limitation(s), (in consideration of the claim as a whole): “wherein the natural language processing is associated with: determining, using natural language processing, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase; and generating an output, wherein the output indicates a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time”. The relevant prior art of record, Wolf et al, (US-PGPUB 20200272660) discloses a computing system comprising: a processor, (see at least: Abstarct, and Par. 0021, the system may include at least one processor), configured to: obtain surgical video data comprising a plurality of images, (see at least: Par. 0107, the video of the surgical procedure may be recorded by an image capture device, such as a camera, in an operating room or in a cavity of a patient, where the video of a surgical procedure may include any series of still images that were captured during and are associated with the surgical procedure, [i.e., obtain surgical video data, “video of a surgical”, comprising a plurality of images, “series of still images”]); perform video footage, “i.e., the video footage implicitly relates to surgical video data”, [i.e., perform processing, “training machine learning model”, on an input, “video footage”, to associate the plurality of images with a plurality of surgical activities, “implicitly associating the video footage with surgical procedures, surgical phases, intraoperative events, and/or event characteristics”, wherein the input is at least one image in the surgical video data, “the video footage at input implicitly comprises to at least one image in the surgical video data”]); generate, based at least in part on the performed Par. 0160, analyzing the video footage to identify the video footage location associated with at least one of the surgical events or the surgical phase, [i.e., generate, based at least in part on the performed processing, “implicit by analyzing the video footage of region of interest of patient”, a prediction result, “identify the video footage location”]’ and from Par. 0155, the video footage location may refer to a time index or timestamp, a time range, a particular starting time and/or ending time, [i.e., wherein the prediction result, “the identify the video footage location”, is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, “the video footage location may refer to …starting time and ending time”]); However, Wolf et al fails to teach or suggest, either alone or in combination with the other cited references, wherein the natural language processing is associated with: determining, using natural language processing, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase; and generating an output, wherein the output indicates a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time A further prior art of record, Ravine, (US-PGPUB 20210099505) discloses performing the natural language processing on at least one image in the video data, (see at least: Par. 0118-0119, where a transformer machine-learning language model could be trained on videos and video summaries in a multi-modal fashion (e.g., being trained on text and videos at the same time) to output video summaries in an end-to-end single process; and on Par. 0117, summarization analysis techniques may be trained to detect and consider additional factors such as scene transitions and salient visual activity to determine portions of a target video that should be included in a summarized video, [i.e., performing the natural language processing, “implicit by using transformer machine-learning language model”, on at least one image in the video data, “implicitly by using one or more videos or video images”, to associate the plurality of images with a plurality of activities, “salient visual activity to determine portions of a target video that should be included in a summarized video”]); generating, based at least in part on the performed natural language processing, a prediction result, wherein the prediction result is configured to indicate a start time and an end time of the plurality of surgical activities in the surgical video data, (see at least: Par. 0119-0120, users can create video summaries with a manual summarization tool to allow the model to learn what users would like summarized, and then (2) the model can be used to generate multiple versions of video summaries that are then presented to users, with an editing tool that allows the users to choose which summarization is more accurate and/or provide direct feedback to the model, where the user my specify the time interval, and the model attempts to determine which information is more relevant and creates an even more compact summarized version that meets the time specification, [i.e., generating, based at least in part on the performed natural language processing, “implicit by using transformer machine-learning language model”, a prediction result, “generating summarized video”, wherein the prediction result, “the summarized video”, is configured to indicate a start time and an end time of the plurality of activities in the video data, “user’s specifying the time interval, which implicitly includes start time and end time”]). However, while disclosing training transformer machine-learning language model on videos and video summaries in a multi-modal fashion, to create a summarized video version that meets the user’s time specification; Ravine fails to teach or suggest, either alone or in combination with the other cited references, wherein the natural language processing is associated with: determining, using natural language processing, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase; and generating an output, wherein the output indicates a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time. The prior art of record, Venkataraman et al, (US-Patent 11,205,508) discloses determining, using machine learning model, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase, (see at least: col. 2, lines 35-46, and col. 11, lines 48-52, two consecutive phases of the set of predefined phases can be separated by an identifiable “phase boundary” in the surgical videos, which indicates the end of a current phase and the beginning of the next phase in the surgical procedure; and from col. 16, lines 48-63, training machine learning classifiers 522 … to generate iteratively more accurate phase boundaries for video segments 518, [i.e., determining, using “trained machine learning classifiers associated with machine learning descriptors 534”, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase, “two consecutive phases of the set of predefined phases can be separated by an identifiable “phase boundary”]); and generating an output, wherein the output indicates a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time, (see at least: col. 2, lines 39-41, detects the phase boundary by detecting an initial appearance of a surgical tool as an indicator of the beginning of a given phase; and from col. 11, lines 48-52, identifiable “phase boundary” in the surgical videos, which indicates the end of a current phase and the beginning of the next phase in the surgical procedure, [i.e., indicating a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time, “first start time and second time are implicitly the initial appearance of a surgical tool in the first and second surgical phases”]). However, Venkataraman fails to teach or suggest, either alone or in combination with the other cited references, wherein the natural language processing is associated with: determining, using natural language processing, a phase boundary associated with the plurality of surgical activities, wherein the phase boundary indicates a boundary between a first surgical phase and a second surgical phase; and generating an output, wherein the output indicates a first surgical phase start time, a first surgical phase end time, a second surgical phase start time, and a second surgical phase end time. With respect to claim 8, the prior art of record, alone or in reasonable combination, does not teach or suggest, the following underlined limitation(s), (in consideration of the claim as a whole): “wherein the natural language processing is associated with identifying an idle period, wherein the idle period is associated with inactivity during the surgical procedure; generating an output, wherein the output indicates an idle start time and an idle end time; and refining the prediction result based on the identified idle period” The prior art of record Wolf, Ravine, Venkataraman, indicated above with respect to claim 7, apply also to claim 8, but none, either alone or in combination, teach or suggest all the claimed limitations of claim 8 A further prior art of record, Donhowe et al, (US-Patent 11,974,813) discloses identifying an idle period, wherein the idle period is associated with inactivity during the surgical procedure, (see at least: col. 3, lines 63-66, a surgical procedure analysis system is able to identify delays that occurred during the surgical procedure by determining idle periods in the surgical procedure; and from col. 7, lines 15-18, idle period can also include a time period when the robotic surgical system 102 is active but the tools on the robotic surgical system 102 are inactive. Further, col. 12, lines 25-29, the surgical procedure analysis system 104 may employ a machine learning model to identify the sequence of events that are associated with the delay based on the records in the surgical procedure data 114, [i.e., identifying an idle period, “determining idle periods in the surgical procedure implicitly using machine learning model”, wherein the idle period is associated with inactivity during the surgical procedure, “time period where the robotic surgical system 102 are inactive”]); generating an output, wherein the output indicates an idle start time and an idle end time, (see at least: col. 12, lines 25-29, the surgical procedure analysis system 104 may employ a machine learning model to identify the sequence of events that are associated with the delay based on the records in the surgical procedure data 114, [i.e., the sequence of event associated with idle implicitly include an idle start time and an idle end time]); and refining the prediction result based on the identified idle period, (see at least: col. 5, lines 11-15, the computing environment 100 further includes a server device 110 configured to present the recommendations to a surgeon 122 or other medical personal and to update the procedure setup plan 112 for future related surgical procedures, [i.e., refining the prediction result based on the identified idle period, “update the procedure setup plan 112 for future related surgical procedures, implicitly based on the identified idle period”]). However, Donhowe fails to teach or suggest, either alone or in combination with the other cited references, that the natural language processing is associated with identifying an idle period, wherein the idle period is associated with inactivity during the surgical procedure; generating an output, wherein the output indicates an idle start time and an idle end time; and refining the prediction result based on the identified idle period The prior art of record, Thurimella, (US-Patent 11,340,887), discloses identifying an idle period, wherein the idle period is associated with inactivity during an event, (see at least: col. 2, lines 33-65, during operation of the control unit, at least one idle time interval is detected, in which at least one software module of the control unit is currently not required, “i.e., idle time interval is associated with the control unit inactivity”); and generating an output, wherein the output indicates an idle start time and an idle end time, (col. 3, lines 61-64, analysis device determines the idle time interval on the basis of a machine learning method and/or a predictive analytics method by using historical operating pattern, which the machine learning method can be implemented on the basis of an artificial neural network; and from col. 10, lines 19-22, wherein predicting the idle time interval comprises predicting a start time and an end time of the idle time interval based at least in part on the at least one historical operating pattern, [i.e., generating an output, wherein the output indicates an idle start time and an idle end time, “predicting a start time and an end time of the idle time interval, based implicitly on training the machine learning with historical operating pattern]); and refining the prediction result based on the identified idle period, (see at least: col. 2, lines 33-37, after the idle time interval is predicted by the analysis device (i.e., ML), the software update is then started at the beginning of the idle time interval) However, Thurimella fails to teach or suggest, either alone or in combination with the other cited references, that the natural language processing is associated with identifying an idle period, wherein the idle period is associated with inactivity during the surgical procedure; generating an output, wherein the output indicates an idle start time and an idle end time; and refining the prediction result based on the identified idle period Regarding claim 9, claim 9 is in condition for allowance based at least on its dependency on claim 8. Regarding claim 18, claim 18 recites substantially similar limitations as set forth in claim 7. As such, claim 18 is in condition for allowance, for similar reasons. Regarding claim 19, claim 19 recites substantially similar limitations as set forth in claim 9. As such, claim 19 is in condition for allowance, for similar reasons. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMARA ABDI whose telephone number is (571)272-0273. The examiner can normally be reached 9:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMARA ABDI/Primary Examiner, Art Unit 2668 09/10/2026
Read full office action

Prosecution Timeline

Oct 13, 2023
Application Filed
Nov 14, 2025
Non-Final Rejection mailed — §103
Feb 17, 2026
Response Filed
Apr 24, 2026
Final Rejection mailed — §103
Jul 24, 2026
Request for Continued Examination
Jul 28, 2026
Response after Non-Final Action
Sep 16, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749345
MONITORING AND ANALYZING BODY LANGUAGE WITH MACHINE LEARNING, USING ARTIFICIAL INTELLIGENCE SYSTEMS FOR IMPROVING INTERACTION BETWEEN HUMANS, AND HUMANS AND ROBOTS
2y 7m to grant Granted Sep 29, 2026
Patent 12749346
FINGER ENCODING BASED POSE CLASSIFICATION
2y 6m to grant Granted Sep 29, 2026
Patent 12749307
CAMERA APPARATUS AND METHOD OF ENHANCED FOLIAGE DETECTION
2y 2m to grant Granted Sep 29, 2026
Patent 12743792
SYSTEMS AND METHODS FOR IMAGE PROCESSING
2y 11m to grant Granted Sep 22, 2026
Patent 12728441
ROBOTIC REPAIR CONTROL SYSTEMS AND METHODS
3y 6m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
76%
With Interview (-7.3%)
2y 6m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 840 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month