Prosecution Insights
Last updated: September 17, 2026
Application No. 19/072,002

METHOD AND SYSTEM TO HIGHLIGHT VIDEO SEGMENTS IN A VIDEO STREAM

Final Rejection §103
Filed
Mar 06, 2025
Priority
Mar 16, 2022 — continuation of 12/273,604
Examiner
JONES, HEATHER RAE
Art Unit
2481
Tech Center
2400 — Computer Networks
Assignee
Istreamplanet Co. LLC
OA Round
2 (Final)
69%
Grant Probability
Favorable
3-4
OA Rounds
1y 10m
Est. Remaining
74%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
525 granted / 763 resolved
+10.8% vs TC avg
Moderate +6% lift
Without
With
+5.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
17 currently pending
Career history
788
Total Applications
across all art units

Statute-Specific Performance

§101
7.3%
-32.7% vs TC avg
§103
62.1%
+22.1% vs TC avg
§102
20.3%
-19.7% vs TC avg
§112
1.3%
-38.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 763 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claims 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 8-11, and 15-18 are rejected under 35 U.S.C. 103 as being unpatentable over Dareddy et al. (U.S. Patent Application Publication 2020/0186897) in view of Quennesson (U.S. Patent Application Publication 2018/0025078) in view of Jeong et al. (U.S. Patent Application Publication 2007/0294716). Regarding claim 1, Dareddy et al. discloses a computer-implemented method executed by an electronic device in a video streaming platform (Fig. 1), the computer-implemented method comprising: selecting, by one or more processors, a set of video segments for training one or more machine-learning models (Fig. 6; paragraph [0111] – at step S602, previously generated video game data is received – the previously generated video game data may comprise at least video and audio signals generated during previous playing of the video game; paragraph [0113] – in some examples, the video and audio may be received together as a video file and need to be separated into respective signals); extracting, by the one or more processors, via one or more neural networks, at least one feature for each video segment in the set of video segments (Fig. 6; paragraph [0114] – at step S604, the method comprises generating feature representations of the signals in the previously generated video game data; paragraph [0115] – in some examples, generating the feature representations of the previously generated video signals may involve inputting at least some of the RGB or YUV video frames in the previously generated video signal into a pre-trained model, such as e.g. DenseNet, ResNet, MobileNet, etc.); applying, by the one or more processors, one or more machine-learning models to the set of video segments to determine at least one context corresponding to the at least one feature from each video segment of the set of video segments (Fig. 6; paragraph [0117] – at step S606, the feature representations generated for each signal are clustered into respective clusters using unsupervised learning – each cluster corresponds to content (for a given signals) that has been identified as being similar in some with respect with other content in that cluster; paragraph [0118] – in some examples, this involves using k-means clustering or mini batch k-means clustering to sort the feature representations generated for each signal into respective clusters; paragraph [0126] – the step of clustering the RGB or YUV video frames and audio frames may also involve generating model files to enable the feature representations in each cluster to be labelled; paragraph [0127] – the training method further comprises a step S608 of manually labelling each cluster (for a given signal) with a label indicating an event associated with that cluster); training, by the one or more processors, the one or more machine-learning models based on the at least one context corresponding to the at least one feature from each video segment of the set of video segments (Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example), wherein the training includes generating at least one confidence level for each of the one or more machine-learning models (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data); selecting, by the one or more processors, at least one of the one or more machine-learning models with a predetermined accuracy (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data); and storing, by the one or more processors, the at least one selected machine- learning model as a recommended machine-learning model (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; once the model is ready for deployment, the system will remember to utilize it). However, Dareddy et al. fails to explicitly disclose applying, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments; and selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold. Referring to the Quennesson reference, Quennesson discloses a computer-implemented method executed by an electronic device in a video streaming platform, the computer implemented method comprising: applying, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments (paragraph [0041] – the video classifier 179 may include a machine learning algorithm trained to classify a small portion (also referred to as video segment) of the video broadcast stream, e.g., a few seconds, into one or more classifications – a classification may be a characteristic of the content of the video segment – examples of possible classifications include, but are not limited to, selfie, noise, computer screen, driving, talk, stage, sports, food, outdoors, animals, humans, texture, graphics, interior, nature, construction, unplugged, art, NSFW, etc. – in some examples, classifications may include sub-classes such as Sports.Skiing, Sports.Baseball, Sports.Golf, etc. – in some examples, the classifications may be hierarchical – in some examples, the classifications may be domain specific – the classifications can be manually curated but may be refined with the help of the system 100; paragraph [0042] – the classifier 179 may be a LSTM (“Long Short-term memory”) classifier that has undergone supervised training – in some examples, the classifier 179 may be one classifier able to classify a video segment of the video broadcast stream into one or several classifications – in other examples, the classifier 179 may represent a group of classifiers, each classifier trained to classify a video segment of the video broadcast stream as either in or not in a specific classification – in either case, the classifier 179 may provide a confidence score for each class for a video segment of the video broadcast stream, the confidence score representing how confident the model is that the video segment is correctly classified as in or not in the classification – the classifications and confidence scores may be stored for each video segment of the video broadcast stream, e.g., as part of broadcast metadata 166; paragraph [0044] – the video highlight creator 180 may create video highlights 181 for a video broadcast stream based on the output of the video classifier 179 and class proportion data 162 – for example, the video highlight creator 180 may use classification data (e.g., the classifications and confidence scores) from the video classifier 179 in view of the class proportion data 162 to determine which part of the video broadcast stream is interesting or uncommon – the class proportion data 162 may be generated automatically). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had applied, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments as disclosed by Quennesson in the method disclosed by Dareddy et al. in order to help determine which segments are considered interesting enough to be included in the highlight video segments. However, Dareddy et al. in view of Quennesson still fails to disclose selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold. Referring to the Jeong et al. reference, Jeong et al. discloses a computer-implemented method executed by an electronic device in a video streaming platform (paragraph [0017] – the input data stream may be a sports video stream), the computer-implemented method comprising: training, by the one or more processors, the one or more machine-learning models based on the at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the training includes generating at least one confidence level for each of the one or more machine-learning models (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); and storing, by the one or more processors, the at least one selected machine-learning model as a recommended machine-learning model (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had selected at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold as disclosed by Jeong et al. in the method disclosed by Dareddy et al. in view of Quennesson in order to provide a method for detecting highlights found within the video content with the highest accuracy possible. Regarding claim 2, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 1 including that wherein training the one or more machine-learning models further comprises, for each machine-learning model of the one or more machine-learning models: generating, by the one or more processors, a set of rankings for the at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein each ranking of the set of rankings includes the at least one confidence level (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level); comparing, by the one or more processors, the set of rankings and the at least one feature for each video segment in the set of video segments (Dareddy et al.: Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; Jeong et al.: paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); and based on the comparing, determining, by the one or more processors, an adjustment to one or more parameters values of at least one machine-learning model of the one or more machine-learning models to improve a context determination (Dareddy et al.: Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; Jeong et al.: paragraph [0085] – in operation S280, the online model may be updated, since the detected event data may be a sample meeting the online model standard – in this instance, the online model may be updated by using a weighted average value of a current online model and a detected event sample – further, the online model may be updated by using a median value of the current online model and the detected event sample – still further, the online model may be updated by using a Gaussian mixture of the current online model and the detected event sample, noting that the alternative embodiments are equally available; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Regarding claim 3, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 1 including that the computer-implemented method further comprising: analyzing, by the one or more processors, the at least one context corresponding to the at least one feature from each video segment of the set of video segments to determine the one or more machine-learning models (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0021] and [0022] – training for video data; paragraph [0023] – training for audio data; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level); and selecting, by the one or more processors, the one or more machine-learning models for training (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0021] and [0022] – training for video data; paragraph [0023] – training for audio data; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Regarding claim 4, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 1 including that wherein extracting the at least one feature for each video segment in the set of video segments is based on applying a visual filter and an acoustic filter to the video segment in the set of video segments (Dareddy et al.: paragraphs [0122] and [0124] – video and audio filters; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: paragraph [0021] – the training of the online model may include segmenting video data of the detected event onto frames according to minimum units when the detected event detected by the offline model is the video data, selectively assigning and generating clusters for the online model by analyzing and generating clusters for the online model by analyzing the minimum units, and selecting a cluster for generating and to be implemented model, from the selectively assigned and generated clusters, and generating the online model with at least the selected cluster; paragraph [0022] – the selectively assigning and the generating includes calculating a difference value between at least one preexisting cluster and a newly calculated cluster based upon the detected event, assigning data of the newly calculated cluster to the at least one preexisting cluster when the difference value meets a difference threshold, and generating at least one new cluster for the data of the newly calculated cluster at least when the difference value does not meet the difference threshold or no preexisting cluster exists; paragraph [0023] – the training may include calculating an audio energy value of an audio frame, when the detected event detected by the offline model is the audio frame, calculating an average energy by using a preexisting calculated audio energy value and the calculated audio energy value for the detected event, and extracting a corresponding recording level, and updating the online model with the extracted recording level; filtering during the segmentation process will refine the segmentation results). Regarding claim 8, Dareddy et al. discloses a non-transitory machine-readable storage medium that provides instructions that, if executed by a processor, will cause the processor to perform operations comprising: selecting, by one or more processors, a set of video segments for training one or more machine-learning models (Fig. 6; paragraph [0111] – at step S602, previously generated video game data is received – the previously generated video game data may comprise at least video and audio signals generated during previous playing of the video game; paragraph [0113] – in some examples, the video and audio may be received together as a video file and need to be separated into respective signals); extracting, by the one or more processors, via one or more neural networks, at least one feature for each video segment in the set of video segments (Fig. 6; paragraph [0114] – at step S604, the method comprises generating feature representations of the signals in the previously generated video game data; paragraph [0115] – in some examples, generating the feature representations of the previously generated video signals may involve inputting at least some of the RGB or YUV video frames in the previously generated video signal into a pre-trained model, such as e.g. DenseNet, ResNet, MobileNet, etc.); applying, by the one or more processors, one or more machine-learning models to the set of video segments to determine at least one context corresponding to the at least one feature from each video segment of the set of video segments (Fig. 6; paragraph [0117] – at step S606, the feature representations generated for each signal are clustered into respective clusters using unsupervised learning – each cluster corresponds to content (for a given signals) that has been identified as being similar in some with respect with other content in that cluster; paragraph [0118] – in some examples, this involves using k-means clustering or mini batch k-means clustering to sort the feature representations generated for each signal into respective clusters; paragraph [0126] – the step of clustering the RGB or YUV video frames and audio frames may also involve generating model files to enable the feature representations in each cluster to be labelled; paragraph [0127] – the training method further comprises a step S608 of manually labelling each cluster (for a given signal) with a label indicating an event associated with that cluster); training, by the one or more processors, the one or more machine-learning models based on the at least one context corresponding to the at least one feature from each video segment of the set of video segments (Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example), wherein the training includes generating at least one confidence level for each of the one or more machine-learning models (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data); selecting, by the one or more processors, at least one of the one or more machine-learning models with a predetermined accuracy (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data); and storing, by the one or more processors, the at least one selected machine- learning model as a recommended machine-learning model (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; once the model is ready for deployment, the system will remember to utilize it). However, Dareddy et al. fails to explicitly disclose applying, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or a source type of a video segment corresponding to the set of video segments; and selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold. Referring to the Quennesson reference, Quennesson discloses a computer-implemented method executed by an electronic device in a video streaming platform, the computer implemented method comprising: applying, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments (paragraph [0041] – the video classifier 179 may include a machine learning algorithm trained to classify a small portion (also referred to as video segment) of the video broadcast stream, e.g., a few seconds, into one or more classifications – a classification may be a characteristic of the content of the video segment – examples of possible classifications include, but are not limited to, selfie, noise, computer screen, driving, talk, stage, sports, food, outdoors, animals, humans, texture, graphics, interior, nature, construction, unplugged, art, NSFW, etc. – in some examples, classifications may include sub-classes such as Sports.Skiing, Sports.Baseball, Sports.Golf, etc. – in some examples, the classifications may be hierarchical – in some examples, the classifications may be domain specific – the classifications can be manually curated but may be refined with the help of the system 100; paragraph [0042] – the classifier 179 may be a LSTM (“Long Short-term memory”) classifier that has undergone supervised training – in some examples, the classifier 179 may be one classifier able to classify a video segment of the video broadcast stream into one or several classifications – in other examples, the classifier 179 may represent a group of classifiers, each classifier trained to classify a video segment of the video broadcast stream as either in or not in a specific classification – in either case, the classifier 179 may provide a confidence score for each class for a video segment of the video broadcast stream, the confidence score representing how confident the model is that the video segment is correctly classified as in or not in the classification – the classifications and confidence scores may be stored for each video segment of the video broadcast stream, e.g., as part of broadcast metadata 166; paragraph [0044] – the video highlight creator 180 may create video highlights 181 for a video broadcast stream based on the output of the video classifier 179 and class proportion data 162 – for example, the video highlight creator 180 may use classification data (e.g., the classifications and confidence scores) from the video classifier 179 in view of the class proportion data 162 to determine which part of the video broadcast stream is interesting or uncommon – the class proportion data 162 may be generated automatically). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had applied, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments as disclosed by Quennesson in the method disclosed by Dareddy et al. in order to help determine which segments are considered interesting enough to be included in the highlight video segments. However, Dareddy et al. in view of Quennesson still fails to disclose selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold. Referring to the Jeong et al. reference, Jeong et al. discloses a computer-implemented method executed by an electronic device in a video streaming platform (paragraph [0017] – the input data stream may be a sports video stream), the computer-implemented method comprising: training, by the one or more processors, the one or more machine-learning models based on the at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the training includes generating at least one confidence level for each of the one or more machine-learning models (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); and storing, by the one or more processors, the at least one selected machine-learning model as a recommended machine-learning model (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had selected at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold as disclosed by Jeong et al. in the method disclosed by Dareddy et al. in view of Quennesson in order to provide a method for detecting highlights found within the video content with the highest accuracy possible. Regarding claim 9, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 8 including that wherein training the one or more machine-learning models further comprises, for each machine- learning model of the one or more machine-learning models: generating, by the one or more processors, a set of rankings for the at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein each ranking of the set of rankings includes the at least one confidence level (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level); comparing, by the one or more processors, the set of rankings and the at least one feature for each video segment in the set of video segments (Dareddy et al.: Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; Jeong et al.: paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); and based on the comparing, determining, by the one or more processors, an adjustment to one or more parameters values of at least one machine-learning model of the one or more machine-learning models to improve a context determination to improve a context determination (Dareddy et al.: Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; Jeong et al.: paragraph [0085] – in operation S280, the online model may be updated, since the detected event data may be a sample meeting the online model standard – in this instance, the online model may be updated by using a weighted average value of a current online model and a detected event sample – further, the online model may be updated by using a median value of the current online model and the detected event sample – still further, the online model may be updated by using a Gaussian mixture of the current online model and the detected event sample, noting that the alternative embodiments are equally available; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Regarding claim 10, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 8 including that the operations further comprising: analyzing, by the one or more processors, the at least one context corresponding to the at least one feature from each video segment of the set of video segments to determine the one or more machine-learning models (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0021] and [0022] – training for video data; paragraph [0023] – training for audio data; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level); and selecting, by the one or more processors, the one or more machine-learning models for training (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0021] and [0022] – training for video data; paragraph [0023] – training for audio data; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Regarding claim 11, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 8 including that wherein extracting the at least one feature for each video segment in the set of video segments is based on applying a visual filter and an acoustic filter to the video segment in the set of video segments (Dareddy et al.: paragraphs [0122] and [0124] – video and audio filters; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: paragraph [0021] – the training of the online model may include segmenting video data of the detected event onto frames according to minimum units when the detected event detected by the offline model is the video data, selectively assigning and generating clusters for the online model by analyzing and generating clusters for the online model by analyzing the minimum units, and selecting a cluster for generating and to be implemented model, from the selectively assigned and generated clusters, and generating the online model with at least the selected cluster; paragraph [0022] – the selectively assigning and the generating includes calculating a difference value between at least one preexisting cluster and a newly calculated cluster based upon the detected event, assigning data of the newly calculated cluster to the at least one preexisting cluster when the difference value meets a difference threshold, and generating at least one new cluster for the data of the newly calculated cluster at least when the difference value does not meet the difference threshold or no preexisting cluster exists; paragraph [0023] – the training may include calculating an audio energy value of an audio frame, when the detected event detected by the offline model is the audio frame, calculating an average energy by using a preexisting calculated audio energy value and the calculated audio energy value for the detected event, and extracting a corresponding recording level, and updating the online model with the extracted recording level; filtering during the segmentation process will refine the segmentation results). Regarding claim 15, Dareddy et al. discloses a computer system, comprising: a memory having processor-readable instructions stored therein (Figs. 1 and 5); and one or more processors configured to access the memory and execute the processor-readable instructions (Figs. 1 and 5), which when executed by the one or more processors configures the one or more processors to perform a plurality of functions, including functions for: selecting, by the one or more processors, a set of video segments for training one or more machine-learning models (Fig. 6; paragraph [0111] – at step S602, previously generated video game data is received – the previously generated video game data may comprise at least video and audio signals generated during previous playing of the video game; paragraph [0113] – in some examples, the video and audio may be received together as a video file and need to be separated into respective signals); extracting, by the one or more processors, via one or more neural networks, at least one feature for each video segment in the set of video segments (Fig. 6; paragraph [0114] – at step S604, the method comprises generating feature representations of the signals in the previously generated video game data; paragraph [0115] – in some examples, generating the feature representations of the previously generated video signals may involve inputting at least some of the RGB or YUV video frames in the previously generated video signal into a pre-trained model, such as e.g. DenseNet, ResNet, MobileNet, etc.); applying, by the one or more processors, one or more machine-learning models to the set of video segments to determine at least one context corresponding to the at least one feature from each video segment of the set of video segments (Fig. 6; paragraph [0117] – at step S606, the feature representations generated for each signal are clustered into respective clusters using unsupervised learning – each cluster corresponds to content (for a given signals) that has been identified as being similar in some with respect with other content in that cluster; paragraph [0118] – in some examples, this involves using k-means clustering or mini batch k-means clustering to sort the feature representations generated for each signal into respective clusters; paragraph [0126] – the step of clustering the RGB or YUV video frames and audio frames may also involve generating model files to enable the feature representations in each cluster to be labelled; paragraph [0127] – the training method further comprises a step S608 of manually labelling each cluster (for a given signal) with a label indicating an event associated with that cluster); training, by the one or more processors, the one or more machine-learning models based on the at least one context corresponding to the at least one feature from each video segment of the set of video segments (Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example), wherein the training includes generating at least one confidence level for each of the one or more machine-learning models (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data); selecting, by the one or more processors, at least one of the one or more machine-learning models with a predetermined accuracy (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data); and storing, by the one or more processors, the at least one selected machine- learning model as a recommended machine-learning model (Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; once the model is ready for deployment, the system will remember to utilize it). However, Dareddy et al. fails to explicitly disclose applying, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments; and selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold. Referring to the Quennesson reference, Quennesson discloses a computer system, comprising: applying, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments (paragraph [0041] – the video classifier 179 may include a machine learning algorithm trained to classify a small portion (also referred to as video segment) of the video broadcast stream, e.g., a few seconds, into one or more classifications – a classification may be a characteristic of the content of the video segment – examples of possible classifications include, but are not limited to, selfie, noise, computer screen, driving, talk, stage, sports, food, outdoors, animals, humans, texture, graphics, interior, nature, construction, unplugged, art, NSFW, etc. – in some examples, classifications may include sub-classes such as Sports.Skiing, Sports.Baseball, Sports.Golf, etc. – in some examples, the classifications may be hierarchical – in some examples, the classifications may be domain specific – the classifications can be manually curated but may be refined with the help of the system 100; paragraph [0042] – the classifier 179 may be a LSTM (“Long Short-term memory”) classifier that has undergone supervised training – in some examples, the classifier 179 may be one classifier able to classify a video segment of the video broadcast stream into one or several classifications – in other examples, the classifier 179 may represent a group of classifiers, each classifier trained to classify a video segment of the video broadcast stream as either in or not in a specific classification – in either case, the classifier 179 may provide a confidence score for each class for a video segment of the video broadcast stream, the confidence score representing how confident the model is that the video segment is correctly classified as in or not in the classification – the classifications and confidence scores may be stored for each video segment of the video broadcast stream, e.g., as part of broadcast metadata 166; paragraph [0044] – the video highlight creator 180 may create video highlights 181 for a video broadcast stream based on the output of the video classifier 179 and class proportion data 162 – for example, the video highlight creator 180 may use classification data (e.g., the classifications and confidence scores) from the video classifier 179 in view of the class proportion data 162 to determine which part of the video broadcast stream is interesting or uncommon – the class proportion data 162 may be generated automatically). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had applied, by the one or more processors, the one or more machine-learning models to the set of video segments to automatically determine at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the at least one context identifies an event type or source type of a video segment corresponding to the set of video segments as disclosed by Quennesson in the system disclosed by Dareddy et al. in order to help determine which segments are considered interesting enough to be included in the highlight video segments. However, Dareddy et al. in view of Quennesson still fails to disclose selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold. Referring to the Jeong et al. reference, Jeong et al. discloses a computer system, comprising: training, by the one or more processors, the one or more machine-learning models based on the at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein the training includes generating at least one confidence level for each of the one or more machine-learning models (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); selecting, by the one or more processors, based on the at least one context corresponding to the at least one feature from each video segment corresponding to the set of video segments, at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); and storing, by the one or more processors, the at least one selected machine-learning model as a recommended machine-learning model (Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had selected at least one of the one or more machine-learning models with the at least one confidence level above a confidence level threshold as disclosed by Jeong et al. in the system disclosed by Dareddy et al. in view of Quennesson in order to provide a method for detecting highlights found within the video content with the highest accuracy possible. Regarding claim 16, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 15 including that wherein training the one or more machine- learning models further comprises, for each machine-learning model of the one or more machine-learning models: generating, by the one or more processors, a set of rankings for the at least one context corresponding to the at least one feature from each video segment of the set of video segments, wherein each ranking of the set of rankings includes the at least one confidence level (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level); comparing, by the one or more processors, the set of rankings and the at least one feature for each video segment in the set of video segments (Dareddy et al.: Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; Jeong et al.: paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream); and based on the comparing, determining, by the one or more processors, an adjustment to one or more parameters values of at least one machine-learning model of the one or more machine-learning models to improve a context determination (Dareddy et al.: Fig. 6; paragraph [0138] – each model may be trained in iterative manner – for example, after each epoch (pass through a training set), a final version of the model may be saved if it is determined that model performed better than it did for a previous iteration – in some cases, a given model may be determined as being ready for deployment when it produces sufficiently accurate results for an unseen input set of training data; Jeong et al.: paragraph [0085] – in operation S280, the online model may be updated, since the detected event data may be a sample meeting the online model standard – in this instance, the online model may be updated by using a weighted average value of a current online model and a detected event sample – further, the online model may be updated by using a median value of the current online model and the detected event sample – still further, the online model may be updated by using a Gaussian mixture of the current online model and the detected event sample, noting that the alternative embodiments are equally available; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Regarding claim 17, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 15 including that the computer system further comprises: analyzing, by the one or more processors, the at least one context corresponding to the at least one feature from each video segment of the set of video segments to determine the one or more machine-learning models (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0021] and [0022] – training for video data; paragraph [0023] – training for audio data; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level); and selecting, by the one or more processors, the one or more machine-learning models for training (Dareddy et al.: Fig. 6; paragraph [0135] – the method comprises a step S610 of inputting the feature representations of each signal and the corresponding labels, into respective machine learning models – the machine learning models are trained using supervised deep learning – the models are said to have been trained using supervised learning because the training data was at least partially generated using unsupervised learning; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: Fig. 7; paragraph [0021] and [0022] – training for video data; paragraph [0023] – training for audio data; paragraph [0107] – referring to Fig. 7, the apparatus for detecting a real time event 700 may include a confidence test unit 710, a first event detection unit 720, an online model training unit 730, and a second event detection unit 740, for example; paragraph [0108] – the confidence test unit 710 may test a confidence for an online model, as calculated based on a sports video stream – specifically, the confidence test unit 710 may calculate the confidence for the online model, as calculated based on the sports video stream, compare the calculated confidence for the online model with a threshold, and test the confidence for the online model; paragraph [0109] - the first event detection 720 may detect the event by using an offline model for the sports video stream, when the confidence for the online model is lower than the threshold; paragraph [0110] – the online model training unit 730 may further train the online model based on the event detected by the offline model; paragraph [0111] – when the event detected by using the offline model is video data, the online model training unit 730 may segment the video data into minimum units, e.g., frames or pixels, assign or generate clusters by analyzing the segmented minimum units, selects a cluster, that may be used to generate the implemented model, from the clusters and generates/updates the online model based upon detected events; paragraph [0112] – when the offline model detected event is audio data, the online model may calculate an audio energy value and the currently calculated audio energy value, extract a recording level, and update the online model by using the extracted recording level; paragraph [0113] – accordingly, the confidence of the online model can be improved by the updating of the online model after the training of the online model through the online model training unit 730; paragraph [0114] – once the confidence of the online model is sufficiently high, e.g., greater than the threshold, the second event detection unit 740 may be used for detecting events based on the online model for the sports video stream). Regarding claim 18, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 15 including that wherein extracting the at least one feature for each video segment in the set of video segments is based on applying a visual filter and an acoustic filter to the video segment in the set of video segments (Dareddy et al.: paragraphs [0122] and [0124] – video and audio filters; paragraph [0136] – the feature representations of the video frames and their corresponding labels may be input to a multi-class classification algorithm (corresponding to the video machine learning model) – the video machine learning model may then be trained to determine a relationship between feature representations of RGB or YUV frames and descriptive labels associated with those frames – the video machine learning model may comprise a neural network, such as a convolutional or recurrent neural network, for example; paragraph [0137] – the feature representations of the audio frames and the corresponding labels may be input to a corresponding audio machine learning model, or a binary classification model such as a Gradient Boosting Trees, Random Forests, Support Vector Machine algorithm, for example; Jeong et al.: paragraph [0021] – the training of the online model may include segmenting video data of the detected event onto frames according to minimum units when the detected event detected by the offline model is the video data, selectively assigning and generating clusters for the online model by analyzing and generating clusters for the online model by analyzing the minimum units, and selecting a cluster for generating and to be implemented model, from the selectively assigned and generated clusters, and generating the online model with at least the selected cluster; paragraph [0022] – the selectively assigning and the generating includes calculating a difference value between at least one preexisting cluster and a newly calculated cluster based upon the detected event, assigning data of the newly calculated cluster to the at least one preexisting cluster when the difference value meets a difference threshold, and generating at least one new cluster for the data of the newly calculated cluster at least when the difference value does not meet the difference threshold or no preexisting cluster exists; paragraph [0023] – the training may include calculating an audio energy value of an audio frame, when the detected event detected by the offline model is the audio frame, calculating an average energy by using a preexisting calculated audio energy value and the calculated audio energy value for the detected event, and extracting a corresponding recording level, and updating the online model with the extracted recording level; filtering during the segmentation process will refine the segmentation results). Claims 5-7, 12-14, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Dareddy et al. in view of Quennesson in view of Jeong et al. as applied to claims 1, 8, and 15 above, and further in view of Matsuoka et al. (U.S. Patent Application Publication 2023/0060753). Regarding claim 5, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 1, but fails to disclose that wherein extracting the at least one feature for each video segment in the set of video segments is based on a language transformation that converts audio data in the video segment into text. Referring to the Matsuoka et al. reference, Matsuoka et al. discloses a computer-implemented method executed by an electronic device, the computer-implemented method comprises: wherein extracting the at least one feature for each video segment in the set of video segments is based on a language transformation that converts audio data in the video segment into text (paragraph [0163] – the natural language processor may include one or more machine-learning models configured to process natural language input in any particular form or format – for example, one such machine-learning model may be trained to convert audio into the equivalent text; paragraph [0218] – the one or more machine-learning models may include one or more natural language processors that can convert audio and/or video to text, parse text to derive a semantic meaning such as an interest or intent). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had used a language transformation that converts audio data in the video segment into text to help extract at least one feature for each video segment as disclosed by Matsuoka et al. in the method disclosed by Dareddy et al. in view of Quennesson in view of Jeong et al. in order to help analyze the video by deriving a semantic meaning of the spoken words. Regarding claim 6, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 1, but fails to disclose that wherein the set of video segments includes data from at least one contemporaneous secondary source. Referring to the Matsuoka et al. reference, Matsuoka et al. discloses a computer-implemented method executed by an electronic device, the computer-implemented method comprises: wherein the set of video segments includes data from at least one contemporaneous secondary source (paragraph [0223] – a next set of machine-learning models in the hierarchy may include machine-learning models configured to process application data and/or data from third-party services, such as, but not limited to calendar, email, direct messaging services (e.g., SMS or the like), social media services, music streaming services, video streaming services, to-do lists, shopping lists, and/or any other application executing on a device associated with the member that may be useable). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had the set of video segments include data from at least one contemporaneous secondary source as disclosed by Matsuoka et al. in the method disclosed by Dareddy et al. in view of Quennesson in view of Jeong et al. in order to include more information that will help determine the details of the video content. Regarding claim 7, Dareddy et al. in view of Quennesson in view of Jeong et al. in view of Matsuoka et al. discloses all of the limitations as previously discussed with respect to claims 1 and 6 including that wherein the at least one contemporaneous secondary source includes social media data and broadcast message data (Matsuoka et al.: paragraph [0223] – a next set of machine-learning models in the hierarchy may include machine-learning models configured to process application data and/or data from third-party services, such as, but not limited to calendar, email, direct messaging services (e.g., SMS or the like), social media services, music streaming services, video streaming services, to-do lists, shopping lists, and/or any other application executing on a device associated with the member that may be useable). Regarding claim 12, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 8, but fails to disclose that wherein extracting the at least one feature for each video segment in the set of video segments is based on a language transformation that converts audio data in the video segment into text. Referring to the Matsuoka et al. reference, Matsuoka et al. discloses a computer-implemented method executed by an electronic device, the computer-implemented method comprises: wherein extracting the at least one feature for each video segment in the set of video segments is based on a language transformation that converts audio data in the video segment into text (paragraph [0163] – the natural language processor may include one or more machine-learning models configured to process natural language input in any particular form or format – for example, one such machine-learning model may be trained to convert audio into the equivalent text; paragraph [0218] – the one or more machine-learning models may include one or more natural language processors that can convert audio and/or video to text, parse text to derive a semantic meaning such as an interest or intent). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had used a language transformation that converts audio data in the video segment into text to help extract at least one feature for each video segment as disclosed by Matsuoka et al. in the method disclosed by Dareddy et al. in view of Quennesson in view of Jeong et al. in order to help analyze the video by deriving a semantic meaning of the spoken words. Regarding claim 13, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 8, but fails to disclose that wherein the set of video segments includes data from at least one contemporaneous secondary source. Referring to the Matsuoka et al. reference, Matsuoka et al. discloses a computer-implemented method executed by an electronic device, the computer-implemented method comprises: wherein the set of video segments includes data from at least one contemporaneous secondary source (paragraph [0223] – a next set of machine-learning models in the hierarchy may include machine-learning models configured to process application data and/or data from third-party services, such as, but not limited to calendar, email, direct messaging services (e.g., SMS or the like), social media services, music streaming services, video streaming services, to-do lists, shopping lists, and/or any other application executing on a device associated with the member that may be useable). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had the set of video segments include data from at least one contemporaneous secondary source as disclosed by Matsuoka et al. in the method disclosed by Dareddy et al. in view of Quennesson in view of Jeong et al. in order to include more information that will help determine the details of the video content. Regarding claim 14, Dareddy et al. in view of Quennesson in view of Jeong et al. in view of Matsuoka et al. discloses all of the limitations as previously discussed with respect to claims 8 and 13 including that wherein the at least one contemporaneous secondary source includes social media data and broadcast message data (Matsuoka et al.: paragraph [0223] – a next set of machine-learning models in the hierarchy may include machine-learning models configured to process application data and/or data from third-party services, such as, but not limited to calendar, email, direct messaging services (e.g., SMS or the like), social media services, music streaming services, video streaming services, to-do lists, shopping lists, and/or any other application executing on a device associated with the member that may be useable). Regarding claim 19, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 15, but fails to disclose that wherein extracting the at least one feature for each video segment in the set of video segments is based on a language transformation that converts audio data in the video segment into text. Referring to the Matsuoka et al. reference, Matsuoka et al. discloses a computer system, comprising: wherein extracting the at least one feature for each video segment in the set of video segments is based on a language transformation that converts audio data in the video segment into text (paragraph [0163] – the natural language processor may include one or more machine-learning models configured to process natural language input in any particular form or format – for example, one such machine-learning model may be trained to convert audio into the equivalent text; paragraph [0218] – the one or more machine-learning models may include one or more natural language processors that can convert audio and/or video to text, parse text to derive a semantic meaning such as an interest or intent). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had used a language transformation that converts audio data in the video segment into text to help extract at least one feature for each video segment as disclosed by Matsuoka et al. in the system disclosed by Dareddy et al. in view of Quennesson in view of Jeong et al. in order to help analyze the video by deriving a semantic meaning of the spoken words. Regarding claim 20, Dareddy et al. in view of Quennesson in view of Jeong et al. discloses all of the limitations as previously discussed with respect to claim 15, but fails to disclose that wherein the set of video segments includes data from at least one contemporaneous secondary source. Referring to the Matsuoka et al. reference, Matsuoka et al. discloses a computer system, comprising: wherein the set of video segments includes data from at least one contemporaneous secondary source (paragraph [0223] – a next set of machine-learning models in the hierarchy may include machine-learning models configured to process application data and/or data from third-party services, such as, but not limited to calendar, email, direct messaging services (e.g., SMS or the like), social media services, music streaming services, video streaming services, to-do lists, shopping lists, and/or any other application executing on a device associated with the member that may be useable). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to have had the set of video segments include data from at least one contemporaneous secondary source as disclosed by Matsuoka et al. in the system disclosed by Dareddy et al. in view of Quennesson in view of Jeong et al. in order to include more information that will help determine the details of the video content. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HEATHER R JONES whose telephone number is (571)272-7368. The examiner can normally be reached Mon. - Fri.: 9:00am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William Vaughn can be reached at (571)272-3922. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HEATHER R JONES/Primary Examiner, Art Unit 2481 September 4, 2026
Read full office action

Prosecution Timeline

Mar 06, 2025
Application Filed
Mar 27, 2026
Non-Final Rejection mailed — §103
Jun 26, 2026
Response Filed
Sep 10, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12730504
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND OPERATION MICROSCOPE SYSTEM
3y 1m to grant Granted Sep 08, 2026
Patent 12726680
SYSTEMS AND METHODS TO PROVIDE ADAPTIVE PLAY SETTINGS
5y 5m to grant Granted Sep 01, 2026
Patent 12717838
UTILIZING TEXT-BASED INPUT TO DYNAMIC SELECT AND PRESENT SPECIFIC PORTIONS OF AUDIOVISUAL CONTENT TO USER
2y 5m to grant Granted Aug 25, 2026
Patent 12713092
4K CONTENT MEMORY FRAGMENTATION VIEWING COUNTERMEASURES
2y 5m to grant Granted Aug 18, 2026
Patent 12711769
LANGUAGE INSTRUCTED TEMPORAL LOCALIZATION IN VIDEOS
2y 0m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
69%
Grant Probability
74%
With Interview (+5.5%)
3y 5m (~1y 10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 763 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month