Prosecution Insights
Last updated: October 02, 2026
Application No. 18/219,652

INTELLIGENT VIDEO AGGREGATION AND CUSTOMIZATION

Non-Final OA §103§112
Filed
Jul 08, 2023
Examiner
STEVENS, ROBERT
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
430 granted / 529 resolved
+21.3% vs TC avg
Moderate +12% lift
Without
With
+11.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
8 currently pending
Career history
544
Total Applications
across all art units

Statute-Specific Performance

§101
22.7%
-17.3% vs TC avg
§103
46.2%
+6.2% vs TC avg
§102
8.0%
-32.0% vs TC avg
§112
18.6%
-21.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 529 resolved cases

Office Action

§103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 20 is rejected under 35 U.S.C. § 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Regarding dependent claim 20: Lines 3-4: It is unclear what is meant by “determine a recall confidence”. Both the specification and the claims merely use this terminology, but provide no guidance as to how the metes and bounds of such terminology is to be established. As such, this confidence is essentially an arbitrary value – i.e., the language means something different to each reader/implementer of the claim. Therefore, the scope of the claim is ambiguous. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 4, 9, 12 and 17 are rejected under 35 U.S.C. §103 as being unpatentable over Croitoru et al (US Patent Application Publication No. 2024/0355119, hereafter referred to as “Croitoru”) in view of Chhaya et al (US Patent Application Publication No. 2023/0290146, hereafter referred to as “Chhaya”). Regarding independent claim 1: Croitoru teaches An apparatus comprising: a processor configured to: extract a snippet of video from among a plurality of snippets of video within the video based on the search topic, (See Croitoru paragraphs 0044-0046 discussing an exemplary segmentation technique for identifying a video segment based upon identified video topics requested by a user via a query.) and output the snippet of video via a user interface of a user device. (See Croitoru Fig. 3 #360 and paragraph 0048 teaching the presentation of video segments to the user. See also, Fig. 1 #114, #116, #118 showing exemplary presentation devices.) Although Croitoru discusses issuing queries for videos (see paragraph 0043), Croitoru does not explicitly teach the remaining limitations as claimed. Chhaya, though, teaches identify a video that is related to a search topic based on one or more keywords associated with the video, (See Chhaya paragraph 0024 teaching an exemplary user keyword search to locate a video related to a topic of interest.) It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Chhaya for the benefit of Croitoru, because to do so provided a designer with options for implementing a system the increased the efficiency for viewing video content, as taught by Chhaya in paragraph 0023. These references were all applicable to the same field of endeavor, i.e., management and processing of multimedia data. Regarding claim 4: Croitoru teaches wherein the processor is further configured to convert audio from the video into text via a speech-to-text converter, and execute a natural language processing (NLP) model on the text to generate a transcript of the video. (See Croitoru paragraphs 0037-0038 and 0044 teaching the ability to generate a transcript/text reflecting audio associated with a video, using an automatic speech recognition component and a natural language processing model.) Claims 9, 12 and 17 are substantially similar to claims 1, 4 and 1, respectively, and therefore likewise rejected. Claims 2, 10 and 18 are rejected under 35 U.S.C. §103 as being unpatentable over Croitoru et al (US Patent Application Publication No. 2024/0355119, hereafter referred to as “Croitoru”) in view of Chhaya et al (US Patent Application Publication No. 2023/0290146, hereafter referred to as “Chhaya”) and Sezan et al (US Patent Application Publication No. 2005/0131727, hereafter referred to as “Sezan”). Regarding claim 2: Croitoru in view of Chhaya does not explicitly teach the remaining limitations as claimed. Sezan, though, teaches wherein the processor is further configured to customize the snippet of video based on previous browsing history of the user device prior to outputting the snippet of video via the user interface. (See Sezan Abstract teaching the use of browsing history in the management of user viewing preferences for media data.) It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Sezan for the benefit of Croitoru in view of Chhaya, because to do so provided a designer with options for implementing a system that facilitated the selection of multimedia programs that were of interest to a user, as taught by Sezan in paragraph 0045. These references were all applicable to the same field of endeavor, i.e., management and processing of multimedia data. Claim 10 is substantially similar to claim 2, and therefore likewise rejected. Claim 18 is substantially similar to claim 2, and therefore likewise rejected. Claims 3, 11 and 19 are rejected under 35 U.S.C. §103 as being unpatentable over Croitoru et al (US Patent Application Publication No. 2024/0355119, hereafter referred to as “Croitoru”) in view of Chhaya et al (US Patent Application Publication No. 2023/0290146, hereafter referred to as “Chhaya”) and McLaughlin et al (US Patent Application Publication No. 2013/0036201, hereafter referred to as “McLaughlin”). Regarding claim 3: Croitoru in view of Chhaya does not explicitly teach the remaining limitations as claimed. McLaughlin, though, teaches wherein the processor is further configured to play the snippet of video via a video player embedded within a browser on the user device. (See McLaughlin paragraph 0061 discussing the ability to playback video clips in a browser-based video player.) It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of McLaughlin for the benefit of Croitoru in view of Chhaya, because to do so provided a designer with options for implementing an efficient and compliant video playback capability, as taught by McLaughlin in paragraph 0061. These references were all applicable to the same field of endeavor, i.e., management and processing of multimedia data. Claim 11 is substantially similar to claim 3, and therefore likewise rejected. Claim 19 is substantially similar to claim 3, and therefore likewise rejected. Claims 5 and 13 are rejected under 35 U.S.C. §103 as being unpatentable over Croitoru et al (US Patent Application Publication No. 2024/0355119, hereafter referred to as “Croitoru”) in view of Chhaya et al (US Patent Application Publication No. 2023/0290146, hereafter referred to as “Chhaya”) and Frost et al (US Patent Application Publication No. 2023/0154184, hereafter referred to as “Frost”). Regarding claim 5: Croitoru in view of Chhaya does not explicitly teach the remaining limitations as claimed. Frost, though, teaches wherein the processor is further configured to execute a topic modeling algorithm on the transcript of the video to identify a plurality of snippets of video within the video corresponding to a plurality of different search topics, respectively, and store the plurality of snippets with a plurality of metadata identifying the plurality of different search topics, respectively. (See Frost paragraph 0025 teaching the use of a machine language topic modeling algorithm to process transcribed audio and to identify topics/scenes in a video, making use of metadata such as topic information and timestamps. See also Fig. 3 showing the use of storage devices and Fig. 1 #116 showing use of a database storage element). It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Frost for the benefit of Croitoru in view of Chhaya, because to do so provided a designer with options for implementing a system to automatically update video compilations, as taught by Frost in the Abstract. These references were all applicable to the same field of endeavor, i.e., management and processing of multimedia data. Claim 13 is substantially similar to claim 5, and therefore likewise rejected. Claims 6-7 and 14-15 are rejected under 35 U.S.C. §103 as being unpatentable over Croitoru et al (US Patent Application Publication No. 2024/0355119, hereafter referred to as “Croitoru”) in view of Chhaya et al (US Patent Application Publication No. 2023/0290146, hereafter referred to as “Chhaya”) and Blong et al (US Patent Application Publication No. 2017/0220869, hereafter referred to as “Blong”). Regarding claim 6: Croitoru in view of Chhaya does not explicitly teach the remaining limitations as claimed. Blong, though, teaches wherein the processor is further configured to output identifiers of a plurality of different video snippets and a plurality of topics corresponding to the plurality of different video snippets via the user interface. (See Blong Fig. 1 showing clips with associated metadata, and Fig. 7B / paragraph 0051 teaching an interface for displaying clips and their identifiers.) It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Blong for the benefit of Croitoru in view of Chhaya, because to do so provided a designer with options for implementing a mechanism for the efficient creation of high quality video supercuts/compilations, as taught by Blong in the Abstract. These references were all applicable to the same field of endeavor, i.e., management and processing of multimedia data. Regarding claim 7: Croitoru in view of Chhaya does not explicitly teach the remaining limitations as claimed. Blong, though, teaches wherein the processor is further configured to receive inputs via the user interface selecting two or more video snippets from among the plurality of different video snippets, aggregate content from the two or more video snippets to generate a custom video, and play the custom video via the user interface of the user device. (See Blong Abstract in the context of Fig. 1 teaching the ability of a user to browse a repository of video clips using a supercut tool, and to create a “supercut” of ordered clips. See also paragraph 0052 discussing the playback of the supercut.) Claims 14-15 are substantially similar to claims 6-7, respectively, and therefore likewise rejected. Claims 8 and 16 are rejected under 35 U.S.C. §103 as being unpatentable over Croitoru et al (US Patent Application Publication No. 2024/0355119, hereafter referred to as “Croitoru”) in view of Chhaya et al (US Patent Application Publication No. 2023/0290146, hereafter referred to as “Chhaya”), Blong et al (US Patent Application Publication No. 2017/0220869, hereafter referred to as “Blong”) and Chen et al (US Patent Application Publication No. 2016/0014482, hereafter referred to as “Chen”). Regarding claim 8: Croitoru in view of Chhaya and Blong does not explicitly teach the remaining limitations as claimed. Chen, though, teaches wherein the processor is further configured to identify a contribution of each video snippet from among the two or more video snippets to the custom video, and store the identified contributions in storage. (See Chen paragraph 0165 discussing the use of a settings menu user interface that can adjust a weighting attributed to a video segment, it having been implied that while / for the menu to be displayed the contribution values were stored [i.e., in a register, cache, etc.].) It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Chen for the benefit of Croitoru in view of Chhaya and Blong, because to do so provided a designer with options for implementing a system that enabled ordering and formation of compilations based on the importance of video segments, as taught by Chen in the Abstract. These references were all applicable to the same field of endeavor, i.e., management and processing of multimedia data. Claim 16 is substantially similar to claim 8, and therefore likewise rejected. Claim 20 is rejected under 35 U.S.C. §103 as being unpatentable over Croitoru et al (US Patent Application Publication No. 2024/0355119, hereafter referred to as “Croitoru”) in view of Chhaya et al (US Patent Application Publication No. 2023/0290146, hereafter referred to as “Chhaya”) and Bao et al (US Patent Application Publication No. 2014/0122705, hereafter referred to as “Bao”). Regarding claim 20: Croitoru in view of Chhaya does not explicitly teach the remaining limitations as claimed. Bao, though, teaches wherein the processor is further configured to perform comparing the content of the extracted video snippet to a browsing history of the user device to identify similar content as the extracted video snippet, determine a recall confidence of the similar content, and store the recall confidence in storage. (See Bao paragraph 0003 discussing the conventional use of browsing history to determine related video clips, in the context of Fig. 1 #34 showing storage. See also, Fig 3 #S307 teaching the determination of a similarity degree.) It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains to apply the teachings of Bao for the benefit of Croitoru in view of Chhaya, because to do so provided a designer with options for implementing a system that recognizes previous user interactive behavior for recommending or displaying video data, as taught by Bao in paragraph 0003. These references were all applicable to the same field of endeavor, i.e., management and processing of multimedia data. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Relevance is provided in at least the Abstract of each cited document. Non-Patent Literature Cardillo, Daniela, et al., “The art of video MashUp: supporting creative users with an innovative and smart application”, Multimedia Tools and Applications, Vol. 53, © Springer Science+Business Media, LLC, February 4, 2010, pp. 1-23. In this paper, we describe the development of a new and innovative tool of video mashup. This application is an easy to use tool of video editing integrated in a cross-media platform; it works taking the information from a repository of videos and puts into action a process of semi-automatic editing supporting users in the production of video mashup. Doing so it gives vent to their creative side without them being forced to learn how to use a complicated and unlikely new technology. The users will be further helped in building their own editing by the intelligent system working behind the tool: it combines semantic annotation (tags and comments by users), low level features (gradient of color, texture and movements) and high level features (general data distinguishing a movie: actors, director, year of production, etc.) to furnish a pre-elaborated editing users can modify in a very simple way. (pages 1-2, Abstract). On a lower level, scenes can subsequently be segmented into a sequence of basic video segments named shots. While scene did exist before video, on the other hand shots originate with the invention of motion cameras and are considered to be the longest continuous frame sequences that come from a single camera take (i.e., what the camera images in an uninterrupted run [17]). Shots sharing common perceptual low-level characteristics can then be clustered together into higher entities called groups (or clusters) of shots (see Fig. 2). Examples of groups are visually similar shots, or video segments sharing the same camera motion. Finally, on the bottom level of the hierarchy, one or more key-frames can be extracted from shots as static significant paradigms of the shot visual content. The video repository used for video mash-up contains video material segmented at different levels of the hierarchy and the related metadata (that are HLF = High-level Features, MLF = Mid-level features, LLF = Low-level features, MA = Manual annotations). In particular, in order to create new content, the following entities can be retrieved from the repository: Full length videos; Clips of any length uploaded by users; logical story units; and, shots and groups of similar shots. (paragraph spanning pages 1-2). processing the user’s criteria of selection in an intelligent way in order to return only the clip really interesting for him following the preferences saved in the user’s profile; (page 14, bulleted item: “rules of selection”). Saini, Mukesh, et al., “MoViMash: Online Mobile Video Mashup”, MM ‘12, Nara, Japan, October 29 – November 2, 2012, pp. 139-148. With the proliferation of mobile video cameras, it is becoming easier for users to capture videos of live performances and socially share them with friends and public. As an attendee of such live performances typically has limited mobility, each video camera is able to capture only from a range of restricted viewing angles and distance, producing a rather monotonous video clip. At such performances, however, multiple video clips can be captured by different users, likely from different angles and distances. These videos can be combined to produce a more interesting and representative mashup of the live performances for broadcasting and sharing. The earlier works select video shots merely based on the quality of currently available videos. In real video editing process, however, recent selection history plays an important role in choosing future shots. In this work, we present MoViMash, a framework for automatic online video mashup that makes smooth shot transitions to cover the performance from diverse perspectives. Shot transition and shot length distributions are learned from professionally edited videos. Further, we introduce view quality assessment in the framework to filter out shaky, occluded, and tilted videos. To the best of our knowledge, this is the first attempt to incorporate history-based diversity measurement, state-based video editing rules, and view quality in automated video mashup generations. (page 139, Abstract). Nixon, Lyndon, et al., “Video Lectures Mashup – remixing learning materials for topic-centered learning across collections”, OCW Global Conference, Ljubljana, Slovenia, November 2014, 15 pages. In this paper, we introduce the VideoLecturesMashup, which presents re-mixes of learning materials from the VideoLectures.NET portal based on shared topics across different lectures. Learners need more efficient access to teaching on specific topics which could be part of a larger lecture (focused on a different topic) and occur across lectures from different collections in distinct domains. Current e-learning video portals cannot address this need, either to quickly dip into a shorter part focused on a specific topic of a longer lecture or to explore what is taught about a certain topic easily across collections. Through application of media technologies promoted by the MediaMixer project1 – semantic annotation and media fragment URIs – we have implemented a first demo of VideoLecturesMashup. (1st page, Abstract). Song, Hao, et al., “Extracting Key Segments of Videos for Event Detection by Learning from Web Sources”, IEEE Transactions on Multimedia, Vol. 20, No. 5, May 2018, pp. 1088-1100. In this paper, we present a novel approach of extracting the key segments for event detection in unconstrained videos. The key segments are automatically extracted by transferring the knowledge learned from Web images and Web videos to consumer videos. We propose an adaptive latent structural support vector machine model, where the locations of key segments in videos are regarded as latent variables due to the unavailability of the ground truth of key-segment locations in training data. In order to alleviate the time-consuming and labor expensive manual annotation of huge amounts of training videos, a large number of loosely labeled Web images as well as videos are collected from the Web sources. Additionally, a limited number of labeled consumer videos are utilized to guarantee the precision of the model. Considering the semantic diversity of key segments, we learn a set of concepts as the semantic description of key segments and explore the temporal information of concepts to capture the sequential relations between the segments. The concepts are automatically discovered by using Web images and videos with their associated tags and description sentences. Comprehensive experiments on the Columbia’s consumer video and the TRECVID 2014 Multimedia Event Detection datasets demonstrate that our method outperforms the state-of-the-art methods. (page 1088, Abstract). The framework of the proposed method. We automatically discover concepts by learning from Web images and videos with their associated tags and description sentences. Then the Web sources are used to train the basic concept SVM classifiers. The knowledge learned from the Web sources is transferred to adapt to an optimal target classifier. Also, we explore the temporal relationship to extract the key segments of a video. A discriminative model is learned by using an adaptive latent structural SVM model for high level event classification. (page 1089, Fig. 1). An event video is represented as a mid-level concept semantic vector. And all the relationships are built on this concept representations. Different from these segment discovery methods, the concepts in our framework are automatically discovered. Our method introduces a latent structural SVM framework of event detection which utilizes large amounts of loosely labeled Web sources and a few labeled training target videos. Each segment corresponds to one concept, the segment discovery and event detection are processed simultaneously. (page 1090. Section “B. Extracting Segments of Videos for Event Detection”). Rautiainen, Mika, et al., “Analysing the Performance of Visual, Concept and Text Features in Content-Based Video Retrieval”, MIR ‘04, New York, NY, October 15-16, 2004, pp. 197-204. In this study semantic concepts (or semantic features) are textually defined conceptual terms that are detected from a video with a certain confidence. Concepts describe the objects and events in a video sequence trying to enhance the information obtained by automatic speech recognition transcripts or video OCR. Principally for each concept there is a trained semantic concept detector. As a result of the detection procedure, a set of confidence estimates is obtained from every video sequence for each concept. Concept search engine considers confidence values as features that are being compared against the user defined concept list of the query. (page 199, 1st paragraph of section “2.2 Concept Search Engine”) US Patent Application Publications Lee 2008/0212932 A video management method and system based on a topic, and a video search method based on a topic. The video management system includes: a video topic management unit extracting a topic from information associated with a video, storing the extracted topic, and monitoring the video associated with the stored topic; a video storage management unit storing the video and information associated with the video, and managing a storage space based on the topic which is stored in the video topic management unit; and a topic video management unit generating a topic video of the monitored video when the video associated with the stored topic is monitored, and providing a search function and a navigation function of the video which is stored in the video storage management unit. (Abstract). The user history database stores a history about the user's manipulation. The history analysis unit 244 analyzes a major pattern based on the history about the user's manipulation. Specifically, the history analysis unit 244 performs an analysis process to extract a user preference topic. In this case, the user preference topic may be analyzed based on a topic which a user has recently performed a search on, a topic which is discussed in a broadcast program that the user frequently views, and the like. Also, when a plurality of users exists, such as with a television, the history analysis unit 244 may manage a history for each user. Also, the history analysis unit 244 may classify and analyze the user preference topic for each user. (paras 0049-0050). Ziai 2024/0320958 As discussed above, some search methods are completely manual and require users to extract digital video clips by hand. (para 0005). In some examples, the classification category prediction displays within the digital content understanding graphical user interface include a playback window loaded within a training digital video clip corresponding to the classification category prediction display. The classification category prediction displays can further include a title of a digital video from which the displayed training digital video clip came, and an option to positively acknowledge or negatively acknowledge the same training digital video clip. Generating the classification category prediction displays within the digital content understanding graphical user interface can further include sorting the classification category prediction displays into high levels of confidence and low levels of confidence and updating the classification category prediction displays within the digital content understanding graphical user interface according to the high levels of confidence and the low levels of confidence. (para 0008). Chhaya 2023/0290146 In one example, a search is performed to locate digital videos that pertain to a particular topic. A digital video sharing system, for instance, supports a keyword search to locate digital videos of interest, e.g., “how to change a tire.” In response, a search result is received that includes representations of digital videos that correspond to the keyword search. A user input is then received that selects representations to specify a plurality of digital videos that are to be used to generate a digital document. In this example, the user input in this instance is used to select digital videos that involve changing a tire and avoid digital videos merely describing tires, tire reviews, and so on that are not related to a topic of interest. (para 0024). The digital document generation system 118 begins by locating action clips. The action clips includes frames that depict actions from the plurality of digital videos 114 (block 906). To do so, an extraction module 218 is employed that includes a key clip extraction module 220 and a transcription module 222. The key clip extraction module 220 is configured to extract key clips 224 of frames 210 from the digital videos 114, e.g., as collection of a predetermined number of frames, using machine learning and object recognition to detect inclusion of entities in frame groupings, and so forth. The transcription module 222 is configured to perform speech-to-text techniques that are usable to convert the digital audio 212 into text as part of the transcript 226, e.g., by accessing an application programming interface (API) of a digital service (block 908). (para 0049). The action clips 230 are then received as an input by a sequence generation module 234 to generate action sequences 236 from the digital videos 114 (block 912). The action sequences 236 describe sequences of verbs 240 in respective digital videos, e.g., one or more per video. Thus, the action sequences 236 capture an order of “what has occurred” in the respective digital video 114. The sequence generation module 234 employs a verb detection module 238 that is configured to detect verbs 240 from the action clips 230 and/or transcript 226. In one example, this is performed using bidirectional transformers trained for natural language understanding for semantic-role labeling, which can be fine-tuned for a specific domain. The verb detection module 238, for instance, is configurable using machine learning to process frames of the action clips 230 identified by the action detection module 228 to detect which verbs are usable to describe “what is going on” in the frames of the action clips 230. In another example, the verb detection module 238 identifies the verbs 240 from portions of the transcript 226 that correspond to the action clips 230. Semantic role labeling analyzes natural language sentences to extract information about “who did what to whom, when, where, and how.” Because each verb in the transcript 226 does not correspond to an action for an instruction (e.g., “keep” in “keep stirring the soup”), a domain-specific dictionary is employable by the sequence generation module 234 as a filter to remove potentially misleading text. The dictionary is also learnable from an enterprise corpus, author-defined input, and so on. (paras 0051-0052). Patterson 2021/0272599 Provided are systems for automatic video processing that employ machine learning models to process input video and understand user video content in a semantic and cultural context. This recognition enables the processing system to recognize interesting temporal events, and build narrative video sequences automatically, for example, by linking or interleaving temporal events or other content with film-based categorizations. In further embodiments, the implementation of the processing system is adapted to mobile computing platforms which can be distributed as an “app” within various app stores. In various example, the mobile apps turn everyday users into professional videographers. In further embodiments, music selection and dialog based editing can likewise be automated via machine learning models to create dynamic and interest professional quality video segments. (Abstract). Various embodiments, are configured to meet the needs of casual and professional users, with a computational editing pipeline that extracts relevant semantic concepts from static frames and video clips, recombines clips into film idioms that form the grammar of cinematic editing, and interact with users to collaboratively create engaging short films. FIG. 11 illustrates a high-level overview of a multi-stage film idiom recognition and generation network. As discussed above, various implementations are tailored to execute even in the context of a mobile device platform with associated computation and power limitations. In some embodiment, first network (e.g., a CNN) can be used to identify important video segments and project them into a semantic embedding space. The semantic embedding space can be used to identify concepts, film idiom, cinematic categories, etc. The identification can be executed by another neural network (e.g., LTSM) and/or by the first network. The identification of concepts, film idiom, cinematic categories, etc., can also identify options narrative progression, and the best candidates for such narrative, for example, using a LTSM network. Based on the identification of a “best” narrative sequence an output video can be presented to end users. Based on any further edits, additional feedback can be created to refine the machine learning implementation. (para 0076). Additional functions provided by the system include dialog-based editing. In some embodiments, processing system can take as an input a script with line of dialog (e.g., 1302), a set of input takes (e.g., 1304 video clips), and a set of editing rules (e.g., 1306 “start wide,” no “long” shots, and intensify emotion, among other option), to generate a final video output 1310. Given the inputs, the processing system can align dialog and clips by extracting and matching concepts (e.g., matching semantic information associated with clips (e.g., as discussed above), and then employing the semantic information to order and/or select specific clips for an output video, for example, as shown in FIG. 13. (para 0087). Denoue 2017/0371496 Example implementations described herein are directed to systems and methods for representing meeting content. Such implementations may involve processing an online presentation for one or more media segments, extracting information from the one or more media segments indicative of one or more relationships between one or more participants of the online presentation and generating an interface for the online presentation, the interface indicative of the one or more relationships between the one or more participants of the online presentation. Through such example implementations, online presentations can be indexed and an interface can be generated for the online presentation that allows for content of the presentation to be searchable. (Abstract). To provide an interface for browsing such meetings, related art implementations have provided a search mechanism, wherein if the meeting is properly indexed using speech to text and optical character recognition (OCR), a search interface can return snippets (e.g., key frames) extracted from videos, allowing users to quickly extract relevant parts of a video meeting. In such related art search systems, the source streams can be subdivided using both a speaker segmentation and topic information derived from the speech transcripts in a multi-level video segmentation. (paras 0004-0005). Blong 2017/0220869 The creation of a supercut is described using techniques to allow users to efficiently create high quality supercuts. A video clip repository may include a number of video clips. The video clip repository may allow users to browse and view video clips in the repository. A supercut creation tool may operate to identify, based on comparison of search criteria received from a user to the set of tags, video clips, from the set of video clips, that are relevant to the search criteria; determine, based on scores of the video clips, an ordering of the video clips; and generate a supercut of the video clips as a single video corresponding to the video clips and arranged in the determined order. (Abstract). As an example of the operation of the supercut creation tool, a user may input one or more search terms, or select one or more categories, relating to video clips that the user is interested in potentially including in a supercut (at 1.1, “search criteria”). For example, the user may enter the names of particular actors, movie titles, directors, or other information. The supercut creation tool may search the video clip repository, such as by searching the metadata associated with the video clips, to determine video clips relevant to the user's search (at 1.2, “obtain relevant video clips”). In some implementations, the supercut creation tool may automatically select video clips and select and an arrangement of the video clips for the supercut (at 1.3, “output supercut with automatically ordered video clips”). In one implementation, the video clips stored by the video clip repository may be associated with a score that quantifies the quality or popularity of each video clip. In one implementation, the supercut creation tool, when automatically arranging the video clips in the supercut, may insert the highest scoring video clip as the first video clip in the supercut and the second highest scoring video clip as the last video clip in the supercut. In this implementation, putting the highest scoring video clip as the first video clip may to tend to maximize the ability of the supercut to grab the viewer's attention and putting the second highest scoring video clip as the last video clip may increase the likelihood that a viewer of the supercut may be motivated to share or otherwise recommend the supercut. (para 0014). Burkitt 2012/0254917 Disclosed embodiment providing for the capture of video content. The video content is segmented in real-time into clips by topic, and those clips are delivered as customized queues of video items relevant to the consumer according to their interests as aggregated from their social graph data and manual entry. (Abstract). The systems and methods disclosed in this specification capture video content, segment the content in real time, can sort the video content into clips by topic, and can delivers those clips as a customized queue of video items relevant to users according to their interests, as determined by their social graph data and/or their manual interest profile configurations. The disclosed systems and methods can segment long-form video content into smaller clips based on the topic of the content. This enables the generation of a searchable index of short, highly relevant video results. The created index is not only useful to the user, but also provides advertisers a deeper context against which relevant advertising can be selected. In addition to enabling `directed search` results for keyword lookups in the index, the disclosed systems and methods enables providing recommendations in the form of a custom video queue or other organization means. These recommendations may be based on users' interests. The catalogue of user interests is built by aggregating user input, social network graph information, and usage feedback. (paras 0010-0012). Begeja 2003/0163815 A method for extracting multimedia content segments, such as electronic clips or "eClips," from a source of video or other multimedia content. The extraction is based on individual preferences such as key terms and/or phrases as well as content source, which a user may identify in a profile. User profiles can be stored in a service platform and continually checked against new content in the system. When matches are found between a user profile and the content, the service platform can alert the user that segments have been identified and extracted. The user may then view/play these automatically provided segments (eClips). In addition, the eClips service is capable of stitching the clips of diverse sources together, providing an automatically generated multimedia experience that revolves around the user's provided profile. (Abstract). In another embodiment, the invention relates to a system for delivering a customized video presentation comprising video clips to a user. The system may include a video capture device operable to receive a plurality of video inputs, as well as a video database operable to store the plurality of video inputs and text associated with the video inputs. The system may also include a video server operable to search the video inputs within the video database in accordance with a user criterion and based on the text. The video server may be further operable to extract from the video inputs video clips corresponding to the user criterion and combine the video clips into a customized video presentation for delivery to the user. (para 0014). Once the content is captured and recorded, it can be segmented, analyzed and/or classified, and thereafter stored on a platform. For example, the content can be broken down into its component parts, such as video, audio and/or text. The text can include, for example, closed-captioning text associated with the original transmission, text generated from an audio portion by speech recognition software, or a transcription of the audio portion created before or after the transmission. In the latter case, it becomes possible to utilize the invention in conjunction with executive speeches, conferences, corporate training, business TV, advertising, and many other sources of video which do not typically have available an associated textual basis for searching the video. Having obtained or generated the text, it can then be used as a basis for searching the multimedia content. In particular, the text provides the basis for an exemplary methodology for overcoming the above-identified problems associated with searching video in the prior art. That is, if a user wishes to search the stored content for video segments relevant to the President of the United States discussing a particular topic, then the President's name and the associated topic can be searched for within the text associated with the video segments. Whenever the President's name and the associated topic are located, an algorithm can be used to determine which portion of an entire video file actually pertains to the desired content and should therefore be extracted for delivery to the user. Thus, if a video file comprises an entire news broadcast about a number of subjects, the user will receive only those portions of the broadcast, if any, that pertain to the President and the particular topic desired. For example, this could include segments in which the President talks about the topic, or segments in which another talks about the topic and the President's position. Once the pertinent segments of the broadcast have been appropriately extracted, for a given user, they can be stitched together for continuous delivery to that user. (paras 0026-0028). Chen 2016/0014482 In many embodiments, the process of generating a personalized playlist is treated as a maximum coverage problem. A maximum coverage problem typically involves a number of sets of elements, where the sets of elements can intersect (i.e. a single element can belong to multiple sets). Solving a maximum coverage problem involves finding the fixed number of elements that cover the largest number of sets of elements. In the context of generating a personalized playlist, the elements are the video segments and video segments that relate to the same content are treated as belonging to the same set. Therefore, the concept of content coverage can be used to refer to the amount of different content covered by a set of video segments. As noted above, video segments can be compared to determine whether the content is related or unrelated. In the context of news stories, many embodiments attempt to span the major news stories of the day and an objective function for solving the maximum coverable problem can be weighted by a linear combination of several personalization factors. These factors can include (but are not limited to) explicit preferences specified by a user, personal information provided by the user and/or obtained from secondary sources including (but not limited to) online social networks, and implicit preferences obtained by analyzing a user's viewing history. Information concerning implicit preferences may be derived by analyzing a user's viewing history with respect to playlists generated by a playlist generation server system. In other embodiments, implicit preferences can be derived from additional sources of information including (but not limited to) a user's browsing activity (especially with respect to online articles relevant to video segment content), activity within an online social network, and/or viewing history with respect to video and/or audio content provided by one or more additional services. (para 0136). As can readily be appreciated, processes similar to those described above with respect to FIG. 24B can be utilized to create summaries of individual video segments, to annotate a given video segment with relevant video content from other video segments, and/or other content from sources associated with one or more video segments identified as relevant to the given video segment. Furthermore, any of a variety of processes can be utilized to identify and score individual video clips extracted from a video segment for the purpose of combining video clips. (para 0175). Scoring metrics can be any value assigned to a video clip that can represent the relative importance and/or relevance of a video clip as compared to other video clips with respect to a specific topic and/or subject. A process for scoring and selecting video clips is illustrated in FIG. 24E. Variety of key features can be extracted (2472) from video clips including, but not limited to, visual, textual, and audio data, or any other feature as appropriate to the requirements of specific applications. Scoring data be generated (2474) for each video clip based upon the extracted key features. Importance of a video clip can be determined based upon key features. (paras 0181-0182). In many embodiments, video clips can be ordered (2478) to enhance the quality of the video summary sequence. In some embodiments, ordering can be based on one or more scores assigned to video clips. Ordering can be determined prior to, during, or after video clips are extracted from video segments. In many embodiments, ordering video clips places the video clips with the highest scores at the beginning of the video summary sequence. In other embodiments, video clips with the highest scores are placed at the end of the video summary sequence. As can be readily appreciated, any ordering of video clips can be used as appropriate to the requirements of specific applications in accordance with embodiments of the invention. Although specific processes are described above with respect to the generation of video summary sequences, any of a variety of techniques can be utilized extract and select video clips from one or more video segments, score video clips, and order video clips as appropriate to the requirements of specific applications in accordance with embodiments of the invention. (paras 0188-0189). Otsubo 2010/0229078 FIG. 4 is a diagram illustrating an example in which the content display control apparatus 1 extracts a plurality of topics from recorded content 40 stored in the content database 8, and the plurality of topics are utilized as video clips. (para 0061). Logan 2005/0061232 Typically, the metadata used for program scanning would use metadata that was different than that used to divide the show into segments, as the index metadata would be too course to allow for the extraction of the short, pithy highlights best suited for this function. The metadata used in the scanning function, and the resulting video clips extracted, would have to be carefully selected so as not to give away too much of the plot or game result in the case of sports, as discussed later under topic "Plot Leakage." The content used in either program-level or segment-level scanning could be customized to a specific viewer, client device, or VOD system. (paras 0354-0356). Sezan 2005/0131727 A method of using a system with at least one of audio, image, and a video comprises a plurality of frames comprising the steps of providing a usage preferences description where the usage preference description includes at least one of a browsing preferences description, a filtering preferences description, a search preferences description, and a device preferences description. The browsing preferences description relates to a user's viewing preferences. The filtering and search preferences descriptions relate to at least one of (1) content preferences of the at least one of audio, image, and video, (2) classification preferences of the at least one of audio, image, and video, (3) keyword preferences of the at least one of audio, image, and video, and (4) creation preferences of the at least one of audio, image, and video. The device preferences description relates to user's preferences regarding presentation characteristics. A usage history description is provided where the usage preference description includes at least one of a browsing history description, a filtering history description, a search history description, and a device usage history description. The browsing history description relates to a user's viewing preferences. The filtering and search history descriptions relate to at least one of (1) content usage history of the at least one of audio, image, and video, (2) classification usage history of the at least one of audio, image, and video, (3) keyword usage history of the at least one of audio, image, and video, and (4) creation usage history of the at least one of audio, image, and video. The device usage history description relates to user's preferences regarding presentation characteristics. The usage preferences description and the usage history description are used to enhance system functionality. (Abstract). McLaughlin 2013/0036201 As a practical example, consider a online sports blog publisher who would like to present his viewing audience with a single basketball highlights composite video from four existing internet based video clips. For the sake of simplicity, we will assume that he will be using "Video Clip 1" (VC1) and "Video Clip 4" (VC4) in their entirety, but would like to specify an "out" point on "Video Clip 2" (VC2) before its end, and an "in" point on "Video Clip 3" (VC3) after its beginning. Suppose that this blog is focusing on one particular athlete, and that his preferred "out" point on VC2 occurs just after the player scores and the ball goes through the net, and the preferred "in" point for VC2 occurs just after the player receives an inbound pass. For both the "in" and "out" points, the timing window for the cut is about 0.1 seconds. Playlist Alternative A (200). Using existing playlist tools (e.g., http://embedr.com), it is possible to string together multiple video clips to play back sequentially in a single browser playback window using one instance of the browser's video player. This embodiment is illustrated 250 by "Playlist Alternative A" 200. In this situation, VC1 260 will play in its entirety, and when it is done, VC2 270 is loaded into the browser's video player, and it is played in its entirety, and similarly followed by VC3 280 and VC4 290. All that is required of the video player API for this embodiment is the ability to programmatically load, unload, and start the videos. (paras 0060-0061). Croitoru 2024/0355119 In various embodiments, the automatic annotation component 250 can combine the outputs of the automatic speech recognition (ASR) component 230 and the natural language processing (NLP) model 240 to automatically generate a transcript of the audio track of a video and apply the transcript to the video component of the video as annotation of the activities occurring in the scenes of the video. The annotation can be added to the video to provide labels for the events occurring within the video scenes. The use of ASR can reduce or eliminate the use of human annotated labels for training and/or video segment identification. Instead, weak supervision from video transcripts can be leveraged. (para 0038). At operation 320, the video segment identification system 130 can obtain the video 305 from the user or from a database based on a specification by the user, for example, the query may include the name of a video, where the video may be available on a database, without the user uploading the specific video to the video segment identification system 130. The video segment identification system 130 may obtain the video 305 from the user 110, where the user 110 uploads the video from a source available to the user or from local storage on a user devise 112, 114, 116, 118, 119. At operation 330, a transcript can be generated by the video segment identification system 130, where the transcript can be generated from the audio track of the video using automatic speech recognition. The transcript can include the detectable words spoken by a character or presenter in the video translated into text. In various embodiments, a natural language processing component 240 can analyze the transcript to improve the interpretation of the content of the transcript, where the natural language processing component 240 can identify topics and descriptions discussed in the audio track. The transcript text can be used to annotate the video. At operation 340, identified topics and descriptions in the transcript can be associated with portions of the original video, where the portions of the video can include scenes associated with identified topic(s). The automatic annotation component 250 may associate identified topics and descriptions in the transcript with portions of the video. The automatic annotation component 250 may be trained to identify whether or not a portion of a video contains visual content related to transcript topics or descriptions. For example, during a video portion, a presenter my digress and discuss other topics or personal information, such as the weather, a pet, or a child, that is unrelated to the visual material and context of the video portion being displayed at the same time. The automatic annotation component 250 may determine that the topic(s) in the transcript are unrelated to the visual objects present in the video at the concurrent time. At operation 350, the video portions that are identified as relating to the topics requested by the user in the query and identified in the transcript can be extracted from the original video 305 to generate a video segment 370. A video segment can be identified by a start time and an end time, where, for example, a start time may be identified by the initial reference to a topic in the associated transcript concurrently with the video displaying objects or scenes related to the topic, whereas the end time may be identified by the objects or scenes used for identifying the start time no longer appearing in the video. The reappearance of the objects or scenes relating to the query topic can cause another video segment having different start and end times to be identified. The video segment(s) 370 having the identified start and end time(s) can be generated from the original video 305. (paras 0043-0046). Frost 2023/0154184 For example, based on the identified entities and objects at a given timestamp of a video/episode that may be fed into the machine learning model, the recap generator program 108A, 108B may identify an entity, such as character A, and an object such as a sword. Furthermore, based on the transcribed audio in a time range that may include the given timestamp, the recap generator program 108A, 108B may identify sounds and dialogue that may correspond to a battle. As such, using the topic modeling algorithms and change in topic algorithms associated with the machine learning model, the recap generator program 108A, 108B may determine that, for a timestamp and range of 43:46-44:00 in the video/episode, character A is engaging in a battle. As such, the recap generator program 108A, 108B may determine that the time of 43:46-44:04 represents a topic that includes a battle scene/clip involving character A. As such, the recap generator program 108A, 108B may use the topic groupings for each video/episode, as well as the correlated and identified entities/objects and transcribed audio, to identify a respective scene and/or clip in the video/episode. In turn, and as depicted at 228 in FIG. 2A, the recap generator program 108A, 108B may generate a list of scenes, whereby the list of scenes may be categorized according to the topics and may also include the entities, objects, and timestamp data. Therefore, each time a video/episode is released and/or added to the episodic collection of videos/episodes, the recap generator program 108A, 108B may process the video/episode according to the operational flowchart 200A described in FIG. 2A and thereby generate a list of scenes corresponding to the topics and entities covered in each video/episode. For example, and as depicted at 228, the recap generator program 108A, 108B may process a received video/episode and thereby generate a list of scenes/topics for that received video/episode, whereby each topic (i.e. Topic A, Topic B . . . Topic n) may also include corresponding data on the entities and objects associated with the topic as well as timestamp data. (para 0025). Bao 2014/0122705 At present, networks have become a common medium for people to access, browse, store, and exchange information on a daily basis. From the perspective of an end user, interaction with the network information may be performed through a site on the network (or simply called "website"). With the development of network technology, more and more sites can mine and study user features, for example, interactive habits, preferences, interests, etc., using a technology such as data analysis, and on this basis, provide personalized and/or customized information service to the users. For example, a video service network can infer from a user's browsing history and previous interactive behaviors which type of information the user potentially prefers, and recommend or display video clips related to this type of information in an eye-catching way. (para 0003). Dontcheva 2020/0334290 For instance, the interactive computing environment may access a video search engine application programming interface (API) and identify videos that are relevant to the video search queries. In an example, a searching subsystem further identifies relevant portions of the video search results to the video search query by comparing the caption tracks of the video search results to the search terms provided by the video search query. In one or more examples, the relevant portions of the video search results are ranked and presented to a user based on a number of matches identified between terms of the video search query (e.g., including context information of an active tool) and text in the caption tracks. (para 0020). US Patents Chang 6,741,655 Object-oriented methods and systems for permitting a user to locate one or more video objects from one or more video clips over an interactive network are disclosed. The system includes one or more server computers (110) comprising storage (111) for video clips and databases of video object attributes, a communications network (120), and a client computer (130). The client computer contains a query interface to specify video object attribute information, including motion trajectory information (134), a browser interface to browse through stored video object attributes within the server computers, and an interactive video player. (Abstract). Systems and methods for generating video summary sequences in accordance with embodiments of the invention are illustrated. An embodiment of the method of the invention includes obtaining a set of annotated video segments using a video summarization system, extracting a set of video clips from the set of annotated video segments based upon clipping cues using the video summarization system, where a video clip in the set of video clips includes at least one key feature and metadata describing the length of the video clip, generating scoring data using a video summarization system, wherein the scoring data includes at least one scoring metric for each video clip in the set of video clips, where the at least one scoring metric describes the at least one key feature of each video clip utilized to determine the relative importance of each video clip within the set of video clips, selecting a subset of the set of video clips based on the generated scoring data such that the sum of the lengths of the video clips in the selected subset of video clips is within a predefined range of lengths using the video summarization system, determining a sequence of at least a subset of video clips from the selected subset of video clips using the video summarization system, generating a video summary sequence including the selected subset of video clips in the determined sequence using the video summarization system, and providing the generated video summary sequence in response to a request for a video summary sequence using the video summarization system. (para 0006). Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner ROBERT STEVENS whose telephone number is (571) 272-4102. The examiner can normally be reached Mon - Fri 6:00 - 2:30. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amy Ng can be reached on (571) 270-1698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ROBERT STEVENS/Primary Examiner, Art Unit 2164 September 16, 2026
Read full office action

Prosecution Timeline

Jul 08, 2023
Application Filed
Nov 21, 2023
Response after Non-Final Action
Sep 18, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748988
DYNAMIC PROTOTYPE LEARNING FRAMEWORK FOR NON-HOMOPHILOUS GRAPHS
3y 6m to grant Granted Sep 29, 2026
Patent 12737407
SYSTEM FOR HELPING OPERATOR TO QUESTION HELP-SEEKER
1y 10m to grant Granted Sep 15, 2026
Patent 12737421
INTERPRETABLE FEATURE DISCOVERY WITH GRAMMAR-BASED BAYESIAN OPTIMIZATION
1y 9m to grant Granted Sep 15, 2026
Patent 12694065
APPARATUS AND A METHOD FOR HEURISTIC RE-INDEXING OF STOCHASTIC DATA TO OPTIMIZE DATA STORAGE AND RETRIEVAL EFFICIENCY
1y 8m to grant Granted Jul 28, 2026
Patent 12651026
APPARATUS AND METHOD FOR OPTIMAL ZONE STRATEGY SELECTION
1y 7m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
93%
With Interview (+11.9%)
2y 11m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 529 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month