Prosecution Insights
Last updated: August 16, 2026
Application No. 19/236,765

ORGANIZING MEDIA CONTENT ITEMS UTILIZING DETECTED SCENE TYPES

Non-Final OA §101
Filed
Jun 12, 2025
Priority
Dec 08, 2022 — provisional 63/386,628 +1 more
Examiner
BIBBEE, JARED M
Art Unit
2161
Tech Center
2100 — Computer Architecture & Software
Assignee
Dropbox Inc.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
1y 10m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
538 granted / 670 resolved
+25.3% vs TC avg
Moderate +14% lift
Without
With
+13.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
12 currently pending
Career history
680
Total Applications
across all art units

Statute-Specific Performance

§101
15.8%
-24.2% vs TC avg
§103
52.3%
+12.3% vs TC avg
§102
17.7%
-22.3% vs TC avg
§112
4.7%
-35.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 670 resolved cases

Office Action

§101
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 4-8, 11-14, and 17-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. As to independent claims 1, 8, and 14: At Step 1: The claims are directed to a “method”, “medium”, and “system” and thus directed to a statutory category. At Step 2A, Prong One: The claims recite the following limitations directed to an abstract idea: “generating embedding data for the digital video by analyzing video data of the digital video and transcript data of the digital video” as drafted recites a mental process. One can mentally evaluate/judge a video and then create embedding data using pen and paper. “analyzing the embedding data for the digital video to determine mappings between video segments from the digital video and scene types” as drafted recites a mental process. One can mentally evaluate a video and the data associated with it to determine a mapping to a scene type. At Step 2A, Prong Two: The claims recite the following additional elements: That the method and system are performed by a “a computer containing a processor”, “a server system”, “a non-transitory computer-readable medium”, “display”, and “graphical user interface” which is a high-level recitation of a generic computer components and represents mere instructions to apply on a computer as in MPEP 2106.05(f), which does not provide integration into a practical application. “accessing a digital video” is insignificant extra-solution activity. This limitation is recited as receiving data (i.e. mere data gathering). This does not provide integration into a practical application. “providing, for display within a graphical user interface, the video segments from the digital video visually organized based on the scene types” is insignificant extra-solution activity. This limitation is recited as mere outputting of data or providing/presenting data. Specifically, Applicant’s specification [0022], [0023], and [0028] states that the content organization system presents the content of the video files as collection objects that portray video segments grouped by scene type. This does not provide integration into a practical application. Viewing the additional limitations together and the claim as a whole, nothing provides integration into a practical application. At Step 2B: The conclusions for the mere implementation using a computer are carried over and do not provide significantly more. With respect to the “accessing” and “providing” identified as extra-solution activity in Step 2A Prong 2, when re-evaluated as Step 2B this limitation is well-understood, routine, and conventional and remains insignificant extra-solution activity. See MPEP 2106.05(d)(II) “i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network).” To the extent this is a request for “input” and “output” on records that is well-understood, routine and conventional. See MPEP 2106.05(d)(II) “iii. Electronic recordkeeping, Alice Corp. Pty. Ltd. v. CLS Bank Int'l, 573 U.S. 208, 225, 110 USPQ2d 1984 (2014) (creating and maintaining "shadow accounts"); Ultramercial, 772 F.3d at 716, 112 USPQ2d at 1755 (updating an activity log).” Looking at the claims as a whole does not change this conclusion and the claim is ineligible. As to dependent claims 4-7, 11-13, and 17-20: At Step 1: The claims are directed to a “method”, “medium”, and “system” and thus directed to a statutory category. At Step 2A, Prong One: The claims recite the following limitations directed to an abstract idea: “analyzing the digital video to determine a magnitude of visual change for a frame transition between a first frame and a second frame of the digital video; and determining a scene change between the first frame and the second frame of the digital video based on the magnitude of visual change for the frame transition” as drafted recites a mental process. One can mentally evaluate/judge a magnitude of change between frames. “determining mappings between the video segments from the digital video and the scene types comprises determining a mapping between at least one video segment and the user-defined scene type” as drafted recites a mental process. One can mentally evaluate a video and the data associated with it to determine a mapping to a scene type. “merging a first video segment and a second video segment of the video segments of the digital video based on determining that the first video segment and the second video segment are mapped to a common scene type” as drafted recites certain methods of organizing human activity. One can group video segments based on a common scene. “generating a summary video for the digital video by: selecting first video segment mapped to a first scene type; selecting a second video segment mapped to a second scene type; and merging the first video segment with the second video segment to generate the summary video” as drafted recites a mental process. One can mentally evaluate/judge a video in order to create a summary using pen and paper. One can also mentally judge/evaluate a video and group video segments in order to create a summary for the grouping using pen and paper. At Step 2A, Prong Two: The claims recite the following additional elements: That the method and system are performed by a “a computer containing a processor”, “a server system”, “a non-transitory computer-readable medium”, “display”, and “graphical user interface” which is a high-level recitation of a generic computer components and represents mere instructions to apply on a computer as in MPEP 2106.05(f), which does not provide integration into a practical application. “generating a video segment from the digital video based on the determined scene change” is insignificant extra-solution activity. This limitation is recited as mere outputting of data or providing/presenting data. Specifically, Applicant’s specification [0022], [0023], and [0028] states that the content organization system presents the content of the video files as collection objects that portray video segments grouped by scene type. This does not provide integration into a practical application. Viewing the additional limitations together and the claim as a whole, nothing provides integration into a practical application. At Step 2B: The conclusions for the mere implementation using a computer are carried over and do not provide significantly more. With respect to the “generating” identified as extra-solution activity in Step 2A Prong 2, when re-evaluated as Step 2B this limitation is well-understood, routine, and conventional and remains insignificant extra-solution activity. See MPEP 2106.05(d)(II) “i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network).” To the extent this is a request for “input” and “output” on records that is well-understood, routine and conventional. See MPEP 2106.05(d)(II) “iii. Electronic recordkeeping, Alice Corp. Pty. Ltd. v. CLS Bank Int'l, 573 U.S. 208, 225, 110 USPQ2d 1984 (2014) (creating and maintaining "shadow accounts"); Ultramercial, 772 F.3d at 716, 112 USPQ2d at 1755 (updating an activity log).” Looking at the claims as a whole does not change this conclusion and the claim is ineligible. Allowable Subject Matter Claims 2-3, 9-10, and 15-16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Specifically, claims 2, 9, and 15 recite “generating the embedding data for the digital video by analyzing the video data of the digital video and the transcript data of the digital video comprises: utilizing a machine learning model to generate a first set of word vector embeddings from the video data and a second set of word vector embeddings from the transcript data; and generating the embedding data by fusing the first set of word vector embeddings with the second set of word vector embeddings”, which would overcome the 101 abstract idea rejections associated with independent claims 1, 8, and 14 above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. PATEL et al (US 20230162502 A1) - a video segmentation and object tracking in video scenes. A video segmentation system can receive a text input and an input video scene from a user of an intended edit to the input video scene. The video segmentation system can extract text features such as an object and an intended edit action from the text input. The video segmentation system determines a corresponding video edit for the object in the video scene. The video segmentation system identifies keyframes in the video scene that include the object and perform image segmentation and frame referring segmentation. The video segmentation system clusters groups of keyframes that include the object based on a number of keyframes and a threshold proximity of the group of keyframes. The video segmentation system fuses the image segmentation with the frame referring segmentation to produce an output set of fusion masks for the object in the input video. The video segmentation system applies the set of fusion masks to the input video scene. The video segmentation system outputs a masked video scene that includes the fusion masks as applied to the input video scene. BAUGHMAN et al (US 20200092610 A1) - Systems and methods for dynamically providing customized versions of video content are disclosed. In embodiments, a method comprises: analyzing a video to determine content and context of portions of the video; assigning one or more content editing categories to the portions of the video based on the analyzing; determining an unwanted scene of the video based on the one or more content editing categories and user profile data of a viewer; determining a style component of the unwanted scene based on context of the unwanted scene and the user profile data; generating custom content to replace the unwanted scene of the video based on an acceptable portion of content corresponding to the unwanted scene and the style component; editing the video to replace the unwanted scene of the video with the custom content to produce an edited video including the custom content; and providing the edited video to the viewer. Ishtiaq et al (US 20150082349 A1) - A method receives video content and metadata associated with video content. The method then extracts features of the video content based on the metadata. Portions of the visual, audio, and textual features are fused into composite features that include multiple features from the visual, audio, and textual features. A set of video segments of the video content is identified based on the composite features of the video content. Also, the segments may be identified based on a user query. Bedi et al (US 10945040 B1) - The present disclosure relates to methods, systems, and non-transitory computer-readable media for generating a topic visual element for a portion of a digital video based on audio content and visual content of the digital video. For example, the disclosed systems can generate a map between words of the audio content and their corresponding timestamps from the digital video and then modify the map by associating importance weights with one or more of the words. Further, the disclosed systems can generate an additional map by associating words embedded in one or more video frames of the visual content with their corresponding timestamps. Based on these maps, the disclosed systems can identify a topic for a portion of the digital video (e.g., a portion currently previewed on a computing device), generate a topic visual element that includes the topic, and provide the topic visual element for display on a computing device. Zhiwen (US 20220114204 A1) - A method includes accessing an audiovisual composition comprising a target video segment and a source video segment. The method also includes, in response to presence of the target video segment and the source video segment in the audiovisual composition: accessing a first keyword associated with the source video segment; and calculating a first relevance score for the first keyword relative to the target video segment based on a temporal position of the source video segment in the audiovisual composition and a temporal position of the target video segment in the audiovisual composition; accessing a textual query comprising the first keyword. The method additionally includes: generating a first query result based on the textual query, the first query result comprising the target video segment based on the first relevance score; and at a native composition application, rendering a representation of the first query result. Lussier et al (US 20110030031 A1) - Methods and systems for receiving, processing and organizing video. Organizational tools are provided that allow users the capacity to solicit, mine, clip, aggregate, organize, and search submitted footage. These tools include: a set of electronic folders, a media clipper and a media submit portal. Studios, projects, folders and subfolders exist in a hierarchical relationship in order to arrange media. The hierarchy of folders created by producers may be made accessible by the public for the purpose of submitting media to a particular project. Various content creators may upload video to the system, may create electronic video clips from the uploaded video by selecting subportions of that video, and may submit the electronic video clips to a specific folder which is associated with a project. Producers may view various folders to select submitted video clips for use in a project. Patluri et al (US 11869240 B1) - Systems and techniques are generally described for semantically segmenting videos. In various examples, a selection of a first video may be received. A first query to segment the first video into segments related to a first category of content may be received. A first plurality of segments related to the first category may be determined. In some examples, time code data representing the first plurality of segments may be sent to a remote computing device, wherein a video player of the remote computing device is effective to play the first plurality of segments based at least in part on the time code data. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to JARED M BIBBEE whose telephone number is (571)270-1054. The examiner can normally be reached Monday-Thursday 8AM-6PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, APU MOFIZ can be reached at 5712724080. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JARED M BIBBEE/Primary Examiner, Art Unit 2161
Read full office action

Prosecution Timeline

Jun 12, 2025
Application Filed
Jul 22, 2026
Non-Final Rejection mailed — §101 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705242
SYSTEM AND METHOD FOR CACHING OBJECT DATA IN A CLOUD DATABASE SYSTEM
2y 10m to grant Granted Aug 11, 2026
Patent 12694053
Multi-Image Search
1y 4m to grant Granted Jul 28, 2026
Patent 12681948
Data Warehouse System-Based Data Processing Method and Data Warehouse System
1y 1m to grant Granted Jul 14, 2026
Patent 12670211
DETERMINING QUERY COMPLEXITY IN VIDEO QUESTION ANSWERING
1y 5m to grant Granted Jun 30, 2026
Patent 12664126
SUGGESTING CONTENT ITEMS TO BE ACCESSED BY A USER
1y 3m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
94%
With Interview (+13.7%)
3y 0m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 670 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month