Prosecution Insights
Last updated: October 04, 2026
Application No. 18/001,174

METHOD AND SYSTEM FOR SELECTING HIGHLIGHT SEGMENTS

Non-Final OA §103§112
Filed
Dec 08, 2022
Priority
Jun 08, 2020 — EU 20178817.1 +1 more
Examiner
MILLER, RONDE LEE
Art Unit
2663
Tech Center
2600 — Communications
Assignee
Dropbox Inc.
OA Round
3 (Non-Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
28 granted / 38 resolved
+11.7% vs TC avg
Strong +22% interview lift
Without
With
+22.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
15 currently pending
Career history
57
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
51.5%
+11.5% vs TC avg
§102
18.4%
-21.6% vs TC avg
§112
17.9%
-22.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 38 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. The Applicant’s Remarks filed 29 April 2026 have been received and considered. Claims 1 – 22 remain pending. Claims 1, 3 – 4, 6, 10, 16 – 17, and 19 have been amended. Claims 1 – 22, all of the claims still pending in this application, remain rejected. Response to Applicant’s Remarks Applicant’s arguments were filed 29 April 2026 regarding amendments to independent claims 1 and 16. Applicant argues that Han fails to teach “receiving user data associated with a user, the user data comprising a reference video segment associated with a user.”. Applicant further argues that Han fails to teach “generating a score for each feature vector based on comparing each feature vector associated with a local neighborhood with the one or more reference feature vectors from the user data.” However, Applicant has not provided any rationale or evidence to support this. Applicant states Han appears to “generally recite an interface with a display for viewing videos and other content, and describes a client device that allows a user to interact with and consume digital content, such as video highlights. Notably, [0019] does not describe receiving any "user data" in the sense recited in the currently amended claims, much less "user data comprising a reference video segment associated with the user," or "generating one or more reference feature vectors from the reference video segment." Rather, [0019] merely describes conventional user interaction with a device (e.g., viewing, selecting, and consuming content).”. Examiner disagrees with the Applicant’s previously mentioned arguments. In the Specification, Applicant defines the user data as: In some embodiments, the user data can be indicative of a user's preference for video segments. That is, the user data preferably reflects a given user's likes and dislikes. For example, if a user likes to see close up of football goals, several video segments depicting those can be used as user data. The more personalised and detailed user data is made, the better the resulting predicted highlight segment will reflect the user's preferences. In some embodiments, the user data can comprise at least one reference video segment. In some such embodiments, the reference video segment can be selected by the user. In some such embodiments, the method can further comprise receiving a plurality of user-selected video segments indicative of user preference and generating user data based on them. If a given user inputs their own preferences by specifically submitting video segments they have previously enjoyed, the resulting personalised highlight segment can be fairly accurate and exhibit a large degree of personalisation, since the user may know best what they would prefer. In other embodiments where the user data comprises reference video segments, the reference video segment can be automatically generated based on a user's video viewing habits. This can be useful, as the user may not want to manually input their likes/dislikes for video segments, and may instead prefer to let them be automatically collected and/or curated. In other such embodiments, the reference video segment can be generated based on viewing habits of a plurality of reference users. There may be a database storing a plurality of video segments reflecting average user preferences. The user data can be generated or created based on this database, provided some further information is known about the given user (so that relevant members of the database can be selected to serve as user data). The Examiner maintains that Han does indeed teach these limitations, specifically in the Paragraph 19 of the specification, that the Applicant also pointed to. Han teaches the user can record video, consume content, browse websites, etc…with the user being able to interact with the device for functions such as viewing, selecting, and viewing sports clips. Inputs, clicks, personally recorded videos, and selecting preferred clips would be considered user data or user preference. If applicant believes that the user data should be defined in a different manner, Examiner recommends that the Applicant amend the claim to specifically define what should be interpreted as user data while also pointing to the section in the specification that enables this interpretation. Han also teaches the preferred clips (in this case various sports clips or reference videos) being categorized into classes and generating category pair-wise feature vectors, which are used to then generate a score for each frame of the clip. The video used by the trained feature model also requires user data/user input (reference video data) for comparison when detecting highlights in other user selected videos. Han receives user data, when user data is needed in order to determine highlights in a video that is preferred by respective user, i.e specifically focusing on sports various clips in this case. Applicant also argues that Han does not teach “generating one or more reference feature vectors from the reference video segment" or "generating a score for each feature vector based on comparing each feature vector associated with a local neighborhood with the one or more reference feature vectors from the user data by determining a distance between each feature vector associated with a local neighborhood and a reference feature vector from the one or more reference feature vectors.”. Examiner also disagrees with this remark. Applicant defines local neighborhood as being at least one frame in a sequence of frames. Han teaches each video frame of an input video receiving a highlight score using the process mentioned above and further elaborated in Paragraphs [0043 – 0046] which correlates to Figures 7A – 7C, showing how the determined highlight scores of each frame of the video are displayed to the user device. Never the less, in view of the Applicant’s arguments previously mentioned, the previously applied prior art rejections are withdrawn. Applicant's arguments are rendered moot in view of the new grounds of rejection set forth below. Claim Rejections - 35 USC § 112 Claims 16 – 22 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 16, claim 16 is an independent system claim for selecting highlight segments. However, the recited claim language is merely a series of steps with no apparent “system”. It fails to mention the necessary components needed to function properly (i.e. memory storing a program, processor, display, etc.). It is unclear what the system comprises without these necessary components, yet alone display results to a user without mentioning the use of a display. Therefore, claim 16 has been rejected. Claims 17 – 22 are rejected by virtue of their dependency on claim 16. Claim 22 recites the limitation "configured to display" in the claim language. There is insufficient antecedent basis for this limitation in the claim (Refer to the above 112b rejection of claim 16). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 – 22 are rejected under 35 U.S.C. 103 as being unpatentable over US Publication No. 2016/0292510 A1 to Han et al. (hereinafter Han) in view of US Publication No. 2017/0109584 A1 to Yao et al. (hereinafter Yao). Examiner note: Although no longer required by in the claim language, Applicant uses a trained model to generate feature vectors. The training of this model is defined in the specification: “On the other hand, training one model per user is inefficient and requires large amounts of personal information which is typically not available. To overcome these limitations, we present a global ranking model which can condition on a particular user's interests. Rather than training one model per user, our model is personalized via its inputs, which allows it to effectively adapt to its predictions, given only a few user-specific examples. To train this model, we create a large- scale dataset of users and the GIFs they created, giving us an accurate indication of their interests.”. Han also uses a model trained on a large dataset (clips from various sports, as example), further disclosed in the following rejections. This is a key component needed to understand the Examiner’s reasons to combine. Claim 1 Regarding Claim 1, an independent method claim, Han teaches a computer-implemented method for selecting a highlight segment (Abstract). Although it’s implied that the dataset used to train the model would be that of the same interest of the user(s) “The training phase 510 has two sub-phases: feature model training based on a large corpus of video training data 502, and highlight detection model training based on a subset 504 of the large corpus of video training data.”, Paragraph [0048], Han does not explicitly teach or suggest receiving user, the user data comprising a reference video segment associated with the user. However, Zao teaches receiving user, the user data comprising a reference video segment associated with the user (“User interface module 214 can interact with I/O interfaces(s) 110. User interface module 214 can present a graphical user interface (GUI) at I/O interface 110. GUI can include features for allowing a user to interact with training module 208, highlight detection module 210, video output module 212 or components of video highlight engine 128. Features of the GUI can allow a user to train neural network(s), select video for analysis and view summarization of analyzed video at consumer device 104.”), where it would be obvious to one skilled in the art that the user can provide the training model with a plurality of preferred videos (user reference video segments associated with the user) as training material. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Han to incorporate training the model with a plurality of user preferred video segments (user data), as disclosed by Yao. The suggestion/motivation for doing so would have been to allow the training model to generate reference feature vectors that are more tailored to a user’s preferred viewing material for better video highlight selections. Han, in view of Yao, further teaches generating one or more reference feature vectors from the reference video segment ("The feature training module 310 classifies the sports videos stored in the video database 132 into different classes and generates feature vectors associated with each class of the sports videos.", Paragraph [0029], wherein these would be the “reference” feature vectors associated with each video segment in their respective classes. receiving a sequence of frames corresponding to video content (“The interface module 410 also receives an input video received by the client device, e.g., a mountain biking activity recorded by a mobile phone of a user or streamed from a video streaming service, and stores the received input video in the frame buffer 402.”, Paragraph [0041]); selecting a local neighborhood for each frame of the sequence of frames, each local neighborhood comprising at least one frame from the sequence of frames (“The feature extraction module 420 extracts visual features from the frames of the input video.”, Paragraph [0042]), wherein each frame of the input video’s sequence of frames is its own local neighborhood; converting each local neighborhood into a feature vector (“The features from the convolution layers are normalized and combined, e.g., by liner embedding, to generate a feature vector for the frame of the sports video.”, Paragraph [0042]), wherein this is implicitly done for each frame on the input video; generatingthe one or more reference feature vectors from the user data by determining a distance between each feature vector associated with a local neighborhood and a reference feature vector from the one or more reference feature vectors (“To detect video highlights in a video frame of the input video, the highlight detection module 430 applies the highlight detection model trained by the training module 136 to the feature vector associated with the video frame. In one embodiment, the highlight detection module 430 compares the feature vector with pair-wise frame feature vectors to determine the similarity between the feature vector associated with the video frame and the feature vector of the pair-wise frame feature vectors representing a video highlight. For example, the highlight detection module 430 computes a Euclidean distance between the feature vector associated with the video frame and the pair-wise feature vector representing a video highlight. Based on the comparison, the highlight detection module 430 computes a highlight score for the video frame.”, Paragraph [0043]). generating at least one highlight segment comprising a plurality of frames from the sequence of frames based on evaluating each score for each feature vector (“The highlight detection module 430 repeats the similar detection process to each video frame of the input video and generates a highlight score for each video frame of the input video.” Paragraph [0044]; “In the example shown in FIG. 7B, the horizontal axis of the graphical user interface shows the frame identification 720 of the video frames of the input video; the vertical axis shows the corresponding highlight scores 710 of the video frames of the input video. The example shown in FIG. 7B further shows a graph of highlight scores for 6 identified videos frames, i.e., 30.sup.th, 60.sup.th, 90.sup.th, 120.sup.th, 150.sup.th and 180.sup.th frame, of the input video, where the 60.sup.th frame has the highest highlight score 730 and the video segment between the 30.sup.th frame and 60.sup.th frame is likely to represent a video highlight of the input video.”, Paragraph [0046]); and providing the highlight segment to the user for display (Figure 7C; “The video segment between the 30.sup.th frame and 60.sup.th frame is presented as a video highlight predicted by the highlight detection module 430 to the users of the client device.”, Paragraph [0046]). PNG media_image1.png 681 501 media_image1.png Greyscale Claim 2 Regarding Claim 2, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1. Han, in view of Yao, further teaches further comprising generating and maintaining a database of video segments and selecting at least one video segment as the user data based on at least one characteristic associated with the user (“The model training module 320 stores the trained video highlight detection model and pair-wise frame features in the model database 134.”, Paragraph [0037]), where the user trained model would incorporate preferred user video highlights, in this case, the user’s interest (characteristic) in various sports clips. Claim 3 Regarding Claim 3, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1. Han, in view of Yao, further teaches wherein receiving the user data comprises receiving (Rejected as applied to claim 1), wherein it is obvious to one skilled in the art that each video segment (reference video segment) used to train the model by the user would be video segments indicative of the user’s preference, i.e sports clips highlights. Claim 4 Regarding Claim 4, dependent on claim 1, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 1. Han, in view of Yao, further teaches receiving the user data comprises receiving (Rejected as applied to claim 1); and generating the one or more reference feature vectors comprises generating a plurality of reference feature vectors from the plurality of reference video segments (Rejected as applied to claim 1). Claim 5 Regarding Claim 5, dependent on claim 4, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 4. Han, in view of Yao, further teaches wherein the plurality of reference video segments are indicative of different user preferences of the user and wherein the plurality of reference video segments are grouped into sets, each set indicative of a particular user preference, and wherein each set is converted into a distinct user data subset comprising a subset of the plurality of reference feature vectors associated with the plurality of reference video segments forming part of it (“Based on the training, the feature training module 310 classifies the sports videos stored in the video database 132 into different classes. For example, the sports videos stored in the video database 132 are classified by the feature training module 310 into classes, e.g., cycling, American football, soccer, table tennis/ping pong, tennis and basketball. “, Paragraph [0033]; “Based on the training, the feature training module 310 generates frame-based feature vectors associated with each class of the sports video.”, Paragraph [0034]). Claim 8 Regarding Claim 8, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1. Han, in view of Yao, further teaches prior to selecting the local neighborhood for each frame, generating at least one segment, each segment comprising at least one frame of the sequence of frames (Rejected as applied to claim 1), where each frame of the sequence of frames is its own segment and its own neighborhood. Claim 9 Regarding Claim 9, dependent on claim 8, Han, in view of Yao, teaches the invention as claimed in claim 8. Han, in view of Yao, further teaches wherein each local neighborhood is comprised within a single segment (Rejected as applied to claim 8). Claim 10 Regarding Claim 10, dependent on claim 4, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 4. Han, in view of Yao, further teaches wherein generating the score for each feature vector comprises: comparing each feature vector associated with a local neighborhood with each reference feature vector of the[[a]] plurality of reference feature vectors associated with the user data (“To detect video highlights in a video frame of the input video, the highlight detection module 430 applies the highlight detection model trained by the training module 136 to the feature vector associated with the video frame. In one embodiment, the highlight detection module 430 compares the feature vector with pair-wise frame feature vectors to determine the similarity between the feature vector associated with the video frame and the feature vector of the pair-wise frame feature vectors representing a video highlight.”, Paragraph [0043]); and assigning scores to each local neighborhood based on a difference with respect to closest matching of each feature vector and the plurality of reference feature vectors (“Based on the comparison, the highlight detection module 430 computes a highlight score for the video frame.”, Paragraph [0043]; “The highlight detection module 430 repeats the similar detection process to each video frame of the input video and generates a highlight score for each video frame of the input video. A larger highlight score of a video frame indicates a higher likelihood that the video frame has a video highlight than another video frame having a smaller highlight score.”, Paragraph [0044]). Claim 11 Regarding Claim 11, dependent on claim 5, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 5. Han, in view of Yao, further teaches wherein generating the score for each feature vector further comprises determining which distinct user data subset is closest to each feature vector and assigning it a value based on a comparison between the subset of the plurality of reference feature vectors and each feature vector (Rejected as applied to claim 10). Claim 13 Regarding Claim 13, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1. Han, in view of Yao, further teaches constructing the highlight segment, wherein the highlight segment comprises at least one local neighborhood (Figure 7A – 7C; “In the example shown in FIG. 7B, the horizontal axis of the graphical user interface shows the frame identification 720 of the video frames of the input video; the vertical axis shows the corresponding highlight scores 710 of the video frames of the input video. The example shown in FIG. 7B further shows a graph of highlight scores for 6 identified videos frames, i.e., 30.sup.th, 60.sup.th, 90.sup.th, 120.sup.th, 150.sup.th and 180.sup.th frame, of the input video, where the 60.sup.th frame has the highest highlight score 730 and the video segment between the 30.sup.th frame and 60.sup.th frame is likely to represent a video highlight of the input video. The video segment between the 30.sup.th frame and 60.sup.th frame is presented as a video highlight predicted by the highlight detection module 430 to the users of the client device.”, Paragraph [0046]). Claim 14 Regarding Claim 14, dependent on claim 13, Han, in view of Yao, teaches the invention as claimed in claim 13. Han, in view of Yao, further teaches wherein constructing the highlight segment comprises evaluating each score for each feature vector corresponding to each frame and their neighboring frames and identifying a plurality of neighboring frames with an average best score (Rejected as applied to claim 13). Claim 15 Regarding Claim 15, dependent on claim 13, Han, in view of Yao, teaches the invention as claimed in claim 13. Han, in view of Yao, further teaches constructing a plurality of highlight segments, each comprising a plurality of frames selected from the sequence of frames, and corresponding to a plurality of distinct neighboring frames with an average highest score (Rejected as applied to claim 13), wherein the client would be able to see all of the video segments that had the highest scores. Claim 16, an independent system claim, is rejected for the same reasons as applied to claim 1. Claims 17 – 22 are rejected for the same reasons as applied to the above claims. Allowable Subject Matter Claims 6 – 7 and 12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ronde Miller whose telephone number is (703) 756-5686 The examiner can normally be reached Monday-Friday 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Gregory Morse can be reached on (571) 272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RONDE LEE MILLER/Examiner, Art Unit 2663 /GREGORY A MORSE/Supervisory Patent Examiner, Art Unit 2698
Read full office action

Prosecution Timeline

Show 3 earlier events
Oct 22, 2025
Examiner Interview Summary
Oct 28, 2025
Response Filed
Feb 17, 2026
Final Rejection mailed — §103, §112
Mar 27, 2026
Interview Requested
Apr 08, 2026
Examiner Interview Summary
Apr 29, 2026
Request for Continued Examination
May 05, 2026
Response after Non-Final Action
Aug 10, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749180
SYSTEMS AND METHODS FOR THE AUTOMATED DETECTION OF CEREBRAL MICROBLEEDS USING 3T MRI
3y 7m to grant Granted Sep 29, 2026
Patent 12749302
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, COMPUTER PROGRAM, AND SENSOR APPARATUS
3y 2m to grant Granted Sep 29, 2026
Patent 12740547
METHOD AND DEVICE FOR NON-DESTRUCTIVE SORTING OF FERTILIZATION INFORMATION OF HATCHING EGGS BEFORE INCUBATION
2y 3m to grant Granted Sep 22, 2026
Patent 12725233
AUTOMATIC DETERMINATION OF THE PRESENCE OF BURN-IN OVERLAY IN VIDEO IMAGERY
2y 8m to grant Granted Sep 01, 2026
Patent 12694506
DEFECT DETECTING DEVICE AND METHOD
3y 1m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
96%
With Interview (+22.1%)
2y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 38 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month