DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . A request for continued examination under 37 CFR 1.114, including
the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application
is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has
been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR
1.114.
The Applicant’s Remarks filed 29 April 2026 have been received and considered.
Claims 1 – 22 remain pending.
Claims 1, 3 – 4, 6, 10, 16 – 17, and 19 have been amended.
Claims 1 – 22, all of the claims still pending in this application, remain rejected.
Response to Applicant’s Remarks
Applicant’s arguments were filed 29 April 2026 regarding amendments to independent claims 1 and 16. Applicant argues that Han fails to teach “receiving user data associated with a user, the user data comprising a reference video segment associated with a user.”. Applicant further argues that Han fails to teach “generating a score for each feature vector based on comparing each feature vector associated with a local neighborhood with the one or more reference feature vectors from the user data.” However, Applicant has not provided any rationale or evidence to support this. Applicant states Han appears to “generally recite an interface with a display for viewing videos and other content, and describes a client device that allows a user to interact with and consume digital content, such as video highlights. Notably, [0019] does not describe receiving any "user data" in the sense recited in the currently amended claims, much less "user data comprising a reference video segment associated with the user," or "generating one or more reference feature vectors from the reference video segment." Rather, [0019] merely describes conventional user interaction with a device (e.g., viewing, selecting, and consuming content).”. Examiner disagrees with the Applicant’s previously mentioned arguments. In the Specification, Applicant defines the user data as:
In some embodiments, the user data can be indicative of a user's preference for video segments. That is, the user data preferably reflects a given user's likes and dislikes. For example, if a user likes to see close up of football goals, several video segments depicting those can be used as user data. The more personalised and detailed user data is made, the better the resulting predicted highlight segment will reflect the user's preferences.
In some embodiments, the user data can comprise at least one reference video segment. In some such embodiments, the reference video segment can be selected by the user. In some such embodiments, the method can further comprise receiving a plurality of user-selected video segments indicative of user preference and generating user data based on them. If a given user inputs their own preferences by specifically submitting video segments they have previously enjoyed, the resulting personalised highlight segment can be fairly accurate and exhibit a large degree of personalisation, since the user may know best what they would prefer.
In other embodiments where the user data comprises reference video segments, the reference video segment can be automatically generated based on a user's video viewing habits. This can be useful, as the user may not want to manually input their likes/dislikes for video segments, and may instead prefer to let them be automatically collected and/or curated.
In other such embodiments, the reference video segment can be generated based on viewing habits of a plurality of reference users. There may be a database storing a plurality of video segments reflecting average user preferences. The user data can be generated or created based on this database, provided some further information is known about the given user (so that relevant members of the database can be selected to serve as user data).
The Examiner maintains that Han does indeed teach these limitations, specifically in the Paragraph 19 of the specification, that the Applicant also pointed to. Han teaches the user can record video, consume content, browse websites, etc…with the user being able to interact with the device for functions such as viewing, selecting, and viewing sports clips. Inputs, clicks, personally recorded videos, and selecting preferred clips would be considered user data or user preference. If applicant believes that the user data should be defined in a different manner, Examiner recommends that the Applicant amend the claim to specifically define what should be interpreted as user data while also pointing to the section in the specification that enables this interpretation. Han also teaches the preferred clips (in this case various sports clips or reference videos) being categorized into classes and generating category pair-wise feature vectors, which are used to then generate a score for each frame of the clip.
The video used by the trained feature model also requires user data/user input (reference video data) for comparison when detecting highlights in other user selected videos. Han receives user data, when user data is needed in order to determine highlights in a video that is preferred by respective user, i.e specifically focusing on sports various clips in this case.
Applicant also argues that Han does not teach “generating one or more reference feature vectors from the reference video segment" or "generating a score for each feature vector based on comparing each feature vector associated with a local neighborhood with the one or more reference feature vectors from the user data by determining a distance between each feature vector associated with a local neighborhood and a reference feature vector from the one or more reference feature vectors.”. Examiner also disagrees with this remark. Applicant defines local neighborhood as being at least one frame in a sequence of frames. Han teaches each video frame of an input video receiving a highlight score using the process mentioned above and further elaborated in Paragraphs [0043 – 0046] which correlates to Figures 7A – 7C, showing how the determined highlight scores of each frame of the video are displayed to the user device.
Never the less, in view of the Applicant’s arguments previously mentioned, the previously applied prior art rejections are withdrawn. Applicant's arguments are rendered moot in view of the new grounds of rejection set forth below.
Claim Rejections - 35 USC § 112
Claims 16 – 22 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 16, claim 16 is an independent system claim for selecting highlight segments. However, the recited claim language is merely a series of steps with no apparent “system”. It fails to mention the necessary components needed to function properly (i.e. memory storing a program, processor, display, etc.). It is unclear what the system comprises without these necessary components, yet alone display results to a user without mentioning the use of a display. Therefore, claim 16 has been rejected.
Claims 17 – 22 are rejected by virtue of their dependency on claim 16.
Claim 22 recites the limitation "configured to display" in the claim language. There is insufficient antecedent basis for this limitation in the claim (Refer to the above 112b rejection of claim 16).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 – 22 are rejected under 35 U.S.C. 103 as being unpatentable over US Publication No. 2016/0292510 A1 to Han et al. (hereinafter Han) in view of US Publication No. 2017/0109584 A1 to Yao et al. (hereinafter Yao).
Examiner note: Although no longer required by in the claim language, Applicant uses a trained model to generate feature vectors. The training of this model is defined in the specification: “On the other hand, training one model per user is inefficient and requires large amounts of personal information which is typically not available. To overcome these limitations, we present a global ranking model which can condition on a particular user's interests. Rather than training one model per user, our model is personalized via its inputs, which allows it to effectively adapt to its predictions, given only a few user-specific examples. To train this model, we create a large- scale dataset of users and the GIFs they created, giving us an accurate indication of their interests.”. Han also uses a model trained on a large dataset (clips from various sports, as example), further disclosed in the following rejections. This is a key component needed to understand the Examiner’s reasons to combine.
Claim 1
Regarding Claim 1, an independent method claim, Han teaches a computer-implemented method for selecting a highlight segment (Abstract).
Although it’s implied that the dataset used to train the model would be that of the same interest of the user(s) “The training phase 510 has two sub-phases: feature model training based on a large corpus of video training data 502, and highlight detection model training based on a subset 504 of the large corpus of video training data.”, Paragraph [0048], Han does not explicitly teach or suggest receiving user, the user data comprising a reference video segment associated with the user.
However, Zao teaches receiving user, the user data comprising a reference video segment associated with the user (“User interface module 214 can interact with I/O interfaces(s) 110. User interface module 214 can present a graphical user interface (GUI) at I/O interface 110. GUI can include features for allowing a user to interact with training module 208, highlight detection module 210, video output module 212 or components of video highlight engine 128. Features of the GUI can allow a user to train neural network(s), select video for analysis and view summarization of analyzed video at consumer device 104.”), where it would be obvious to one skilled in the art that the user can provide the training model with a plurality of preferred videos (user reference video segments associated with the user) as training material.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Han to incorporate training the model with a plurality of user preferred video segments (user data), as disclosed by Yao. The suggestion/motivation for doing so would have been to allow the training model to generate reference feature vectors that are more tailored to a user’s preferred viewing material for better video highlight selections.
Han, in view of Yao, further teaches generating one or more reference feature vectors from the reference video segment ("The feature training module 310 classifies the sports videos stored in the video database 132 into different classes and generates feature vectors associated with each class of the sports videos.", Paragraph [0029], wherein these would be the “reference” feature vectors associated with each video segment in their respective classes.
receiving a sequence of frames corresponding to video content (“The interface module 410 also receives an input video received by the client device, e.g., a mountain biking activity recorded by a mobile phone of a user or streamed from a video streaming service, and stores the received input video in the frame buffer 402.”, Paragraph [0041]);
selecting a local neighborhood for each frame of the sequence of frames, each local neighborhood comprising at least one frame from the sequence of frames (“The feature extraction module 420 extracts visual features from the frames of the input video.”, Paragraph [0042]), wherein each frame of the input video’s sequence of frames is its own local neighborhood;
converting each local neighborhood into a feature vector (“The features from the convolution layers are normalized and combined, e.g., by liner embedding, to generate a feature vector for the frame of the sports video.”, Paragraph [0042]), wherein this is implicitly done for each frame on the input video;
generatingthe one or more reference feature vectors from the user data by determining a distance between each feature vector associated with a local neighborhood and a reference feature vector from the one or more reference feature vectors (“To detect video highlights in a video frame of the input video, the highlight detection module 430 applies the highlight detection model trained by the training module 136 to the feature vector associated with the video frame. In one embodiment, the highlight detection module 430 compares the feature vector with pair-wise frame feature vectors to determine the similarity between the feature vector associated with the video frame and the feature vector of the pair-wise frame feature vectors representing a video highlight. For example, the highlight detection module 430 computes a Euclidean distance between the feature vector associated with the video frame and the pair-wise feature vector representing a video highlight. Based on the comparison, the highlight detection module 430 computes a highlight score for the video frame.”, Paragraph [0043]).
generating at least one highlight segment comprising a plurality of frames from the sequence of frames based on evaluating each score for each feature vector (“The highlight detection module 430 repeats the similar detection process to each video frame of the input video and generates a highlight score for each video frame of the input video.” Paragraph [0044]; “In the example shown in FIG. 7B, the horizontal axis of the graphical user interface shows the frame identification 720 of the video frames of the input video; the vertical axis shows the corresponding highlight scores 710 of the video frames of the input video. The example shown in FIG. 7B further shows a graph of highlight scores for 6 identified videos frames, i.e., 30.sup.th, 60.sup.th, 90.sup.th, 120.sup.th, 150.sup.th and 180.sup.th frame, of the input video, where the 60.sup.th frame has the highest highlight score 730 and the video segment between the 30.sup.th frame and 60.sup.th frame is likely to represent a video highlight of the input video.”, Paragraph [0046]); and
providing the highlight segment to the user for display (Figure 7C; “The video segment between the 30.sup.th frame and 60.sup.th frame is presented as a video highlight predicted by the highlight detection module 430 to the users of the client device.”, Paragraph [0046]).
PNG
media_image1.png
681
501
media_image1.png
Greyscale
Claim 2
Regarding Claim 2, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1.
Han, in view of Yao, further teaches further comprising generating and maintaining a database of video segments and selecting at least one video segment as the user data based on at least one characteristic associated with the user (“The model training module 320 stores the trained video highlight detection model and pair-wise frame features in the model database 134.”, Paragraph [0037]), where the user trained model would incorporate preferred user video highlights, in this case, the user’s interest (characteristic) in various sports clips.
Claim 3
Regarding Claim 3, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1.
Han, in view of Yao, further teaches wherein receiving the user data comprises receiving (Rejected as applied to claim 1), wherein it is obvious to one skilled in the art that each video segment (reference video segment) used to train the model by the user would be video segments indicative of the user’s preference, i.e sports clips highlights.
Claim 4
Regarding Claim 4, dependent on claim 1, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 1.
Han, in view of Yao, further teaches receiving the user data comprises receiving (Rejected as applied to claim 1);
and generating the one or more reference feature vectors comprises generating a plurality of reference feature vectors from the plurality of reference video segments (Rejected as applied to claim 1).
Claim 5
Regarding Claim 5, dependent on claim 4, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 4.
Han, in view of Yao, further teaches wherein the plurality of reference video segments are indicative of different user preferences of the user and wherein the plurality of reference video segments are grouped into sets, each set indicative of a particular user preference, and wherein each set is converted into a distinct user data subset comprising a subset of the plurality of reference feature vectors associated with the plurality of reference video segments forming part of it (“Based on the training, the feature training module 310 classifies the sports videos stored in the video database 132 into different classes. For example, the sports videos stored in the video database 132 are classified by the feature training module 310 into classes, e.g., cycling, American football, soccer, table tennis/ping pong, tennis and basketball. “, Paragraph [0033]; “Based on the training, the feature training module 310 generates frame-based feature vectors associated with each class of the sports video.”, Paragraph [0034]).
Claim 8
Regarding Claim 8, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1.
Han, in view of Yao, further teaches prior to selecting the local neighborhood for each frame, generating at least one segment, each segment comprising at least one frame of the sequence of frames (Rejected as applied to claim 1), where each frame of the sequence of frames is its own segment and its own neighborhood.
Claim 9
Regarding Claim 9, dependent on claim 8, Han, in view of Yao, teaches the invention as claimed in claim 8.
Han, in view of Yao, further teaches wherein each local neighborhood is comprised within a single segment (Rejected as applied to claim 8).
Claim 10
Regarding Claim 10, dependent on claim 4, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 4.
Han, in view of Yao, further teaches wherein generating the score for each feature vector comprises: comparing each feature vector associated with a local neighborhood with each reference feature vector of the[[a]] plurality of reference feature vectors associated with the user data (“To detect video highlights in a video frame of the input video, the highlight detection module 430 applies the highlight detection model trained by the training module 136 to the feature vector associated with the video frame. In one embodiment, the highlight detection module 430 compares the feature vector with pair-wise frame feature vectors to determine the similarity between the feature vector associated with the video frame and the feature vector of the pair-wise frame feature vectors representing a video highlight.”, Paragraph [0043]); and
assigning scores to each local neighborhood based on a difference with respect to closest matching of each feature vector and the plurality of reference feature vectors (“Based on the comparison, the highlight detection module 430 computes a highlight score for the video frame.”, Paragraph [0043]; “The highlight detection module 430 repeats the similar detection process to each video frame of the input video and generates a highlight score for each video frame of the input video. A larger highlight score of a video frame indicates a higher likelihood that the video frame has a video highlight than another video frame having a smaller highlight score.”, Paragraph [0044]).
Claim 11
Regarding Claim 11, dependent on claim 5, Han, in view of Yao, teaches (As Best Understood) the invention as claimed in claim 5.
Han, in view of Yao, further teaches wherein generating the score for each feature vector further comprises determining which distinct user data subset is closest to each feature vector and assigning it a value based on a comparison between the subset of the plurality of reference feature vectors and each feature vector (Rejected as applied to claim 10).
Claim 13
Regarding Claim 13, dependent on claim 1, Han, in view of Yao, teaches the invention as claimed in claim 1.
Han, in view of Yao, further teaches constructing the highlight segment, wherein the highlight segment comprises at least one local neighborhood (Figure 7A – 7C; “In the example shown in FIG. 7B, the horizontal axis of the graphical user interface shows the frame identification 720 of the video frames of the input video; the vertical axis shows the corresponding highlight scores 710 of the video frames of the input video. The example shown in FIG. 7B further shows a graph of highlight scores for 6 identified videos frames, i.e., 30.sup.th, 60.sup.th, 90.sup.th, 120.sup.th, 150.sup.th and 180.sup.th frame, of the input video, where the 60.sup.th frame has the highest highlight score 730 and the video segment between the 30.sup.th frame and 60.sup.th frame is likely to represent a video highlight of the input video. The video segment between the 30.sup.th frame and 60.sup.th frame is presented as a video highlight predicted by the highlight detection module 430 to the users of the client device.”, Paragraph [0046]).
Claim 14
Regarding Claim 14, dependent on claim 13, Han, in view of Yao, teaches the invention as claimed in claim 13.
Han, in view of Yao, further teaches wherein constructing the highlight segment comprises evaluating each score for each feature vector corresponding to each frame and their neighboring frames and identifying a plurality of neighboring frames with an average best score (Rejected as applied to claim 13).
Claim 15
Regarding Claim 15, dependent on claim 13, Han, in view of Yao, teaches the invention as claimed in claim 13.
Han, in view of Yao, further teaches constructing a plurality of highlight segments, each comprising a plurality of frames selected from the sequence of frames, and corresponding to a plurality of distinct neighboring frames with an average highest score (Rejected as applied to claim 13), wherein the client would be able to see all of the video segments that had the highest scores.
Claim 16, an independent system claim, is rejected for the same reasons as applied to claim 1.
Claims 17 – 22 are rejected for the same reasons as applied to the above claims.
Allowable Subject Matter
Claims 6 – 7 and 12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ronde Miller whose telephone number is (703) 756-5686 The examiner can normally be reached Monday-Friday 8:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Gregory Morse can be reached on (571) 272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RONDE LEE MILLER/Examiner, Art Unit 2663
/GREGORY A MORSE/Supervisory Patent Examiner, Art Unit 2698