Prosecution Insights
Last updated: October 04, 2026
Application No. 19/324,761

VIDEO SUMMARIZATION METHOD, APPARATUS, COMPUTER DEVICE, COMPUTER-READABLE STORAGE MEDIUM AND COMPUTER PROGRAM PRODUCT

Non-Final OA §101§103§112
Filed
Sep 10, 2025
Priority
Oct 29, 2024 — CN 202411527505.3
Examiner
FAN, HUA
Art Unit
2426
Tech Center
2400 — Computer Networks
Assignee
Shanghai Hode Information Technology Co. Ltd.
OA Round
1 (Non-Final)
70%
Grant Probability
Favorable
1-2
OA Rounds
2y 10m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
549 granted / 787 resolved
+11.8% vs TC avg
Strong +21% interview lift
Without
With
+21.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
31 currently pending
Career history
810
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
40.4%
+0.4% vs TC avg
§102
18.2%
-21.8% vs TC avg
§112
21.3%
-18.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 787 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This office action is in response to communication filed 9/10/2025. Claims 1-20 are pending for examination, the rejection cited as stated below. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter as follows. Claim 20 defines a “a computer program product, comprising a computer program.” The body of the claim lacks definite structure indicative of a physical apparatus. Therefore, the claim as a whole appears to be nothing more than a “system” of software elements, thus defining functional descriptive material per se. Functional descriptive material may be statutory if it resides on a non-transitory “computer-readable medium or computer-readable memory”. The claim(s) indicated above lack structure, and do not define a non-transitory computer readable medium and are thus non-statutory for that reason (i.e., “When functional descriptive material is recorded on some computer-readable medium it becomes structurally and functionally interrelated to the medium and will be statutory in most cases since use of technology permits the function of the descriptive material to be realized” – Guidelines Annex IV). The scope of the presently claimed invention encompasses products that are not necessarily computer readable, and thus NOT able to impart any functionality of the recited program. The examiner suggests: Adding structure to the body of the claim that would clearly define a statutory apparatus. Any amendment to the claim should be commensurate with its corresponding disclosure. Claim Rejections - 35 USC § 112 4. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. 5. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. 6. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. a) Claim 1 recites “the video content summary” which lacks sufficient antecedent basis, because none of the preceding limitation(s) recites “a video content summary”. It is unclear whether and how the claimed “the video content summary” relates to the previously recited “a plurality of video content summaries”. Applicant is required to clarify. For the sake of the examination, Examiner interprets as any relationship and any video content summary. Claims 2-20 are similarly rejected. b) Claim 11 recites “the method according to claim 1, wherein the displaying a video summary area in response to selecting the video summary control comprises: obtaining the video summary list and the subtitle list in response to selecting the video summary control; when the video summary list and the subtitle list are obtained, displaying the video summary area, and displaying the video summary list and the subtitle list by using the video summary area, and when the video summary list and the subtitle list are not obtained, obtaining a basic content list of the target video, wherein the basic content list comprises one or more of manuscript information, a comment, and a bullet-screen comment; and displaying the video summary area, and displaying the basic content list by using the video summary area”. The recited “the subtitle list” lacks sufficient antecedent basis, because neither this claim nor the parent claim 1 recites “a subtitle list”. Applicant is required to clarify. For the sake of the examination, Examiner presumes that the claimed limitations read: “the method according to claim 1, wherein the displaying a video summary area in response to selecting the video summary control comprises: obtaining the video summary list in response to selecting the video summary control; when the video summary list are obtained, displaying the video summary area, and displaying the video summary list by using the video summary area, and when the video summary list are not obtained, obtaining a basic content list of the target video, wherein the basic content list comprises one or more of manuscript information, a comment, and a bullet-screen comment; and displaying the video summary area, and displaying the basic content list by using the video summary area.” Claim Rejections - 35 USC § 103 7. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 8. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 9. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 10. Claims 1, 10-12, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over LANZA et al (US 2024/0357060 here after LANZA) in view of Hu et al (CN116992079A, Google Patent translation is relied upon, hereafter Hu) and Chen (CN 114845152A, Google Patent translation is relied upon). As to claim 1, LANZA discloses a video summarization method, comprising: displaying a playback page, wherein the playback page comprises a player and a video summary control, the player is configured to play a target video, and the target video comprises a plurality of video clips (Fig. 6; [0071], “FIG. 6 shows an example of a video on the left, and the system interface on the right, including information relating to the video”; [0159], “Video section 602 of the interface shows the video as it is playing. System interface section 604 shows information that is captured and collected relating to video shown in section 602. Interface 604 shows up on the screen when a user engages the visual video summarization system. System indicator 606 is visible when the system has been engaged, to show the user that the system is analyzing the video for slides, DIVIs and other content”; [0015], “The video visual summarization system disclosed herein has the ability to summarize videos into concise visual slides along with associated comments, summaries and/or links to the original video”. See [0127]-[0136] for video clips/segments); displaying a video summary area, wherein the video summary area comprises a first display control (Fig. 6; [0159], “System interface section 604 shows information that is captured and collected relating to video shown in section 602. Interface 604 shows up on the screen when a user engages the visual video summarization system. System indicator 606 is visible when the system has been engaged”; [0160], “Pulldown selection area 608 allows the viewer to select a notebook for filing this video summary” and Fig. 6, wherein the selectable “Fun stuff” option in the pulldown section area 608 is equivalent to a first display control); and displaying a video summary list in the video summary area in response to selecting the first display control, wherein the video summary list comprises a plurality of video content summaries (Fig. 6; and [0160], “Pulldown selection area 608 allows the viewer to select a notebook for filing this video summary”, such as the “fun stuff” notebook as shown in Fig. 6, which displays a video summary list in the 604 video summary area in response to the user’s selecting “fun stuff” notebook, wherein the list comprises a plurality of video content summaries, e.g., multiple 610 slides each corresponding to a video clip. See [0160], “These slides may be regular slides, build slides, DIVI builds or other types of slides/builds”; [0161], “The viewer is also able to add comments in comment section 614. In some instances, the system will automatically add comments based on the written and/or audio text of the video. In some instances, the system may use AI to summarize the written and/or audio text of the video”); wherein the video content summary comprises a text summary and a picture summary of a corresponding video clip, the picture summary comprises a frame image of the corresponding video clip (see 112 rejection and Examiner’s interpretation stated therein. See Fig. 6, wherein each component 610 comprises an image and text(s). See [0161], “The viewer is also able to add comments in comment section 614. In some instances, the system will automatically add comments based on the written and/or audio text of the video. In some instances, the system may use AI to summarize the written and/or audio text of the video”). But does not expressly disclose that the displaying of the video summary area is in response to selecting the video summary control, and that the plurality of video content summaries are in a one-to-one correspondence with the plurality of video clips, or that the frame image is determined based on the text summary. Hu discloses determining a frame image based on text summary (abstract; and claims 1-3). Before the effective filing date of the invention, it would have been obvious for an ordinary skilled in the art to combine LANZA with Hu. The suggestion/motivation of the combination would have been to select important video frames (Hu, abstract). Chen discloses a concept that displaying video summary are is in response to selecting a video summary control (abstract, “the target video is divided to obtain the plurality of video fragments, the playing control comprising the plurality of sub-controls is drawn based on the starting time of each video fragment, the summary information of each video fragment is rendered on each sub-control”; page 10, “In other embodiments, while playing the target video, the second user may call the display of the play control through some triggering operations, optionally including but not limited to: at least one of clicking a video playing interface, double clicking the video playing interface, pressing a bottom area of the video playing interface, a voice instruction or a gesture instruction is clicked, and the type of the triggering operation is not specifically limited in the embodiment of the present application”). Chen also discloses that a plurality of video content summaries are in a one-to-one correspondence with a plurality of video clips (see citation above, e.g., abstract, “the target video is divided to obtain the plurality of video fragments, the playing control comprising the plurality of sub-controls is drawn based on the starting time of each video fragment, the summary information of each video fragment is rendered on each sub-control”). Before the effective filing date of the invention, it would have been obvious for an ordinary skilled in the art to combine LANZA with Chen. The suggestion/motivation of the combination would have been to allow user an option to view video summaries with progress information (Chen, abstract and page 10). As to claim 12, see similar rejection to claim 1, wherein a processor and a memory are implied in order to perform the “video summarization method” as required by claim 1. Also see LANZA, [0198], “processor”; [0199], “memory”. As to claim 20, see similar rejection to claim 1, wherein a processor and a computer program product is implied in order to perform the “video summarization method” as required by claim 1. Also see LANZA,[0205], “software”. As to claim 10, LANZA in view of Hu and Chen discloses the claimed invention substantially as discussed in claim 1. Chen further discloses display a playback progress bar of the target video (abstract, “the target video is divided to obtain the plurality of video fragments, the playing control comprising the plurality of sub-controls is drawn based on the starting time of each video fragment, the summary information of each video fragment is rendered on each sub-control”; page 7, “the playback control is a target stripe and the plurality of child controls are segments of the target stripe. Optionally, the target strip is also referred to as a outline progress bar, the target strip may be an elongated strip, or the target strip may also be an annular strip, and the shape of the target strip is not specifically limited in the embodiments of the present application” page 10, “Fig. 10 is a schematic diagram of a playing control of a target video provided in an embodiment of the present application, please refer to fig. 10, where an outline progress bar 1001 (also called a playing control) is displayed in a bottom region of a video playing interface 1000, summary information (also called outline content) of a corresponding video segment is displayed on each segment of the outline progress bar 1001, a ratio of a length of each segment to a length of the outline progress bar is equal to a ratio of a segment duration of each video segment to a total duration of the video. in addition, a background color of the outline progress bar 1001 may change along with the playing progress to indicate a current playing progress of the target video, and the outline progress bar 1001 completely replaces a conventional playing progress bar to control a video playing progress and intuitively displays the outline content of each video segment. Fig. 11 is a schematic diagram of a target video playing control provided in an embodiment of the present application, please refer to fig. 11, a traditional progress bar 1101 and an outline progress bar 1102 are respectively displayed in a bottom region of a video playing interface 1100”), and the video summary area further comprises a progress bar linkage switch (Page 10, “In other embodiments, while playing the target video, the second user may call the display of the play control through some triggering operations, optionally including but not limited to: at least one of clicking a video playing interface, double clicking the video playing interface, pressing a bottom area of the video playing interface, a voice instruction or a gesture instruction is clicked, and the type of the triggering operation is not specifically limited in the embodiment of the present application”); and the method further comprises: displaying a plurality of video summary nodes on the playback progress bar in response to turning on the progress bar linkage switch, wherein the plurality of video summary nodes are in one-to-one correspondence with the plurality of video clips (see citation above, e.g., Page 10, “In other embodiments, while playing the target video, the second user may call the display of the play control through some triggering operations, optionally including but not limited to: at least one of clicking a video playing interface, double clicking the video playing interface, pressing a bottom area of the video playing interface, a voice instruction or a gesture instruction is clicked, and the type of the triggering operation is not specifically limited in the embodiment of the present application”. Also see claim 1, “acquiring a plurality of start moments and a plurality of summary information of a plurality of video segments in a target video, wherein each video segment corresponds to one start moment and one summary information; based on the plurality of start moments, acquiring a playing control of the target video, wherein the playing control is used for controlling the playing progress of the target video and comprises a plurality of sub-controls respectively corresponding to the plurality of video clips; and respectively displaying the plurality of summary information on the plurality of sub-controls”; Claim 2, “The method of claim 1, wherein the playing control is a target stripe, the plurality of sub-controls are a plurality of segments of the target stripe”); determining a corresponding video clip in response to selecting one of the video summary nodes (see citation in rejection to the preceding limitations. Also see claim 11, “wherein after displaying the plurality of sub-controls included in the playback control in the video playback interface, the method further comprises: and responding to the triggering operation of any sub-control, and playing the video clip corresponding to the any sub-control”); obtaining a corresponding video content summary based on the corresponding video clip (claim 1, “acquiring a plurality of start moments and a plurality of summary information of a plurality of video segments in a target video, wherein each video segment corresponds to one start moment and one summary information; based on the plurality of start moments, acquiring a playing control of the target video, wherein the playing control is used for controlling the playing progress of the target video and comprises a plurality of sub-controls respectively corresponding to the plurality of video clips; and respectively displaying the plurality of summary information on the plurality of sub-controls”); and displaying the corresponding video content summary based on a display location of the selected video summary node (abstract, “According to the method and the device, the target video is divided to obtain the plurality of video fragments, the playing control comprising the plurality of sub-controls is drawn based on the starting time of each video fragment, the summary information of each video fragment is rendered on each sub-control”). Before the effective filing date of the invention, it would have been obvious for an ordinary skilled in the art to combine LANZA in view of Hu and Chen with the further disclosure of Chen. The suggestion/motivation of the combination would have been so that the man-machine interaction efficiency is improved (Chen, abstract). As to claim 11, LANZA in view of Hu and Chen discloses the method according to claim 1, wherein the displaying a video summary area in response to selecting the video summary control comprises: obtaining the video summary list and the subtitle list in response to selecting the video summary control (see 112 rejection and Examiner’s interpretation stated therein. See LANZA, see citation in rejection to claim 1, e.g., Fig. 6; [0160], “Pulldown selection area 608 allows the viewer to select a notebook for filing this video summary”); when the video summary list and the subtitle list are obtained, displaying the video summary area (see 112 rejection and Examiner’s interpretation stated therein. See LANZA, see citation in rejection to claim 1, e.g., Fig. 6; [0160], “Pulldown selection area 608 allows the viewer to select a notebook for filing this video summary” such as the “fun stuff” notebook as shown in Fig. 6, which displays a video summary list in the 604 video summary area in response to the user’s selecting “fun stuff” notebook, wherein the list comprises the video summary list, e.g., multiple 610 slides each corresponding to a video clip. See [0160], “These slides may be regular slides, build slides, DIVI builds or other types of slides/builds”; [0161], “The viewer is also able to add comments in comment section 614. In some instances, the system will automatically add comments based on the written and/or audio text of the video. In some instances, the system may use AI to summarize the written and/or audio text of the video”), and displaying the video summary list and the subtitle list by using the video summary area (see 112 rejection and Examiner’s interpretation stated therein. See LANZA, see citation in rejection to claim 1, e.g., Fig. 6; [0160], “Pulldown selection area 608 allows the viewer to select a notebook for filing this video summary”), and when the video summary list and the subtitle list are not obtained, obtaining a basic content list of the target video, wherein the basic content list comprises one or more of manuscript information, a comment, and a bullet-screen comment; and displaying the video summary area, and displaying the basic content list by using the video summary area (see 112 rejection and Examiner’s interpretation stated therein. See LANZA, as cited in the preceding limitations, wherein the summary generated by AI is optional (see [0161], “The viewer is also able to add comments in comment section 614. In some instances, the system will automatically add comments based on the written and/or audio text of the video. In some instances, the system may use AI to summarize the written and/or audio text of the video” therefore in the instances wherein the system does not use AI to summarize the written and/or audio text of the video, i.e., when the summary list is not obtained, then only the basis content list such as comments are displayed, see Fig. 6). 11. Claims 2-5 and 13-16 are rejected under 35 U.S.C. 103 as being unpatentable over LANZA in view of Hu and Chen, as applied to claim 1 above, and further in view of Shen et al (CN118102050A, Google Patent translation is relied upon, hereafter Shen). As to claim 2, LANZA in view of Hu and Chen discloses the method according to claim 1, wherein the video content summary further comprises a first time jump control (LANZA, [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”); and the first time jump control and the text summary are obtained by performing the following operations: obtaining text information of the target video, wherein the text information comprises a plurality of subtitles of the target video (LANZA, [0029]-[0051], e.g., [0034], “obtaining auditory or textual information from the one or more images”; [0046], “presenting a video notebook may generally comprise presenting one or more video summaries upon a display to a user as a corresponding thumbnail image, wherein each of the one or more video summaries provide one or more searchable keywords corresponding to a content of the video summary, and wherein creation of each of the one or more video summaries comprises: obtaining one or more images from a video as the video is played, locating a presence of a shape or text from each of the one or more images, determining whether the shape or text corresponds to a prior shape or text from a prior base image, determining whether each of the one or more images comprises a corresponding slide, presenting the one or more images as one or more slides upon an interface displayed to the user, providing a timestamp upon each of the one or more slides, whereby selection of the timestamp by the user plays the video at a location which correlates to the timestamp within the video, and presenting the one or more slides including the timestamp to the user”. See claim 8, “The summary is generated as a transcript”. See Hu, claim 1, “step 1, acquiring frame characteristic representation of a video: for a video subtitle data set D= { V, Y }, wherein V represents a video set and Y represents an English subtitle sentence set corresponding to each video in the video set V”; see also claim 2); and wherein one text summary corresponds to one video clip (Chen, see citation in rejection to claim 1, e.g., abstract, “the target video is divided to obtain the plurality of video fragments, the playing control comprising the plurality of sub-controls is drawn based on the starting time of each video fragment, the summary information of each video fragment is rendered on each sub-control”), each text summary has a corresponding first time jump control, and the first time jump control is configured to locate a corresponding video clip of the text summary (LANZA, [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”; See also Chen, page 10, “when the second user clicks the segment of the outline progress bar, the second user can directly jump to the start time of the corresponding video segment and play the video segment”), LANZA also discloses obtain a plurality of text summaries by using a machine learning model (see [0161], “In some instances, the system may use AI to summarize the written and/or audio text of the video”), but does not expressly disclose inputting the text information into a pre-trained language model or that each subtitle has a corresponding time identifier. Shen disclose inputting the text information into a pre-trained language model (claim 10, “A method according to any one of claims 1-3, wherein the text processing model is a large-scale language model”; page 7, “In step S203, a task input corresponding to the summary generating task may be configured, and at least a portion of subtitles and a corresponding subtitle timestamp included in one of the plurality of packets and the configured task input are input together into the text processing model, so as to obtain a first processing result of the packet output by the text processing model, that is, at least one candidate segment summary of the packet and a candidate segment timestamp of each of the at least one candidate segment summary. According to some embodiments, the text processing model may be a large scale language model (Large Language Model, LLM). A large-scale language model is an artificial intelligence system trained on large data sets, aimed at understanding, generating, answering questions, etc., human language-related task”) and that each subtitle has a corresponding time identifier (claim 1, “A method for generating a video summary, the method comprising: determining a plurality of subtitles in a target video, the plurality of subtitles each having a subtitle timestamp”). Before the effective filing date of the invention, it would have been obvious for an ordinary skilled in the art to combine LANZA in view of Hu and Chen, with Shen. The suggestion/motivation of the combination would have been to utilize a text processing model to obtain video summaries (Shen, page 4). As to claim 13, see similar rejection to claim 2. As to claim 3, LANZA in view of Hu, Chen and Shen discloses the method according to claim 2, further comprising: generating a jump instruction in response to selecting a first time jump control of one of the video content summaries; and executing the jump instruction, wherein the jump instruction is configured to instruct the player to enable the target video to jump to a location indicated by the first time jump control to play (LANZA, [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”). As to claim 14, see similar rejection to claim 3. As to claim 4, LANZA in view of Hu, Chen and Shen discloses the method according to claim 3, wherein the video summary area further comprises a second display control (LANZA, Fig. 6, [0160], “Pulldown selection area 608 allows the viewer to select a notebook for filing this video summary”, wherein another selectable notebook other than “fun stuff” is equivalent to a second display control, which displays a video summary list in the 604 video summary area in response to the user’s selecting the other selectable notebook such as another one of the notebooks shown in Fig. 8); and the method further comprises: displaying a subtitle list in the video summary area in response to selecting the second display control, wherein the subtitle list comprises a plurality of subtitles of the target video and a second time jump control corresponding to each subtitle, and the second time jump control is configured to locate a corresponding video clip of the subtitle (see citation above and LANZA, [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”. Also See Padhi, as cited in rejection to claim 3 regarding subtitles). As to claim 15, see similar rejection to claim 4. As to claim 5, LANZA in view of Hu, Chen and Shen discloses the method according to claim 4, further comprising: generating a jump instruction in response to selecting a second time jump control of one of the subtitles; and executing the jump instruction, wherein the jump instruction is configured to instruct the player to enable the target video to jump to a location indicated by the second time jump control to play (Lanza, [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”. Also See Shen, as cited in rejection to claim 1 regarding subtitles and a location indicated by the timestamp). As to claim 16, see similar rejection to claim 5. 12, Claims 6-7 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over LANZA in view of Hu, Chen and Shen, as applied to claim 4 above, and further in view of GOSHEN et al (WO 2023/215260A1, Google Patent translation is relied upon, hereafter GOSHEN). As to claim 6, LANZA in view of Hu, Chen and Shen discloses the method according to claim 4, further comprising: determining a corresponding video clip based on the selected second time jump control (LANZA, see citation in rejection to claim 1, e.g., Fig. 6, a timestamp 616 on each slide; [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”); determining a corresponding video content summary based on the corresponding video clip (LANZA, see citation in rejection to claim 1, e.g., Fig. 6; and [0161], “The viewer is also able to add comments in comment section 614. In some instances, the system will automatically add comments based on the written and/or audio text of the video. In some instances, the system may use AI to summarize the written and/or audio text of the video”); but does not expressly disclose generating a first list synchronization instruction based on the corresponding video content summary; and executing the first list synchronization instruction, wherein the first list synchronization instruction is configured to instruct the video summary area to adjust the video summary list, so as to display the corresponding video content summary at a predetermined location of the video summary list. GOSHEN discloses generating a first list synchronization instruction based on corresponding video content summary; and executing the first list synchronization instruction, wherein the first list synchronization instruction is configured to instruct the video summary area to adjust a video summary list, so as to display the corresponding video content summary at a predetermined location of the video summary list ([0105], “The user interface may also include controls for navigating relative to the generated summary snippets and/or the corresponding transcript or source audio/video file. For example, the user interface may include a first scroll bar associated with the summary snippets and a second scroll bar associated with the formatted transcript. These scroll bars may enable vertical navigation (e.g., scrolling) relative to the snippets and/or transcript. In some embodiments, the navigation of the snippets and transcript may be linked such that changing the position of the first scroll bar to move locations of summary snippets in window 1430 will cause movement of the transcript in window 1440. For example, if a change in position of the first scroll bar results in removal of a snippet from window 1430, a section of transcript corresponding to the removed snippet may also be removed from window 1440. Similarly, if a change in position of the first scroll bar results in the appearance of a new snippet (or portion of a new snippet) into window 1430, a section of transcript corresponding to the new snippet may also be added/shown in window 1440”). Before the effective filing date of the invention, it would have been obvious for an ordinary skilled in the art to combine LANZA in view of Hu, Chen and Shen with GOSHEN. The suggestion/motivation of the combination would have been to respond to link the navigation of the snippets and transcript such that changing the position of the first scroll bar to move locations of summary snippets in window 1430 will cause movement of the transcript in window 1440 (GOSHEN, [0105]). As to claim 17, see similar rejection to claim 6. As to claim 7, LANZA in view of Hu, Chen, Shen and GOSHEN discloses the method according to claim 4, further comprising: determining a corresponding video clip based on the selected first time jump control (LANZA, see citation in rejection to claim 1, e.g., Fig. 6, a timestamp 616 on each slide; [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”); determining a subtitle of the corresponding video clip based on the corresponding video clip (LANZA, see citation in rejection to claim 1, e.g., Fig. 6; and [0161], “The viewer is also able to add comments in comment section 614. In some instances, the system will automatically add comments based on the written and/or audio text of the video. In some instances, the system may use AI to summarize the written and/or audio text of the video”); generating a second list synchronization instruction based on the subtitle of the corresponding video clip; and executing the second list synchronization instruction, wherein the second list synchronization instruction is configured to instruct the video summary area to adjust the subtitle list, so as to display the subtitle of the corresponding video clip at a predetermined location of the subtitle list (GOSHEN, ([0105], “The user interface may also include controls for navigating relative to the generated summary snippets and/or the corresponding transcript or source audio/video file. For example, the user interface may include a first scroll bar associated with the summary snippets and a second scroll bar associated with the formatted transcript. These scroll bars may enable vertical navigation (e.g., scrolling) relative to the snippets and/or transcript. In some embodiments, the navigation of the snippets and transcript may be linked such that changing the position of the first scroll bar to move locations of summary snippets in window 1430 will cause movement of the transcript in window 1440. For example, if a change in position of the first scroll bar results in removal of a snippet from window 1430, a section of transcript corresponding to the removed snippet may also be removed from window 1440. Similarly, if a change in position of the first scroll bar results in the appearance of a new snippet (or portion of a new snippet) into window 1430, a section of transcript corresponding to the new snippet may also be added/shown in window 1440”). As to claim 18, see similar rejection to claim 7. 13. Claims 8-9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over LANZA in view of Hu, Chen and Shen, as applied to claim 2 above, and further in view of Swaminathan et al (US 10650245 hereafter Swaminathan). As to claim 8, LANZA in view of Hu, Chen and Shen discloses the method according to claim 2, wherein the frame image is determined based on the text summary (Hu, see citation in rejection to claim 1, e.g., Hu, abstract), and the method further comprises: determining a corresponding video clip based on the first time jump control of the text summary (see citation in rejection to claim 1, e.g., LANZA, Fig. 6, a timestamp 616 on each slide; and [0162], “Also shown within system interface 604 is timestamp 616, as well as a play icon next to the timestamp. This timestamp is associated with the adjacent slide. If the viewer clicks on the icon next to the timestamp, the video will go to that point and start playing from that point-the point associated with the captured slide”); extracting a plurality of frame images based on the corresponding video clip (see citation in rejection to claim 1, e.g., LANZA, [0029], See also [0033], “In another aspect of the method for visually summarizing a video, obtaining the one or more images may comprise obtaining the one or more images from one or more build images to create the corresponding slide comprising a build sequence slide”); determining text feature vectors of the plurality of frame images, and an image feature vector (see Hu, abstract; and page 3, “step 1, acquiring frame characteristic representation of a video; for a video subtitle data set D= { V, Y }, wherein V represents a video set and Y represents an English subtitle sentence set corresponding to each video in the video set V; processing any ith video in the video set V by adopting a visual encoder of the CLIP model to obtain a frame characteristic representation F of the ith video i ={f i,1, f i,2,...,f i,n ,..,f i,N -a }; wherein f i,n Representing an nth frame characteristic representation in an ith video, wherein N represents the total frame number of the video i”; “step 2, acquiring the characteristic representation of the caption: text encoder adopting CLIP model is used for English caption sentence Y corresponding to ith video in pairs i ={y i,1,1 ,...,y i,1,W ;y i,m,1 ,y i,m,2 ,...,y i,m,t ,...,y i,m,W ;y i,M,1,...,y i,M,W Processing to obtain English caption text vector T corresponding to video i i ={t i,1 ,t i,2 ,...,t i,m ,..,t i,M -wherein y i,m,t Representing the t-th word, t, in the mth subtitle sentence corresponding to the ith video i,m Representing an mth subtitle vector in an English subtitle sentence corresponding to the ith video; w represents the total number of words”); determining a picture-text similarity corresponding to the each frame image based on the text feature vector and the image feature vector corresponding to the frame image (see Hu, abstract; and page 3, steps 1-2 and “step 3, obtaining the characteristic representation f of the nth frame in the ith video by using the formula (1) i,n And caption text vector T i Average similarity s (f) i,k ,T i ) And represents f as an nth frame feature of video i i,n Automated scoring of (a)”); and determining a target frame image based on the picture-text similarity that are corresponding to the each frame image, and determining the target frame image as the picture summary (see Hu, abstract; page 3, “step 3, obtaining the characteristic representation f of the nth frame in the ith video by using the formula (1) i,n And caption text vector T i Average similarity s (f) i,k ,T i )And represents f as an nth frame feature of video i i,n Automated scoring of (a)In the formula (1), tr represents vector transposition”). But does not expressly an aesthetic score that are corresponding to each frame image and that the aesthetic score is based on to determine a target frame image. Swaminathan discloses an aesthetic score corresponding to a frame image and that the aesthetic score is based on to determine a target frame image (see abstract, “utilize an aesthetics neural network to determine aesthetics scores for frames of a digital video and a relevancy neural network to generate importance scores for frames of the digital video. Utilizing the aesthetic scores and relevancy scores, the disclosed systems can select a subset of frames and apply a generative reconstructor neural network to create a digital video reconstruction. By comparing the digital video reconstruction and the original digital video, the disclosed systems can accurately identify representative frames and flexibly generate a variety of different digital video summaries”). Before the effective filing date of the invention, it would have been obvious for an ordinary skilled in the art to combine LANZA in view of Hu, Chen and Shen with Swaminathan. The suggestion/motivation of the combination would have been to select a subset of frames utilizing aesthetics scores (Swaminathan, abstract). As to claim 19, see similar rejection to claim 8. As to claim 9, LANZA in view of Hu, Chen, Shen and Swaminathan discloses the method according to claim 8, wherein the determining text feature vectors of the plurality of frame images, and an image feature vector and an aesthetic score that are corresponding to each frame image comprises: inputting the text summary into a pre-trained generative embeddings model, so as to obtain the text feature vectors by using the pre-trained generative embeddings model (Hu, page 4, “step 2, acquiring the characteristic representation of the caption: using the CLIP model English caption sentence Y corresponding to ith video in the text encoder pair of (2) i ={y i,1,1 ,...,y i,1,W ;y i,m,1 ,y i,m,2 ,...,y i,m,t,..., y i,m,W ;y i,M,1 ,...,yi,M,W Processing to obtain caption text vector T corresponding to video i i ={t i,1 ,t i,2 ,...,t i,m ,..,t i,M -wherein y i,m,t Representing the t-th word, t, in the mth subtitle sentence corresponding to the ith video i,m Representing the mth subtitle vector in the corresponding english subtitle sentence in the ith video, m=20, w=30 in this embodiment”, wherein the CLIP’s text encoder is a pre-trained text embedding language model that produces a text feature vector from the textual content/summary, functioning as a pre-trained model that generates a text embedding vector from input text, therefore is equivalent to the claimed text-vector step); inputting the plurality of frame images into a pre-trained contrastive language-image model, so as to obtain the image feature vector corresponding to the each frame image by using the pre-trained contrastive language-image model (Hu, page 4, “step 1, acquiring frame characteristic representation of a video: for a video subtitle data set D= { V, Y }, wherein V represents a video set and Y represents an English subtitle sentence set corresponding to each video in the video set V; processing any ith video in the video set V by adopting a visual encoder of the CLIP model to obtain a frame characteristic representation F of the ith video i ={f i,1, f i,2,...,f i,n ,..,f i,N -a }; wherein f i,n Representing the nth frame feature representation in the ith video, N representing the total frame number of video i, in this embodiment, n=12”, wherein C:IP is the pre-trained contrastive language-image model); and inputting the plurality of frame images into a pre-trained aesthetic scoring model, so as to obtain the aesthetic score corresponding to the each frame image by using the pre-trained aesthetic scoring model (Swaminathan, abstract). Prior Art Cited but not Applied in the Rejection 15. Below is a list of prior art reference(s) cited but not applied in the rejection: a) DELGO et al (US 2010/0070483), disclosing most relevant and salient visual appearances depicted in the videos are presented to the user, both for the sake of summarizing the video content for the users to “see before they watch”, as well as for allowing the users to further refine their video search result set according to the most relevant and salient video content. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUA FAN whose telephone number is (571)270-5311. The examiner can normally be reached on 9-6. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nasser Goodarzi, can be reached at (571) 272-4195. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HUA FAN/Primary Examiner, Art Unit 2426
Read full office action

Prosecution Timeline

Sep 10, 2025
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12701154
AUTOMATED DELIVERY OF ALERTS WITH CONFIRMATION OF RECEIPT
6y 1m to grant Granted Aug 04, 2026
Patent 12665931
TECHNIQUES FOR DYNAMIC CLIENT-SIDE TRAFFIC ROUTING WITH SERVER-SIDE CONTROL
2y 3m to grant Granted Jun 23, 2026
Patent 12652241
PROTOCOL INDEPENDENT MULTICAST (PIM) ACROSS TRANSPORT NETWORK
2y 5m to grant Granted Jun 09, 2026
Patent 12627728
GRAPHICALLY INTEGRATING SENSOR DATA THROUGH EDGE DEVICES
2y 1m to grant Granted May 12, 2026
Patent 12615179
CONNECTIVITY FAILURE SOLUTIONS FOR CONTAINER PLATFORMS
2y 5m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
70%
Grant Probability
91%
With Interview (+21.2%)
3y 11m (~2y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 787 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month