Prosecution Insights
Last updated: August 13, 2026
Application No. 18/572,175

VIDEO PROCESSING METHOD, APPARATUS AND ELECTRONIC DEVICE

Non-Final OA §103
Filed
Dec 19, 2023
Priority
Aug 17, 2022 — CN 202210987415.7 +1 more
Examiner
TOPGYAL, GELEK W
Art Unit
2481
Tech Center
2400 — Computer Networks
Assignee
Beijing Zitiao Network Technology Co., Ltd.
OA Round
2 (Non-Final)
59%
Grant Probability
Moderate
2-3
OA Rounds
11m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 59% of resolved cases
59%
Career Allowance Rate
363 granted / 614 resolved
+1.1% vs TC avg
Strong +19% interview lift
Without
With
+18.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
11 currently pending
Career history
649
Total Applications
across all art units

Statute-Specific Performance

§101
6.7%
-33.3% vs TC avg
§103
56.9%
+16.9% vs TC avg
§102
24.2%
-15.8% vs TC avg
§112
3.3%
-36.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 614 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Information Disclosure Statement The information disclosure statement (IDS) submitted on 10/2/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Response to Arguments Applicant’s arguments with respect to claim(s) 1, 3-11, 13-14 and 18-23 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3-4, 13-14 and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over More et al. (US 2019/0303403) in view of Clifton et al. (US2008/0215979). Regarding claim 1, More teaches a video processing method, comprising: obtaining a video edit draft, the video edit draft being used to record an initial multimedia material and indication information of a first video editing operation for the initial multimedia material, the initial multimedia material comprising a video material and/or an image material (Fig. 2, steps S1005 and S1010, results in obtaining a video edit draft including portions of video segments/clips that meet the claimed initial multimedia material used for the video edit draft. The claimed indication information refers to the one or more edits performed by the user on the video segments/clips used for the generation of the video edit draft as further discussed in paragraphs 138-156); extracting the initial multimedia material from the video edit draft (paragraphs 138-156 teaches accessing video files that are used in the video edit draft such that the system is able to perform future edits as discussed); determining, based on the initial multimedia material, a target edit template that matches the initial multimedia material, the target edit template being used to record indication information of a second video editing operation (Fig. 2, steps S1020-S1025 and paragraphs 138-156 teaches wherein analysis is performed on user actions, registration attributes, performance data, etc. to generate a recommendation for editing the draft video file); and processing, based on the target edit template, the initial multimedia material according to the second video editing operation, to obtain a first target video having an editing effect of the second video editing operation (Fig. 2, step S1030 and paragraphs 97 and 138-156 teaches the claimed processing the initial multimedia material by editing the draft to generate a newer customized video file in accordance with the recommendation). While More in paragraphs 183 and 189 teaches tags associated with the initial video content used in generating recommendations for generating the edits for the draft video, fails to explicitly teach the use of content tags indicating material type/quantity/duration, and therefore fails to teach the following, but Clifton teaches: wherein determining, based on the initial multimedia material, the target edit template that matches the initial multimedia material comprises: analyzing content of the initial multimedia material to obtain a content tag corresponding to the initial multimedia material, the content tag being configured to indicate characteristics of the initial multimedia material, including a material type, a material quantity, or a material duration of the initial multimedia material; and determining, based on the content tag, the target edit template that matches the initial multimedia material (Figs. 2A-2B, wherein metadata stores information about the video analysis and audio analysis. The analysis results in the visual media items and audio media items being processed to “create and store metadata” based on the results. Firstly, examiner notes that the limitation of the content tag is presented as an alternative language. However, examiner will address all three content “tags” with the material type being differentiable between visual media items (video or images) see paragraphs 122 for video media input and paragraphs 66-68 where certain “key images” are input. Secondly, with regards to the “quantity”, paragraphs 66-67 states “the user could upload a number of digital visual media items” or “list of audiovisual works” to be input into the system. With regards to the “duration”, certain portions/”shots” can be deemed key shots and therefore, certain lengths of input visual media items are prioritized to be included in the output. Specific metadata of “length” is also included in the metadata. Altogether, based on these data, the system is able to select one or modules as discussed in paragraph 93). Therefore, It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Clifton into the system of More, such that More’s template selection also accommodates for specific metadata relating to the material’s type/quantity/duration because said incorporation allows for the benefit of creating automatic video output based not only implicit but inferred metadata of the input content (abstract and paragraph 93). Regarding claim 3, More teaches the claimed wherein analyzing content of the initial multimedia material to obtain a content tag corresponding to the initial multimedia material comprises: processing the initial multimedia material with the first model to obtain the content tag (paragraphs 183 and 189 teaches tags associated with the initial video content used in generating recommendations for generating the edits for the draft video); wherein the first model is obtained by learning with a plurality of sets of samples, and the plurality of sets of samples comprise sample multimedia materials and sample content tags corresponding to the sample multimedia materials (paragraphs 183 and 189 teaches tags associated with the initial video content used in generating recommendations for generating the edits for the draft video. Paragraph 177 teaches wherein a grouping model is trained using attributed weight, tags and points, which as discussed above are with regards to the initial video content. Training in general requires multiple sets of data inputs, which in this case would meet the input requirement). Regarding claim 4, Clifton teaches the claimed comprises: determining, based on the content tag, the target edit template from a plurality of preset edit templates if: the content tag indicates that a number of images of the initial multimedia material is greater than or equal to a first threshold, or (examiner notes the alternative language, Figs. 2A-2B, wherein metadata stores information about the video analysis and audio analysis. The analysis results in the visual media items and audio media items being processed to “create and store metadata” based on the results. Firstly, examiner notes that the limitation of the content tag is presented as an alternative language. However, examiner will address all three content “tags” with the material type being differentiable between visual media items (video or images) see paragraphs 122 for video media input and paragraphs 66-68 where certain “key images” are input, meaning certain number of images may be selected as part of a “key image” or “key shot”, which includes a plurality of frames that makes up the video frames. Secondly, with regards to the “quantity”, paragraphs 66-67 states “the user could upload a number of digital visual media items” or “list of audiovisual works” to be input into the system. With regards to the “duration”, certain portions/”shots” can be deemed key shots and therefore, certain lengths of input visual media items are prioritized to be included in the output. Specific metadata of “length” is also included in the metadata. Altogether, based on these data, the system is able to select one or modules as discussed in paragraph 93); the content tag indicates that a number of videos of the initial multimedia material is greater than or equal to a second threshold, or (examiner notes the alternative language, Figs. 2A-2B, wherein metadata stores information about the video analysis and audio analysis. The analysis results in the visual media items and audio media items being processed to “create and store metadata” based on the results. Firstly, examiner notes that the limitation of the content tag is presented as an alternative language. However, examiner will address all three content “tags” with the material type being differentiable between visual media items (video or images) see paragraphs 122 for video media input and paragraphs 66-68 where certain “key images” are input, meaning certain number of images may be selected as part of a “key image” or “key shot”, which includes a plurality of frames that makes up the video frames. Secondly, with regards to the “quantity”, paragraphs 66-67 states “the user could upload a number of digital visual media items” or “list of audiovisual works” to be input into the system, meaning a base of at least one video clip or a plurality of video media items are selected based on the input. With regards to the “duration”, certain portions/”shots” can be deemed key shots and therefore, certain lengths of input visual media items are prioritized to be included in the output. Specific metadata of “length” is also included in the metadata. Altogether, based on these data, the system is able to select one or modules as discussed in paragraph 93) the content tag indicates that a video duration of the initial multimedia material is greater than or equal to a third threshold (examiner notes the alternative language, Figs. 2A-2B, wherein metadata stores information about the video analysis and audio analysis. The analysis results in the visual media items and audio media items being processed to “create and store metadata” based on the results. Firstly, examiner notes that the limitation of the content tag is presented as an alternative language. However, examiner will address all three content “tags” with the material type being differentiable between visual media items (video or images) see paragraphs 122 for video media input and paragraphs 66-68 where certain “key images” are input, meaning certain number of images may be selected as part of a “key image” or “key shot”, which includes a plurality of frames that makes up the video frames. Secondly, with regards to the “quantity”, paragraphs 66-67 states “the user could upload a number of digital visual media items” or “list of audiovisual works” to be input into the system, meaning a base of at least one video clip or a plurality of video media items are selected based on the input. With regards to the “duration of the initial multimedia material”, certain portions/”shots” can be deemed key shots and therefore, certain lengths of input visual media items, being above a base threshold length, are prioritized to be included in the output. Specific metadata of “length” is also included in the metadata. Altogether, based on these data, the system is able to select one or modules as discussed in paragraph 93). Device claim 13 and NTCRM claim 14 are rejected for the same reasons as discussed in claim 1 above. Furthermore, paragraphs 67 and 74 teaches a computer readable storage medium storing executable instructions to implement the methodology in claim 1. Device claim 18 is rejected for the same reasons as discussed in claim 3. Claim 19 is rejected for the same reasons as discussed in claim 4. Claims 5-6 and 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over More et al. (US 2019/0303403)) in view of Clifton et al. (US2008/0215979) and further in view of Novikoff et al. (US 2018/0068019). Regarding claim 5, More and Clifton teaches the claimed wherein obtaining the video edit draft comprises: if the video edit application enables the automatic draft video creation function, obtaining the video edit draft at an ending of editing of the initial multimedia material (as discussed in claim 1 above, wherein Fig. 2, step S1030 and paragraphs 97 and 138-156 teaches the claimed processing the initial multimedia material by editing the draft to generate a newer customized video file in accordance with the recommendation.) However, while More teaches the automatic editing when enabled, fails to explicitly teach a button/User Interface to turn it off/on and therefore, fails, but Novikoff teaches that if a video edit application does not enable an automatic draft video creation function, displaying a first page which is a function setting page (See Figs. 6E, wherein a user is presented with the option to create movie 672 prior to the automatic movie creation being initiated); in response to a trigger operation on a first control in the first page, enabling the automatic draft video creation function and obtaining the video edit draft, the first control being a switch control of the automatic draft video creation function (Fig. 6E allows the user to choose the option to turn on automatic editing function, resulting in the processing of the automatic video draft creation). It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Novikoff into the system of More and Clifton because said incorporation allows for the benefit of improving user engagement by reducing time necessary required for editing. Regarding claim 6, while More and Clifton fails to teach the following, Novikoff teaches the claimed wherein displaying the first page comprises: in response to a touch operation on a second control in a second page, displaying the first page, the second page being used to present the first target video, the second control being used to jump to the first page (Figs. 6E, wherein second page presents first target video, selecting second control 673 allows jumping to Figs. 6D, 6E or 6G). It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Novikoff into the system of More because said incorporation allows for the benefit of improving user engagement by reducing time necessary required for editing. Claims 20-21 are rejected for the same reasons as discussed in claims 5-6, respectively. Claims 7-8 and 22-23 are rejected under 35 U.S.C. 103 as being unpatentable over More et al. (US 2019/0303403)) in view of Clifton et al. (US2008/0215979) and further in view of Rav-Acha et al. (US 9,554,111). Regarding claim 7, More and Clifton teaches the claimed as discussed in claim 1 above, however fails to, but Rav-Acha teaches wherein, after obtaining the first target video, the method further comprises: obtaining a target image of the first target video (col. 6, line 58-67); generating a first cover of the first target video based on the target image (Fig. 31, teaches displaying a video result using representative images); and displaying the first target video in a second page based on the first cover (Fig. 31 teaches wherein generated video content can be played as visual entities on a second page, wherein the visual entities are altered/edited media stream). It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Rav-Acha into the system of More and Clifton because such an incorporation allows for the benefit of improving the production of edited videos (Rav: col. 19, lines 50-55). Regarding claim 8, More teaches the claimed as discussed in claim 1 above, however fails to, but Rav-Acha teaches wherein generating the first cover of the first target video based on the target image comprises: obtaining a video content of the first target video (col. 6, line 58-67); determining a target text based on the video content (col. 6, line 58-67, text acquired about image); and generating the first cover of the first target video based on the target image and the target text (Fig. 31 and col. 6, line 58-67 shows representative images displayed with corresponding text). It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Rav-Acha into the system of More and Clifton because such an incorporation allows for the benefit of improving the production of edited videos (Rav: col. 19, lines 50-55). Claims 22-23 are rejected for the same reasons as discussed in claims 7-8, respectively. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over More et al. (US 2019/0303403)) in view of Clifton et al. (US2008/0215979) further in view of Rav-Acha et al. (US 9,554,111) and further in view of Wu (CN 113 589 982) (using US2024/0289141 as translation). Regarding claim 9, More teaches the claimed as discussed in claims 1 and 7 above, however fails to teach, but Wu teaches wherein displaying the first cover in the second page comprises: obtaining a creation time instant of the first object video (paragraphs 37-38 at least teaches presenting a sequence of draft resources based on a swiping action, resulting in the subsequent draft resource being played. Each draft is therefore indicative of different times when created); displaying, based on the creation time instant, the first cover corresponding to the first target video in the second page (paragraphs 37-38 at least teaches presenting a sequence of draft resources starting with a first draft resource and based on a swiping action resulting in the subsequent draft resource being played). It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Wu into the proposed combination of More, Clifton and Rav-Acha because said incorporation allows for the benefit of improving the user experience by displaying related draft versions simply. Claims 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over More et al. (US 2019/0303403)) in view of Clifton et al. (US2008/0215979) and further in view of Rav-Acha et al. (US 9,554,111) and further in view of Wu (CN 113 589 982) (using US2024/0289141 as translation) and further in view of Svendsen et al. (US2014/0255009). Regarding claim 10, More, Clifton, Rav-Acha and Wu teaches the claimed as discussed in claims 1 and 7 above, however, while Rav-Acha, More and Wu both teach that the video output is viewable and further edits can be made (Rav: Fig. 6F: Edit movie 678, etc.) fails to explicitly teach, but Svendsen teaches wherein, after displaying the first cover in the second page, the method further comprises: in response to a touch operation on the first cover, playing the first target video of the first cover, the first target video comprising a template replacement control (Figs. 6 and 7, teaches playing a target video that has a particular theme/template implemented); in response to a touch operation on the template replacement control, displaying template controls corresponding to a plurality of preset edit templates (Figs. 6 and 7, teaches wherein the template can be changed via a touch screen (para 29)); in response to a touch operation on the template controls, replacing the target edit template of the first target video (Fig. 7 teaches wherein the template can be changed via a touch screen (para 29)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Svendsen into the proposed combination of More, Clifton and Rav-Acha because said incorporation allows for the benefit of improving the user experience by allowing easier changes to their desired output. Regarding claim 11, More, Clifton and Rav-Acha teaches the system as discussed in claim 10 above, however fails to teach, but Svendsen teaches the claimed wherein replacing the target edit template of the first target video in response to a touch operation on the template controls comprises: obtaining a first edit template, the first edit template being used to record indication information of a third video editing operation (Figs. 6 and 7, teaches wherein the template is first applied to the video output. Additional changes are made through the user interface in selecting editing operations like transitions, adding media/audio, etc. in area 604 of the interface); and replacing the second video editing operation in the first target video with the third video editing operation to obtain the second target video, the second target video having an editing effect of the third video editing operation (paragraphs 6-7 and 29 at least teaches wherein changes to the output video is made). It would have been obvious to one of ordinary skill in the art before the effective filing date of the current application to incorporate the teachings of Svendsen into the proposed combination of More, Clifton and Rav-Acha because said incorporation allows for the benefit of improving the user experience by allowing easier changes to their desired output. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GELEK W TOPGYAL whose telephone number is (571)272-8891. The examiner can normally be reached M-F (9:30-6 PST). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William Vaughn can be reached at 571-272-3922. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GELEK W TOPGYAL/ Primary Examiner, Art Unit 2481
Read full office action

Prosecution Timeline

Dec 19, 2023
Application Filed
Oct 02, 2025
Non-Final Rejection mailed — §103
Jan 04, 2026
Response Filed
May 05, 2026
Final Rejection mailed — §103
Jul 06, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700428
TRIM PASS METADATA PREDICTION IN VIDEO SEQUENCES USING NEURAL NETWORKS
1y 8m to grant Granted Aug 04, 2026
Patent 12693536
SYSTEM AND METHOD FOR PRESENTING IMAGE CONTENT ON MULTIPLE DEPTH PLANES BY PROVIDING MULTIPLE INTRA-PUPIL PARALLAX VIEWS
2y 9m to grant Granted Jul 28, 2026
Patent 12687999
DISPLAY APPARATUS, DOOR BODY, AND CABINET BODY
2y 2m to grant Granted Jul 21, 2026
Patent 12688876
Highlight Video Generation
1y 7m to grant Granted Jul 21, 2026
Patent 12682932
Systems and methods for automatically generating a video production
2y 2m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
59%
Grant Probability
78%
With Interview (+18.6%)
3y 7m (~11m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 614 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month