Prosecution Insights
Last updated: August 17, 2026
Application No. 19/033,358

Multimedia Data Generating Method, Apparatus, Electronic Device, Medium, and Program Product

Final Rejection §103§DP
Filed
Jan 21, 2025
Priority
Oct 28, 2021 — CN 202111266196.5 +2 more
Examiner
ZHAO, DAQUAN
Art Unit
2484
Tech Center
2400 — Computer Networks
Assignee
Beijing Zitiao Network Technology Co., Ltd.
OA Round
2 (Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
1y 2m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
806 granted / 1044 resolved
+19.2% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
24 currently pending
Career history
1064
Total Applications
across all art units

Statute-Specific Performance

§101
11.7%
-28.3% vs TC avg
§103
46.8%
+6.8% vs TC avg
§102
18.5%
-21.5% vs TC avg
§112
13.1%
-26.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1044 resolved cases

Office Action

§103 §DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claim(s) s 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 10-11 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Lee (US 2023/0308731) and further in view of Angquist et al (US 2019/0104259). For claim 1, Lee teaches a multimedia data generating method, comprising: receiving text information (e.g. paragraph 73, figure 4: “speech input and text conversion”); displaying the text information and acquiring a first reading speech of the text information (e.g. figure 4: “I went to a nice beach and saw seals and nice boats on the sandy rocks of the beach”); obtaining a video image matched with the text information (e.g. paragraphs 74-75, figure 4, images match beach , seals boats and the sandy rock on the beach); and generating a first multimedia data based on the video image and the first reading speech and displaying the first multimedia data (e.g. paragraphs 76-77, figure 4: “Synthesized and converted video multimedia content); wherein each of the plurality of first multimedia segments comprises a first video segment including a video image matched with a first text segment (e.g. paragraphs 74-75, figure 4, images match beach , seals boats and the sandy rock on the beach)and a first speech segment including a reading speech of the first text segment (e.g. figure 5, paragraph 75: the video resource matching unit 130 matches caption and font resources to the sentence information as the element information and matches the audio made by converting the sentence information into speech as a sound resource to the sentence information. Furthermore, the video resource matching unit 130 matches animation information to the sentence information). Lee does not further teach: wherein, the first multimedia data comprises a plurality of first multimedia segments, the plurality of first multimedia segments corresponding to a plurality of first text segments included in the text information, respectively, the first video segment and the first speech segment being displayed in editing tracks. Angquist et al teach: wherein, the first multimedia data comprises a plurality of first multimedia segments, the plurality of first multimedia segments corresponding to a plurality of first text segments included in the text information, respectively, the first video segment and the first speech segment being displayed in editing tracks (e.g. figure 1: Media Editor Graphical User Interface comprises Video clip 107, audio clip 108 (or audio clip 112), caption 105). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the GUI of Angquist et al into the teaching of Lee to display video, audio and subtitle to improve editing efficiency by allowing editor to accurately specify the starting and/or ending point of a media object in the timeline (e.g. paragraph 79, Angquist et al). Claims 10 and 19 are rejected for the same reasons as discussed in claim 1 above, wherein Lee teaches various processes performed by processor (e.g. paragraph 21). For claims 2, 11 and 20, Angquist et al teach converting the text information into speech data in response to a multimedia synthesis operation; and generating second multimedia data based on the text information and the speech data, and displaying the second multimedia data; wherein the second multimedia data comprise the speech data and the video image matched with the text information; the second multimedia data comprise a plurality of second multimedia segments, the plurality of second multimedia segments corresponding to a plurality of second text segments included in the text information, respectively; each of the plurality of second multimedia segments comprises a second video segment including a video image matched with a second text segment and a second speech segment including a reading speech of the second text segment (e.g. figure 4 shows “Speech input and text conversion”, matching video image to the text such as “beach” is matched with the image of a “[U.S.A] beach resource”, and then synthesized and converted video multimedia content, also see discussion of claim 1 above.). Claims 3, 4, 6, 8-9, 12, 13, 15 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Lee and Angquist et al, as applied to claims 1-2, 10-11 and 19-20, and further in view of Rodriguez et al (US 8744239). For claims 4 and 13, Angquist et al teach displaying the first text segment corresponding to the first speech segment, and acquiring a read segment of the first text segment; and displaying the read segment in an area corresponding to the first speech segment (e.g. figure 1 audio segment 112, video clip, subtitle, caption are display on the same column or align in the same column). Lee and Angquist et al do not further disclose deleting, in response to a rerecording operation for the first speech segment, the first speech segment. Rodriguez et al teach deleting, in response to a rerecording operation for the first speech segment, the first speech segment (.g. column 28, lines 6-16: “text is delete”). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Rodriguez et al into the teaching of Lee and Angquist et al to delete speech segment to improve storage efficiency. For claims 6 and 15, Lee and Angquist et al do not teach moving, in response to a speech segment sliding operation and sliding of a first cursor pointing to the first reading speech till the first speech segment, a second cursor pointing to the text information till the first text segment. Rodriguez et al teach moving, in response to a speech segment sliding operation and sliding of a first cursor pointing to the first reading speech till the first speech segment, a second cursor pointing to the text information till the first text segment (e.g. figure 1, column 7, lines 14-19). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Rodriguez et al into the teaching of Lee and Angquist et al to correlate the scripted words with the voice of the narrator to reduce the number of take the narrator has to performs (e.g. column 1, lines 33-46) to improve convenience for the narrator. For claims 9 and 18, Angquist et al teach wherein the generating a first multimedia data based on the text information and the first reading speech and displaying the first multimedia data comprises: generating the first multimedia data based on the text information and the fourth reading speech, and displaying the first multimedia data (e.g. figure 1 audio segment 112, video clip, subtitle, caption are display on the same column or align in the same column). Lee and Angquist et al do not further disclose subjecting, after the first reading speech is acquired, the first reading speech to voice change processing and/or speed change processing to obtain a fourth reading speech. Rodriguez et al teach subjecting, after the first reading speech is acquired, the first reading speech to voice change processing and/or speed change processing to obtain a fourth reading speech (e.g. column 7, lines 20-30). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Rodriguez et al into the teaching of Lee and Angquist et al to correlate the scripted words with the voice of the narrator to reduce the number of take the narrator has to performs (e.g. column 1, lines 33-46) to improve convenience for the narrator. For claims 3 and 12, Lee teaches acquiring a second reading speech of the text information; and generating a third multimedia data based on the text information and the second reading speech, and displaying the third multimedia data to overwrite the second multimedia data; wherein the third multimedia data comprise the second reading speech and the video image matched with the text information; the third multimedia data comprise a plurality of third multimedia segments, the plurality of third multimedia segments corresponding to a plurality of third text segments included in the text information, respectively; each of the plurality of third multimedia segments includes a third video segment including a video image matched with a third text segment and a third target speech segment including a reading speech of the third text segment (e.g. figure 1 audio segment 112, video clip, subtitle, caption are display on the same column or align in the same column; figure 4 matching text with video image). Lee and Angquist et al do not further disclose displaying, in response to a recording trigger operation after the second multimedia data are generated, the text information. Rodriguez et al teach displaying, in response to a recording trigger operation for the text information, the text information (e.g. column 7, lines 50-58). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Rodriguez et al into the teaching of Lee and Angquist et al to correlate the scripted words with the voice of the narrator to reduce the number of take the narrator has to performs (e.g. column 1, lines 33-46) to improve convenience for the narrator. For claims 8 and 17, Lee teaches editing, after the first multimedia data are generated and displayed, the text information in response to an edit operation for the text information to obtain a modified target text information; updating the first reading speech based on the target modified reading speech of the modified text information to obtain a third reading speech; and generating a fourth multimedia data based on the modified text information and the third reading speech, and displaying the fourth multimedia data; wherein the fourth multimedia data comprise the third reading speech and the video image matched with the text information; the fourth multimedia data comprise a plurality of fourth multimedia segments, the plurality of fourth multimedia segments corresponding to a plurality of fourth text segments included in the text information, respectively; a fourth target each of the plurality of fourth multimedia segments comprises a fourth target video segment including a video image matched with a fourth text segment and a fourth target speech segment including a reading speech of the fourth text segment. (e.g. figure 1 audio segment 112, video clip, subtitle, caption are display on the same column or align in the same column; figure 4 matching text with video image). Lee and Angquist et al do not further displaying, in response to a recording trigger operation for the modified text information, the target modified text information and acquiring a reading speech of the modified text information. Rodriguez et al teaches displaying, in response to a recording trigger operation for the modified text information, the target modified text information and acquiring a reading speech of the modified text information (e.g. column 7, lines 50-58). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Rodriguez et al into the teaching of Lee and Angquist et al to correlate the scripted words with the voice of the narrator to reduce the number of take the narrator has to performs (e.g. column 1, lines 33-46) to improve convenience for the narrator. Claims 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Lee and Angquist et al, as applied to claims 1-2, 10-11 and 19, and further in view of Tokunaka et al (US 2008/0002949). For claims 7 and 16, Lee and Angquist et al do not further disclose prominently displaying a text segment currently read by the user while acquiring the first reading speech. Tokunaka et al teach prominently displaying a text segment currently read by the user while acquiring the first reading speech (e.g. figure 6, paragraph 132, Bar IB). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Tokunaka et al into the teaching of Lee and Angquist et al to improve editing efficiency. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-4, 6-13, and 15-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-18 of U.S. Patent No. 12,206,955 B2 in view of Rodriquez et al (US 8,744,239). Claims 1-4, 6-13, and 15-20 of the instant application corresponds to claims 1-18 of the Patent, respectively. Independent claims 1, 9 and 17 of the Patent do not further specify the first target video segment and the first target speech segment is displayed in editing tracks. Rodriguez et al teach the first target video segment and the first target speech segment is displayed in editing tracks (e.g. figure 1, video and text are displayed on the same screen. Column 4, line 62-column 5, line 5: The composite display area 130 includes multiple tracks that span a timeline 160, and displays one or more graphical representations of media clips in the composite presentation. As shown, the composite display area 130 displays a music clip representation 165 and a video clip representation 170.). It would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teaching of Rodriguez et al into claims 1-18 of the Patent to correlate the scripted words with the voice of the narrator to reduce the number of take the narrator has to performs (e.g. column 1, lines 33-46) to improve convenience for the narrator. Allowable Subject Matter Claims 5 and 14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAQUAN ZHAO whose telephone number is (571)270-1119. The examiner can normally be reached M-Thur: 7:00 am-5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Thai Tran can be reached on 571-272-7382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. Email: daquan.zhao1@uspto.gov. Phone: (571)270-1119 /DAQUAN ZHAO/Primary Examiner, Art Unit 2484
Read full office action

Prosecution Timeline

Jan 21, 2025
Application Filed
Mar 05, 2026
Non-Final Rejection mailed — §103, §DP
Jun 05, 2026
Response Filed
Jul 21, 2026
Final Rejection mailed — §103, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706478
SYSTEMS AND METHODS FOR MONITORING EQUIPMENT
3y 1m to grant Granted Aug 11, 2026
Patent 12699908
KNOWLEDGE MANAGEMENT SYSTEM AND KNOWLEDGE MANAGEMENT METHOD
3y 3m to grant Granted Aug 04, 2026
Patent 12700754
SYSTEMS AND METHODS FOR MONITORING EQUIPMENT
3y 1m to grant Granted Aug 04, 2026
Patent 12700430
ELECTRONIC DEVICE, METHOD, AND COMPUTER READABLE STORAGE MEDIUM FOR EDITING VIDEO
1y 10m to grant Granted Aug 04, 2026
Patent 12694907
METHOD, APPARATUS, DEVICE, AND STORAGE MEDIUM FOR CONTENT INTERACTION
1y 8m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
92%
With Interview (+14.5%)
2y 9m (~1y 2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1044 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month