Prosecution Insights
Last updated: October 02, 2026
Application No. 19/042,319

METHOD, APPARATUS, DEVICE AND STORAGE MEDIUM FOR GENERATING MEDIA CONTENT

Non-Final OA §103§112
Filed
Jan 31, 2025
Priority
May 21, 2024 — CN 202410635356.6
Examiner
CHOKSHI, PINKAL R
Art Unit
2425
Tech Center
2400 — Computer Networks
Assignee
Beijing Youzhuju Network Technology Co., Ltd.
OA Round
2 (Non-Final)
61%
Grant Probability
Moderate
2-3
OA Rounds
1y 9m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
317 granted / 519 resolved
+3.1% vs TC avg
Strong +29% interview lift
Without
With
+29.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
26 currently pending
Career history
544
Total Applications
across all art units

Statute-Specific Performance

§101
4.8%
-35.2% vs TC avg
§103
65.1%
+25.1% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
13.2%
-26.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 519 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claim 1 have been considered but are moot because the arguments do not apply in view of newly found reference Brooks being used in the current rejection. Furthermore, newly added limitations to claims 1, 12, and 20 now invoke 112(a) rejection. See the new rejection below. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claims 1, 12, and 20 recite limitations “the image is a last frame of the first segment, and the second segment is a temporal extension of the first segment starting from the image”, the examiner is unable to find support for this limitation in the originally filed spec and/or drawings. Paragraphs (0023, 0033) of the specification disclose how the extension button is used to generate a new segment. However, nowhere in the spec and/or drawing discloses the image is a last frame of the first segment, and the second segment is a temporal extension of the first segment starting from the image. Application fails to provide adequate support required by 112(a) or the first paragraph of 35 USC 112 for above mentioned negative limitation in the detailed description. Applicant to provide support for this limitation. Claims 2-11 and 13-19 are rejected due to their dependency on the independent claims 1 and 12, respectively. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 8-16, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over US PG Pub 2024/0273796 to Edson (“Edson”) in view of US PG Pub 2014/0075335 to Hicks (“Hicks”) and US PG Pub 2025/0259362 to Brooks (“Brooks”). Regarding claim 1, “A method for generating media content” reads on the technique for generating animated image file using an image and text instructions (abstract) disclosed by Edson and represented in Fig. 3. As to “comprising: presenting a configuration interface based on a selection of first media content, wherein the configuration interface comprises an input control for inputting a prompt item” Edson discloses (¶0029, ¶0036, ¶0062) that the AIG module receives an image selection on the user interface as represented in Fig. 4D; (¶0063-¶0064) in response to user selecting the image file, the user interface displays a text prompt field that enables the user to input instructions in natural-language that may be used to affect edits to the selected image as represented in Fig. 4F. As to “acquiring the prompt item via the input control” Edson discloses (¶0063-¶0064) that the user inputs a text prompt instructing the model to change make the changes for the image file selected. As to “outputting second media content, wherein the second media content comprises…a second segment, … and the second segment is based on an image in the first media content and the acquired prompt item” Edson discloses (¶0066-¶0068, ¶0070, claim 1) that the system generates the animated image file based on the selected image file and the user inputted text prompt as represented in Fig. 4F-H. Edson meets all the limitations of the claim except “outputting a second media content, wherein the second media content comprises a first segment and a second segment, the first segment corresponds to the first media content.” However, Hicks discloses (¶0087-¶0089) that the system generates and displays edited image along with the original image as represented in Fig. 4E. Therefore, it would have been obvious to one of the ordinary skills in the art before the effective filing date of the invention to modify Edson’s system by generating second media content comprising first/original segment and second/edited segment as taught by Hicks in order to visually see the difference in the original image and the edited image on the screen (Hicks - ¶0020). Combination of Edson and Hicks meets all the limitations of the claim except “the image is a last frame of the first segment, and the second segment is a temporal extension of the first segment starting from the image.” However, Brooks discloses (¶0095) that the input prompt includes a prompt video as the visual media and the text instructs visual media generative response engine to generate an extended video in a time-forward dimension from the prompt video; the prompt requests that the output video extend the input video, the input frames would be located first, and the noisy frames that will be generated into new frames to extend the input video would be located subsequent to the input frames; a video responsive to the prompt that extended the input video with additional frames. Therefore, it would have been obvious to one of the ordinary skills in the art before the effective filing date of the invention to modify Edson and Hicks’ systems by outputting the second segment as an extension after the last frame of the first segment as taught by Brooks in order to create high-quality videos and longer-form outputs of any modality (Brooks - ¶0006). Regarding claim 2, “The method according to claim 1, wherein presenting the configuration interface comprises: receiving a selection of a generation mode; and presenting the configuration interface corresponding to the generation mode” Edson discloses (¶0060-¶0061) that the user selects to generate an animated image using one of the selection options on the screen as represented in Figs. 4B-4D. Regarding claim 3, “The method according to claim 1, wherein the configuration interface corresponds to a first generation mode, the configuration interface further comprises an image control, and the method further comprises: acquiring at least one reference image via the image control; and generating the second segment based on the image, the acquired prompt item and the at least one reference image” Edson discloses (¶0060-¶0061) that the user selects to generate an animated image using one of the selection options on the screen; (¶0029, ¶0036, ¶0062) the AIG module receives an image selection on the user interface as represented in Fig. 4D; (¶0063-¶0064) in response to user selecting the image file, the user interface displays a text prompt field that enables the user to input instructions in natural-language that may be used to affect edits to the selected image; (¶0066-¶0068, ¶0070, claim 1) the system generates the animated image file based on the selected image file and the user inputted text prompt as represented in Fig. 4F-H. Regarding claim 4, “The method according to claim 3, wherein the second segment comprises at least one video frame corresponding to the at least one reference image” Edson discloses (¶0066-¶0068, ¶0070, claim 1) that the system generates the animated image file based on the selected image file and the user inputted text prompt; (¶0076-¶0079) each of the frames extracted from an animated image file may be compiled and rendered into a single image and utilized as input for a fine-tuned model. Regarding claim 5, “The method according to claim 4, wherein a location of the at least one video frame in the second segment is determined based on a configuration operation of a user” Edson discloses (¶0047, ¶0050) the generative AI model may apply one or more computer vision algorithms to identify positions or angles of the features of each layer; the generative AI model outputs a natural-language description of the identified feature positions and angles based on the user input text instructions. Regarding claim 8, “The method according to claim 1, wherein the configuration interface further comprises a an input component, and the method further comprises: acquiring at least one media parameter via the input component, such that the second segment is generated further based on the at least one media parameter” Edson discloses (¶0050) that the user may input text instructions to change an angle or position of a described feature. The approved or modified natural-language text output of the generative AI model may be re-used by the same generative AI model or inputted into another model to generate the animated image file Regarding claim 9, “The method according to claim 8, wherein the at least one media parameter comprises at least one of: a first media parameter, used to indicate a motion amplitude of the segment to be generated; a second media parameter, used to indicate lens information of the segment to be generated; and a third media parameter, used to indicate proportional information of the segment to be generated” Edson discloses (¶0088, ¶0078-¶0079) that the user subsequently instructs the model to modify the image using a natural-language text prompt, such as “Move the cat's left paw to X degree” as represented in Fig. 5. Regarding claim 10, “The method according to claim 1, further comprising: displaying a candidate prompt item corresponding to generating the first media content at the input control for inputting the prompt item in the configuration interface; and determining the modified candidate prompt item as the acquired prompt item in response to the candidate prompt item being modified” Edson discloses (¶0040) that the AIG module generates a storyboard based on the text instructions. The storyboard may represent text descriptions or questions of features, e.g., visual elements, actions, changes, events, or the like, of the animated image described in the text instructions. In one embodiment, the storyboard may be generated by the generative AI model using an input prompt that includes the text instructions of the user and instructions to the generative AI model to generate questions about the inputted text instructions of the use. Regarding claim 11, “The method according to claim 1, further comprising: presenting at least one piece of media content generated based on a group of parameters; and receiving the selection of the first media content from the at least one piece of media content” Edson discloses (¶0061-¶0063) that the user is provided with multiple options to generate an animated image based on the selection of an image file as represented in Fig. 4. Regarding claim 12, see rejection similar to claim 1. Furthermore, Edson discloses (¶0005, ¶0097-¶0099) that the CRM device stores instruction executed by the processor to perform above mentioned method. Regarding claim 13, see rejection similar to claim 2. Regarding claim 14, see rejection similar to claim 3. Regarding claim 15, see rejection similar to claim 4. Regarding claim 16, see rejection similar to claim 5. Regarding claim 19, see rejection similar to claim 8. Regarding claim 20, see rejection similar to claim 1. Furthermore, Edson discloses (¶0005, ¶0097-¶0099) that the CRM device stores instruction executed by the processor to perform above mentioned method. Claims 6-7 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Edson in view of Hicks and Brooks as applied to claims 1-3 and 12 above, and further in view of US PG Pub 2023/0368534 to Vachhani (“Vachhani”). Regarding claim 6, “The method according to claim 3, wherein generating the second segment based on the image, the acquired prompt item and the at least one reference image comprises…generating the second segment…” Edson discloses (¶0060-¶0061) that the user selects to generate an animated image using one of the selection options on the screen; (¶0029, ¶0036, ¶0062) the AIG module receives an image selection on the user interface as represented in Fig. 4D; (¶0063-¶0064) in response to user selecting the image file, the user interface displays a text prompt field that enables the user to input instructions in natural-language that may be used to affect edits to the selected image; (¶0066-¶0068, ¶0070, claim 1) the system generates the animated image file based on the selected image file and the user inputted text prompt as represented in Fig. 4F-H. However, combination of Edson, Hicks, and Brooks does not explicitly teach “determining a reference start frame of the segment to be generated based on the image; determining a reference end frame of the segment to be generated based on the at least one reference image; and generating the second segment based on the reference start frame, the reference end frame and the acquired prompt item.” Vachhani discloses (abstract) a method for generating a segment of a video by the device by identifying a context associated with the video; (¶0075) the device analyzes the parameter in every frame of the video, where the parameter includes a subject, an environment, an action of the subject, an object, etc., by determining in the frame of an occurrence of a change in the parameter and generates the segment of the video comprising the frame at which there is a change in the parameter as a temporal boundary of the at least one segment. Therefore, it would have been obvious to one of the ordinary skills in the art before the effective filing date of the invention to modify Edson, Hicks, and Brooks’ systems by determining reference start/end frame to generate second segment as taught by Vachhani in order to identify and generate intelligent temporal segments for improving temporal consistency (Vachhani - ¶0008). Regarding claim 7, “The method according to claim 1, wherein the configuration interface corresponds to a second generation mode, and the method further comprises: determining a reference start frame of the segment to be generated based on the image; and generating the second segment based on the reference start frame and the acquired prompt item” combination of Edson and Vachhani teaches this limitation, where Edson discloses (¶0060-¶0061) that the user selects to generate an animated image using one of the selection options on the screen; (¶0029, ¶0036, ¶0062) the AIG module receives an image selection on the user interface as represented in Fig. 4D; (¶0063-¶0064) in response to user selecting the image file, the user interface displays a text prompt field that enables the user to input instructions in natural-language that may be used to affect edits to the selected image; (¶0066-¶0068, ¶0070, claim 1) the system generates the animated image file based on the selected image file and the user inputted text prompt as represented in Fig. 4F-H, and Vachhani discloses (abstract) a method for generating a segment of a video by the device by identifying a context associated with the video; (¶0075) the device analyzes the parameter in every frame of the video, where the parameter includes a subject, an environment, an action of the subject, an object, etc., by determining in the frame of an occurrence of a change in the parameter and generates the segment of the video comprising the frame at which there is a change in the parameter as a temporal boundary of the at least one segment. Therefore, it would have been obvious to one of the ordinary skills in the art before the effective filing date of the invention to modify Edson, Hicks, and Brooks’ systems by determining reference start/end frame to generate second segment as taught by Vachhani in order to identify and generate intelligent temporal segments for improving temporal consistency (Vachhani - ¶0008). Regarding claim 17, see rejection similar to claim 6. Regarding claim 18, see rejection similar to claim 7. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PINKAL R CHOKSHI whose telephone number is (571)270-3317. The examiner can normally be reached Monday - Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BRIAN T PENDLETON can be reached at (571)272-7527. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PINKAL R CHOKSHI/Primary Examiner, Art Unit 2425
Read full office action

Prosecution Timeline

Jan 31, 2025
Application Filed
Feb 13, 2026
Non-Final Rejection mailed — §103, §112
May 13, 2026
Response Filed
Jul 14, 2026
Final Rejection mailed — §103, §112
Sep 11, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748820
CLASSIFICATION DEVICE AND CLASSIFICATION METHOD
3y 9m to grant Granted Sep 29, 2026
Patent 12744956
DISPLAY DEVICE
2y 6m to grant Granted Sep 22, 2026
Patent 12732654
Systems and Methods for Customizing Channel Programming
2y 5m to grant Granted Sep 08, 2026
Patent 12726676
WIRELESS TRANSMISSION DEVICE, DISPLAY DEVICE AND DATA TRANSMISSION METHOD AND STREAM PROCESSING METHOD THEREOF
1y 10m to grant Granted Sep 01, 2026
Patent 12713112
SYSTEM AND A METHOD FOR GENERATING AND DISTRIBUTING MULTIMEDIA CONTENT
3y 2m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
61%
Grant Probability
90%
With Interview (+29.2%)
3y 5m (~1y 9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 519 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month