Prosecution Insights
Last updated: August 14, 2026
Application No. 18/796,330

CONTENT GENERATION METHOD, COMPUTER DEVICE, AND STORAGE MEDIUM

Final Rejection §102
Filed
Aug 07, 2024
Priority
Sep 15, 2023 — CN 202311199451.8
Examiner
BARHAM, RYAN ALLEN
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Beijing Zitiao Network Technology Co., Ltd.
OA Round
2 (Final)
56%
Grant Probability
Moderate
3-4
OA Rounds
4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 56% of resolved cases
56%
Career Allowance Rate
9 granted / 16 resolved
-5.7% vs TC avg
Strong +54% interview lift
Without
With
+53.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
24 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
2.4%
-37.6% vs TC avg
§103
49.6%
+9.6% vs TC avg
§102
44.9%
+4.9% vs TC avg
§112
2.4%
-37.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 16 resolved cases

Office Action

§102
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Maschmeyer (US 20240320444 A1). Maschmeyer was filed on the same day as the claimed invention (09/15/2023); however, it claims domestic priority over US Provisional Application US 63491321 (filed 03/21/2023) and US Provisional Application US 63501526 (filed 05/11/2023). Thus, it was effectively filed before the effective filing date of the claimed invention. Regarding claim 1, Maschmeyer teaches a content generation method, comprising: acquiring a text (par. 0023: “In some implementations, the user edits may comprise at least one of: deletion of a portion of an output; replacement of a portion of an output; or addition of text or image.”), wherein description content for the role comprises content describing an appearance, a behavior, an identity, or emotions of the role (par. 0078: “More particularly, the input prompt may include a description of content that is requested to be generated by the generative AI model 112. By way of example, the input prompt may indicate various desired features, properties, or requirements for the requested content.”), and description content for the scenario comprises content describing a type, a scenery, a layout, time, weather, or an event of the scenario (par. 0092: “The first text prompt is an initial prompt that is supplied by the user to the generative model. In particular, the first text prompt may be or include a command/request to generate new content. For example, the first text prompt may include a description of content that is requested to be generated. Additionally, or alternatively, the first text prompt may indicate certain features, properties, or requirements for the requested content.”); generating at least one prompt word based on the text, wherein the at least one prompt word comprises at least one of a role prompt word corresponding to the role or a scenario prompt word corresponding to the scenario (par. 0044: “Upon receiving an indication from the user that they are finished making and/or configuring their selections and edits of the original outputs, the system is configured to modify the original prompt in accordance with the user selections/edits and the optional adherence weight. The modified prompt is then provided back to the model (e.g., LLM) for generation, and the returned output is presented as the final generation to the user.”); inputting the at least one prompt word into a content generation model, and generating at least one preview image based on the at least one prompt word, wherein prompt words associated with different preview images are at least partially different (par. 0045: “In some implementations, the system may provide a preview of the final generation based on the user's selections/edits on an ongoing basis. The “preview” may itself be a partial output (e.g., based on a reduced prompt) or otherwise a complete output of the LLM.”); in response to a first modification operation on a prompt word associated with a first preview image, inputting a modified prompt word into the content generation model, and generating a new preview image corresponding to the first preview image, wherein the first preview image is any one of the at least one preview image (par. 0044, as above); and generating multimedia content corresponding to the text based on the new preview image (par. 0030: “The output/result may comprise new content such as text, images, audio, and the like.”). Regarding claim 2, Maschmeyer teaches the content generation method according to claim 1, wherein the generating a prompt word based on the text comprises: splitting the text to obtain a plurality of text segments, wherein any text segment comprises: at least part of a first description content for the role, and/or at least part of a second description content for the scenario (par. 0063: “An example of how the transformer 50 may process textual input data is now described. Input to a language model (whether transformer-based or otherwise) typically is in the form of natural language as may be parsed into tokens. It should be appreciated that the term “token” in the context of language models and NLP has a different meaning from the use of the same term in other contexts such as data security. Tokenization, in the context of language models and NLP, refers to the process of parsing textual input (e.g., a character, a word, a phrase, a sentence, a paragraph, etc.) into a sequence of shorter segments that are converted to numerical representations referred to as tokens (or “compute tokens”). Typically, a token may be an integer that corresponds to the index of a text segment (e.g., a word) in a vocabulary dataset.”); and for each text segment of the plurality of text segments, performing semantic analysis on each text segment to obtain a prompt word corresponding to each text segment (par. 0064: “ An embedding 60 is a learned numerical representation (such as, for example, a vector) of a token that captures some semantic meaning of the text segment represented by the token 56. The embedding 60 represents the text segment corresponding to the token 56 in a way such that embeddings corresponding to semantically-related text are closer to each other in a vector space than embeddings corresponding to semantically-unrelated text.”). Regarding claim 3, Maschmeyer teaches the content generation method according to claim 2, wherein the inputting the prompt word into a content generation model, and generating at least one frame of preview image corresponding to the text, comprises: inputting prompt words respectively corresponding to the plurality of text segments into the content generation model, to obtain a preview image corresponding to each text segment (par. 0064: “In general, the token sequence that is inputted to the transformer 50 may be of any length up to a maximum length defined based on the dimensions of the transformer 50 (e.g., such a limit may be 2048 tokens in some LLMs). Each token 56 in the token sequence is converted into an embedding vector 60 (also referred to simply as an embedding). An embedding 60 is a learned numerical representation (such as, for example, a vector) of a token that captures some semantic meaning of the text segment represented by the token 56. The embedding 60 represents the text segment corresponding to the token 56 in a way such that embeddings corresponding to semantically-related text are closer to each other in a vector space than embeddings corresponding to semantically-unrelated text.”). Regarding claim 4, Maschmeyer teaches the content generation method according to claim 3, wherein the in response to a first modification operation on a prompt word associated with a first preview image, inputting a modified prompt word into the content generation model, and generating a new preview image corresponding to the first preview image comprises: in response to a first modification operation on a prompt word associated with any text segment, inputting a modified prompt word corresponding to the any text segment into the content generation model, and generating a new preview image corresponding to the any text segment (par. 0080: “The content generation engine 114 receives outputs that are produced by the generative AI model 112. Each output comprises content that is generated based on an input prompt to the generative AI model 112. The input prompt may be an initial prompt that is supplied by a user, or it may be a modified prompt derived by the content generation engine 114 based on revising an initial prompt. The outputs of the generative AI model 112 may be provided to the content generation engine 114 for further processing and refining, for example, to obtain a final content output.”). Regarding claim 5, Maschmeyer teaches the content generation method according to claim 4, further comprising: determining an associated preview image from other preview images other than the first preview image based on the modified prompt word corresponding to the any text segment (par. 0081: “The content generation engine 114 enables users to customize AI-generated content. More particularly, the content generation engine 114 supports selectively combining portions from different outputs of the generative AI model 112. A user can select portions from one or more of the outputs that are desired to be included in a final content output.”); and modifying the associated preview image based on the modified prompt word corresponding to the any text segment, to obtain a new preview image corresponding to the associated preview image (par. 0081: “In at least some implementations, the content generation engine 114 may iteratively perform the steps of receiving user selections of desired portions from outputs of the generative AI model, deriving modified input prompts based on the user selections, and providing the modified prompts to the generative AI model 112.”). Regarding claim 6, Maschmeyer teaches the content generation method according to claim 1, wherein before the generating multimedia content corresponding to the text based on the new preview image, the content generation method further comprises: generating caption information corresponding to the text (par. 0066: “Conceptually, the decoder 54 is designed to map the features represented by the feature vectors 62 into meaningful output, which may depend on the task that was assigned to the transformer 50. For example, if the transformer 50 is used for a translation task, the decoder 54 may map the feature vectors 62 into text output in a target language different from the language of the original tokens 56.”), and/or determining a timbre corresponding to the text (par. 0124: “The output of the model, i.e., a generated product description, may be displayed in the output display area 520. FIG. 5A shows a first text output 510a that is generated based on information inputted in the input field 502. A different product description may be generated if the user selects the “Try Again” button 506a. The user can change the language tone/style for the product description using the drop-down menu 504.”); and the generating multimedia content corresponding to the text based on the new preview image comprises: generating the multimedia content corresponding to the text based on the new image and at least one selected from a group consisting of the caption information and the timbre (par. 0124, as above). Regarding claim 7, Maschmeyer teaches the content generation method according to claim 6, wherein the determining a timbre corresponding to the text comprises: determining a sound feature of the role based on the text, and matching a corresponding timbre for the role based on the sound feature (par. 0100: “In some implementations, the second text prompt may be generated by adding, to the first text prompt, supplementary text that describes the selection(s). The supplementary text may comprise text describing properties of the user-selected portions. By way of example, the supplementary text may include text specifying one or more of a content type, location (e.g., absolute and/or relative location), size, etc. for each of the user-selected portions. The properties may be automatically determined by the computing system, for example, based on parsing the outputs and the user-selected portions using text and/or image processing algorithms.”); or receiving a timbre determined by a user from a plurality of candidate timbres (par. 0124, as above in claim 6 rejection). Regarding claim 8, Maschmeyer teaches the content generation method according to claim 1, further comprising: acquiring painting style information and/or image ratio information of the preview image (par. 0031: “The generative AI trained on such a data set is then able to take an input prompt in text form, which may include suggested topics, features, styles or other suggestions, and provide an output image that reflects, at least to some degree, the input prompt.”); and the inputting the prompt word into a content generation model and generating at least one frame of preview image corresponding to the text comprises: inputting the prompt word and at least one selected from the group consisting of the painting style information and the image ratio information into the content generation model, and generating the at least one frame of preview image corresponding to the text (par. 0031, as above). Regarding claim 9, Maschmeyer teaches the content generation method according to claim 1, wherein before the generating at least one frame of preview image corresponding to the text, the method further comprises: obtaining appearance feature information corresponding to the role, wherein the appearance feature information is obtained by performing role feature analysis on the text, and/or receiving the appearance feature information corresponding to the role input by a user (par. 0079: “In some implementations, a user-supplied prompt may be processed by the system 100 to generate a suitable prompt for inputting to the generative AI model 112.”); inputting the appearance feature information into the content generation model, to obtain a role image of the role (par. 0078: “An input prompt supplied by a user is received from the user device 120 via a network 150. In the context of content generation, the input prompt may be or include a command/request to generate content of a specific type. The input prompt may comprise text, images, audio, and/or other forms of unstructured data.”); and the inputting the prompt word into a content generation model, and generating at least one frame of preview image corresponding to the text, comprising: inputting the prompt word and the role image into the content generation model to generate the at least one frame of preview image corresponding to the text (par. 0079, as above). Regarding claim 10, Maschmeyer teaches the content generation method according to claim 9, further comprising: in response to a second modification operation on the appearance feature information corresponding to the role, generating a new role image of the role based on a modified appearance feature information (par. 0086: “The generative AI model 112 may iteratively produce new outputs based on modified prompts provided by the content generation engine 114, and the user may select an output from one of the iterations as the final content output.”); determining a preview image corresponding to the role from a plurality of preview images (par. 0121: “In some implementations, the computing system presents a preview of a final content output based on the user selection and editing input (operation 414).”); and modifying the role in the preview image corresponding to the role based on the new role image, to obtain a second preview image (par. 0126: “A second input prompt may be obtained by modifying the first input prompt, and the second input prompt is provided to the generative AI model as part of a request to generated the final text output 510c. The second input prompt may include, at least, a description of the user selections from and/or edits of the previously generated descriptions.”); the generating multimedia content corresponding to the text based on the new preview image comprising: generating the multimedia content corresponding to the text based on the second preview image and the new preview image (par. 0086, as above). Response to Arguments Applicant’s arguments, see Remarks, filed 05/19/2026, with respect to the rejection(s) of claim(s) 1-20 under Tao (US 20190325626 A1) have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Maschmeyer (US 20240320444 A1). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to RYAN A BARHAM whose telephone number is (571)272-4338. The examiner can normally be reached Mon-Fri, 8:30am-5pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu, can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /RYAN ALLEN BARHAM/Examiner, Art Unit 2613 /XIAO M WU/Supervisory Patent Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Aug 07, 2024
Application Filed
Mar 05, 2026
Non-Final Rejection mailed — §102
May 19, 2026
Response Filed
Jun 18, 2026
Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12657785
SIMULATING SHUTTER ROLLING EFFECT
2y 4m to grant Granted Jun 16, 2026
Patent 12646187
METHOD AND DEVICE FOR ALIGNING LASER POINT CLOUD AND IMAGE BASED ON DEEP LEARNING
2y 5m to grant Granted Jun 02, 2026
Patent 12639935
Visual Analytics Framework for Explainable Data Slicing-Based Model Validation
2y 5m to grant Granted May 26, 2026
Patent 12633031
STOCHASTIC TEXTURE FILTERING
2y 4m to grant Granted May 19, 2026
Patent 12564345
MEDICAL APPARATUS, AND IMAGE GENERATION METHOD FOR VISUALIZING TEMPORAL TRENDS OF BIOMAGNETIC DATA ON AN ORGAN MODEL
2y 10m to grant Granted Mar 03, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
56%
Grant Probability
99%
With Interview (+53.8%)
2y 4m (~4m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 16 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month