Prosecution Insights
Last updated: October 02, 2026
Application No. 19/068,044

METHOD AND DEVICE FOR GENERATING SYNTHETIC VIDEO DATA FROM A TEXT PROMPT

Non-Final OA §103
Filed
Mar 03, 2025
Priority
Mar 19, 2024 — EU 24 16 4422.8
Examiner
PUNTIER, CHRIS ALEJANDRO
Art Unit
Tech Center
Assignee
Robert Bosch GmbH
OA Round
1 (Non-Final)
95%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 95% — above average
95%
Career Allowance Rate
40 granted / 42 resolved
+35.2% vs TC avg
Moderate +7% lift
Without
With
+6.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
10 currently pending
Career history
48
Total Applications
across all art units

Statute-Specific Performance

§101
5.5%
-34.5% vs TC avg
§103
74.2%
+34.2% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 42 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Allowable Subject Matter Claim 6,7 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1,2,8,9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen(Chen, Haoxin, et al. "Videocrafter1: Open diffusion models for high-quality video generation." arXiv preprint arXiv:2310.19512 (2023).), in view of Hong(Hong, Susung, et al. "Direct2v: Large language models are frame-level directors for zero-shot text-to-video generation." arXiv preprint arXiv:2305.14330 (2023).) Regarding claim 1, Chen discloses A method for generating synthetic video data from a text prompt, including for providing video data for training and/or testing and/or verifying and/or validating a machine learning model, the method comprising the following steps:providing an input text prompt descriptive of content of the video data to be generated(at Fig. 1“We have open-sourced two diffusion models for video generation in VideoCrafter1. The Text-to-Video (T2V) model takes a text prompt as input and generates a video accordingly” This explicitly discloses receiving a text prompt as the input and generating a video according to that prompt.);generating a text embedding for each of the at least two text sub- prompts(Chen discloses at Section 3.1 “Attention(Q,K,V) = softmax QKT √ d ·V,where (5) Q=W(i) Q ·φi(zt),K = W(i) K ·ϕ(y),V = W(i) V ·ϕ(y). (6) φi(zt) ∈ RN× di ϵ represents spatially flattened tokens of video latent, ϕ denotes the Clip text encoder, and y is the input text prompt” Alongside figure 4, this explains that the U-Net backbone features are processed with the text and image embeddings via a dual cross-attention layer. This expressly teaches converting an input text prompt into a CLIP-based text representation that is then used by the video diffusion model.);and generating synthetic video data by a Video Diffusion Model based on the generated text embeddings(Chen discloses in Section 3.1 “The VideoCrafter T2V model is a Latent Video Diffusion Model (LVDM) [24] consisting of two key components: a video VAE and a video latent dif fusion model, as illustrated in Fig. 3.” Chen also states in connection with Figure 4 that the U-Net features are processed with “the text and image embeddings via a dual cross-attention layer.” This maps to the claim element because Chen expressly uses a latent video diffusion model, feeds the text-derived embedding into that model through cross-attention, and decodes the denoised latent into a generated video.) However, Chen does not disclose decomposing the provided text prompt into at least two text sub- prompts by a large language model. Hong does disclose decomposing the provided text prompt into at least two text sub- prompts by a large language model(Hong at paragraph 3 in the introduction states “To this end, we use LLM directors to divide user inputs into separate prompts for each frame, essentially separating static and dynamic elements within the user prompts” This maps directly to the claim element because the reference starts with a single prompt, processes it with a LLM and divide that prompt into multiple separate frame=level prompts.); It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to incorporate the teachings of Hong into the teachings of Chen in order to allow for multiple LLM-generated prompts to provide more detailed temporal guidance for generating video content that changes over time. Regarding claim 2, the combination of Chen and Hong disclose all the elements of claim 1 as discussed above. The combination also discloses wherein the text sub-prompts decompose the text prompt to describe sequential visual states of the content to be generated as the video data, wherein the sequential visual states of the content are to be represented in the generated video data by at least two frames(Hong discloses at paragraphs 3-4 of the introduction “In contrast to images, which can be described by one or a few sentences, videos contain sequences of time-varying actions and contexts, requiring much more de scriptive information … To address these limitations, we devise a novel frame work that leverages instruction-tuned large language models (LLMs) (Ouyang et al., 2022; Wei et al., 2021), such as GPT-4 (OpenAI, 2023) and PaLM2 (Google, 2023), for generating frame-by-frame descriptions in video creation from a single abstract user prompt. Starting from the analysis of LLMs’ ability to generate time-varying frame-level directions, we propose a method called DirecT2V, which enables zero-shot video creation by utilizing carefully designed task prompts tailored for instruction-tuned LLMs. To this end, we use LLM directors to divide user inputs into separate prompts for each frame, essentially separating static and dynamic elements within the user prompts” This maps closely to the “text sub-promps” corresponding to Hongs teaching of separate prompts for each frame, and those prompts describe the changing visual content at successive times. Hong’s Figure 3 also makes this concrete, generating distinct descriptions for multiple frames progressing from the initial scene. The successive frame descriptions show the cisual scene changing over time.) The rationale from claim 1 is incorporated herein. Claims 8 and 9, which are similar in scope to claim 1, thus rejected under the same rationale. Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen as modified by Hong as applied to claim 1 above, and further in view of Ungureanu(US-20250131604-A1). Regarding claim 3, the combination of Chen and Hong disclose all the elements of claim 1 as discussed above. However, the combination does not fully disclose wherein the method further comprises interpolating between two adjacent generated text embeddings to derive a text embedding for another intermediate frame of the to be generated video data. Ungureanu does disclose (para.[0047] “ In some cases, prompt encoder 310 interpolates between the prompt embedding and the expanded prompt embedding based on the diversity input 315 to generate an interpolated embedding. For instance, given a high diversity input (e.g., a 1.0 on a scale from 0 to 1.0), the interpolated embedding might be, or closely resemble, the expanded prompt embedding. Conversely, with a lower diversity input, the interpolated embedding might be, or closely resemble, the prompt embedding.” Chen teaches text embeddings used to condition Hong teaches successive frame-level text conditions and Hong teaches successive frame-level text conditions. Ungureanu teaches the known technique of interpolating between two text embeddings to obtain an intermediary embedding. Applying the interpolation technique to adjacent frame-level embeddings would predictably provide smoother intermediate semantic conditioning for intervening video frames. It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to incorporate the teachings of Ungureanu into the combinations of teachings of Chen and Hong in order to derive an intermediary text embedding representing semantic content between the two adjacent conditions and provide intermediate conditioning for an intervening video frame. Claim(s) 4,5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen as modified by Hong as applied to claim 1 above, and further in view of Qi (Qi, Chenyang, et al. "Fatezero: Fusing attentions for zero-shot text-based video editing." 2023 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2023.) Regarding claim 4, the combination of Chen and Hong disclose all the elements of claim 1 as discussed above. The combination also discloses wherein the Video Diffusion Model includes convolutional layers, at least one spatial transformer and at least one temporal transformer(Chen at Section 3.1 para. 3 teaches “As illustrated in Fig.3, the denoising U-Net is a 3D U-Net architecture consisting of a stack of basic spatial-temporal blocks with skip connections. Each block comprises convolutional layers, spatial transformers (ST), and temporal transformers (TT)…”), wherein the method further comprises the following steps:extracting an attention map form the temporal transformer(Chen also teaches this in the section Denoising U-Net. It states that each block includes temporal transformers and mathematically defines the temporal-attention operation. This establishes that temporal attention weights are generated internally and the attention weights are dependent on the temporal transformers are indicated by Fig. 3 and 4.); However, the combination does not disclose providing a delta attention map for regularization of the extracted attention map; and regularizing the attention map based on the delta attention map. Qi does disclose disclose providing a delta attention map for regularization of the extracted attention map; and regularizing the attention map based on the delta attention map(Chen teaches a Video diffusion model having temporal transformers that generate temporal attention as discussed above. While Qi teaches extracting spatial temporal attention maps and subsequently modifying an editing stage attention map based on a separately supplied stored attention map by attention fusion. Fig.2 and Fig.3 of Qi visually show the source self-attention + editing self-attention + blended self attention and the caption explains that the stored inversion maps are fused with the editing stage maps at each denoising step. While Qi does not call it a “delta attention map” the stored source attention map performs the analogous role of a n auxiliary attention map supplied to constrain the current attention map.) It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to incorporate the teachings of Qi into the combinations of teachings of Chen and Hong in order to constrain the temporal attention distribution and improve temporal consistency. Regarding claim 5, the combination of Chen, Hong, and Qi disclose all the elements of claim 4 as discussed above. Qi also discloses wherein the regularizing of the attention map includes: (i) transforming the extracted attention map and the delta attention map based on an affine transformation; or (ii) summing the extracted attention map with the delta attention map multiplied by a scaling factor(Qi teaches in the section Attention Map Blending “stfused = Mt⊙stedit +(1−Mt)⊙ stsrc .” This explains that the editing stage attention map and the stored inversion-stage attention map are blended with the binary mask to form a fused attention map. This is conceptually analogous to the limitation, where one attention map is combined by summation with another attention map after weighting.) The rationale from claim 4 is incorporated herein. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRIS ALEJANDRO PUNTIER whose telephone number is (703)756-1893. The examiner can normally be reached M-F 7:30-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CHRIS ALEJANDRO PUNTIER/ Examiner, Art Unit 2616 /DANIEL F HAJNIK/ Supervisory Patent Examiner, Art Unit 2616
Read full office action

Prosecution Timeline

Mar 03, 2025
Application Filed
Sep 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744012
LOCAL DIMMING PROCESSING ALGORITHM AND CORRECTION SYSTEM
3y 1m to grant Granted Sep 22, 2026
Patent 12725340
METHOD, DEVICE, AND PROGRAM PRODUCT FOR GENERATING AVATAR ANIMATION
2y 3m to grant Granted Sep 01, 2026
Patent 12705801
SYSTEMS AND METHODS FOR PERSONALIZED IMAGE GENERATION
2y 6m to grant Granted Aug 11, 2026
Patent 12694594
ENCODER, DECODER AND SCENE DESCRIPTION DATA SUPPORTING MULTIPLE ANIMATIONS AND/OR MOVEMENTS FOR AN OBJECT
3y 1m to grant Granted Jul 28, 2026
Patent 12682568
GENERATING COMPLETE THREE-DIMENSIONAL SCENE GEOMETRIES USING MACHINE LEARNING
3y 0m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
95%
Grant Probability
99%
With Interview (+6.9%)
2y 4m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 42 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month