DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 1 is objected to because of the following informalities: Claim 1 recites, “receiving a command to generate the visual media the text input received by the prompt editor.” Appropriate correction is required.
The Examiner recommends amending the claim to state, “receiving a command to generate the visual media based on the text input received by the prompt editor.”
Claims 14 and 18 recite a similar limitation, and these claims are objected to for the same reason.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 3 and 15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 3 recites, “provide at least one noisy frame corresponding to the duration, the resolution, or the aspect ratio to a diffusion-transformer layer to converted in to the visual media.” It appears that “to converted in to the visual media” is grammatically incorrect. It is unclear what has been converted, or whether “in to” should be “into.” It is unclear what is required to associate the claimed “noisy frame”/”diffusion-transformer layer” to the claimed “visual media.”
For the purposes of art rejection, the Examiner is reading the limitation as “provide at least one noisy frame corresponding to the duration, the resolution, or the aspect ratio to a diffusion-transformer layer to generate the visual media.”
Claim 15 is substantially similar to Claim 3, and Claim 15 is rejected for the similar reason.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 4-5, 7-11, 14, and 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over Martindale (“How to make AI videos from just text - Digital Trends”)
Regarding Claim 1, Martindale teaches A method comprising:
presenting a graphical user interface including a prompt editor (“Using Pika is as straightforward as typing your prompt into the text box at the bottom of the main page and pressing the Enter key. The prompt and the options you select can have a big impact on what the tool ends up outputting back to you.” Martindale p. 2.);
receiving at least a text input into the prompt editor as part of an input prompt (“Using Pika is as straightforward as typing your prompt into the text box at the bottom of the main page and pressing the Enter key.” Martindale p. 2.), wherein the text input is a prompt that describes a visual media to be generated by a visual media generative response engine (“Give Pika a subject, shot type, description of the scene, and a lighting effect. You can also choose a style like anime, cinematic, or pixel art.” Martindale p. 3.);
receiving a command to generate the visual media the text input received by the prompt editor (“Using Pika is as straightforward as typing your prompt into the text box at the bottom of the main page and pressing the Enter key.” Martindale p. 2.);
prior to generating the visual media (“Write or upload your prompt, but before you hit Enter, select the Video options icon -- it looks like the four disconnected corners of a square.” Martindale p. 4.), determining at least one of a duration, resolution, or aspect ratio in which to generate the visual media ( Martindale p. 4:
PNG
media_image1.png
686
660
media_image1.png
Greyscale
, where “aspect ratio” and “frames per second” could be set based on a graphical user interface.), wherein the visual media generative response engine is capable of generating the visual media in multiple durations (“Select the three-dot menu icon under the video, and select Add 4s. This will add your prompt to the input box, along with your video as a base and a command to add an additional four seconds.” Martindale p. 7.), resolutions (temporal resolution, which is frames per second of 8-24), and aspect ratios (Aspect ratio of 16:9, 9:16, 1:1, and 5:2);
receiving the visual media generated based on the prompt, the visual media was generated in the duration, the resolution, or the aspect ratio that was determined (
“Select the three-dot menu icon under the video, and select Add 4s. This will add your prompt to the input box, along with your video as a base and a command to add an additional four seconds.” Martindale p. 7.
temporal resolution, which is frames per second of 8-24
Aspect ratio of 16:9, 9:16, 1:1, and 5:2).
With respect to receiving a command, the Examiner mapped the claim feature to the command generated and received after “pressing the Enter key.” However, Martindale is not explicit.
The Examiner takes an Office Notice that it would have been well-known in the art that a command to execute a program could be initiated and received after “pressing the Enter key.” The benefits of combining this well-known knowledge would have made it convenient for a user to interact with the computing system. One click is quick and familiar to the user.
Regarding Claim 2, Martindale further teaches The method of claim 1, wherein the input prompt includes visual media and the text input (
“Select the three-dot menu icon under the video, and select Add 4s. This will add your prompt to the input box, along with your video as a base and a command to add an additional four seconds.” Martindale p. 7.
“You can also select the Image or video button and attach an image or video to provide additional context or data for the AI to use.” Martindale p. 3.).
Regarding Claim 4, Martindale further teaches The method of claim 2,
wherein the input prompt includes an image as the visual media and the text input instructs to generate a video from the image (
“Pika is an incredible artificial intelligence tool that takes text and images and creates impressive, believable, and often beautiful multi-second video clips.” Martindale p. 1.
“You can also select the Image or video button and attach an image or video to provide additional context or data for the AI to use.” Martindale p. 3.), and
the video responsive to the input prompt includes the image as part of the video responsive to the input prompt (“Use an image you generated in a separate text-to-image generator: Midjourney and Dall-E are great tools for creating a base model for Pika to animate. Try importing something you made there. Need help? Check out our guide on how to use Dall-E.” Martindale p. 3.).
Regarding Claim 5, Martindale further teaches The method of claim 2, wherein the input prompt includes a prompt video as the visual media and the text input instructs to generate an extended video in a time-forward or time-backward dimension from the prompt video, and the extended video responsive to the input prompt includes the prompt video with additional frames (
“Select the three-dot menu icon under the video, and select Add 4s. This will add your prompt to the input box, along with your video as a base and a command to add an additional four seconds.” Martindale p. 7. The additional seconds correspond to additional frames. Each inserted/added frame is either time-forward or time-backward with respect to an existing neighboring frame.).
Regarding Claim 7, Martindale further teaches The method of claim 2, wherein the input prompt includes a prompt video (“your video as a base”) as the visual media and the text input (“your prompt”) instructs to modify an aspect of the prompt video (“Add 4s”), and the video responsive to the input prompt includes the prompt video modified as instructed by the input prompt (“Select the three-dot menu icon under the video, and select Add 4s. This will add your prompt to the input box, along with your video as a base and a command to add an additional four seconds.” Martindale p. 7.).
Regarding Claim 8, Martindale further teaches The method of claim 2, wherein the graphical user interface further includes at least one of an aspect ratio input control or a resolution input control (Martindale p. 4:
PNG
media_image1.png
686
660
media_image1.png
Greyscale
, where “aspect ratio” and “frames per second”/temporal resolution could be set based on a graphical user interface),
wherein the determining of the at least one of the aspect ratio or the resolution in which to generate the visual media is determined based on explicit input (button for “Aspect ratio”; slider for “Frames per second”/temporal resolution) provided using the aspect ratio input control or the resolution input control (as shown in the figure.).
Regarding Claim 9, Martindale further teaches The method of claim 2, wherein the at least one of the aspect ratio or the resolution in which to generate the visual media is determined from an inference derived from the text input (
“Use negative prompts: If you find Pika keeps outputting certain video types you don't like, or they have something in them you would rather wasn't there, try using the negative prompts in the Parameters menu (the two lines with circles at opposite ends). You can input things like, ‘low resolution,’ or ‘morphing,’ or ‘blurry background,’ depending on what you're looking for.” Martindale p. 3.
The resolution is inferred from the text input “low resolution.”).
Regarding Claim 10, Martindale further teaches The method of claim 2, wherein the graphical user interface includes a visual media upload button that enables an input visual media to be included as part of the prompt (“You can also select the Image or video button and attach an image or video to provide additional context or data for the AI to use.” Martindale p. 3.).
Regarding Claim 11, Martindale further teaches The method of claim 2, wherein the graphical user interface further includes a duration input control, wherein the duration input control is effective to control a duration of the generated visual media (“Select the three-dot menu icon under the video, and select Add 4s. This will add your prompt to the input box, along with your video as a base and a command to add an additional four seconds.” Martindale p. 7.).
Claim 14 is substantially similar to Claim 1. The rejections analyses based on Marindale for Claim 1 are also applied to Claim 14. In addition, Claim 14 recites, “A system comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, configure the system to: . . . .” The Examiner takes an Official Notice that Pika is used on a computer that comprises “at least one processor; and a memory storing instructions that, when executed by the at least one processor, configure the system to” execute Pika. Pika is a well-known software product. The benefits of combining this well-known knowledge would have been to quickly process data and provide a user with better user experience.
Claim 16 is substantially similar to Claims 2+4. The rejections analyses based on Marindale for Claims 2+4 are also applied to Claim 16.
Claim 17 is substantially similar to Claim 2+7. The rejections analyses based on Marindale for Claim 2+7 are also applied to Claim 17.
Claim 18 is substantially similar to Claim 1. The rejections analyses based on Marindale for Claim 1 are also applied to Claim 18. In addition, Claim 18 recites, “A non-transitory computer-readable storage medium comprising instructions that when executed by at least one processor, cause the at least one processor to: . . . .” The Examiner takes an Official Notice that Pika is used on a computer that comprises “at least one processor; and a memory/non-transitory computer-readable storage medium storing instructions that, when executed by the at least one processor, configure the system to” execute Pika. Pika is a well-known software product. The benefits of combining this well-known knowledge would have been to quickly process data and provide a user with better user experience.
Regarding Claim 19, Martindale further teaches The non-transitory computer-readable storage medium of claim 18, wherein the input prompt includes visual media and the text input, wherein the graphical user interface further includes a duration input control, wherein the duration input control is effective to control a duration of the generated visual media (“Select the three-dot menu icon under the video, and select Add 4s. This will add your prompt to the input box, along with your video as a base and a command to add an additional four seconds.” Martindale p. 7.).
Claims 3 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Martindale as applied to Claim 1 or 14, in further view of Ma et al. (“Latte: Latent Diffusion Transformer for Video Generation”).1
Regarding Claim 3, Martindale teaches The method of claim 1.
Martindale does not explicitly disclose; however, Ma teaches further comprising:
providing at least one noisy frame (Ma 3.1 Preliminary of Latent Diffusion Models: “The key processes: diffusion and denoising. The diffusion process gradually introduces Gaussian noise into the latent code z, . . . .”) corresponding to the duration, the resolution, or the aspect ratio (
PNG
media_image2.png
286
592
media_image2.png
Greyscale
, where the number of video frames correspond to duration, height and width correspond to aspect ratio and resolution) to a diffusion-transformer layer (Ma Fig. 2, where “Each block depicted in light orange represents a Transformer block”) to converted in to the visual media (Ma Fig. 2: The pipeline of Latte for video generation.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Ma’s diffusion models with primary reference Marindale. One of ordinary skill in the art would be motivated to improve the quality of the media output. Ma states, “The prior work on the CNN-based video diffusion model proposes a joint image-video training strategy that greatly improves the quality of the generated videos (Ho et al, 2022). We explore whether this training strategy can also improve the performance of the Transformer-based video diffusion model.” Ma 3.3.4.
Claim 15 is substantially similar to Claim 3. The rejections analyses based on Marindale in view of Ma for Claim 3 are also applied to Claim 15.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Martindale as applied to Claim 2, in further view of WILSON et al. (US 20250200843 A1).
Martindale teaches The method of claim 2.
Martindale does not explicitly disclose; however, Wilson teaches wherein the input prompt includes at least two prompt videos as the visual media and the text input (340; after the combination of Martindale and Wilson, the text input would be similar to Martindale’s text input) instructs to create a blended video (360) that blends from a first of the at least two prompt videos (310) to a second of the at least two prompt videos (320), and the blended video (360) responsive to the input prompt includes aspects of the at least two prompt videos ( Wilson:
PNG
media_image3.png
474
726
media_image3.png
Greyscale
“As shown in FIG. 3A, control interface 200 is shown below an image 310 from a video signal received from a computing device of a first user, and an image 320 from a video signal received from a computing device of a second user. One of the users configures the control interface using a ‘generate from canvas’ composition mode and enters a ‘brainstorming’ activity with a ‘Hawaii vacation’ theme.” Wilson ¶ 41.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Wilson’s blending technique with primary reference Marindale. One of ordinary skill in the art would be motivated to integrate videos into one screen so that others could conveniently observe both videos simultaneous. This could be suitable for certain scenarios, e.g., virtual office.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Martindale as applied to Claim 2, in further view of HuggingFace (“huggingface/diffusers”).2
Regarding Claim 12, Marindale teaches The method of claim 2, wherein the graphical user interface further includes various controls (
PNG
media_image1.png
686
660
media_image1.png
Greyscale
, which the controls include those for “Aspect ratio” and “Frames per second.”).
Marindale does not explicitly discloses; however, HuggingFace teaches
a number of generations input control, wherein the number of generations input control is effective to cause the visual media generative response engine to generate more than one visual media (
PNG
media_image4.png
42
532
media_image4.png
Greyscale
After Marindale and HuggingFace are combined, Marindale’s graphical user interface is used to set “num_images_per_prompt.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine HuggingFace’s “num_images_per_prompt” with primary reference Marindale. One of ordinary skill in the art would be motivated to provide better services to a user. When a user could specify the number of candidate images/videos that the user receives, the user would have more choices if needed.
Claim 13 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Martindale as applied to Claim 2 or 18, in further view of Kotaru et al. (US 20240330589 A1).
Regarding Claim 13, Martindale further teaches The method of claim 2, wherein the graphical user interface includes a prompt enhancement option (the Parameters menu (the two lines with circles at opposite ends); enhancement through “negative prompts”), wherein, when selected (when Parameters menu selected), the prompt enhancement option is effective to input the prompt “input things like, ‘low resolution,’ or ‘morphing,’ or ‘blurry background,’” and the inputs are language input to be processed by a language model) to generate an enhanced prompt (prompt enhanced by negative prompt) and to replace the prompt with the enhanced prompt in the prompt editor (the replacement causes the prompt to change from
PNG
media_image5.png
220
402
media_image5.png
Greyscale
to
PNG
media_image6.png
208
306
media_image6.png
Greyscale
)(
“Use negative prompts: If you find Pika keeps outputting certain video types you don't like, or they have something in them you would rather wasn't there, try using the negative prompts in the Parameters menu (the two lines with circles at opposite ends). You can input things like, "low resolution," or "morphing," or "blurry background," depending on what you're looking for.” Martindale p. 3.).
Martindale does not explicitly disclose; however, Kotaru teaches input the prompt to a language modelto generate an enhanced prompt (
“ The prompt engineering approach involves augmenting the user's query with relevant text from a database before feeding it to the large language model. This approach allows for the system to account for updates or modifications in the database. On the other hand, fine-tuning requires a reasonable number (e.g., a couple hundred) of new training samples to be effective.” Kotaru ¶ 65. Here, a language model is used to enhance the language of the user’s query/prompt.
“Prompt engineering is a technique used in language models to fine-tune the model's output for a specific task by providing tailored prompts as inputs to the model. Prompt engineering involves crafting a specific prompt that elicits the desired response from the model. The prompt can include various elements, such as keywords, context, and formatting, and can be optimized using various techniques such as grid search or reinforcement learning. The goal is to create a prompt that provides the right amount of information to the model without being too prescriptive, allowing the model to generate accurate and relevant output.” Kotaru ¶ 52.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Kotaru’s prompt engineering with primary reference Marindale. One of ordinary skill in the art would be motivated to acquire better results from machine learning model. “Prompt engineering is a technique used in language models to fine-tune the model's output for a specific task by providing tailored prompts as inputs to the model. Prompt engineering involves crafting a specific prompt that elicits the desired response from the model.” Kotaru ¶ 52.
Claim 20 is substantially similar to Claim 2+13. The rejections analyses based on Marindale in view of Kotaru for Claim 2+13 are also applied to Claim 20.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZHENGXI LIU whose telephone number is (571)270-7509. The examiner can normally be reached M-F 9 AM - 5 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at 571-272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZHENGXI LIU/Primary Examiner, Art Unit 2611
1 Included in Applicant’s filed IDS.
2 https://web.archive.org/web/20221021050108/https://github.com/huggingface/diffusers/blob/main/src/diffusers/pipelines/stable_diffusion/pipeline_stable_diffusion_inpaint.py