DETAILED ACTION
Claims 1-20 are pending in this application, claims 1-20 have been examined under the priority date of 09/11/2024 in accordance with applicant’s claim to the benefit of the previously filed provisional application.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 01/15/2026 and 08/13/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
Client device in claims 2, 11, 12, and 18.
Processing device in claim 10.
Memory component in claim 10.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 5-7, and 17-19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Fang (US 20250292472 A1).
Regarding claim 1 Fang discloses: A computer-implemented method comprising:
generating, from an image-to-video request comprising a digital image, a set of image tokens from the digital image (Fang, [0011]-[0012] a machine learning model is trained to complete a user initiated image-to-video task (image to video request) [0019] images are taken into the model transformer to generate patches of the input images [0020] multiple tokens are generated from each input by the transformer);
generating a set of anchor tokens from the set of image tokens by adding a timestep embedding to the set of image tokens that indicates that the set of anchor tokens are fully denoised (Fang, [0036] a sequence of grounding and visual tokens are generated for each frame in the sequence to be used in video generation [0038] tokens are generated for frames at timesteps marked by “i”, indicating a timestep is added to the image tokens described in [0036], [0046] the video data may be encoded into corresponding text prompt embeddings where these embedding predict the noise strength added [0051] the model performs iterative denoising using a latent diffusion model to reverse the noise addition to the variables, where timesteps and text embeddings (anchor tokens) are used to determine whether the tokens are denoised);
PNG
media_image1.png
478
374
media_image1.png
Greyscale
PNG
media_image2.png
66
370
media_image2.png
Greyscale
(Fang, [0051])
generating combined tokens from the set of anchor tokens and noised tokens that are generated from noise (Fang, [0046] the model takes the latent codes and generates noised latent codes (noised tokens));
PNG
media_image3.png
312
466
media_image3.png
Greyscale
(Fang, [0046])
generating, utilizing a diffusion transformer model to process the combined tokens, denoised tokens (Fang, [0051] a latent diffusion model is used to denoise the combined tokens, where the noise added inputs (combined tokens) are input into the model, then the noise is removed to output the denoised latent inputs (denoised token));
and generating a digital video comprising at least a portion of the digital image based on the denoised tokens (Fang, [0031] the system generates multiple frames from the encodings/tokens to generate a full video from the frames [0086] the video is generated from the denoised latent embeddings (output tokens)).
Regarding claim 2 Fang discloses; The computer-implemented method of claim 1, further comprising:
receiving, from a client device at inference time, the image-to-video request to generate the digital video comprising the digital image and the image-to-video request indicates that the digital image is to be portrayed in the digital video (Fang, [0013]- [0014] the system allows a user to selectively animate an input frame portion of an input frame to an Image-to-video model, where the user indicates the motion trajectory of the object in the image they wish to animate, [0035] a condition or initial frame is provided by the user to be animated indicating the frame is to be included in the video),
wherein the image-to-video request indicates that the digital image is to be included as a first frame of a sequence of frames, an intermediate frame of the sequence of frames, or a final frame of the sequence of frames (Fang, [0080]-[0082] the input includes a first image frame and a second image frame, where the trajectory of the objects motion to be animated in video form is shown, meaning that a first frame is a first location or specified as the first frame in the sequence).
Regarding claim 3 Fang discloses; The computer-implemented method of claim 1, wherein generating the set of image tokens comprises: generating, from the digital image of the image-to-video request, an embedding that represents the digital image (Fang, [0033] the autoencoder takes the image and maps it from pixel space to generate a latent embedding of the input image);
and generating, utilizing a tokenization model to break down the digital image into a plurality of image patches, the set of image tokens from the embedding (Fang, [0019] vision transformers may be used to generate patches of the input image feature map embeddings, [0020] for each input feature map tokens may be generated, [0037] the model generates tokens corresponding to multiple image patches/regions).
Regarding claim 5 Fang discloses; The computer-implemented method of claim 1, wherein generating the digital video comprises generating a sequence of frames from the denoised tokens (Fang, claim 1, the denoised output tokens are used to generate a video from the input image frame, [0086] the video is generated from the denoised latent embeddings (output tokens)),
wherein the digital video includes the digital image as at least one of a portion of a frame of the sequence of frames, one or more keyframes in the sequence of frames, or one or more motion frames in the sequence of frames (Fang, [0029] the input image or series of input images are used to generate the video frame sequence, [0035] – [0036] an initial frame is use to generate a motion sequence of subsequent frames (one or more motion frames generated from the conditional digital image frame.).
Regarding claim 6 Fang discloses; The computer-implemented method of claim 1, further comprising training the diffusion transformer model by: generating, from a frame of a sequence of training frames, a training embedding that represents the frame (Fang, [0019]-[0020] the model encoder takes the input image and generates a set of embeddings, which are then used to generate tokens from the image, [0033] the autoencoder takes the image and maps it from pixel space to generate a latent embedding of the input image);
generating, utilizing a tokenization model, a set of training tokens from the training embedding (Fang, [0019]- [0020] the model encoder takes the input image and generates a set of embeddings, which are then used to generate tokens from the image);
and generating, utilizing the tokenization model, noised training tokens from the sequence of training frames that does not include the frame (Fang, [0021]-[0022] there is a token for each embedding in the sequence which is used to train the transformer (tokenization model), [0046] as part of the training process the latent embeddings for the input image along with the text (tokens) are noised).
Regarding claim 7 Fang discloses; The computer-implemented method of claim 6, further comprising: generating training anchor tokens by concatenating timestep embeddings to the set of training tokens to indicate that the set of training tokens are fully denoised (Fang, [0038] visual tokens (anchor tokens) (vi) are generate at multiple timesteps (denoted as i), where, [0051] the input latent information from the encoder/visual tokens/anchor tokens are denoised at multiple timesteps);
and generating, utilizing the diffusion transformer model to process the training anchor tokens and the noised training tokens, denoised training tokens (Fang, [0051] a latent diffusion model is used to denoise the combined tokens, where the noise added inputs (combined tokens) are input into the model, then the noise is removed to output the denoised latent inputs (denoised token), [0016] notes that the models may be trained or re-trained on data as more becomes available, therefore the examiner is interpreting this as meaning all data generation steps may also be used as training data).
Regarding claim 17 Fang discloses; A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising (Fang, [0052] the system has a processing component coupled to a memory):
generating, from an image-to-video request comprising a digital image, a set of image tokens from the digital image (Fang, [0011]-[0012] a machine learning model is trained to complete a user initiated image-to-video task (image to video request) [0019] images are taken into the model transformer to generate patches of the input images [0020] multiple tokens are generated from each input by the transformer);
generating a set of anchor tokens from the set of image tokens by adding a timestep embedding to the set of image tokens that indicates that the set of anchor tokens are fully denoised (Fang, [0036] a sequence of grounding and visual tokens are generated for each frame in the sequence to be used in video generation [0038] tokens are generated for frames at timesteps marked by “i”, indicating a timestep is added to the image tokens described in [0036], [0046] the video data may be encoded into corresponding text prompt embeddings where these embedding predict the noise strength added [0051] the model performs iterative denoising using a latent diffusion model to reverse the noise addition to the variables, where timesteps and text embeddings (anchor tokens) are used to determine whether the tokens are denoised);
generating combined tokens from the set of anchor tokens and noised tokens that are generated from noise (Fang, [0046] the model takes the latent codes and generates noised latent codes (noised tokens));
generating, utilizing a diffusion transformer model to process the combined tokens, denoised tokens (Fang, [0051] a latent diffusion model is used to denoise the combined tokens, where the noise added inputs (combined tokens) are input into the model, then the noise is removed to output the denoised latent inputs (denoised token));
and generating a digital video comprising at least a portion of the digital image based on the denoised tokens (Fang, [0031] the system generates multiple frames from the encodings/tokens to generate a full video from the frames [0086] the video is generated from the denoised latent embeddings (output tokens)).
Regarding claim 18 Fang discloses; The non-transitory computer-readable medium of claim 17, wherein the operations further comprise: receiving, from a client device at inference time, the image-to-video request to generate the digital video comprising the digital image and the image-to-video request indicates that the digital image is to be portrayed in the digital video (Fang, [0013]- [0014] the system allows a user to selectively animate an input frame portion of an input frame to a Image-to-video model, where the user indicates the motion trajectory of the object in the image they wish to animate, [0035] a condition or initial frame is provided by the user to be animated indicating the frame is to be included in the video),
wherein the image-to-video request indicates that the digital image is to be included as a first frame of a sequence of frames, an intermediate frame of the sequence of frames, or a final frame of the sequence of frames (Fang, [0080]-[0082] the input includes a first image frame and a second image frame, where the trajectory of the objects motion to be animated in video form is shown, meaning that a first frame is a first location or specified as the first frame in the sequence).
Regarding claim 19 Fang discloses; The non-transitory computer-readable medium of claim 17, wherein the operations further comprise training the diffusion transformer model by adding noise to a subset of frames of a sequence of frames of a training video (Fang, [0046] each frame of the video data has noise added to it in latent space),
wherein the subset of frames does not include one or more frames that are anchor frames in the training video (Fang, [0046] the frame noise addition adds noise to the sequence of frames [0035] where the sequence of frames is generated from the input condition frame (anchor frame) but does not include the anchor frame itself).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
2. Claims 4 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Fang (US 20250292472 A1) in view of Zhang (US 20250175679 A1).
Regarding claim 4 Fang fails to disclose; The computer-implemented method of claim 1, wherein generating the denoised tokens comprises: initializing the noised tokens by sampling a random level of noise from a noise distribution;
and removing noise from the noised tokens according to the set of anchor tokens utilizing the diffusion transformer model.
However, in the same field of endeavor, Zhang teaches;
wherein generating the denoised tokens comprises initializing the noised tokens by sampling a random level of noise from a noise distribution (Zhang, [0030] random noise is added to latent representations of a text prompt and a latent representation of an image to generate a noised latent representation (noised token));
and removing noise from the noised tokens according to the set of anchor tokens utilizing the diffusion transformer model (Zhang, [0033] the diffusion denoising model iteratively denoises the combined latent inputs based on timesteps T, where the timesteps in combination with the visual latent variables (tokens) are functionally equivalent to the anchor tokens).
The combination of Fang and Zhang would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of random noise to the latent data allows the model to learn a set variable amount to reverse or reconstruct. (Zhang, [0029]- [[0037])
Regarding claim 8 the combination of Fang and Zhang teaches; The computer-implemented method of claim 7, further comprising: generating, utilizing a detokenization model, denoised training embeddings from the denoised training tokens (Fang, [0020] the tokens which are output by the encoder are taken in by the decoder of the transformer (detokenization model) and are then used to generate output embeddings, therefore [0051] the denoising done in latent space using an autoencoder would denoise the tokens to generate denoised tokens in latent space and output denoised latent embeddings since the autoencoder is an encoder-decoder pair which takes the text token inputs and outputs then as denoised latent embeddings via the decoder);
and comparing the denoised training embeddings with embeddings generated from the sequence of training frames prior to tokenization (Zhang, [0035] the denoised latent input is compared with the input latent representation);
and determining a measure of loss from comparing the denoised training embeddings with the embeddings generated from the sequence of training frames prior to tokenization to modify parameters of the diffusion transformer model (Zhang, [0035] the denoised latent input is compared with the input latent representation and [0036] a loss is computed based on this to fine tune the model).
The combination of Fang and Zhang would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the loss function of Zhang allows for comparison of the reconstructed or generated frames with the originals for improvement of the accuracy of the model. (Zhang, [0029]- [[0037])
Claims 9-16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Fang (US 20250292472 A1) in view of Azarian (US 20250166236 A1)
Regarding claim 9 Fang discloses; The computer-implemented method of claim 1, wherein generating the digital video comprises: generating, from a first pass of the combined tokens through the diffusion transformer model, a conditional token output from the denoised tokens (Fang, [0046] the model takes the latent codes and generates noised latent codes (noised tokens), [0051] where the noised latent codes (combined tokens) are conditioned by the text embeddings, therefore when the latent diffusion model denoises these, the generated denoised output (conditional token output) is conditioned using a condition variable (condition token));
Fang fails to teach; generating, from a second pass of additional combined tokens through the diffusion transformer model, an unconditional token output from additional denoised tokens;
generating a final token output by combining the conditional token output and the unconditional token output;
and generating the digital video comprising at least the portion of the digital image based on the final token output.
However, in the same field of endeavor, Azarian teaches; generating, from a second pass of additional combined tokens through the diffusion transformer model, an unconditional token output from additional denoised tokens (Azarian, [0051] a model generates an original noise prediction in the denoising based on a conditional probability input from an image, where the conditional probability is applied to the generate tokens in [0050] to output a conditional token output as notes in equation 1 (first pass), further [0052] an unconditional score is also generated an used to scale the noise based on this score according to equation 2 (second pass), [0054] both the conditional and the unconditional scores are applied to the tokens to generate conditional and unconditional tokens respectively);
generating a final token output by combining the conditional token output and the unconditional token output (Azarian, [0053]-[0054] the model generates a modified noise score based on both the conditional and unconditional generated scores and applies both to the tokens to base an output on the combination of the conditional and unconditional tokens as noted in equations 4 and 5, where these tokens are denoised by the model to generate a final output);
and generating the digital video comprising at least the portion of the digital image based on the final token output (Azarian, [0041]-[0042] the predicted noise is used to denoise the latent image representation and then generate an output latent image representation frame, where the denoised tokens are based on a combination of the conditional and unconditional token scoring output as described in [0050]-[0054], where [0063]-[0064] notes that the output image frames may generate a video output).
The combination of Fang and Azarian would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of the combined conditional and unconditional token information allows the model to generate images which are more aligned with a specific class or prompt. (Azarian, [0050]- [0055])
Regarding claim 10 the combination of Fang and Azarian teaches; A system comprising: a memory component (Fang, [0052] the system has a processing component coupled to a memory);
and one or more processing devices coupled to the memory component, the one or more processing devices to perform operations comprising (Fang, [0052] the system has a processing component coupled to a memory):
generating a set of image tokens from a digital image as part of an image-to-video request (Fang, [0011]-[0012] a machine learning model is trained to complete a user initiated image-to-video task (image to video request) [0019] images are taken into the model transformer to generate patches of the input images [0020] multiple tokens are generated from each input by the transformer);
generating, from a first pass of combined tokens through a trained diffusion transformer model, a conditional token output (Fang, [0046] the model takes the latent codes and generates noised latent codes (noised tokens), [0051] where the noised latent codes (combined tokens) are conditioned by the text embeddings, therefore when the latent diffusion model denoises these, the generated denoised output (conditional token output) is conditioned using a condition variable (condition token)),
wherein the combined tokens comprise a set of anchor tokens from the set of image tokens and noised tokens (Fang, [0036] a sequence of grounding and visual tokens are generated for each frame in the sequence to be used in video generation [0038] tokens are generated for frames at timesteps marked by “i”, indicating a timestep is added to the image tokens described in [0036], [0046] the video data may be encoded into corresponding text prompt embeddings where these embedding predict the noise strength added [0051] the model performs iterative denoising using a latent diffusion model to reverse the noise addition to the variables, where timesteps and text embeddings (anchor tokens) are used to determine whether the tokens are denoised);
generating, from a second pass of additional combined tokens through the trained diffusion transformer model, an unconditional token output (Azarian, [0051] a model generates an original noise prediction in the denoising based on a conditional probability input from an image, where the conditional probability is applied to the generate tokens in [0050] to output a conditional token output as notes in equation 1 (first pass), further [0052] an unconditional score is also generated an used to scale the noise based on this score according to equation 2 (second pass), [0054] both the conditional and the unconditional scores are applied to the tokens to generate conditional and unconditional tokens respectively),
wherein the additional combined tokens comprise the set of anchor tokens and additional noised tokens (Azarian, [0044] and [0045] the system generates the outputs and tokens by iterating through at multiple timesteps, therefore each token has a corresponding timestep, indicating that the inputs are anchor tokens);
generating a final token output by combining the conditional token output and the unconditional token output (Azarian, [0053]-[0054] the model generates a modified noise score based on both the conditional and unconditional generated scores and applies both to the tokens to base an output on the combination of the conditional and unconditional tokens as noted in equations 4 and 5, where these tokens are denoised by the model to generate a final output);
and generating a digital video comprising at least a portion of the digital image based on the final token output (Azarian, [0041]-[0042] the predicted noise is used to denoise the latent image representation and then generate an output latent image representation frame, where the denoised tokens are based on a combination of the conditional and unconditional token scoring output as described in [0050]-[0054], where [0063]-[0064] notes that the output image frames may generate a video output).
The combination of Fang and Azarian would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of the combined conditional and unconditional token information allows the model to generate images which are more aligned with a specific class or prompt. (Azarian, [0050]- [0055])
Regarding claim 11 the combination of Fang and Azarian teaches The system of claim 10, wherein the operations comprise receiving, from a client device at inference time, the image-to-video request to generate the digital video and a text prompt that indicates that the digital image is to be portrayed in the digital video (Fang, [0013]- [0014] the system allows a user to selectively animate an input frame portion of an input frame to an Image-to-video model, where the user indicates the motion trajectory of the object in the image they wish to animate, [0035] a condition or initial frame is provided by the user to be animated indicating the frame is to be included in the video).
Regarding claim 12 the combination of Fang and Azarian teaches The system of claim 10, wherein the operations comprise: receiving, from a client device, the image-to-video request that comprises a conditional prompt for the trained diffusion transformer model to include in the digital video and an unconditional prompt for the digital video (Fang, [0019] and [0020] the system takes as input an image and generates a vector input of the image (functionally equivalent to an unconditional input prompt, per [0029] of Fang) and as well as [0034] a condition image and a text prompt to be used for conditioning (conditional prompt)),
wherein the conditional prompt indicates that the digital image is to be portrayed as at least one of an initial frame, an intermediate frame, or a subset of frames in the digital video (Fang, [0080]-[0082] the condition input includes a first image frame and a second image frame, where the trajectory of the objects motion to be animated in video form is shown, meaning that a first frame is a first location or specified as the first frame in the sequence).
Regarding claim 13 the combination of Fang and Azarian teaches; The system of claim 10, wherein generating the combined tokens comprises: generating text tokens from a conditional prompt of a text prompt of the image-to-video request (Fang, [0030] a set of text prompt tokens are generated from a conditional input to the model);
generating an image embedding from the digital image (Fang, [0033] the autoencoder takes the image and maps it from pixel space to generate a latent embedding of the input image);
initializing the noised tokens from a noise distribution (Fang, [0046] the model takes the latent codes and generates noised latent codes (noised tokens), where the noise distribution is given in paragraph [0046] as well);
and combining the text tokens, the image embedding, and the noised tokens (Fang, [0046] the noised latent codes are combined with the text prompt tokens and the embedded video/image data).
Regarding claim 14 the combination of Fang and Azarian teaches The system of claim 13, further comprising generating the set of anchor tokens from the image embedding by: generating, utilizing a tokenization model to break down the digital image into a plurality of image patches, the set of image tokens from the image embedding (Fang, [0019] vision transformers may be used to generate patches of the input image feature map embeddings, [0020] for each input feature map tokens may be generated, [0037] the model generates tokens corresponding to multiple image patches/regions);
and adding a timestep embedding to the set of image tokens from the image embedding, wherein the timestep embedding indicates to the trained diffusion transformer model that the set of image tokens are fully denoised (Fang, [0033] the system uses iterative noise diffusion guided by timestamps to conduct denoising in latent space, [0046] the video data may be encoded into corresponding text prompt embeddings where these embedding predict the noise strength added [0051] the model performs iterative denoising using a latent diffusion model to reverse the noise addition to the variables, where timesteps and text embeddings (anchor tokens) are used to determine whether the tokens are denoised).
Regarding claim 15 the combination of Fang and Azarian teaches; The system of claim 10, wherein generating the additional combined tokens comprises: generating text tokens from an unconditional prompt of a text prompt of the image-to-video request (Fang, [0049]-[0050] the model may be prompted with fixed background filling text prompts (unconditional text prompt), where tokens are generated from this to fill the background conditions from this prompt);
generating an image embedding from the digital image (Fang, [0033] the autoencoder takes the image and maps it from pixel space to generate a latent embedding of the input image);
initializing the noised tokens from a noise distribution (Fang, [0046] the model takes the latent codes and generates noised latent codes (noised tokens), where the noise distribution is given in paragraph [0046] as well);
and combining the text tokens, the image embedding, and the noised tokens (Fang, [0046] the noised latent codes are combined with the text prompt tokens and the embedded video/image data);
Regarding claim 16 the combination of Fang and Azarian teaches; The system of claim 10, wherein combining the conditional token output and the unconditional token output comprises: interpolating, utilizing a classifier free guidance model, between the conditional token output and the unconditional token output (Azarian, [0052] the system uses classifier free guidance to generate an estimated modified noise prediction which is used in generating the final denoise latent/token output using both the conditional and unconditional model scores),
wherein interpolating comprises a guidance scale that encourages the trained diffusion transformer model to generate the digital video based on the conditional token output (Azarian, [0051]-[0052] the model scales the noise predication based on the conditional and unconditional scores which is used to encourage the model to generate frames/images which are more aligned with a conditional class);
and generating the final token output based on the interpolation of the classifier free guidance model (Azarian, [0053]-[0054] the model generates a modified noise score based on both the conditional and unconditional generated scores and applies both to the tokens to base an output on the combination of the conditional and unconditional tokens as noted in equations 4 and 5, where these tokens are denoised by the model to generate a final output, [0041]-[0042] the predicted noise is used to denoise the latent image representation and then generate an output latent image representation frame, where the denoised tokens are based on a combination of the conditional and unconditional token scoring output as described in [0050]-[0054], where [0063]-[0064] notes that the output image frames may generate a video output).
The combination of Fang and Azarian would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of the combined conditional and unconditional token information allows the model to generate images which are more aligned with a specific class or prompt. (Azarian, [0050]- [0055])
Regarding claim 20 the combination of Fang and Azarian teaches; The non-transitory computer-readable medium of claim 17, wherein generating the digital video comprises:
generating, from a first pass of combined tokens through a trained diffusion transformer model, a conditional token output (Fang, [0046] the model takes the latent codes and generates noised latent codes (noised tokens), [0051] where the noised latent codes (combined tokens) are conditioned by the text embeddings, therefore when the latent diffusion model denoises these, the generated denoised output (conditional token output) is conditioned using a condition variable (condition token)),
wherein the combined tokens comprise a set of anchor tokens from the set of image tokens and noised tokens (Fang, [0033] the system uses iterative noise diffusion guided by timestamps to conduct denoising in latent space, [0046] the video data may be encoded into corresponding text prompt embeddings where these embedding predict the noise strength added [0051] the model performs iterative denoising using a latent diffusion model to reverse the noise addition to the variables, where timesteps and text embeddings (anchor tokens) are used to determine whether the tokens are denoised);
generating, from a second pass of additional combined tokens through the diffusion transformer model, an unconditional token output from additional denoised tokens (Azarian, [0051] a model generates an original noise prediction in the denoising based on a conditional probability input from an image, where the conditional probability is applied to the generate tokens in [0050] to output a conditional token output as notes in equation 1 (first pass), further [0052] an unconditional score is also generated an used to scale the noise based on this score according to equation 2 (second pass), [0054] both the conditional and the unconditional scores are applied to the tokens to generate conditional and unconditional tokens respectively),
wherein the additional combined tokens comprise the set of anchor tokens and additional noised tokens (Azarian, [0044] and [0045] the system generates the outputs and tokens by iterating through at multiple timesteps, therefore each token has a corresponding timestep, indicating that the inputs are anchor tokens);
generating a final token output by combining the conditional token output and the unconditional token output (Azarian, [0053]-[0054] the model generates a modified noise score based on both the conditional and unconditional generated scores and applies both to the tokens to base an output on the combination of the conditional and unconditional tokens as noted in equations 4 and 5, where these tokens are denoised by the model to generate a final output);
and generating a digital video comprising at least a portion of the digital image based on the final token output (Azarian, [0041]-[0042] the predicted noise is used to denoise the latent image representation and then generate an output latent image representation frame, where the denoised tokens are based on a combination of the conditional and unconditional token scoring output as described in [0050]-[0054], where [0063]-[0064] notes that the output image frames may generate a video output).
The combination of Fang and Azarian would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of the combined conditional and unconditional token information allows the model to generate images which are more aligned with a specific class or prompt. (Azarian, [0050]- [0055])
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For a listing of prior art of record, please see the attached PTO-892 Notice of References Cited form.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORDAN M ELLIOTT whose telephone number is (703)756-5463. The examiner can normally be reached M-F 8AM-5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.M.E./Examiner, Art Unit 2666 /Molly Wilburn/Primary Examiner, Art Unit 2666