Prosecution Insights
Last updated: October 02, 2026
Application No. 18/948,902

MEDIA DATA GENERATION IN ASSOCATION WITH A GENERATIVE MODEL AND AN ADAPTER

Non-Final OA §102§103§112
Filed
Nov 15, 2024
Examiner
RIVERA-MARTINEZ, GUILLERMO M
Art Unit
2677
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
81%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
401 granted / 514 resolved
+16.0% vs TC avg
Minimal +3% lift
Without
With
+3.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
32 currently pending
Career history
547
Total Applications
across all art units

Statute-Specific Performance

§101
5.9%
-34.1% vs TC avg
§103
44.5%
+4.5% vs TC avg
§102
22.6%
-17.4% vs TC avg
§112
25.0%
-15.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 514 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION This Office action is in response to the Application filed on November 15, 2024. An action on the merits follows. Claims 1-20 are pending on the application. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 3 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 3 recites the limitation “the adapter is configured to approximate operation of a second set of one or more layers of the multiple layers of the generative model” in lines 1-3 of the claim. However, it is not clear if the claimed “approximate operation” term recited in lines 1-2 of the claim encompass embodiments corresponding to the claimed “first sampling operation of multiple sampling operations” previously recited in line 7 of claim 1, or if the claimed “approximate operation” term recited in lines 1-2 of claim 3 encompass embodiments corresponding to the claimed “first portion of the first sampling operation” previously recited in line 9 of claim 1, or if the claimed “approximate operation” term recited in lines 1-2 of claim 3 encompass embodiments corresponding to the claimed “second portion of the first sampling operation” previously recited in line 13 of claim 1, for example, or if the claimed “approximate operation” term recited in lines 1-2 of claim 3 encompass embodiments corresponding to another “approximate operation” different from the claimed “first sampling operation of multiple sampling operations” previously recited in line 7 of claim 1, or different from “first portion of the first sampling operation” previously recited in line 9 of claim 1, or different from the claimed “second portion of the first sampling operation” previously recited in line 13 of claim 1, respectively, for example. Additionally, the claimed “approximate operation” term recited in lines 1-2 of the claim is not defined by the claim. Therefore, based on above, the metes and bounds of the claim are not clearly set forth and the examiner cannot clearly determine which elements are encompassed by the claim language, which renders the claim indefinite. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 4-6, 8, 16, and 18-20 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Hinz et al. (US PG Publication No. 2024/0320789 A1), hereafter referred to as Hinz. Regarding claim 1, Hinz discloses a device (Par. [0004-7]: systems and methods for generating a high-resolution image… method, apparatus, and non-transitory computer readable medium for high-resolution image generation are described… apparatus and system for high-resolution image generation are described; Par. [0033-37]: systems and methods for generating a high-resolution image… system and an apparatus for image generation is described) comprising: a memory (e.g. Fig. 4, No. 410 and Fig. 20, No. 2010; Par. [0007]: apparatus and system include… one or more memory components; Par. [0037]: the system and apparatus include… one or more memory components) configured to store (Par. [0066-69]: FIG. 4 shows an example of an image generation apparatus 400 according to aspects of the present disclosure… image generation apparatus 400 includes… memory unit 410… Memory unit 410 includes one or more memory devices… memory is used to store; Par. [0280-283]: FIG. 20 shows an example of a computing device 2000 according to aspects of the present disclosure…. computing device 2000 includes… memory subsystem 2010… memory subsystem 2010 includes one or more memory devices. Memory subsystem 2010 is an example of, or includes aspects of, the memory unit as described with reference to FIG. 4… memory is used to store): a generative model (Par. [0004-7]: systems and methods for generating a high-resolution image… using a generative adversarial network… using the generative adversarial network to generate a high-resolution image… generating, using a generative adversarial network (GAN), a high-resolution image… a generative adversarial network (GAN)… the GAN trained to generate a high-resolution image; Par. [0028-37]: image generation using a machine learning model… example of a current machine learning model that can generate an image based on a text input is a generative adversarial network (GAN), which is trained to produce a final output by iteratively refining an output of a synthesis network… the GAN trained to generate a high-resolution image; Par. [0096]: Generative adversarial network (GAN) 445 is an example… GAN 445 is trained to generate a high-resolution image; Par. [0168]: machine learning model 900 includes… generative adversarial network (GAN) 925) including multiple layers (Par. [0077-79]: machine learning model 420 includes one or more ANNs… In ANNs, a hidden (or intermediate) layer includes hidden nodes and is located between an input layer and an output layer. Hidden layers perform nonlinear transformations of inputs entered into the network. Each hidden layer is trained to produce a defined output that contributes to a joint output of the output layer of the neural network. Hidden representations are machine-readable data representations of an input that are learned from a neural network's hidden layers and are produced by the output layer. As the neural network's understanding of the input improves as it is trained, the hidden representation is progressively differentiated from earlier iterations… Different layers perform different transformations on their inputs. The initial layer is known as the input layer and the last layer is known as the output layer; Par. [0134-138]: FIG. 6 shows an example of a U-Net 600… example shown includes U-Net 600, input features 605, initial neural network layer 610… down-sampling layer 620… final neural network layer 645… U-Net 600 receives additional input features to produce a conditionally generated output… the additional input features are combined with intermediate features 615 within U-Net 600 at one or more layers; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 applies one or more down-sampling layers followed by up-sampling layers… GAN 925 includes a series of up-sampling convolution layers; Par. [0177-181]: Machine learning model 1000 is an example… machine learning model 1000 includes text encoder 1005, mapping network 1010, and generative adversarial network (GAN) 1015… GAN 1015 includes multiple (e.g., three) down-sampling layers and multiple (e.g., seven) up-sampling layers/units); and an adapter (Par. [0047-48]: image generation apparatus 115 generates an adaptive convolution… image generation apparatus 115 generates the high-resolution image… using the adaptive convolution filter; Par. [0106-107]: GAN 445 includes adaptive convolution component 450… adaptive convolution component 450 generates an adaptive convolution filter… an adaptive convolution filter is a filter that can automatically adjust the filter's parameters based on the input data; Par. [0159-162]: GAN 825 performs an adaptive convolution filter process… the convolution blocks of GAN 825 comprise a series of up-sampling convolution layers… each convolution layer is enhanced with an adaptive convolution filter; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection; Par. [0217-229]: an image generation apparatus… uses a GAN of a machine learning model… the GAN generates an adaptive convolution filter from a bank of convolution filters… the image generation apparatus generates the image based on the adaptive convolution filter… the system generates an adaptive convolution filter… the system generates a high-resolution image corresponding to the low-resolution image based on the adaptive convolution filter… adaptive convolution filter 1425 is used in a convolution pipeline of the GAN; Par. [0243-245]: system performs a convolution process on the feature map based on the adaptive convolution filter… the operations of this step refer to, or may be performed by, a GAN… the learned features of the low-resolution images may be features that the adaptive convolution filter has learned to recognize for a specific task… the GAN is trained to process the feature map using the adaptive convolution filter to predict a high-resolution image); and one or more processors (e.g. Fig. 4, No. 405 and Fig. 20, No. 2005; Par. [0007]: apparatus and system include one or more processors; one or more memory components coupled with the one or more processors; Par. [0037]: system and apparatus include one or more processors; one or more memory components coupled with the one or more processors) configured to (Par. [0066-68]: image generation apparatus 400 includes processor unit 405… noise component 415, machine learning model 420, training component 460, or a combination thereof are implemented as software stored in a memory subsystem and executed by one or more processors… processor unit 405 is configured to execute computer-readable instructions stored in memory unit 410 to perform various functions… processor unit 405 comprises the one or more processors; Par. [0280-283]: computing device 2000 includes processor(s) 2005, memory subsystem 2010… memory subsystem 2010 includes one or more memory devices. Memory subsystem 2010 is an example of, or includes aspects of, the memory unit as described with reference to FIG. 4… memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor to perform various functions): obtain an input image frame (Par. [0006]: obtaining an input image; Par. [0185]: method include obtaining an input image; Par. [0194]: system obtains an input image; Par. [0281]: computing device 2000 includes one or more processors 2005 that can execute instructions stored in memory subsystem 2010 to obtain an input image); for a first sampling operation of multiple sampling operations, perform, based on the input image frame (Par. [0006]: obtaining an input image… generating, using a diffusion model, a low-resolution image based on the input image… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image; Par. [0028-40]: image generation using a machine learning model. Machine learning algorithms build a model based on sample data, known as training data, to make a prediction or a decision in response to an input… machine learning model that can generate an image based on a text input is a generative adversarial network (GAN), which is trained to produce a final output by iteratively refining an output of a synthesis network… and diffusion models… an iterative sampling process used by diffusion models… the diffusion model trained to generate a low-resolution image… and a generative adversarial network (GAN)… trained to generate a high-resolution image based on the low-resolution image… the low-resolution image is generated using multiple iterations of the diffusion model and the high-resolution image is generated using a single iteration of the GAN; Par. [0091-120]: diffusion model 435 is trained to generate a low-resolution image… diffusion model 435 takes the image embedding as input… the low-resolution image is generated using multiple iterations of diffusion model 435… the low-resolution image is generated using multiple iterations of diffusion model 435 and the high-resolution image is generated using a single iteration of GAN 445… Diffusion models function by iteratively adding noise to data during a forward diffusion process and then learning to recover the data by denoising the data during a reverse diffusion process; Par. [0174-185]: GAN 925 applies one or more down-sampling layers followed by up-sampling layers… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection… GAN 925 takes low-resolution image 970 (such as the image output by the diffusion model as described with reference to FIGS. 5 and 12) as input and generates high-resolution image 975 in response… obtaining an input image having a first resolution… using a diffusion model, a low-resolution image based on the input image… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image; for a first sampling operation of multiple sampling operations, perform, based on the input image frame (e.g. systems and methods for image generation using a machine learning model (i.e. a generative model) to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images (i.e. based on the input image frame), for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process (i.e. for a first, second, third… Nth sampling operation of multiple sampling operations, perform, based on the input image frame), as indicated above), for example): a first portion of the first sampling operation via a first set of one or more layers of the multiple layers of the generative model (Par. [0077-79]: machine learning model 420 includes one or more ANNs… In ANNs, a hidden (or intermediate) layer includes hidden nodes and is located between an input layer and an output layer. Hidden layers perform nonlinear transformations of inputs entered into the network. Each hidden layer is trained to produce a defined output that contributes to a joint output of the output layer of the neural network. Hidden representations are machine-readable data representations of an input that are learned from a neural network's hidden layers and are produced by the output layer. As the neural network's understanding of the input improves as it is trained, the hidden representation is progressively differentiated from earlier iterations… Different layers perform different transformations on their inputs. The initial layer is known as the input layer and the last layer is known as the output layer; Par. [0134-138]: FIG. 6 shows an example of a U-Net 600… The example shown includes U-Net 600, input features 605, initial neural network layer 610… down-sampling layer 620… final neural network layer 645… a diffusion model (such as the diffusion model described with reference to FIGS. 4-5) is based on an ANN architecture known as a U-Net… U-Net 600 receives input features 605, where input features 605 include an initial resolution and an initial number of channels, and processes input features 605 using an initial neural network layer 610 (e.g., a convolutional network layer) to produce intermediate features 615… intermediate features 615 are then down-sampled using a down-sampling layer 620 such that down-sampled features 625 have a resolution less than the initial resolution… this process is repeated multiple times, and then the process is reversed. For example, down-sampled features 625 are up-sampled using up-sampling process 630 to obtain up-sampled features 635… U-Net 600 receives additional input features to produce a conditionally generated output… the additional input features include a vector representation of an input prompt… the additional input features are combined with intermediate features 615 within U-Net 600 at one or more layers; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 applies one or more down-sampling layers followed by up-sampling layers; Par. [0177-181]: Machine learning model 1000 is an example… machine learning model 1000 includes text encoder 1005, mapping network 1010, and generative adversarial network (GAN) 1015… GAN 1015 takes style vector 1045 and low-resolution image 1050 as input and applies a down-sampling process followed by an up-sampling process to generate high-resolution image 1055… GAN 1015 includes multiple (e.g., three) down-sampling layers and multiple (e.g., seven) up-sampling layers/units); a first portion of the first sampling operation via a first set of one or more layers of the multiple layers of the generative model (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process, for example, by using different layers (i.e. multiple layers of the generative model) to perform different transformations (i.e. a first, second, third… Nth portion of a first, second, third… Nth sampling operation) on corresponding inputs, including the obtained input images, for example, including hidden (or intermediate) layers between an input layer and the output layer, an initial neural network layer, or multiple down-sampling layers and multiple up-sampling layers (i.e. a first, portion of the first sampling operation via a first set of one or more layers of the multiple layers of the generative model), as indicated above), for example), the first set of one or more layers including a first layer associated with a first resolution (Par. [0006]: obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; Par. [0135-137]: a diffusion model (such as the diffusion model described with reference to FIGS. 4-5) is based on an ANN architecture known as a U-Net… U-Net 600 receives input features 605, where input features 605 include an initial resolution… and processes input features 605 using an initial neural network layer 610 (e.g., a convolutional network layer) to produce intermediate features 615… intermediate features 615 are then down-sampled using a down-sampling layer 620 such that down-sampled features 625 have a resolution less than the initial resolution… this process is repeated multiple times; Par. [0185]: method include obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; Par. [0194-195]: system obtains an input image having a first resolution… noise component generates the input image having the first resolution (e.g., a noisy image) using a forward diffusion process (such as the forward diffusion process described with reference to FIGS. 5 and 12)… the system generates a low-resolution image based on the input image, where the low-resolution image has the first resolution; Par. [0281]: obtain an input image having a first resolution… generate, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; the first set of one or more layers including a first layer associated with a first resolution (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process, for example, by using different layers (i.e. multiple layers of the generative model) to perform different transformations (i.e. a first, second, third… Nth portion of a first, second, third… Nth sampling operation) on corresponding inputs, including the obtained input images, for example, including an initial neural network layer that processes input features having an initial resolution, such as an input image having a first resolution (i.e. a first, second, third… Nth layer associated with a first resolution), for example, or a down-sampling layer that generates down-sampled features that have the first resolution (the initial resolution) or a resolution less than the initial resolution (i.e. a first, second, third… Nth layer associated with a first resolution), as indicated above), for example); and a second portion of the first sampling operation via the adapter, the adapter associated with a second resolution that is different from the first resolution (Par. [0006]: obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; Par. [0047-48]: image generation apparatus 115 generates an adaptive convolution… image generation apparatus 115 generates the high-resolution image… using the adaptive convolution filter; Par. [0159-162]: GAN 825 performs an adaptive convolution filter process… the convolution blocks of GAN 825 comprise a series of up-sampling convolution layers… each convolution layer is enhanced with an adaptive convolution filter; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection; Par. [0185]: method include obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; Par. [0217-229]: an image generation apparatus… uses a GAN of a machine learning model… the GAN generates an adaptive convolution filter from a bank of convolution filters… the image generation apparatus generates the image based on the adaptive convolution filter… the system generates an adaptive convolution filter… the system generates a high-resolution image corresponding to the low-resolution image based on the adaptive convolution filter… adaptive convolution filter 1425 is used in a convolution pipeline of the GAN; Par. [0243-245]: system performs a convolution process on the feature map based on the adaptive convolution filter… the operations of this step refer to, or may be performed by, a GAN… the learned features of the low-resolution images may be features that the adaptive convolution filter has learned to recognize for a specific task… the GAN is trained to process the feature map using the adaptive convolution filter to predict a high-resolution image; Par. [0281]: obtain an input image having a first resolution… generate, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; the first set of one or more layers including a first layer associated with a first resolution… and generate, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; and a second portion of the first sampling operation via the adapter, the adapter associated with a second resolution that is different from the first resolution (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process, for example, by using different layers (i.e. multiple layers of the generative model) to perform different transformations (i.e. a first, second, third… Nth portion of a first, second, third… Nth sampling operation) on corresponding inputs, including the obtained input images, for example, including an adaptive convolution filter process (i.e. the adapter) that automatically adjust convolution filter's parameters based on the input data to generate the high-resolution images (i.e. and a second, third… Nth portion of the first, second, third… Nth sampling operation via the adapter), for example, in which each high-resolution image has a second resolution that is greater than the first resolution (i.e. the adapter associated with a second resolution that is different from the first resolution), as indicated above), for example); and output, based on the multiple sampling operations, one or more output image frames (Par. [0030-34]: machine learning model that can generate an image based on a text input is a generative adversarial network (GAN), which is trained to produce a final output by iteratively refining an output of a synthesis network… image generation system uses a text-conditioned diffusion model to first generate a low-resolution (e.g., 128×128 pixel) RGB image and then uses a text-conditioned GAN model to up-sample the diffusion model's output to a high-resolution (e.g., 1024×1024 pixel) image; Par. [0135-137]: a diffusion model (such as the diffusion model described with reference to FIGS. 4-5) is based on an ANN architecture known as a U-Net… U-Net 600 receives input features 605, where input features 605 include an initial resolution… and processes input features 605 using an initial neural network layer 610 (e.g., a convolutional network layer) to produce intermediate features 615… intermediate features 615 are then down-sampled using a down-sampling layer 620 such that down-sampled features 625 have a resolution less than the initial resolution… this process is repeated multiple times, and then the process is reversed. For example, down-sampled features 625 are up-sampled using up-sampling process 630 to obtain up-sampled features 635… the combination of intermediate features 615 and up-sampled features 635 are processed using final neural network layer 645 to produce output features 650; Par. [0174]: GAN 925 applies one or more down-sampling layers followed by up-sampling layers… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection as described with reference to FIGS. 13-14, followed by attention layers. Thus, GAN 925 takes low-resolution image 970 (such as the image output by the diffusion model as described with reference to FIGS. 5 and 12) as input and generates high-resolution image 975 in response; Par. [0199-200]: system generates a high-resolution image based on the low-resolution image using a generative adversarial network (GAN)… the GAN takes the output of the diffusion model (e.g., the low-resolution image or an embedding of the low-resolution image) as input and generates the high-resolution image by up-sampling the low-resolution image or the embedding of the low-resolution image… the GAN generates the high-resolution image by generating a feature map corresponding to the low-resolution image or the low-resolution image embedding and performing convolution processes on the feature map to obtain the high-resolution image; Par. [0224-245]: system generates a high-resolution image corresponding to the low-resolution image based on the adaptive convolution filter… performing the convolution process includes applying the adaptive convolution filter over the feature map… performing the convolution process generates output that captures the learned features of the low-resolution images, and the high-resolution images may be generated based on the output… GAN performs a convolution process on the feature map based on the adaptive convolution filter… the GAN is trained to process the feature map using the adaptive convolution filter to predict a high-resolution image; output, based on the multiple sampling operations, one or more output image frames (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. based on the multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images (i.e. one or more output image frames) based on the low-resolution images generated by the iterative sampling process (i.e. output, based on the multiple sampling operations, one or more output image frames), as indicated above), for example). Regarding claim 4, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), wherein the one or more processors are configured to, for a second sampling operation of the multiple sampling operations, perform, based on the input image frame: perform the second sampling operations via the multiple layers of the generative model (Par. [0077-79]: machine learning model 420 includes one or more ANNs… In ANNs, a hidden (or intermediate) layer includes hidden nodes and is located between an input layer and an output layer. Hidden layers perform nonlinear transformations of inputs entered into the network. Each hidden layer is trained to produce a defined output that contributes to a joint output of the output layer of the neural network. Hidden representations are machine-readable data representations of an input that are learned from a neural network's hidden layers and are produced by the output layer. As the neural network's understanding of the input improves as it is trained, the hidden representation is progressively differentiated from earlier iterations… Different layers perform different transformations on their inputs. The initial layer is known as the input layer and the last layer is known as the output layer; Par. [0134-138]: FIG. 6 shows an example of a U-Net 600… The example shown includes U-Net 600, input features 605, initial neural network layer 610… down-sampling layer 620… final neural network layer 645… a diffusion model (such as the diffusion model described with reference to FIGS. 4-5) is based on an ANN architecture known as a U-Net… U-Net 600 receives input features 605, where input features 605 include an initial resolution and an initial number of channels, and processes input features 605 using an initial neural network layer 610 (e.g., a convolutional network layer) to produce intermediate features 615… intermediate features 615 are then down-sampled using a down-sampling layer 620 such that down-sampled features 625 have a resolution less than the initial resolution… this process is repeated multiple times, and then the process is reversed. For example, down-sampled features 625 are up-sampled using up-sampling process 630 to obtain up-sampled features 635… U-Net 600 receives additional input features to produce a conditionally generated output… the additional input features include a vector representation of an input prompt… the additional input features are combined with intermediate features 615 within U-Net 600 at one or more layers; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 applies one or more down-sampling layers followed by up-sampling layers; Par. [0177-181]: Machine learning model 1000 is an example… machine learning model 1000 includes text encoder 1005, mapping network 1010, and generative adversarial network (GAN) 1015… GAN 1015 takes style vector 1045 and low-resolution image 1050 as input and applies a down-sampling process followed by an up-sampling process to generate high-resolution image 1055… GAN 1015 includes multiple (e.g., three) down-sampling layers and multiple (e.g., seven) up-sampling layers/units); for a second sampling operation of the multiple sampling operations, perform, based on the input image frame: perform the second sampling operations via the multiple layers of the generative model (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process, for example, by using different layers (i.e. multiple layers of the generative model) to perform different transformations (i.e. a first, second, third… Nth portion of a first, second, third… Nth sampling operation) on corresponding inputs, including the obtained input images, for example, including hidden (or intermediate) layers between an input layer and the output layer, an initial neural network layer, or multiple down-sampling layers and multiple up-sampling layers (i.e. for a second sampling operation of the multiple sampling operations, perform, based on the input image frame: perform the second sampling operations via the multiple layers of the generative model), as indicated above), for example). Regarding claim 5, claim 4 is incorporated and Hinz discloses the device (Par. [0004-7]), wherein the one or more processors are configured to, for a third sampling operation of the multiple sampling operations, perform, based on the input image frame: a first portion of the third sampling operation via the first set of one or more layers of the multiple layers of the generative model (Par. [0077-79]: machine learning model 420 includes one or more ANNs… In ANNs, a hidden (or intermediate) layer includes hidden nodes and is located between an input layer and an output layer. Hidden layers perform nonlinear transformations of inputs entered into the network. Each hidden layer is trained to produce a defined output that contributes to a joint output of the output layer of the neural network. Hidden representations are machine-readable data representations of an input that are learned from a neural network's hidden layers and are produced by the output layer. As the neural network's understanding of the input improves as it is trained, the hidden representation is progressively differentiated from earlier iterations… Different layers perform different transformations on their inputs. The initial layer is known as the input layer and the last layer is known as the output layer; Par. [0134-138]: FIG. 6 shows an example of a U-Net 600… The example shown includes U-Net 600, input features 605, initial neural network layer 610… down-sampling layer 620… final neural network layer 645… a diffusion model (such as the diffusion model described with reference to FIGS. 4-5) is based on an ANN architecture known as a U-Net… U-Net 600 receives input features 605, where input features 605 include an initial resolution and an initial number of channels, and processes input features 605 using an initial neural network layer 610 (e.g., a convolutional network layer) to produce intermediate features 615… intermediate features 615 are then down-sampled using a down-sampling layer 620 such that down-sampled features 625 have a resolution less than the initial resolution… this process is repeated multiple times, and then the process is reversed. For example, down-sampled features 625 are up-sampled using up-sampling process 630 to obtain up-sampled features 635… U-Net 600 receives additional input features to produce a conditionally generated output… the additional input features include a vector representation of an input prompt… the additional input features are combined with intermediate features 615 within U-Net 600 at one or more layers; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 applies one or more down-sampling layers followed by up-sampling layers; Par. [0177-181]: Machine learning model 1000 is an example… machine learning model 1000 includes text encoder 1005, mapping network 1010, and generative adversarial network (GAN) 1015… GAN 1015 takes style vector 1045 and low-resolution image 1050 as input and applies a down-sampling process followed by an up-sampling process to generate high-resolution image 1055… GAN 1015 includes multiple (e.g., three) down-sampling layers and multiple (e.g., seven) up-sampling layers/units); for a third sampling operation of the multiple sampling operations, perform, based on the input image frame: a first portion of the third sampling operation via the first set of one or more layers of the multiple layers of the generative model (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process, for example, by using different layers (i.e. multiple layers of the generative model) to perform different transformations (i.e. a first, second, third… Nth portion of a first, second, third… Nth sampling operation) on corresponding inputs, including the obtained input images, for example, including hidden (or intermediate) layers between an input layer and the output layer, an initial neural network layer, or multiple down-sampling layers and multiple up-sampling layers (i.e. for a third sampling operation of the multiple sampling operations, perform, based on the input image frame: a first portion of the third sampling operation via the first set of one or more layers of the multiple layers of the generative model), as indicated above), for example); and a second portion of the third sampling operation via the adapter (Par. [0006]: obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; Par. [0047-48]: image generation apparatus 115 generates an adaptive convolution… image generation apparatus 115 generates the high-resolution image… using the adaptive convolution filter; Par. [0159-162]: GAN 825 performs an adaptive convolution filter process… the convolution blocks of GAN 825 comprise a series of up-sampling convolution layers… each convolution layer is enhanced with an adaptive convolution filter; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection; Par. [0185]: method include obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; Par. [0217-229]: an image generation apparatus… uses a GAN of a machine learning model… the GAN generates an adaptive convolution filter from a bank of convolution filters… the image generation apparatus generates the image based on the adaptive convolution filter… the system generates an adaptive convolution filter… the system generates a high-resolution image corresponding to the low-resolution image based on the adaptive convolution filter… adaptive convolution filter 1425 is used in a convolution pipeline of the GAN; Par. [0243-245]: system performs a convolution process on the feature map based on the adaptive convolution filter… the operations of this step refer to, or may be performed by, a GAN… the learned features of the low-resolution images may be features that the adaptive convolution filter has learned to recognize for a specific task… the GAN is trained to process the feature map using the adaptive convolution filter to predict a high-resolution image; Par. [0281]: obtain an input image having a first resolution… generate, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; the first set of one or more layers including a first layer associated with a first resolution… and generate, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; a second portion of the third sampling operation via the adapter (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process, for example, by using different layers (i.e. multiple layers of the generative model) to perform different transformations (i.e. a first, second, third… Nth portion of a first, second, third… Nth sampling operation) on corresponding inputs, including the obtained input images, for example, including an adaptive convolution filter process (i.e. the adapter) that automatically adjust convolution filter's parameters based on the input data to generate the high-resolution images (i.e. and a second, third… Nth portion of the first, second, third… Nth sampling operation via the adapter), for example, in which each high-resolution image has a second resolution that is greater than the first resolution (i.e. the adapter associated with a second resolution that is different from the first resolution), as indicated above), for example). Regarding claim 6, claim 5 is incorporated and Hinz discloses the device (Par. [0004-7]), wherein: the second sampling operation is performed after the first sampling operation; and the third sampling operation is performed after the second sampling operation (Par. [0006]: obtaining an input image… generating, using a diffusion model, a low-resolution image based on the input image… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image; Par. [0028-40]: image generation using a machine learning model. Machine learning algorithms build a model based on sample data, known as training data, to make a prediction or a decision in response to an input… machine learning model that can generate an image based on a text input is a generative adversarial network (GAN), which is trained to produce a final output by iteratively refining an output of a synthesis network… and diffusion models… an iterative sampling process used by diffusion models… the diffusion model trained to generate a low-resolution image… and a generative adversarial network (GAN)… trained to generate a high-resolution image based on the low-resolution image… the low-resolution image is generated using multiple iterations of the diffusion model and the high-resolution image is generated using a single iteration of the GAN; Par. [0091-120]: diffusion model 435 is trained to generate a low-resolution image… diffusion model 435 takes the image embedding as input… the low-resolution image is generated using multiple iterations of diffusion model 435… the low-resolution image is generated using multiple iterations of diffusion model 435 and the high-resolution image is generated using a single iteration of GAN 445… Diffusion models function by iteratively adding noise to data during a forward diffusion process and then learning to recover the data by denoising the data during a reverse diffusion process; Par. [0174-185]: GAN 925 applies one or more down-sampling layers followed by up-sampling layers… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection… GAN 925 takes low-resolution image 970 (such as the image output by the diffusion model as described with reference to FIGS. 5 and 12) as input and generates high-resolution image 975 in response… obtaining an input image having a first resolution… using a diffusion model, a low-resolution image based on the input image… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image; wherein: the second sampling operation is performed after the first sampling operation; and the third sampling operation is performed after the second sampling operation (e.g. systems and methods for image generation using a machine learning model (i.e. a generative model) to generate high-resolution images include an iterative sampling process (i.e. wherein: the second sampling operation is performed after the first sampling operation; and the third sampling operation is performed after the second sampling operation) used by diffusion models trained to generate low-resolution images based on obtained input images (i.e. based on the input image frame), for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process (i.e. for a first, second, third… Nth sampling operation of multiple sampling operations, perform, based on the input image frame), as indicated above), for example). Regarding claim 8, claim 4 is incorporated and Hinz discloses the device (Par. [0004-7]), wherein: the second sampling operation is performed prior to the first sampling operation (Par. [0006]: obtaining an input image… generating, using a diffusion model, a low-resolution image based on the input image… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image; Par. [0028-40]: image generation using a machine learning model. Machine learning algorithms build a model based on sample data, known as training data, to make a prediction or a decision in response to an input… machine learning model that can generate an image based on a text input is a generative adversarial network (GAN), which is trained to produce a final output by iteratively refining an output of a synthesis network… and diffusion models… an iterative sampling process used by diffusion models… the diffusion model trained to generate a low-resolution image… and a generative adversarial network (GAN)… trained to generate a high-resolution image based on the low-resolution image… the low-resolution image is generated using multiple iterations of the diffusion model and the high-resolution image is generated using a single iteration of the GAN; Par. [0091-120]: diffusion model 435 is trained to generate a low-resolution image… diffusion model 435 takes the image embedding as input… the low-resolution image is generated using multiple iterations of diffusion model 435… the low-resolution image is generated using multiple iterations of diffusion model 435 and the high-resolution image is generated using a single iteration of GAN 445… Diffusion models function by iteratively adding noise to data during a forward diffusion process and then learning to recover the data by denoising the data during a reverse diffusion process; Par. [0174-185]: GAN 925 applies one or more down-sampling layers followed by up-sampling layers… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection… GAN 925 takes low-resolution image 970 (such as the image output by the diffusion model as described with reference to FIGS. 5 and 12) as input and generates high-resolution image 975 in response… obtaining an input image having a first resolution… using a diffusion model, a low-resolution image based on the input image… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image; wherein: the second sampling operation is performed prior to the first sampling operation (e.g. systems and methods for image generation using a machine learning model (i.e. a generative model) to generate high-resolution images include an iterative sampling process (i.e. wherein: the second sampling operation is performed prior to the first sampling operation used by diffusion models trained to generate low-resolution images based on obtained input images (i.e. based on the input image frame), for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process (i.e. for a first, second, third… Nth sampling operation of multiple sampling operations, perform, based on the input image frame), as indicated above), for example); and the multiple layers of the generative model include the first layer associated with the first resolution and a second layer associated with the second resolution (Par. [0006]: obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; Par. [0047-48]: image generation apparatus 115 generates an adaptive convolution… image generation apparatus 115 generates the high-resolution image… using the adaptive convolution filter; Par. [0159-162]: GAN 825 performs an adaptive convolution filter process… the convolution blocks of GAN 825 comprise a series of up-sampling convolution layers… each convolution layer is enhanced with an adaptive convolution filter; Par. [0168-174]: machine learning model 900 includes… generative adversarial network (GAN) 925… GAN 925 includes a series of up-sampling convolution layers, where convolution block 930 is enhanced with a sample-adaptive kernel selection; Par. [0185]: method include obtaining an input image having a first resolution… generating, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution… and generating, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; Par. [0217-229]: an image generation apparatus… uses a GAN of a machine learning model… the GAN generates an adaptive convolution filter from a bank of convolution filters… the image generation apparatus generates the image based on the adaptive convolution filter… the system generates an adaptive convolution filter… the system generates a high-resolution image corresponding to the low-resolution image based on the adaptive convolution filter… adaptive convolution filter 1425 is used in a convolution pipeline of the GAN; Par. [0243-245]: system performs a convolution process on the feature map based on the adaptive convolution filter… the operations of this step refer to, or may be performed by, a GAN… the learned features of the low-resolution images may be features that the adaptive convolution filter has learned to recognize for a specific task… the GAN is trained to process the feature map using the adaptive convolution filter to predict a high-resolution image; Par. [0281]: obtain an input image having a first resolution… generate, using a diffusion model, a low-resolution image based on the input image, wherein the low-resolution image has the first resolution; the first set of one or more layers including a first layer associated with a first resolution… and generate, using a generative adversarial network (GAN), a high-resolution image based on the low-resolution image, wherein the high-resolution image has a second resolution that is greater than the first resolution; and the multiple layers of the generative model include the first layer associated with the first resolution and a second layer associated with the second resolution (e.g. systems and methods for image generation using a machine learning model to generate high-resolution images include an iterative sampling process (i.e. a first, second, third… Nth sampling operation of multiple sampling operations) used by diffusion models trained to generate low-resolution images based on obtained input images, for example, and a generative adversarial network (GAN) trained to generate high-resolution images based on the low-resolution images generated by the iterative sampling process, for example, by using different layers (i.e. multiple layers of the generative model) to perform different transformations (i.e. and the multiple layers of the generative model include the first layer associated with the first resolution and a second layer associated with the second resolution) on corresponding inputs, including the obtained input images, for example, including an adaptive convolution filter process (i.e. the adapter) that automatically adjust convolution filter's parameters based on the input data to generate the high-resolution images (i.e. and a second, third… Nth portion of the first, second, third… Nth sampling operation via the adapter), for example, in which each high-resolution image has a second resolution that is greater than the first resolution (i.e. the adapter associated with a second resolution that is different from the first resolution), as indicated above), for example). Regarding claim 16, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), further comprising a modem coupled to the one or more processors, the modem configured to transmit the one or more output image frames to a second device for output by the second device (Par. [0068]: Processor unit 405 includes one or more processors… processor unit 405 includes special purpose components for modem processing, baseband processing, digital signal processing, or transmission processing. According to some aspects, processor unit 405 comprises the one or more processors; Par. [0282-285]: processor(s) 2005 are included in the processor unit… a processor includes special-purpose components for modem processing, baseband processing, digital signal processing, or transmission processing… I/O interface 2020 is controlled by an I/O controller to manage input and output signals for computing device 2000… the I/O controller represents or interacts with a modem, a keyboard, a mouse, a touchscreen, or a similar device. In some cases, the I/O controller is implemented as a component of a processor. In some cases, a user interacts with a device via I/O interface 2020 or via hardware components controlled by the I/O controller). Regarding claim 18, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), wherein the one or more processors are integrated in at least one of a mobile phone, a tablet computer device, a wearable electronic device, a virtual reality headset, a mixed reality headset, an augmented reality headset, or a camera device (Par. [0049-52]: user device 110 is a personal computer, laptop computer, mainframe computer, palmtop computer, personal assistant, mobile device, or any other suitable processing apparatus… image generation apparatus 115 includes a computer implemented network… image generation apparatus 115 also includes one or more processors… image generation apparatus 115 is implemented on a server… the server comprises a general-purpose computing device, a personal computer, a laptop computer, a mainframe computer, a supercomputer, or any other suitable processing apparatus). Regarding claim 19, is a corresponding method claim rejected as applied to the apparatus claim 1 above. Regarding claim 20, Hinz discloses a non-transitory computer-readable medium that stores instructions that are executable by one or more processors to cause the one or more processors to (Par. [0006]: method, apparatus, and non-transitory computer readable medium for high-resolution image generation are described; Par. [0056]: FIG. 2 shows an example of a method 200 for image generation according to aspects of the present disclosure… these operations are performed by a system including a processor executing a set of codes to control functional elements of an apparatus; Par. [0068-69]: processor unit 405 is configured to execute computer-readable instructions stored in memory unit 410 to perform various functions… memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor of processor unit 405 to perform various functions described herein; Par. [0258]: operations are performed by a system including a processor executing a set of codes to control functional elements of an apparatus; Par. [0281-290]: computing device 2000 includes one or more processors 2005 that can execute instructions stored in memory subsystem 2010… a processor is configured to execute computer-readable instructions stored in a memory to perform various functions… memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor to perform various functions described herein… the functions described herein may be implemented in hardware or software and may be executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored in the form of instructions or code on a computer-readable medium… Computer-readable media includes… non-transitory computer storage media… A non-transitory storage medium may be any available medium that can be accessed by a computer. For example, non-transitory computer-readable media can comprise random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disk (CD) or other optical disk storage, magnetic disk storage, or any other non-transitory medium for carrying or storing data or code). The steps of the program further recited in claim 20 correspond to claim 1 when executed and are rejected as applied to apparatus claim 1 above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Hinz, as applied to claim 1 above, in view of Gupta et al. (US PG Publication No. 2024/0155071 A1), hereafter referred to as, in view of Gupta. Regarding claim 2, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), wherein: the generative model has a U-Net architecture (Par. [0092-103]: diffusion model 435 comprises a latent diffusion model… diffusion model 435 comprises a U-Net… the synthesis network of GAN 445 includes an encoder and a decoder with a skip connection in a U-Net architecture. For example, a layer of the decoder is connected to a layer of the encoder by a skip connection in a U-Net architecture), but fails to teach the following as further recited in claim 2. However, Gupta teaches the generative model includes an image-to-video generative model; the generative model includes an image-to-video generative model: or a combination thereof (Par. [0001-5]: an adaptive artificial intelligence (AI) text-video generation model that leverages text-to-image generation… systems and methods for text-to-video generation… method includes receiving a text input, generating a representation frame based on text embeddings of the text input, generating a set of frames based on the representation frame and a first frame rate, interpolating the set of frames based on a second frame rate, generating a first video based on the interpolated set of frames… system configured for text-to-video generation. The system includes one or more processors, and a memory storing instructions which, when executed by the one or more processors, cause the system to perform operations. The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames… a non-transient computer-readable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method(s) for text-to-video generation. The instructions causing the one or more processors to receive a text input, generate a representation frame based on text embeddings of the text input, decoding a set of frames conditioned on the representation frame and a frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames; Par. [0069-73]: representation frame is generated using a text-to-image model… generate a set of frames based on the representation frame and a first frame rate. The first frame rate may be predefined by the system. The set of frames may be combined to generate a low-resolution video… The video generating module 650 is configured to generate a first video with an increased frame rate based on the interpolated set of frames; the generative model includes an image-to-video generative model (e.g. adaptive artificial intelligence (AI) text-video generation model that leverages text-to-image generation includes receiving a text input, generating a representation frame (i.e. an image) based on text embeddings of the text input (i.e. text-image generation), generating a set of frames (i.e. images) based on the representation frame and a first frame rate, interpolating the set of frames based on a second frame rate, and generating a first video based on the interpolated set of frames (i.e. the generative model includes an image-to-video generative model), as indicated above), for example) or a combination thereof (Par. [0021]: text-video generation model decomposes a full temporal U-Net and attention tensors and approximates them in space and time; Par. [0061]: When fine-tuning on masked frame interpolation, an additional 4 channels may be added to the input of the U-Net; or a combination thereof (e.g. adaptive artificial intelligence (AI) text-video generation model that leverages text-to-image generation includes receiving a text input, generating a representation frame (i.e. an image) based on text embeddings of the text input (i.e. text-image generation), generating a set of frames (i.e. images) based on the representation frame and a first frame rate, interpolating the set of frames based on a second frame rate, and generating a first video based on the interpolated set of frames, in which the text-video generation model decomposes a full temporal U-Net (i.e. the generative model has a U-Net architecture or a combination thereof) and attention tensors and approximates them in space and time, as indicated above), for example). Hinz and Gupta are considered to be analogous art because they pertain to i.e. image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify the method, apparatus, and non-transitory computer readable medium for high-resolution image generation (as disclosed by Hinz) with the generative model includes an image-to-video generative model; the generative model includes an image-to-video generative model: or a combination thereof (as taught by Gupta, Abstract, Par. [0001-5, 21, 61, 69-73]) to generate an output video by increasing a resolution of the video based on spatial information, to generate high resolution and frame rate videos with a video decoder, to implement text-video generation, to improve video resolution by considering temporal information and to improve generalization capabilities that generate more coherent video (Gupta, Abstract, Par. [0001-5, 21, 24, 45, 65]). Regarding claim 13, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), but fails to teach the following as further recited in claim 13. However, Gupta teaches wherein the generative model is applied to perform a text-based video generation, a text-based video content editing operation, image-based video generation, a video enhancement operation, video compression, a data augmentation operation, or a combination thereof (Par. [0001-5]: an adaptive artificial intelligence (AI) text-video generation model that leverages text-to-image generation… systems and methods for text-to-video generation… method includes receiving a text input, generating a representation frame based on text embeddings of the text input, generating a set of frames based on the representation frame and a first frame rate, interpolating the set of frames based on a second frame rate, generating a first video based on the interpolated set of frames… system configured for text-to-video generation. The system includes one or more processors, and a memory storing instructions which, when executed by the one or more processors, cause the system to perform operations. The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames… a non-transient computer-readable storage medium having instructions embodied thereon, the instructions being executable by one or more processors to perform a method(s) for text-to-video generation. The instructions causing the one or more processors to receive a text input, generate a representation frame based on text embeddings of the text input, decoding a set of frames conditioned on the representation frame and a frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames; Par. [0069-73]: representation frame is generated using a text-to-image model… generate a set of frames based on the representation frame and a first frame rate. The first frame rate may be predefined by the system. The set of frames may be combined to generate a low-resolution video… The video generating module 650 is configured to generate a first video with an increased frame rate based on the interpolated set of frames; wherein the generative model is applied to perform a text-based video generation, a text-based video content editing operation, image-based video generation, a video enhancement operation, video compression, a data augmentation operation, or a combination thereof (e.g. adaptive artificial intelligence (AI) text-video generation model that leverages text-to-image generation includes receiving a text input, generating a representation frame (i.e. an image) based on text embeddings of the text input (i.e. perform a text-based video generation), generating a set of frames (i.e. images) based on the representation frame and a first frame rate, interpolating the set of frames based on a second frame rate, and generating a first video based on the interpolated set of frames (i.e. image-based video generation), as indicated above), for example). The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 2. Claims 7, 14-15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Hinz, as applied to claim 1 above, in view of Kreis et al. (US PG Publication No. 2024/0171788 A1), hereafter referred to as, in view of Kreis. Regarding claim 7, claim 4 is incorporated and Hinz discloses the device (Par. [0004-7]), but fails to teach the following as further recited in claim 7. However, Kreis teaches wherein a first power consumption of performance of the first sampling stage is less than a second power consumption of performance of the second sampling stage (Par. [0050]: the encoder 104 can map an input image (e.g., the image data 108) from an image or pixel space to a lower dimensional, compressed latent space (e.g., latent tensors, latent representations, latent encoding, and so on) where the image diffusion model 102 can be trained more efficiently in terms of power consumption). Hinz and Gupta are considered to be analogous art because they pertain to i.e. image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify the method, apparatus, and non-transitory computer readable medium for high-resolution image generation (as disclosed by Hinz) with wherein a first power consumption of performance of the first sampling stage is less than a second power consumption of performance of the second sampling stage (as taught by Kreis, Abstract, Par. [0050]) to improve scalability at high resolutions, to improve the computational and memory efficiency over pixel-space Diffusion Models (DMs) by first training a compression model to transform input images into a spatially lower-dimensional latent space of reduced complexity, from which the original data can be reconstructed at high fidelity, to train an image diffusion model more efficiently in terms of power consumption, and to improve efficiency for high-spatial-resolution and high-temporal-resolution video synthesis (Kreis, Abstract, Par. [0001-17, 29, 44, 50, 96]). Regarding claim 14, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), but fails to teach the following as further recited in claim 14. However, Kreis teaches further comprising: one or more cameras coupled to the one or more processors and configured to generate image data associated with the input image frame (Par. [0070]: decoder 106 can enforce temporally coherent reconstructions across frames by applying the latent representations 504 corresponding to the frames as input and outputting frames that are aligned as frames 506 of a video having temporal coherence; Par. [0109]: I/O ports 712 may enable the computing device 700 to be logically coupled to other devices including the I/O components 714, the presentation component(s) 718, and/or other components, some of which may be built in to (e.g., integrated in) the computing device 700. Illustrative I/O components 714 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The computing device 700 may be include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. The I/O components 714 can include the camera 154 for generating images and video); and an input device configured to receive an input and provide the input to the one or more processors, wherein the input includes a request to generate video data including the one or more output image frames based on the image data from the one or more cameras (Par. [0070]: decoder 106 can enforce temporally coherent reconstructions across frames by applying the latent representations 504 corresponding to the frames as input and outputting frames that are aligned as frames 506 of a video having temporal coherence; Par. [0109]: I/O ports 712 may enable the computing device 700 to be logically coupled to other devices including the I/O components 714, the presentation component(s) 718, and/or other components, some of which may be built in to (e.g., integrated in) the computing device 700. Illustrative I/O components 714 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The computing device 700 may be include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. The I/O components 714 can include the camera 154 for generating images and video). The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 7. Regarding claim 15, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), but fails to teach the following as further recited in claim 15. However, Kreis teaches further comprising: one or more cameras coupled to the one or more processors and configured to generate image data associated with the input image frame, wherein the one or more output image frames a is generated by the one or more processors at least partially based on the image data from the one or more cameras (Par. [0070]: decoder 106 can enforce temporally coherent reconstructions across frames by applying the latent representations 504 corresponding to the frames as input and outputting frames that are aligned as frames 506 of a video having temporal coherence; Par. [0109]: I/O ports 712 may enable the computing device 700 to be logically coupled to other devices including the I/O components 714, the presentation component(s) 718, and/or other components, some of which may be built in to (e.g., integrated in) the computing device 700. Illustrative I/O components 714 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The computing device 700 may be include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. The I/O components 714 can include the camera 154 for generating images and video); and a display device coupled to the one or more processors and configured to output the one or more output image frames as video content (Par. [0109-111]: I/O ports 712 may enable the computing device 700 to be logically coupled to other devices including the I/O components 714, the presentation component(s) 718, and/or other components, some of which may be built in to (e.g., integrated in) the computing device 700. Illustrative I/O components 714 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The computing device 700 may be include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. The I/O components 714 can include the camera 154 for generating images and video…presentation component(s) 718 may include a display (e.g., a monitor, a touch screen, a television screen, a heads-up-display (HUD), other display types, or a combination thereof), speakers, and/or other presentation components. The presentation component(s) 718 may receive data from other components (e.g., the GPU(s) 708, the CPU(s) 706, DPUs, etc.), and output the data (e.g., as an image, video, sound, etc.)). The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 7. Regarding claim 17, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), but fails to teach the following as further recited in claim 17. However, Kreis teaches further comprising: a microphone configured to provide an input signal to the one or more processors to cause the one or more processors to generate the one or more output image frames (Par. [0109]: I/O ports 712 may enable the computing device 700 to be logically coupled to other devices including the I/O components 714, the presentation component(s) 718, and/or other components, some of which may be built in to (e.g., integrated in) the computing device 700. Illustrative I/O components 714 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The computing device 700 may be include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. The I/O components 714 can include the camera 154 for generating images and video); a speaker configured to output audio associated with the one or more output image frames (Par. [0111]: presentation component(s) 718 may include a display (e.g., a monitor, a touch screen, a television screen, a heads-up-display (HUD), other display types, or a combination thereof), speakers, and/or other presentation components. The presentation component(s) 718 may receive data from other components (e.g., the GPU(s) 708, the CPU(s) 706, DPUs, etc.), and output the data (e.g., as an image, video, sound, etc.)); or a combination thereof (Par. [0109-111]: I/O ports 712 may enable the computing device 700 to be logically coupled to other devices including the I/O components 714, the presentation component(s) 718, and/or other components, some of which may be built in to (e.g., integrated in) the computing device 700. Illustrative I/O components 714 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The computing device 700 may be include depth cameras, such as stereoscopic camera systems, infrared camera systems, RGB camera systems, touchscreen technology, and combinations of these, for gesture detection and recognition. The I/O components 714 can include the camera 154 for generating images and video…presentation component(s) 718 may include a display (e.g., a monitor, a touch screen, a television screen, a heads-up-display (HUD), other display types, or a combination thereof), speakers, and/or other presentation components. The presentation component(s) 718 may receive data from other components (e.g., the GPU(s) 708, the CPU(s) 706, DPUs, etc.), and output the data (e.g., as an image, video, sound, etc.)). The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 7. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Hinz, as applied to claim 1 above, in view of Zhang et al. (US PG Publication No. 2024/0362830 A1), hereafter referred to as, in view of Zhang. Regarding claim 12, claim 1 is incorporated and Hinz discloses the device (Par. [0004-7]), wherein: the one or more processors (e.g. Fig. 4, No. 405 and Fig. 20, No. 2005; Par. [0007]: apparatus and system include one or more processors; one or more memory components coupled with the one or more processors; Par. [0037]: system and apparatus include one or more processors; one or more memory components coupled with the one or more processors) are configured to (Par. [0066-68]: image generation apparatus 400 includes processor unit 405… noise component 415, machine learning model 420, training component 460, or a combination thereof are implemented as software stored in a memory subsystem and executed by one or more processors… processor unit 405 is configured to execute computer-readable instructions stored in memory unit 410 to perform various functions… processor unit 405 comprises the one or more processors; Par. [0280-283]: computing device 2000 includes processor(s) 2005, memory subsystem 2010… memory subsystem 2010 includes one or more memory devices. Memory subsystem 2010 is an example of, or includes aspects of, the memory unit as described with reference to FIG. 4… memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor to perform various functions) encode, via autoencoder (Par. [0032]: Diffusion first trains an autoencoder to encode/decode images to/from a 64×64 latent space and then trains a diffusion model for the 64×64 latent space. The decoder is used at test time to generate high-resolution images from the 64×64 latent space, which is faster and cheaper than using a diffusion model. However, Stable Diffusion currently only supports a max resolution of 768×768 pixels, and using the autoencoder increases memory cost during training, requires high-resolution images for training, and tightly couples the autoencoder and diffusion model to each other (meaning if the autoencoder changes, the diffusion model has to be trained again)), the input the input image frame to generate a latent representation of the input image frame (Par. [0032]: Diffusion first trains an autoencoder to encode/decode images to/from a 64×64 latent space and then trains a diffusion model for the 64×64 latent space. The decoder is used at test time to generate high-resolution images from the 64×64 latent space, which is faster and cheaper than using a diffusion model. However, Stable Diffusion currently only supports a max resolution of 768×768 pixels, and using the autoencoder increases memory cost during training, requires high-resolution images for training, and tightly couples the autoencoder and diffusion model to each other (meaning if the autoencoder changes, the diffusion model has to be trained again); Par. [0092-97]: diffusion model 435 comprises a latent diffusion model… the generator learns to generate a candidate by mapping information from a latent space to a data distribution of interest; Par. [0117-124]: FIG. 5 shows an example of a guided latent diffusion architecture 500… The example shown includes guided latent diffusion architecture 500, original image 505, pixel space 510, image encoder 515… use a deterministic process so that a same input results in a same output. Diffusion models may also be characterized by whether noise is added to an image itself, or to image features generated by an encoder, as in latent diffusion… the diffusion model is a latent diffusion model; Par. [0206]: an image generation apparatus (such as the image generation apparatus described with reference to FIGS. 1 and 4) uses diffusion processes 1200 to generate a low-resolution image. In some cases, forward diffusion process 1205 adds first noise to an image (or image features in a latent space) to obtain a noisy image (or noisy image features) (e.g., an input image including random noise). In some cases, the image is an image prompt provided by the user or retrieved by the image generation apparatus. In some cases, the image is an initial noisy image including sampled noise. In some cases, reverse diffusion process 1210 removes second noise from the noisy image (or noisy image features in the latent space) to obtain a low-resolution image), but fails to teach the following as further recited in claim 17. However, Zhang teaches a variational autoencoder (VAE) (Par. [0004]: Generating images from text descriptions has been improved with deep generative models, including… variational autoencoders (VAEs)); and wherein the one or more output image frames include fourteen or more image frames associated with the input image frame (Par. [0078-88]: discriminator 300 may be trained to distinguish whether image inputs are synthesized by generator 200 or sampled from real data… textual description may describe a plurality of scenes, and a plurality of video frames of video content corresponding to the respective plurality of scenes may be generated by the trained neural network; Par. [0142]: one or more images can be one or more still images and/or one or more images utilized in video imagery). Hinz and Zhang are considered to be analogous art because they pertain to i.e. image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify the method, apparatus, and non-transitory computer readable medium for high-resolution image generation (as disclosed by Hinz) with a variational autoencoder (VAE) and wherein the one or more output image frames include fourteen or more image frames associated with the input image frame (as taught by Zhang, Abstract, Par. [0004, 78-88, 142]) generate a final high-resolution image, to improve deep generative models by using variational autoencoders (VAEs), and to produce high quality output results in text-to-image generation, using a conditional GAN formulation (Zhang, Abstract, Par. [0002-7]). Allowable Subject Matter Claims 9 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The prior art of record fails to anticipate or render obvious the following limitations as claimed: In view of claim 1 in its entirety, the further limitations of “… receive a first feature output of the first layer for the first sampling operation, the first feature output associated with the first resolution; and receive a second feature output of the second layer for the second sampling operation, the second feature output associated with the second resolution; one or more spatial-temporal modules coupled in series and configured to receive an output of the first convolution module; and a second convolutional module configured to: receive an output of the one or more spatial-temporal modules; and output a third feature output for the first sampling operation, the third feature output associated with the second resolution” as recited in claim 9. Claims 10-11 are dependent upon claim 9. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to GUILLERMO M RIVERA-MARTINEZ whose telephone number is (571) 272-4979. The examiner can normally be reached on 9 am to 5 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached on 571-270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GUILLERMO M RIVERA-MARTINEZ/ Primary Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Nov 15, 2024
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §102, §103, §112
Sep 17, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731397
Method of generating a peripheral image of an aircraft and associated electronic generation device and computer program product
3y 3m to grant Granted Sep 08, 2026
Patent 12717041
METHOD FOR MONITORING A LOADING AREA
2y 7m to grant Granted Aug 25, 2026
Patent 12694563
SYSTEM AND METHOD FOR USING DYNAMIC OBJECTS TO ESTIMATE CAMERA POSE
2y 6m to grant Granted Jul 28, 2026
Patent 12682441
CHARACTERIZATION SYSTEM AND METHOD IMPLEMENTING IMAGE ENHANCEMENT FOR IMPROVED DEFECT DETECTION
4y 4m to grant Granted Jul 14, 2026
Patent 12651392
STATIONARY MULTI-SOURCE AI-POWERED REAL-TIME TOMOGRAPHY (SMART)
2y 9m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
81%
With Interview (+3.3%)
2y 6m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 514 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month