DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Application Status
This office action is responsive to the Application No.:19/019,882 filed on 01/14/2025 (Priority Date: 05/14/2024).
Claims 1-20 are pending and presented for examination.
This action has been made NON-FINAL.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/14/2025 is being considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Gupta, US 20240155071 in view of Rawat, US 20220222940.
Claim 1:
Gupta discloses a processing system in a device (See Gupta Abstract). However, Gupta failed to explicitly recite “pixels.” However, Gou discloses pixels in paragraphs 0014; 0017-0020; 0041; 0046-0047. It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have further modified Gupta by the teachings of GOU to enable improved video generation, while using pixels values, as expressly taught by GOU, more effectively. In addition, both of the references teach features that are directed to analogous art and they are directed to the same field of endeavor, such as video generation. This close relation between both of the refences highly suggests an expectation of success.
As modified:
The combination of Gupta and GOU discloses the following:
a memory (See Gupta Paragraphs 0027; 0030) configured to store machine learning model parameters (See Gupta Paragraphs 0031-0032);
and one or more processors (See Gupta Paragraphs 0027; 0030), coupled to the memory (See Gupta Paragraphs 0027; 0030), configured to:
access a transformed version of image (See Gupta Figure 7, Paragraphs 00041; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) in a machine learning model (See Gupta Paragraphs 0031-0032) trained to provide controllability of generated videos (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030);
generate a spatial version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) based on the transformed version of image (See Gupta Figure 7, Paragraphs 00042; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using a spatial attention component (See Gupta Paragraphs 0001-0005; 0019-0023);
generate a temporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) based on the transformed version of image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using a temporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023);
generate a first spatiotemporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) based on processing at least one of the transformed version of image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the spatial version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), or the temporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using a first spatiotemporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023);
generate an output version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) based on the first spatiotemporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) and at least one of the spatial version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) or the temporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047);
and generate a set of output image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels from the machine learning model (See Gupta Paragraphs 0031-0032) based on the output version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels, wherein the set of output image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels portray motion depicted in prompt video data (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) for the machine learning model (See Gupta Paragraphs 0031-0032).
Claim 2:
The combination of Gupta and GOU discloses wherein: to generate the first spatiotemporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to process the spatial version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using the first spatiotemporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023); to generate the temporal version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to process the spatial version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using the temporal attention component; and to generate the output version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to aggregate the first spatiotemporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) and the temporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047).
Claim 3:
The combination of Gupta and GOU discloses wherein: to generate the first spatiotemporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to process the transformed version of image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using the first spatiotemporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023); and to generate the temporal version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels, the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to: generate an aggregated version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels by aggregating the first spatiotemporal version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) and the spatial version (See Gupta Paragraphs 0001-0005; 0019-0023) of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047); and process the aggregated version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using the temporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023).
Claim 4:
The combination of Gupta and GOU discloses wherein: the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to generate a second spatiotemporal version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) based on processing the aggregated version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) using a second spatiotemporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023); and to generate the output version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to aggregate the second spatiotemporal version of the image (See Gupta Paragraphs 0001-0005; 0019-0023) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) and the temporal version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047).
Claim 5:
The combination of Gupta and GOU discloses wherein: to generate the spatial version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to generate, for each respective frame of the transformed version of image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), a respective spatial self-attention value based on a respective plurality of spatial elements in the respective frame (See Gupta Paragraphs 0001-0005; 0019-0023); to generate the temporal version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels, the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to generate, for each respective spatial element of a version (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) of the image pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) input for the temporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023), a respective temporal self-attention value based on a respective plurality of frames of the version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) input for the temporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023); and to generate the first spatiotemporal version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to generate at least one spatiotemporal self-attention value (See Gupta Paragraphs 0001-0005; 0019-0023) based on a plurality of spatial elements and a plurality of frames of a version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) input for the first spatiotemporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023).
Claim 6:
The combination of Gupta and GOU discloses wherein, to generate the first spatiotemporal version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047), the one or more processors (See Gupta Paragraphs 0027; 0030) are configured to generate a plurality of self-attention values (See Gupta Paragraphs 0001-0005; 0019-0023), each respective self-attention value of the plurality of self-attention values (See Gupta Paragraphs 0001-0005; 0019-0023) being generated based on a respective tubelet of a plurality of tubelets (See GOU Figure 2E; Paragraph 00413) in the version of the image (See Gupta Figure 7, Paragraphs 0004; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) input for the first spatiotemporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023).
Claim 7:
The combination of Gupta and GOU discloses wherein each respective tubelet of the plurality of tubelets (See GOU Figure 2E; Paragraph 00414) comprises at least two spatial elements (See Gupta Figure 7, Paragraphs 00045; 0017-0024; 0030) of the plurality of spatial elements across at least two frames of the plurality of frames of the version of the image (See Gupta Figure 7, Paragraphs 00046; 0017-0024; 0030) pixels input for the first spatiotemporal attention component (See Gupta Paragraphs 0001-0005; 0019-0023).
Claim 8:
The combination of Gupta and GOU discloses wherein a respective size of each respective tubelet of the plurality of tubelets was learned during training of the first spatiotemporal attention component (See GOU Figure 2E; Paragraph 00417).
Claim 9:
The combination of Gupta and GOU discloses wherein the machine learning model comprises a text-to-video machine learning model (See Gupta Paragraphs 0001; 0003-0005; 0031-0035).
Claim 10:
The combination of Gupta and GOU discloses further comprising a camera (See Gupta Paragraph 0026) configured to capture a set of image (See Gupta Figure 7, Paragraphs 00048; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047) that are transformed to generate the transformed version of image (See Gupta Figure 7, Paragraphs 00049; 0017-0024; 0030) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047).
Claim 11:
The combination of Gupta and GOU discloses further comprising a display (See Gupta Paragraphs 0026; 0077; 0089; 0098) configured to display the set of output image (See Gupta Figure 7, Item 790; Paragraph 0088) pixels (See GOU Figure 1; Paragraphs 0014; 0017-0020; 0041; 0046-0047).
Claims 12-20:
Claims 12-20 are rejected on the same basis as claims 1-9.
Pertinent Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US Patent Application Publication No.: 20220405883 relates to image processing, and more particularly, to a system and method for super-resolution image processing.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHEREE N BROWN whose telephone number is (571)272-4229. The examiner can normally be reached M-F 5:30-2:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, SAID BROOME can be reached at (571) 272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHEREE N BROWN/Primary Examiner, Art Unit 2612 August 20, 2026
1 Gupta paragraph 0004 recites “The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames, increase a resolution of the first video based on spatiotemporal information using a first super-resolution model to generate a second video, and generate an output video by increasing a resolution of the second video based on spatial information using a second super-resolution model.”
2 Gupta paragraph 0004 recites “The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames, increase a resolution of the first video based on spatiotemporal information using a first super-resolution model to generate a second video, and generate an output video by increasing a resolution of the second video based on spatial information using a second super-resolution model.”
3 Paragraph 0041 recites “The tubelets are localized to correct for pixel-level foreground-background biases, which are then turned into short spatio-temporal action tubelets that are passed to a classification network to obtain multi-label predictions.”
4 Paragraph 0041 recites “The tubelets are localized to correct for pixel-level foreground-background biases, which are then turned into short spatio-temporal action tubelets that are passed to a classification network to obtain multi-label predictions.”
5 Gupta paragraph 0004 recites “The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames, increase a resolution of the first video based on spatiotemporal information using a first super-resolution model to generate a second video, and generate an output video by increasing a resolution of the second video based on spatial information using a second super-resolution model.”
6 Gupta paragraph 0004 recites “The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames, increase a resolution of the first video based on spatiotemporal information using a first super-resolution model to generate a second video, and generate an output video by increasing a resolution of the second video based on spatial information using a second super-resolution model.”
7 Paragraph 0041 recites “The tubelets are localized to correct for pixel-level foreground-background biases, which are then turned into short spatio-temporal action tubelets that are passed to a classification network to obtain multi-label predictions.”
8 Gupta paragraph 0004 recites “The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames, increase a resolution of the first video based on spatiotemporal information using a first super-resolution model to generate a second video, and generate an output video by increasing a resolution of the second video based on spatial information using a second super-resolution model.”
9 Gupta paragraph 0004 recites “The operations include to receive a text input, generate a representation frame based on text embeddings of the text input, generate a set of frames based on the representation frame and a first frame rate, interpolate the set of frames based on a second frame rate, generate a first video based on the interpolated set of frames, increase a resolution of the first video based on spatiotemporal information using a first super-resolution model to generate a second video, and generate an output video by increasing a resolution of the second video based on spatial information using a second super-resolution model.”