Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: APPARATUS, SYSTEMS, AND METHODS FOR IMAGE DECODING BASED ON CAPTION DATA.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 10, 13-14, 18 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Bilcu (WO 2021170230 A1).
Regarding claim 1, Bilcu discloses a decoder apparatus (Fig. 2, [pg. 11, lines 1-9] Fig. 2 illustrates a schematic representation of a block diagram of an apparatus 200; Fig. 6, [pg. 13, lines 16-19] CNN 601 may transform the image data 600 into a rich representation by embedding the image data into a fixed-length vector which be provided as an input to the decoder RNN 602 which in turn generates the output description 603) comprising:
receiving circuitry to receive caption data indicative of a language-based description for a first image and encoded data representative of the first image (Fig. 4, [pg. 18, lines 1-12] at 401, a textual description of at least one image is obtained, the textual description of the image may be retrieved from a memory or over a network; [pg. 18, lines 14-18] at 402, an auxiliary image associated with the textual description is obtained, the auxiliary image may be a low-quality version of the image, for example, a resolution of the auxiliary image may be lower than a resolution of the image; [pg. 13, lines 12-13] CNN may encode the image into a compact representation, followed by RNN that may generate a corresponding description of the image); and
decoder circuitry comprising one or more trained machine learning models operable to generate a reconstructed image in dependence on the caption data and the encoded data, the reconstructed image having a higher image quality than an image quality associated with the encoded data representative of the first image (Fig. 4, [pg. 18, lines 17-24] at 403, the image is reconstructed based on the textual description and the auxiliary image, for example, the auxiliary image may be used as a drawing canvas, and the textual description may be used to reconstruct the image on the auxiliary image, this may enable that the compressed image may be decompressed in a higher resolution based on the textual description of the image and the significantly lower quality version of the image).
Regarding claim 10, Bilcu discloses the decoder apparatus according to claim 1 as applied above. Bilcu further discloses wherein the decoder circuitry comprises a trained decoder model operable to receive the encoded data and the caption data and generate the reconstructed image in dependence on the encoded data and the caption data (Fig. 4, [pg. 18, lines 17-24] at 403, the image is reconstructed based on the textual description and the auxiliary image, for example, the auxiliary image may be used as a drawing canvas, and the textual description may be used to reconstruct the image on the auxiliary image, this may enable that the compressed image may be decompressed in a higher resolution based on the textual description of the image and the significantly lower quality version of the image; [pg. 18, lines 25-26] the text-to-image algorithm may be trained by randomly sampling sentences and pairing them with images).
Regarding claim 13, Bilcu discloses the decoder apparatus according to claim 1 as applied above. Bilcu further discloses wherein the decoder apparatus is used in association with an encoder apparatus, the encoder apparatus comprising: encoder circuitry comprising a trained encoder model operable to perform lossy compression of the first image to generate the encoded data representative of the first image ([pg. 9, lines 28-30] image compression may be performed by converting the image into a text description of the image with a device…the image is either saved as a single text file using image-to-text function, or a low-resolution version of the image is saved together with the text; [pg. 15, lines 24-26] The device 300 executing the method may achieve a considerably higher compression rate than conventional compression methods );
a trained image captioning model operable to generate the caption data indicative of the language-based description for the first image ([pg. 13, lines 1-5] a textual description of the image may be determined, the conversion to the textual form may implemented using an image-to-text algorithm); and
communication circuitry to communicate the encoded data and the caption data to the decoder apparatus via a network (Fig. 1, [pg. 22, lines 8-11] the device 400 may retrieve the auxiliary image 702 and the modified textual description 703 from the memory).
Regarding claim 14, Bilcu discloses everything claimed as applied above (See rejection of claim 1).
Regarding claim 18, Bilcu discloses everything claimed as applied above (see rejection of claim 1) including further disclosing a non-transitory computer-readable medium comprising computer executable instructions adapted to cause a computer system to perform a method ([pg. 12, lines 1-14] the memory 202 may be any medium, including non-transitory storage media, on which the program code 203 is stored…the program code may comprise instructions which when executed cause the processor, computer, or the like, to perform at least one of the methods described herein).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 2-9, 15-17, 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bilcu (WO 2021170230 A1) in view of Li (Li, H., Gu, J., Koner, R., Sharifzadeh, S., & Tresp, V. (2022). Do DALL-E and Flamingo Understand Each Other?. arXiv preprint arXiv:2212.12249.).
Regarding claim 2, Bilcu discloses the decoder apparatus according to claim 1 as applied above. Bilcu further discloses wherein the decoder circuitry comprises a trained decoder model operable to receive the encoded data and generate the reconstructed image ([pg. 18, lines 14-18] at 402, an auxiliary image associated with the textual description is obtained, the auxiliary image may be a low-quality version of the image, for example, a resolution of the auxiliary image may be lower than a resolution of the image; Fig. 4, [pg. 18, lines 17-24] at 403, the image is reconstructed based on the textual description and the auxiliary image).
Bilcu fails to disclose a trained image captioning model operable to receive the reconstructed image and generate predicted caption data for the reconstructed image.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses a trained image captioning model operable to receive the reconstructed image and generate predicted caption data for the reconstructed image ([pg. 3, Text-Image-Text] we randomly sample a text from the image-text pair dataset and generate N images for each text using SD, then BLIP generates a description for each input image using beam search (see also Figs. 1, 2, 5)).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu including a trained image captioning model operable to receive the reconstructed image and generate predicted caption data for the reconstructed image, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 3, Bilcu in view of Li discloses the decoder apparatus according to claim 2 as applied above. Bilcu fails to disclose wherein the decoder circuitry is operable to output the reconstructed image in dependence on a comparison of the predicted caption data for the reconstructed image and the received caption data.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses wherein the decoder circuitry is operable to output the reconstructed image in dependence on a comparison of the predicted caption data for the reconstructed image and the received caption data (Fig. 2 right side; [pg. 3, Text-Image-Text] we use different methods to calculate the similarity between the input and generated text; Fig. 4, [pg. 4, Insight II] by comparing the generated captions with the input text, we find better images that faithfully depict the text (see also fig. 5, right side)).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu wherein the decoder circuitry is operable to output the reconstructed image in dependence on a comparison of the predicted caption data for the reconstructed image and the received caption data, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 4, Bilcu in view of Li discloses the decoder apparatus according to claim 2 as applied above. Bilcu fails to disclose wherein the trained decoder model is controlled using a set of learned parameters and the trained decoder model is operable to update one or more of the parameters in dependence on a difference between the predicted caption data for the reconstructed image and the received caption data.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses wherein the trained decoder model is controlled using a set of learned parameters and the trained decoder model is operable to update one or more of the parameters in dependence on a difference between the predicted caption data for the reconstructed image and the received caption data ([pg. 5, 4. Method] in the second pipeline, we optimize the SD model using the loss obtained from comparing the generated text by BLIP with the input text (see also Fig. 6)).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu wherein the trained decoder model is controlled using a set of learned parameters and the trained decoder model is operable to update one or more of the parameters in dependence on a difference between the predicted caption data for the reconstructed image and the received caption data, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 5, Bilcu in view of Li discloses the decoder apparatus according to claim 4 as applied above. Bilcu fails to disclose wherein the trained decoder model is operable to update one or more of the parameters of the trained decoder model to update the trained decoder model to compensate for differences between the predicted caption data for the reconstructed image and the received caption data.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses wherein the trained decoder model is operable to update one or more of the parameters of the trained decoder model to update the trained decoder model to compensate for differences between the predicted caption data for the reconstructed image and the received caption data ([pg. 5, 4. Method] in the second pipeline, we optimize the SD model using the loss obtained from comparing the generated text by BLIP with the input text (see also Fig. 6); [pg. 5, Text Generation Stage] the training objective of the image captioning model is to minimize the cross-entropy loss between the ground truth text and the predicted text, defined as equation 1).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu wherein the trained decoder model is operable to update one or more of the parameters of the trained decoder model to update the trained decoder model to compensate for differences between the predicted caption data for the reconstructed image and the received caption data, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 6, Bilcu in view of Li discloses the decoder apparatus according to claim 4 as applied above. Bilcu fails to disclose wherein the trained decoder model is operable to generate another reconstructed image in dependence on the encoded data using updated parameters.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses wherein the trained decoder model is operable to generate another reconstructed image in dependence on the encoded data using updated parameters (Fig. 7-8, [pg. 8, Qualitative Results] figure 8 shows an example of generated images using our model and SD baseline method, the baseline may neglect certain objects like the car, whereas our method reflects the text prompt (i.e. 'our model' includes their optimized parameters, see [pg. 5, 4. Method], vs. the baseline which does not have updated parameters).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu wherein the trained decoder model is operable to generate another reconstructed image in dependence on the encoded data using updated parameters, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 7, Bilcu in view of Li discloses the decoder apparatus according to claim 4 as applied above. Bilcu fails to disclose wherein the trained decoder model and the trained image captioning model are operable together to continue to generate further reconstructed images and to continue to generate further predicted caption data for each of the further reconstructed images and to continue to update the trained decoder model until a predetermined condition is satisfied.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses wherein the trained decoder model and the trained image captioning model are operable together to continue to generate further reconstructed images and to continue to generate further predicted caption data for each of the further reconstructed images and to continue to update the trained decoder model until a predetermined condition is satisfied ([pg. 7, 4.3. Full Training Objective] the parameters of both models are updated at each iteration, enabling a joint improvement of both models; [pg. 5, Text Generation Stage] the training objective of the image captioning model is to minimize this cross-entropy loss between the ground truth text and the predicted text….this loss is further utilized to update the weights of BLIP, preventing its predictions from deviating from the text used in pretraining; [pg. 6, Image Reconstruction Stage] the training procedure involves a forward diffusion process wherein a clean image is progressively destructed by introducing noise in iterative steps (see also Fig. 6)).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu wherein the trained decoder model and the trained image captioning model are operable together to continue to generate further reconstructed images and to continue to generate further predicted caption data for each of the further reconstructed images and to continue to update the trained decoder model until a predetermined condition is satisfied, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 8, Bilcu in view of Li discloses the decoder apparatus according to claim 7 as applied above. Bilcu fails to disclose wherein the predetermined condition comprises one of more of whether a predetermined number of reconstructed images have been generated and whether a difference between respective further predicted caption data and the received caption data is less than a threshold difference.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses wherein the predetermined condition comprises one of more of whether a predetermined number of reconstructed images have been generated and whether a difference between respective further predicted caption data and the received caption data is less than a threshold difference ([pg. 5, Text Generation Stage] the training objective of the image captioning model is to minimize this cross-entropy loss between the ground truth text and the predicted text, see equation 1).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu wherein the predetermined condition comprises one of more of whether a predetermined number of reconstructed images have been generated and whether a difference between respective further predicted caption data and the received caption data is less than a threshold difference, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 9, Bilcu in view of Li discloses the decoder apparatus according to claim 4 as applied above. Bilcu fails to disclose wherein the trained decoder model is operable to update at least some of the parameters in dependence on a loss function computed in dependence on the difference between the predicted caption data and the received caption data.
Li, in a related system from the same field of endeavor of image reconstruction including text captioning (Abstract), discloses wherein the trained decoder model is operable to update at least some of the parameters in dependence on a loss function computed in dependence on the difference between the predicted caption data and the received caption data ([pg. 5, Text Generation Stage] the training objective of the image captioning model is to minimize this cross-entropy loss between the ground truth text and the predicted text, see equation 1).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Li with Bilcu wherein the trained decoder model is operable to update at least some of the parameters in dependence on a loss function computed in dependence on the difference between the predicted caption data and the received caption data, as disclosed by Li, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving text-to-image and image-to-text models and image reconstruction (See Li: [pg. Improvements]).
Regarding claim 15, Bilcu discloses the computer-implemented method of claim 14 as applied above. Bilcu in view of Li further discloses everything claimed as applied above (see rejection of claim 2).
Regarding claim 16, Bilcu in view of Li discloses the computer-implemented method according to claim 15 as applied above. Bilcu in view of Li further discloses everything claimed as applied above (see rejection of claim 3).
Regarding claim 17, Bilcu in view of Li discloses the computer-implemented method of claim 15 as applied above. Bilcu in view of Li further discloses everything claimed as applied above (see rejection of claim 4).
Regarding claim 19, Bilcu discloses the computer-implemented method of claim 18 as applied above. Bilcu in view of Li further discloses everything claimed as applied above (see rejection of claim 2).
Regarding claim 20, Bilcu in view of Li discloses the computer-implemented method according to claim 19 as applied above. Bilcu in view of Li further discloses everything claimed as applied above (see rejection of claim 3).
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bilcu (WO 2021170230 A1) in view of Noguchi (Noguchi, Chihiro, Shun Fukuda, and Masao Yamanaka. "Scene Text Image Super-resolution based on Text-conditional Diffusion Models." arXiv preprint arXiv:2311.09759 (2023).).
Regarding claim 11, Bilcu discloses the decoder apparatus of claim 10 as applied above. Bilcu further discloses wherein the trained decoder model is controlled using a set of learned parameters and having been initially trained using training data captions indicative of language-based descriptions for the image pairs to learn a function for mapping encoded data representative of a lower resolution image and a corresponding caption to a higher resolution image ([pg. 18, lines 24-28] at 403, the image is reconstructed based on the textual description and the auxiliary image, for example, the auxiliary image may be used as a drawing canvas, and the textual description may be used to reconstruct the image on the auxiliary image…the text-to-image algorithm may be trained by randomly sampling sentences and pairing them with images…the algorithm may be trained with a dataset of paired samples of sources and targets).
Bilcu fails to disclose the training data comprising low- and high- resolution image pairs.
Noguchi, in a related system from the same field of endeavor of training image content recognition models based on image pairs (g. 1, Introduction), discloses the training data comprising low- and high- resolution image pairs ([pg. 4, 3.2 LR-HR Paired Text Image Synthesis] synthesizing LR-HR paired text images; [pg. 6, Datasets] when training super-resolver and degrader, we used LR-HR paired text images of TextZoom; [pg. 1, Introduction first paragraph] STISR aims to restore the high-resolution (HR) text images from low resolution (LR) text images).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Noguchi with Bilcu, wherein the training data comprises low- and high- resolution image pairs, as disclosed by Noguchi, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of generating high-quality reconstructed high-resolution images based on the training (See Noguchi: [pg. 8, 7. Conclusions]).
Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bilcu (WO 2021170230 A1) in view of Zeng (US 20250166667 A1).
Note: Zeng (US 20250166667 A1) claims priority to Chinese patent application CN 117544833 B (application number CN202311543966) and was thus effectively filed 11/17/2023. Support for the selections of Zeng relied upon in this office action can be found in the provided copy and translation of CN 117544833 B in at least paragraph [n0119].
Regarding claim 12, Bilcu discloses the decoder apparatus according to claim 1 as applied above. Bilcu fails to disclose wherein the receiving circuitry is operable to receive the encoded data representative of a sequence of video images and caption data indicative of a language-based description for at least some of the sequence of video images, and the decoder circuitry is operable to generate a reconstructed video image sequence in dependence on the caption data and the encoded data.
Zeng, in a related system from the same field of endeavor of image and video generation based on a received text description (Abstract), discloses wherein the receiving circuitry is operable to receive the encoded data representative of a sequence of video images and caption data indicative of a language-based description for at least some of the sequence of video images (Fig. 16, [0115] at block 1610, an image for describing at least any of a head image and a tail image of a target video is received, at block 1620, a text for describing a content of the target video is received), and the decoder circuitry is operable to generate a reconstructed video image sequence in dependence on the caption data and the encoded data (Fig. 16, [0115] at block 1630, the target video is generated based on the images and the text according to a generation model).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to combine Zeng with Bilcu wherein the receiving circuitry is operable to receive the encoded data representative of a sequence of video images and caption data indicative of a language-based description for at least some of the sequence of video images, and the decoder circuitry is operable to generate a reconstructed video image sequence in dependence on the caption data and the encoded data, as disclosed by Zeng, as part of a decoder apparatus for generating reconstructed image data, as disclosed by Bilcu, for the purpose of improving performance of a machine learning model for generating relevant images and text descriptions (See Zeng: [0060], [0072], [0095]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Cho (US 20230153522 A1) discloses image captioning including obtaining an encoded training image and computing a reward function based on an encoded training caption and the encoded training image and updating parameters of a network in response to the function in order to generate improved image captions.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAROLINE DEPALMA whose telephone number is (571)270-0769. The examiner can normally be reached Mon-Thurs 9:00am-4pm Eastern Time.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CAROLINE E. DEPALMA/Examiner, Art Unit 2675
/SJ Park/Primary Examiner, Art Unit 2675