DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 18 is objected to because of the following informalities: “one or more ne or more”. Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 8, 12-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lindsay Sparks et al. (hereinafter Sparks) (US 20220270162 A1, 2022-08-25) in view of Kevin Lin et al. (hereinafter Lin) (“DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design,” 2023-10-23), further in view of Thomas M. Isaacson et al. (hereinafter Isaacson) (US 20180315040 A1, 2018-11-01).
Regarding claim 1, Sparks teaches;
A data processing system comprising: a processor; and a memory storing executable instructions that, when executed, cause the processor alone or in combination with other processors to perform operations of ([0360] a user device for generating a digital gift is provided. It includes: … a processor system; … and a memory system. The memory system includes executable instructions to at least:): receiving, from a first user interface of an application, a first natural language utterance describing digital content to be generated ([0225] User Connect app determines if … voice mode is activated … If the voice mode is activated, then the operations at block 1102 to 1110 are executed [0219] the user device, via the voice interface, asks one or more questions to the gift giver, and the user device at block 1104 receives voice data … process the received voice data to … generate a digital gift); …
receiving a second natural language utterance describing digital wrapping ([0222] generating a digital gift wrapper based on the information provided in blocks 1103, 1104, 1105… [0219] 1103, the user device, via the voice interface, asks one or more questions to the gift giver … 1104 receives voice data … 1105, the user device … process the received voice data to … generate) for the digital content ([0054] The term “digital gift wrapper” herein refers to digital content that precedes the digital gift), the digital wrapping to be presented to a recipient of the digital content ([0055] the digital gift wrapper is sent first to a user device of the gift recipient [0165] the user device 102 displays … a digital gift wrapper), …
and sending the digital content and the digital wrapping to a client device of a recipient with a link ([0319] gift giver's user device … generates and sends a digital gift 105 to a gift recipient … sends a text message that includes a data link (e.g. URL) that links to the digital gift [0223] In the process of receiving the digital gift, the user device of the gift recipient receives, … a digital gift wrapper from a gift giver),
Sparks fails to teach but Lin teaches;
constructing a first prompt ([pg. 7] expanded prompt), using a prompt construction unit ([pg. 7] adopts ChatGPT [65] for prompt expansion), the first prompt being based on the first natural language utterance ([pg. 7] prompt expansion, i.e., converting an input user query into a more detailed text description) instructing a generative model to generate the digital content ([pg. 49] User Input: create a very cute stuffy for my daughter birthday present); providing the first prompt to the generative model to cause the generative model to generate the digital content ([pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations); obtaining the digital content as an output from the generative model ([Abstract] T2I models … generating … images); …
constructing a second prompt ([pg. 7] expanded prompt) based on the second natural language utterance using the prompt construction unit ([pg. 7] adopts ChatGPT [65] for prompt expansion, i.e., converting an input user query into a more detailed text description), the second prompt comprising instructions to the generative model to generate content comprising one or more images representing the digital wrapping ([pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations … [pg. 44] Expanded Prompt: … Pusheen sits at a crafting table, wrapping presents with colorful papers and ribbons); providing the second prompt to the generative model ([pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations) to cause the generative model to generate the … wrapping (See pg. 44 below);
[pg. 44]
PNG
media_image1.png
796
1137
media_image1.png
Greyscale
obtaining the digital wrapping as an output from the generative model (See Lin pg. 44 above);
OBVIOUSNESS TO COMBINE LIN:
Sparks is analogous art to the present disclosure as it pertains to digital gift submission, and Lin is analogous art to the present disclosure as it pertains to prompt expansion and generating digital content using generative models. Sparks teaches receiving natural language utterances that are utilized to create a digital gift and digital wrapping, and sending the digital gift and wrapping to a recipient with a link, while Lin teaches constructing expanded prompts from natural language utterances to provide a more detailed description of image content to be generated, inputting the expanded prompts to generative models to generate corresponding image content, and even teaches that generative models were known to be capable of images depicting gifts and wrapping. Additionally, Lin states that “[Lin, Abstract] Recent T2I models like DALL-E 3 … and others, have demonstrated remarkable capabilities in generating photorealistic images that align closely with textual inputs”, and further states that “[Lin, pg. 56] expanded text prompts are helpful in improving the design fidelity across all experimented T2I models.” Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Sparks’ digital gift generation process according to Lin by converting Sparks’ received natural language utterances to expanded prompts, and using said expanded prompts as input to a text to image model, as taught by Lin, to generate the digital content and the digital wrapping. Such a combination would have resulted in using known prompt expansion and image generation methods taught by Lin to generate the digital content and wrapping based on the received natural language utterances of Sparks, predictably providing generated image content for the digital content and wrapping that more closely reflects the user’s input and has improved design fidelity.
Sparks and Lin fail to explicitly teach but Isaacson teaches;
the digital wrapping providing a virtual experience of receiving and unwrapping a gift comprising the digital content ([0241] the virtual unwrapping event can mimic the experience of unwrapping a gift to view or access the gift's contents) …
the digital wrapping comprising a control, which activated causes the client device to present an animation of the digital wrapping being removed and the digital content to be presented ([0276] in response to receiving an input associated with the separable flap on the image overlay, the computing system 100 generates the animated unwrapping effect to remove the image overlay and reveal the image underlay. The animated unwrapping effect creates an illusion perceived by users as the unwrapping of a wrapped egift) on a user interface of the client device ([0247] The virtual egift 2902 can be rendered on a user interface).
OBVIOUSNESS TO COMBINE ISAACSON:
Isaacson is analogous art to the present disclosure as it pertains to simulating unwrapping a gift to present a digital gift to a user. Sparks already teaches a method of providing a digital gift with digital wrapping to a recipient, where the digital wrapping is presented before the gift, Lin teaches generating images from expanded prompts, such as images of gifts and wrapping, and Isaacson teaches an image overlay (which depicts wrapping) that obscures an image underlay (which depicts a gift) having a control mechanism that allows the user to trigger an image animation (unwrapping event) of the image overlay being removed from the image underlay. Isaacson further teaches that “[0241] the virtual unwrapping event can mimic the experience of unwrapping a gift to view or access the gift's contents, peeling off a label or sticker, tearing off a wrapping paper or covering, etc.” Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the digital gift giving of Sparks as modified by Lin by presenting the wrapping imagery as an image overlay that obscures the underlying digital gift content, and providing a control for the user to trigger an animation of the wrapping being removed from the gift, as taught by Isaacson. Such a combination would have resulted in using Isaacson’s known image overlay and control method to modify the presentation of the digital gift and wrapping of the system of Sparks as modified by Lin, predictably providing the digital gift recipient with an interactive experience that more closely mimics the experience of unwrapping a physical gift.
Regarding claim 8, Sparks teaches;
wherein sending the digital content and the digital wrapping to the client device of the recipient further comprises: generating a link ([0330] The connect server generates a data link to the digital gift) referencing a storage location of the digital content and the digital wrapping ([0329] connect server 103 stores the digital gift … [0304] the digital gift 105 includes a digital gift wrapper … reference to a digital gift may herein include a digital gift wrapper although not explicitly stated), the link when activated causes the client device to access the storage location of the digital content and the digital wrapping and obtain a copy of the digital content and the digital wrapping from the storage location ([0329] connect server 103 stores the digital gift … [0208] The recipient opens the data link to view the digital gift on the web browser [0304] the digital gift 105 includes a digital gift wrapper … reference to a digital gift may herein include a digital gift wrapper although not explicitly stated); and sending the link to the client device of the recipient ([0208] sends a data link to a gift recipient via text, email, a social media app, or messaging app, or a combination thereof. The recipient opens the data link to view the digital gift on the web browser).
Regarding claim 12, Sparks teaches
one or more images representing the digital wrapping and the one or more digital accessories ([0165] the digital gift wrapper 504 is an image, animation, or a video … of a physically wrapped gift box with a ribbon … [0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content … A ribbon 701 overlays the windowed surfaces)
Sparks fails to teach but Lin teaches;
constructing the second prompt further comprises constructing a plurality of second prompts based on the second natural language utterance ([pg. 7] ChatGPT [65] for prompt expansion, i.e., converting an input user query into a more detailed text description … default prompt expansion behavior in ChatGPT … such as generating four prompts sequentially and producing four images), each prompt of the plurality of second prompts instructing the generative model to generate one or more images ([pg. 7] prompt expansion … generating four prompts sequentially and producing four images … [pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations) … and wherein providing the plurality of second prompts to the generative model to cause the generative model to generate the one or more images ([pg. 7] prompt expansion i.e., converting an input user query into a more detailed text description … generating four prompts sequentially and producing four images … [pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations)
OBVIOUSNESS: Using the same reasoning from claim 1.
Regarding claim 13,
Claim 13 is a method claim that is substantially similar to system claim 1, and is rejected using the same reasoning.
Claim(s) 2-7, 14-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sparks (US 20220270162 A1, 2022-08-25) in view of Lin (“DEsignBench: Exploring and Benchmarking DALL-E 3for Imagining Visual Design,” 2023-10-23) further in view of Isaacson (US 20180315040 A1, 2018-11-01) as applied to claim 1 above, further in view of Javed Qadrud-Din et al. (hereinafter Quadrud-Din) (US 11860914 B1)
Regarding claim 2, Sparks teaches
wherein constructing the second prompt further comprises: … the digital wrapping for the digital content ([0054] The term “digital gift wrapper” herein refers to digital content that precedes the digital gift) based on the second natural language utterance ([0222] generating a digital gift wrapper based on the information provided in blocks 1103, 1104, 1105… [0219] 1103, the user device, via the voice interface, asks one or more questions to the gift giver … 1104 receives voice data … 1105, the user device … process the received voice data to … generate)
Sparks, Lin, and Isaacson fail to teach but Qadrud-Din teaches;
accessing a first prompt template from a prompt datastore ([col. 11, ln. 20] Prompt templates may be selected from the prompt library), the first prompt template comprising first instructions to generate [content] ([col. 11, ln. 30-32] a prompt template may include a set of instructions for causing a … model to generate); and customizing the first prompt template based on [input text] ([col. 11, ln. 26-28] a prompt may be determined by supplementing and/or modifying a prompt template based on the input text).
OBVIOUSNESS TO COMBINE QADRUD-DIN: Qadrud-Din is analogous art to the present disclosure as it pertains to storing and customizing prompt templates. Sparks already teaches using natural language utterances to generate the digital wrapping, Lin teaches converting natural language utterances into expanded prompts, and providing the expanded prompts to a generative model to generate corresponding image content, while Qadrud-Din teaches storing, in a prompt datastore/library, prompt templates having instructions for causing a generative model to generate content, and customizing said prompt templates based on input text. One of ordinary skill in the art, before the effective filing date, would have been motivated to apply Qadrud-Din’s prompt template technique to the prompt generation process of Sparks as modified by Lin to provide a structured and reusable mechanism for constructing generative model prompts by selecting stored prompt templates and adapting them to received natural language input. The combination would have amounted to applying Qadrud-Din’s known prompt template construction technique to Lin’s known natural language to generative model prompting process, with the predictable result of constructing prompts for generating Sparks’ digital wrapping by selecting a stored prompt template and customizing the template based on the user’s natural language description of the wrapping.
Regarding claim 3, Sparks teaches;
generate a digital wrapping paper for a virtual gift box ([0222] generating a digital gift wrapper … [0165] the digital gift wrapper 504 is an image, animation, or a video … of a physically wrapped gift box with a ribbon … [0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content)
Sparks fails to teach but Lin teaches;
instruct the generative model to generate ([pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations) [images of wrapping paper and gift boxes] (Note: See page 44 below).
[pg. 44]
PNG
media_image1.png
796
1137
media_image1.png
Greyscale
OBVIOUSNESS: Using the same reasoning from claim 1
Sparks, Lin, and Isaacson fail to teach but Qadrud-Din teaches
wherein the first instructions instruct the generative model to generate ([col. 11, ln. 30-32] a prompt template may include a set of instructions for causing a large language model to generate)
OBVIOUSNESS: Using the same reasoning from claim 2
Regarding claim 4, Sparks teaches;
generate a virtual envelope ([0222] generating a digital gift wrapper … [0165] the digital gift wrapper 504 is an image, animation or video of a physical card in a physical envelope)
Sparks fails to teach but Lin teaches;
generative model to generate ([pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations) [images of an envelope] (Note: See page 84 below, showing outputs of various text to image outputs depicting envelopes. Qualitative comparisons among Ideogram (first image), Firefly 2 (second image), and DALL-E 3 (third image)).
[pg. 84]
PNG
media_image2.png
245
1277
media_image2.png
Greyscale
OBVIOUSNESS: Using the same reasoning from claim 1
Sparks, Lin, and Isaacson fail to teach but Qadrud-Din teaches
wherein the first instructions instruct the generative model to generate ([col. 11, ln. 30-32] a prompt template may include a set of instructions for causing a large language model to generate)
OBVIOUSNESS: Using the same reasoning from claim 2
Regarding claim 5, Sparks teaches;
generate one or more digital accessories ([0181] ribbon) for the digital wrapping based on the second natural language utterance ([0222] generating a digital gift wrapper based on the information provided in blocks 1103, 1104, 1105 … [0219] 1103, the user device, via the voice interface, asks one or more questions to the gift giver … 1104 receives voice data … 1105, the user device … process the received voice data to … generate … [0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content … A ribbon 701 overlays the windowed surfaces)
Sparks fails to teach but Lin teaches;
providing the third prompt to the generative model to cause the generative model to generate ([pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations) the one or more … accessories (Note: See page 44 below)
[pg. 44]
PNG
media_image1.png
796
1137
media_image1.png
Greyscale
Sparks, Lin, and Isaacson fail to teach but Qadrud-Din teaches;
accessing a second prompt template from a prompt datastore ([col. 11, ln. 20] Prompt templates may be selected from the prompt library … [col. 11, ln. 15] One or more prompt templates are determined), the second prompt template comprising first instructions to generate [content] ([col. 11, ln. 30-32] a prompt template may include a set of instructions for causing a … model to generate); customizing the second prompt template based on the [input text] to generate a third prompt (Note: The prompt determined from the modified prompt template constitutes the generated third prompt … [col. 11, ln. 26-28] a prompt may be determined by supplementing and/or modifying a prompt template based on the input text);
OBVIOUSNESS: Sparks already teaches generating digital wrapping based on received voice information and generating a ribbon overlaid on the digital wrapping, Lin teaches converting natural language descriptions, including descriptions of wrapping and ribbons, into expanded prompts and supplying those prompts to a generative model to generate corresponding image content, while Qadrud-Din teaches storing, in a prompt datastore/library, prompt templates having instructions for causing a generative model to generate content, and customizing said prompt templates based on input text. One of ordinary skill in the art, before the effective filing date, would have been motivated to apply Qadrud-Din’s prompt template technique to the prompt generation process of Sparks as modified by Lin to provide a structured and reusable mechanism for constructing generative model prompts by selecting stored prompt templates and adapting them to received natural language input. The combination would have amounted to applying Qadrud-Din’s known prompt template construction technique to Lin’s known natural language to generative model prompting process, with the predictable result of constructing prompts for generating Sparks’ digital accessories by selecting a stored prompt template and customizing the template based on the user’s natural language description of the wrapping.
Regarding claim 6, Sparks teaches;
storing the one or more digital accessories ([0181] ribbon) in a content datastore ([0203] The digital gift wrapper database 902 includes … customized digital gift wrappers … [0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content … A ribbon 701 overlays the windowed surfaces); and associating the one or more digital accessories ([0181] ribbon) with the digital content and the digital wrapping ([0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content … A ribbon 701 overlays the windowed surfaces … [0059] A digital gift wrapper 105, which is associated with the digital gift)
Regarding claim 7, Sparks teaches;
wherein sending the digital content and the digital wrapping to the client device of the recipient ([0319] sends a digital gift 105 to a gift recipient … via text message and user device 101 specifies a gift recipient's phone number … [0304] the digital gift 105 includes a digital gift wrapper … reference to a digital gift may herein include a digital gift wrapper although not explicitly stated) further comprises sending the one or more digital accessories ([0181] ribbon) to the client device of the recipient ([0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content … A ribbon 701 overlays the windowed surfaces … [0223] In the process of receiving the digital gift, the user device of the gift recipient receives, … a digital gift wrapper from a gift giver).
Regarding claims 14-16,
Claims 14-16 are method claims that are substantially similar to system claims 2-4, respectively, and are rejected using the same reasoning.
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sparks (US 20220270162 A1, 2022-08-25) in view of Lin (“DEsignBench: Exploring and Benchmarking DALL-E 3for Imagining Visual Design,” 2023-10-23) further in view of Isaacson (US 20180315040 A1, 2018-11-01) further in view of Qadrud-Din (US 11860914 B1) as applied to claim 5 above, further in view of Nicholas Isaac Kolkin et al. (hereinafter Kolkin) (US 20240135610 A1, 2024-04-25).
Regarding claim 9, Sparks teaches
digital wrapping (using the same reasoning from claim 1)
Sparks fails to explicitly teach but Lin teaches;
receiving a third natural language utterance via the first user interface of the application, the third natural language utterance providing feedback requesting that the generative model modify the [image] ([pg. 7] one may refer to a specific generated image and give an editing instruction, such as “Change the cloth in the second image into the blue color,”); constructing a third natural language prompt ([pg. 7] expanded prompt) based on the third natural language prompt ([pg. 7] prompt expansion, i.e., converting an input user query into a more detailed text description) using the prompt construction unit ([pg. 7] ChatGPT [65] for prompt expansion);
OBVIOUSNESS: Using the same reasoning from claim 1
Sparks, Lin, Isaacson, and Quadrud-Din fail to explicitly teach but Kolkin teaches
and providing the third prompt and the [image] to the generative model to cause the generative model to modify the [image] ([0004] an image generation system receives an original image, … and a target prompt … describing a desired modification to the element … the image generation system generates an output image that depicts the desired modification).
OBVIOUSNESS TO COMBINE KOLKIN:
Kolkin is analogous art to the present disclosure as it pertains to using generative models and natural language inputs to create and modify images. Sparks teaches generating digital wrapping (which can comprise images) based on received natural language, Lin teaches generating expanded prompts from natural language inputs, providing the expanded prompts to a generative model to generate images (even providing example images including wrapping paper), and modifying previously generated images based on additional input, while Kolkin explicitly teaches providing additional natural language input with a previous image to a generative model to modify the previous image. Kolkin states that their image generation system “[0005] generates a more accurate modified image than conventional image generation systems.” Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the digital wrapping generation of Sparks in view of Lin to use Kolkin’s method of providing the previously generated image and corresponding modification prompt to the generative model to modify the image representing the digital wrapping, predictably providing a more accurate image modification system than conventional image generation systems.
Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sparks (US 20220270162 A1, 2022-08-25) in view of Lin (“DEsignBench: Exploring and Benchmarking DALL-E 3for Imagining Visual Design,” 2023-10-23) further in view of Isaacson (US 20180315040 A1, 2018-11-01), further in view of Yuwei Guo et al. (hereinafter Guo) (“ANIMATEDIFF: ANIMATE YOUR PERSONALIZED TEXT-TO-IMAGE DIFFUSION MODELS WITHOUT SPECIFIC TUNING,” 2024-02-08).
Regarding claim 10
Regarding claim 1, Sparks and Lin fail to teach but Isaacson teaches;
generating the animation ([0276] generates the animated unwrapping effect to remove the image overlay and reveal the image underlay)
Sparks, Lin, and Isaacson fail to teach but Guo teaches;
generating [an] animation using an animation model ([Abstract] we present AnimateDiff, a practical framework for animating personalized T2I models … motion module can be inserted into a personalized T2I model to form a personalized animation generator)
OBVIOUSNESS TO COMBINE GUO: Guo is analogous art to the present disclosure as it pertains to text to image generative models, and the generation of animations from imagery produced by such models. Sparks teaches generating digital gifts and digital wrapping that can be represented by images, Lin teaches using text to image models to generate image, including images applicable to gifts and wrapping, Isaacson teaches generating an animation depicting digital wrapping being removed from a digital gift, and Guo teaches augmenting text to image generative models with a motion module to form an animation generator capable of generating animation clips. Guo further teaches that its framework permits, “[Guo, Abstract] animating personalized T2I models without requiring model-specific tuning.” Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement the gift unwrapping animation taught by Isaacson in the combined system of Sparks, Lin and Isaacson using Guo’s model-based animation generation technique. Such a modification would have constituted use of Guo’s known animation generation technique with text to image generated visual content of the combined system, predictably permitting the known gift unwrapping animation of Isaacson to be generated using an animation model while obtaining Guo’s stated benefit of enabling animation of text to input generated content without requiring model specific tuning.
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sparks (US 20220270162 A1, 2022-08-25) in view of Lin (“DEsignBench: Exploring and Benchmarking DALL-E 3for Imagining Visual Design,” 2023-10-23) further in view of Isaacson (US 20180315040 A1, 2018-11-01), further in view of Guo (“ANIMATEDIFF: ANIMATE YOUR PERSONALIZED TEXT-TO-IMAGE DIFFUSION MODELS WITHOUT SPECIFIC TUNING,” 2024-02-08) as applied to claim 10 above, further in view of Yuying Ge et al. (hereinafter Ge) (“Making LLaMA SEE and Draw with SEED Tokenizer”).
Regarding claim 11, Sparks, Lin, Isaacson, and Guo fail to explicitly teach but Ge teaches;
wherein the generative model is a large language model (LLM) ([pg. 3] We further present SEED-LLaMA by equipping the pre-trained LLM [2] with SEED tokenizer … [pg. 8] SEED-LLaMA generates images that are highly correlated with text prompts)
OBVIOUSNESS TO COMBINE GE:
It would have been obvious to one of ordinary skill in the art, before the effective filing date, to implement Lin’s text to image generative model used in the combined system of Sparks and Lin as Ge’s LLM based generative model. Ge expressly teaches that “[Ge, pg. 8] SEED-LLaMA generates images that are highly correlated with text prompts.” Because Lin relies on textual prompts to specify the visual content to be generated, generating images that are highly correlated with those text prompts would improve the correspondence between the requested visual design and the resulting generated image. Such a modification would have predictably permitted the customized digital wrapping of Sparks to be generated from textual prompts using a large language model while obtaining Ge’s benefit of generated images that are highly correlated with input prompts.
Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sparks et al. (US 20220270162 A1, 2022-08-25), in view of Kolkin et al. (hereinafter Kolkin) (US 20240135610 A1, 2024-04-25), further in view of Isaacson (hereinafter Isaacson) (US 20180315040 A1, 2018-11-01).
Regarding claim 17, Sparks teaches;
A data processing system comprising: a processor; and a memory storing executable instructions that, when executed, cause the processor alone or in combination with other processors to perform operations of ([0360] a user device for generating a digital gift is provided. It includes: … a processor system; … and a memory system. The memory system includes executable instructions to at least:): receiving, in a user interface of an application on a client device, a natural language utterance describing digital wrapping ([0222] generating a digital gift wrapper based on the information provided in blocks 1103, 1104, 1105… [0219] 1103, the user device, via the voice interface, asks one or more questions to the gift giver … 1104 receives voice data … 1105, the user device … process the received voice data to … generate) for digital content ([0054] The term “digital gift wrapper” herein refers to digital content that precedes the digital gift), the digital wrapping to be presented to a recipient of the digital content ([0055] the digital gift wrapper is sent first to a user device of the gift recipient);… and sending an indication to the wrapping generation pipeline to send a link to the digital content and the digital wrapping to a recipient ([0319] gift giver's user device … generates and … sends a text message that includes a data link (e.g. URL) that links to the digital gift [0304] the digital gift 105 includes a digital gift wrapper … reference to a digital gift may herein include a digital gift wrapper although not explicitly stated).
Sparks fails to explicitly teach but Kolkin teaches;
sending the digital content and the natural language utterance to a [image] generation pipeline to analyze the digital content and the natural language utterance and generate the [image] based on the digital content and the natural language utterance ([0004] an image generation system receives an original image, … and a target prompt … describing a desired modification to the element … the image generation system generates an output image that depicts the desired modification); obtaining the digital [image] as an output from the wrapping generation pipeline ([0004] the image generation system generates an output image that depicts the desired modification);
OBVIOUSNESS TO COMBINE KOLKIN:
Sparks teaches generating digital content (digital gifts) and corresponding digital wrapping that can be represented using images based on respective natural language utterances, while Kolkin teaches a generative model pipeline which receives an input image and input prompt to generate a new modified image. Kolkin states that their image generation system “[0005] generates a more accurate modified image than conventional image generation systems.” Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the digital wrapping generation of Sparks by providing the image representing the digital content and the natural language utterance associated with the desired digital wrapping as inputs to the generative model of Kolkin, such that the generated digital wrapping is based on both the digital content and the natural language utterance. Such a modification would have predictably applied Kolkin’s known image and text conditioned generation technique to Sparks’ digital wrapping generation while obtaining Kolkin’s stated benefit of generating more accurate modified images for the wrapping generation than conventional image-generation systems.
Sparks and Kolkin fail to teach but Isaacson teaches;
presenting, on the user interface of the application on the client device, the digital wrapping on a second user interface of the client device ( [0197] preview option is a variation in which the system sends a preview to the giver before sending the actual virtual egift … in the preview communication [the giver can] change any settings associated with the scheduled virtual egift … the system can present a graphic or multimedia presentation to the giver illustrating the policy for that virtual egift. Changes to the policy would be shown in the graphic) and controls, which when activated, cause the client device to present an animation of the digital wrapping being removed and the digital content to be presented ([0024] virtual egift including an overlay and underlay image, … configured to perform an animated unwrapping effect that resembles unwrapping a gift, when the recipient provides input, such as selecting a separable flap in the image overlay);
OBVIOUSNESS:
Sparks already teaches a method of providing a virtual gift with digital wrapping to a recipient, while Isaacson teaches presenting an egift having an having a control mechanism that allows the user to trigger an animation of the image overlay (depicting wrapping) being removed from the image underlay (depicting the egift content), and providing a preview of the egift presentation to the gift giver via a preview interface. Isaacson further teaches that “[0241] the virtual unwrapping event can mimic the experience of unwrapping a gift to view or access the gift's contents, peeling off a label or sticker, tearing off a wrapping paper or covering, etc.” Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify the digital gift presentation of Sparks as modified by Kolkin to provide the gift giver with a preview of the wrapped digital gift prior to transmission, wherein the preview includes an image overlay representing the wrapping and an interactive control for triggering an animated removal of the wrapping to reveal the underlying digital gift content, as taught by Isaacson. Such a modification would have predictably allowed the gift giver to preview and verify the appearance and interactive unwrapping presentation of the digital gift before sending it, while preserving Isaacson’s intended immersive gift-unwrapping experience.
Claim(s) 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sparks et al. (US 20220270162 A1, 2022-08-25), in view of Kolkin et al. (hereinafter Kolkin) (US 20240135610 A1, 2024-04-25), further in view of Isaacson (hereinafter Isaacson) (US 20180315040 A1, 2018-11-01) as applied to claim 17 above, further in view of Lin (“DEsignBench: Exploring and Benchmarking DALL-E 3for Imagining Visual Design,” 2023-10-23).
Regarding claim 18, Sparks teaches;
one or more ne or more digital accessories ([0181] ribbon) to enhance the digital wrapping ([0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content … A ribbon 701 overlays the windowed surfaces … [0059] A digital gift wrapper 105, which is associated with the digital gift)
Sparks, Kolkin, and Isaacson fail to explicitly teach but Lin teaches;
the natural language utterance includes instructions to generate one or more ne or more … accessories to enhance the … wrapping (See pg. 44 below)
[pg. 44]
PNG
media_image1.png
796
1137
media_image1.png
Greyscale
OBVIOUSNESS:
Sparks teaches generating customizable digital wrapping for a digital gift, including ribbon accessory, Kolkin teaches using a natural language target prompt to specify desired visual content to be generated, and Lin further demonstrates that text to image models were known to generate depictions of wrapping accessories, including ribbons, in response to textual prompts. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to configure the natural language utterance of the combined system of Sparks, Kolkin, and Isaacson to include an instruction requesting a ribbon accompanying the digital wrapping, as taught by Lin. Such a modification would have predictably allowed the user to specify and generate the ribbon accessory already taught by Sparks using Kolkin’s existing text conditioned generation process, while retaining Kolkin’s benefit of producing output that more closely adheres to the details of the target prompt.
Regarding claim 19, Sparks teaches;
digital wrapping … for a virtual gift box ([0179] system … incorporates the images, animations, videos or sounds (or a combination thereof) into a digital gift wrapper template, to produce a custom digital gift wrapper … [0181] the images, animations or videos could form the 3D surfaces of a gift box. … that respectively include different windowed surfaces 703a, 703b to hold and display visual content … A ribbon 701 overlays the windowed surfaces)
Sparks, Kolkin, and Isaacson fail to explicitly teach but Lin teaches;
wherein the natural language utterance further requests that the wrapping generation pipeline generate a … wrapping paper for a … gift (See pg. 44 below)
[pg. 44]
PNG
media_image1.png
796
1137
media_image1.png
Greyscale
OBVIOUSNESS:
Sparks teaches generating customizable digital wrapping for a digital gift, including digital wrapping associated with a virtual gift box, Kolkin teaches using a natural language target prompt to specify desired visual content to be generated, and Lin further demonstrates that text to image models were known to generate depictions of presents wrapped with wrapping paper, in response to textual prompts. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to configure the natural language utterance of the combined system of Sparks, Kolkin, and Isaacson to request digital wrapping paper for Spark’s virtual gift box, as taught by Lin. Such a modification would have predictably allowed the user to specify and generate wrapping paper for the virtual gift box using Kolkin’s existing text conditioned generation process, while retaining Kolkin’s benefit of production output that more closely adheres to the details of the target prompt.
Regarding claim 20, Sparks teaches;
generate a virtual envelope ([0222] generating a digital gift wrapper … [0165] the digital gift wrapper 504 is an image, animation or video of a physical card in a physical envelope)
Sparks, Kolkin, and Isaacson fail to explicitly teach but Lin teaches;
wherein the natural language utterance further requests that the wrapping generation pipeline generate ([pg. 57] Each T2I model takes the expanded text prompt as input, and generates four image variations) a … envelope (Note: See page 84 below, showing outputs of various text to image outputs depicting envelopes. Qualitative comparisons among Ideogram (first image), Firefly 2 (second image), and DALL-E 3 (third image)).
[pg. 84]
PNG
media_image2.png
245
1277
media_image2.png
Greyscale
OBVIOUSNESS:
Sparks teaches generating customizable digital wrapping for a digital gift, including wrapping in the form of an envelope, Kolkin teaches using a natural language target prompt to specify desired visual content to be generated, and Lin further demonstrates that text to image models were known to generate depictions of envelopes in response to textual prompts. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to configure the natural language utterance of the combined system of Sparks, Kolkin, and Isaacson to request a virtual envelope, as taught by Lin. Such a modification would have predictably allowed the user to specify and generate Spark’s known envelope type digital wrapping using Kolkin’s existing text conditioned generation process, while retaining Kolkin’s benefit of generating visual output responsive to the details specified in the target prompt.
CONCLUSION
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW ALAN CADY/ Examiner, Art Unit 2145
/CESAR B PAULA/ Supervisory Patent Examiner, Art Unit 2145