DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This communication is responsive to the application filed 8/5/2024.
Claims 1-20 are pending with claims 1, 9, and 15 as independent claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Aggarwal et al. (US 2024/0355018, filed 4/20/2023, hereinafter as Aggarwal) in view of Radford et al. (Learning Transferable Visual Models From Natural Language Supervision, published 2021, pages 1-16).
Claim 1. A system comprising:
one or more processors; Aggarwal discloses in [0100] “the components of the mask aware image editing system 102 include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the computing device 900).” (emphasis added),
a non-transitory computer-readable medium storing a program executable by the one or more processors, the program comprising sets of instructions for:
receiving a first natural language input comprising a plurality of words; Aggarwal discloses in [0022 and 0077-0078] “the mask aware image editing system can capture a text prompt and generate a base digital image from the text prompt. The mask aware image editing system can then utilize the base digital image to generate a stylized image from a shape mask… If the style prompt is in text modality, the mask aware image editing system 102 can use an image generator to generate a style image.” (emphasis added) examiner note: the text prompt may be natural language (text) input,
associating the plurality of words with a plurality of controls in a user interface, the plurality of controls comprising visual representations in the user interface corresponding to the plurality of words, wherein the visual representations are configurable in the user interface; Aggarwal discloses in [0085-0089] “the stylized image 518 reflects the base digital image 506 and the shape mask 510 according to the structural edit strength parameter indicated by the structural weight element 514 and the diffusion noising model selection element 512… the mask aware image editing system 102 can take advantage of the structure weight parameter to generate different variations for the typographies generated based on the dominance of the style image. The mask aware image editing system 102 can even use the structure weight as a slider in the user interface to get different variations… if a client device captures a digital image, the mask aware image editing system 102 can automatically generate a stylized image that transforms the captured digital image (e.g., automatically generate a stylized font from the captured digital image).” (emphasis added) examiner note: the shape mask 510 may represent a word and the weight element 514 may represent a control in user interface 502 such that the weight element may be configured, when the strength parameter changes, to at least the style of shape mask 510 as shown in fig. 5,
receiving a first user input modifying a configuration of the visual representations; Aggarwal discloses in [0081 and 0086-0092] “The mask aware image editing system 102 can determine a structural edit strength parameter (i.e., a structural number of steps) based on user interaction with the structural weight element 514… the mask aware image editing system 102 generates a stylized animation corresponding to multiple different structural edit strength parameters (i.e., multiple structural numbers of steps)… the mask aware image editing system 102 can select structural edit strength parameters of 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, and 0.8 (and corresponding structural numbers of steps such as 20, 30, 40, 50, 60, and 80)… the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added) examiner note: the user may modify the configuration of the user interface 502 by interacting with the structural weight element 514,
in response to receiving the first user input, mapping the visual representations to a first set of numeric values, wherein each value of the first set of numeric values is based on the configuration of the visual representations, and wherein changes to the configuration of the visual representations change corresponding numeric values; Aggarwal discloses in [0086-0092] “the mask aware image editing system 102 generates a stylized animation corresponding to multiple different structural edit strength parameters (i.e., multiple structural numbers of steps)… the mask aware image editing system 102 can select structural edit strength parameters of 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, and 0.8 (and corresponding structural numbers of steps such as 20, 30, 40, 50, 60, and 80)… the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added) examiner note: the user may modify the configuration of the user interface 502 by interacting with the structural weight element 514, wherein the user interface comprises visual representations such as base image 506 and textual character 510. Changes to the structural weight slider 514 would change strength parameters or numeric values or weights as shown in fig. 5A-5B,
mapping each numeric value of the first set of numeric values to a first set of predefined natural language terms; Aggarwal discloses in [0086-0092] “the mask aware image editing system 102 generates a stylized animation corresponding to multiple different structural edit strength parameters (i.e., multiple structural numbers of steps)… the mask aware image editing system 102 can select structural edit strength parameters of 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, and 0.8 (and corresponding structural numbers of steps such as 20, 30, 40, 50, 60, and 80)… the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added) examiner note: Changes to the structural weight slider 514 would change strength parameters or numeric values or weights such that the changes would be mapped to the textual character (A) style as shown in fig. 5B, wherein the textual character (A) may represent at least one predefined natural language term,
generating a prompt based on the first set of predefined natural language terms; Aggarwal discloses in [0024 and 0084-0085] “the mask aware image editing system can provide realistic and higher quality results for both image-based style prompts and text-based style prompts… the user interface 502 also includes a generated image element 516. Based on user action with the generate image element 516, the mask aware image editing system 102 can generate a stylized image based on the base digital image 506 and the shape mask 510… FIG. 5B illustrates the user interface 502 that includes a stylized image 518. As shown, the stylized image 518 reflects the base digital image 506 and the shape mask 510 according to the structural edit strength parameter indicated by the structural weight element 514 and the diffusion noising model selection element 512.” (emphasis added) examiner note: user interaction with the generate image element 516 may generate a text-based style prompt,
sending the prompt to [a large language machine learning model]; Aggarwal discloses in [0041 and 0091] “the mask aware image editing system 102 also utilizes a diffusion neural network 208 to generate a stylized image 210 from the mask-segmented image 206. As used herein, the term neural network refers to a machine learning model of interconnected artificial neurons (e.g., organized in layers) that communicate and learn to approximate complex functions and generate outputs… if a client device captures a digital image, the mask aware image editing system 102 can automatically generate a stylized image that transforms the captured digital image (e.g., automatically generate a stylized font from the captured digital image)… The mask aware image editing system 102 generates stylized typography characters from these two base digital images, a constant structural weight, and two different diffusion noising models. As shown, the mask aware image editing system 102 can generate natural realistic stylized images from input text (e.g., a text prompt).” (emphasis added).
Aggarwal does not explicitly disclose a large language machine learning model. However, Radford, in an analogous art, discloses in [p. 2, section 2.2] “In Figure 2 we show that a 63 million parameter transformer language model, which already uses twice the compute of its ResNet50 image encoder, learns to recognize ImageNet classes three times slower than an approach similar to Joulin et al. (2016) that predicts a bag-of-words encoding of the same text.” (emphasis added) examiner note: the transformer language model may be a large language model (LLM).
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Aggarwal with the teaching of Radford because “We demonstrate that a simplified version of ConVIRT trained from scratch, which we call CLIP, for Contrastive Language Image Pre-training, is an efficient and scalable method of learning from natural language supervision. We find that CLIP learns to perform a wide set of tasks during pre-training including OCR, geo-localization, action recognition, and outperforms the best publicly available ImageNet model while being more computationally efficient. We also find that zero-shot CLIP models are much more robust than equivalent accuracy supervised ImageNet models.” Radford [P. 1, Introduction], and
producing, by the large language machine learning model, one or more output images. Aggarwal discloses in [0085] “FIG. 5B illustrates the user interface 502 generated by the mask aware image editing system 102 in response to user interaction with the generate image element 516. In particular, FIG. 5B illustrates the user interface 502 that includes a stylized image 518. As shown, the stylized image 518 reflects the base digital image 506 and the shape mask 510 according to the structural edit strength parameter indicated by the structural weight element 514 and the diffusion noising model selection element 512.” (emphasis added) examiner note: image 518 may be generated based on text-based style prompt as shown in fig. 5B.
Claims 2, 10, and 16. The rejection of the system of claim 1 is incorporated, wherein the visual representations comprise a slider user interface component associated with each word in the plurality of words, wherein different positions of each slider in the slider user interface components change the configuration of the visual representations, and in accordance therewith, change the corresponding numeric values. Aggarwal discloses in [0086-0087] “the mask aware image editing system 102 can select a plurality of structural edit strength parameters (i.e., a plurality of structural numbers of steps). For example, the mask aware image editing system 102 can select structural edit strength parameters of 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, and 0.8 (and corresponding structural numbers of steps such as 20, 30, 40, 50, 60, and 80). For example, for each structural edit strength parameter, the mask aware image editing system 102 can generate a corresponding stylized image. The mask aware image editing system 102 can then combine the stylized images as frames in a stylized animation. Moreover, the mask aware image editing system 102 can provide the stylized animation for display (e.g., as a stylized animation that sequentially displays each of the stylized frames in a loop)… The mask aware image editing system 102 can even use the structure weight as a slider in the user interface to get different variations.” (emphasis added).
Claims 3, 11, and 17. The rejection of the system of claim 1 is incorporated, wherein the visual representations comprise, for each word in the plurality of words, visual representations of the plurality of words in the user interface having associated font sizes, wherein the font sizes are configurable in the user interface, and wherein changes to the font sizes change the configuration of the visual representations, and in accordance therewith, change the corresponding numeric values. Aggarwal discloses in [0089-0090] “if a client device captures a digital image, the mask aware image editing system 102 can automatically generate a stylized image that transforms the captured digital image (e.g., automatically generate a stylized font from the captured digital image)… FIG. 6 illustrates example stylized images (i.e., stylized typography characters) generated from based digital images and shape masks (e.g., typography masks corresponding to the letters A, B, C, D, E, F, and G) in accordance with one or more embodiments. In particular, FIG. 6 illustrates stylized images generated utilizing two different base digital images, a constant structural weight, and two different diffusion noising models. FIG. 6 illustrates example stylized images (i.e., stylized typography characters) generated from based digital images and shape masks (e.g., typography masks corresponding to the letters A, B, C, D, E, F, and G) in accordance with one or more embodiments. In particular, FIG. 6 illustrates stylized images generated utilizing two different base digital images, a constant structural weight, and two different diffusion noising models.” (emphasis added).
Claim 4. The rejection of the system of claim 1 is incorporated, wherein producing the one or more output images comprises generating, in the user interface, a preview, wherein the preview is based on the one or more output images and associated with the first natural language input being displayed in the user interface. Aggarwal discloses in [0085] “FIG. 5B illustrates the user interface 502 generated by the mask aware image editing system 102 in response to user interaction with the generate image element 516. In particular, FIG. 5B illustrates the user interface 502 that includes a stylized image 518. As shown, the stylized image 518 reflects the base digital image 506 and the shape mask 510 according to the structural edit strength parameter indicated by the structural weight element 514 and the diffusion noising model selection element 512.” (emphasis added) examiner note: image 518 may be generated based on text-based style prompt as shown in fig. 5B.
Claims 5, 12, and 18. The rejection of the system of claim 4 is incorporated, the claim further comprising:
receiving, from the large language machine learning model, the one or more output images as the preview; Aggarwal discloses in [0066] “The structural edit strength parameter can indicate structural number of steps, and thus the structural transition step 408, the first set of noising steps 410, the second set of noising steps 412, the first set of denoising steps 414, and the second set of denoising steps 416.” (emphasis added),
in response to receiving the preview, receiving a second user input modifying the configuration of the visual representations in the user interface; Aggarwal discloses in [0066] “The structural edit strength parameter can indicate structural number of steps, and thus the structural transition step 408, the first set of noising steps 410, the second set of noising steps 412, the first set of denoising steps 414, and the second set of denoising steps 416.” (emphasis added),
mapping the configuration to a second set of numeric values by changing corresponding numeric values of the first set of numeric values based on the second user input and the configuration of the visual representations; Aggarwal discloses in [0092] “the mask aware image editing system 102 can also modify structural weights in generating stylized images. FIG. 8, illustrates a plurality of stylized images generated utilizing different structural weights in accordance with one or more embodiments. In particular, the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added),
mapping each numeric value of the second set of numeric values to a second set of predefined natural language terms; Aggarwal discloses in [0092] “the mask aware image editing system 102 can also modify structural weights in generating stylized images. FIG. 8, illustrates a plurality of stylized images generated utilizing different structural weights in accordance with one or more embodiments. In particular, the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added),
updating the prompt based on the second set of predefined natural language terms; Aggarwal discloses in [0092] “the mask aware image editing system 102 can also modify structural weights in generating stylized images. FIG. 8, illustrates a plurality of stylized images generated utilizing different structural weights in accordance with one or more embodiments. In particular, the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added),
sending the prompt to the large language machine learning model;
producing, by the large language machine learning model, one or more new output images; Aggarwal discloses in [0092] “the mask aware image editing system 102 can also modify structural weights in generating stylized images. FIG. 8, illustrates a plurality of stylized images generated utilizing different structural weights in accordance with one or more embodiments. In particular, the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added), and
updating the user interface with an updated preview corresponding to the one or more new output images by replacing the preview corresponding to the one or more output images with the updated preview in the user interface. Aggarwal discloses in [0092] “the mask aware image editing system 102 can also modify structural weights in generating stylized images. FIG. 8, illustrates a plurality of stylized images generated utilizing different structural weights in accordance with one or more embodiments. In particular, the mask aware image editing system 102 generates the stylized images from the shape mask 510 and the base digital image 506 from FIG. 5. Thus, the mask aware image editing system 102 can generate a first stylized image (and/or stylized typography character) utilizing a first structural number of steps (e.g., noising or denoising steps) and a second stylized image (and/or second stylized typography character) utilizing a second structural number of steps (e.g., noising or denoising steps).” (emphasis added).
Claims 6, 7, 13, and 19. The rejection of the system of claim 1 is incorporated, Aggarwal does not explicitly disclose wherein the first set of predefined natural language terms comprise one or more of a plurality of adjectives and/or a plurality of adverbs, and wherein the plurality of adjectives and/or the plurality of adverbs are predefined based on the large language machine learning model. However, Radford, in an analogous art, discloses in [p. 2, section 2.2] “CLIP jointly trains an image encoder and a text encoder to predict the correct pairings of a batch of (image, text) training examples. At test time the learned text encoder synthesizes a zero-shot linear classifier by embedding the names or descriptions of the target dataset’s classes.” (emphasis added) examiner note: the text prompt may comprises textual terms as shown in fig. 1.
Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Aggarwal with the teaching of Radford because “We demonstrate that a simplified version of ConVIRT trained from scratch, which we call CLIP, for Contrastive Language Image Pre-training, is an efficient and scalable method of learning from natural language supervision. We find that CLIP learns to perform a wide set of tasks during pre-training including OCR, geo-localization, action recognition, and outperforms the best publicly available ImageNet model while being more computationally efficient. We also find that zero-shot CLIP models are much more robust than equivalent accuracy supervised ImageNet models.” Radford [P. 1, Introduction].
Claims 8, 14, and 20. The rejection of the system of claim 1 is incorporated, wherein the one or more output images comprises a video and wherein the large language machine learning model is configured to produce the video based on the prompt. Aggarwal discloses in [0020] “the mask aware image editing system can flexibly control the structural fidelity relative to the base digital image in generating a stylized digital image. Indeed, in some implementations, the mask aware image editing system generates animated stylized images by generating different stylized images based on different style weights and then combining the different stylized images as frames in a stylized animation.” (emphasis added) examiner note: the animation may be video.
Claim 9. The claim is directed towards a method to implement the steps of the system of claim 1, therefore, the claim is similarly rejected as claim 1.
Claim 15. The claim is directed towards a non-transitory computer readable medium, storing a program executable by at least one processing unit of a device, to implement the steps of the system of claim 1, therefore, the claim is similarly rejected as claim 1.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO-892.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AHAMED I NAZAR whose telephone number is (571)270-3174. The examiner can normally be reached 10 am to 7 pm Mon-Fri.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Stephen Hong can be reached at 571-272-4124. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AHAMED I NAZAR/Examiner, Art Unit 2178 06/25/2026
/STEPHEN S HONG/Supervisory Patent Examiner, Art Unit 2178