DETAILED ACTION
This FINAL action is in response to Application No. 18/906,870 originally filed 10/04/2024. The amendment presented on 06/15/2026 which provides amendments to claims 1, 5, 9, 13-17, and 19 is hereby acknowledged.
Currently Claim(s) 1-20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-7 and 9-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Liu et al. U.S. Patent Application Publication No. 2024/0153153 A1 hereinafter Liu.
Consider Claim 1:
Liu discloses a computer-implemented method comprising: (Liu, See Abstract.)
receiving a single natural language text input for modifying a digital image, the single natural language text input indicating a first style, a second style, a first region of the digital image, and a second region of the digital image; (Liu, [0003], [0039], [0053], [0091-0093], [0073], “At step 1102, an input is received from a user. The input can include a first input text and a second input text. In some implementations, the input includes at least a third input text. The input texts can include phrases that describe objects, scenes, and/or scenarios. The phrases can further include an artistic phrase describing an artistic style in which to render the image. The input can also include information specifying regions.”)
determining, from the single natural language text input, a first style for modifying a to modify the first region of the digital image using the first style and a second style for modifying a to modify the second region of the digital image using the second style; modifying, using a multi-region style transfer neural network, the digital image by incorporating the first style within the first region and incorporating the second style within the second region; and (Liu, [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
providing the modified digital image for display on a graphical user interface of a client device. (Liu, [0077], “At step 1116, the initial image is updated with an image generated by applying the gradient to the processed image. Steps 1106 through steps 1116 are performed for a predetermined number of iterations. The predetermined number of iterations can vary. In some implementations, the predetermined number of iterations is between 70 to 100 iterations. At step 1118, a final image is outputted. The final image is the current updated initial image after the predetermined number of iterations has been performed.”)
Consider Claim 2:
Liu discloses the computer-implemented method of claim 1, wherein modifying, using the multi-region style transfer neural network, the digital image by incorporating the first style within the first region and incorporating the second style within the second region comprises: generating, using the multi-region style transfer neural network and from the digital image, a first modified digital image that incorporates the first style within the first region; and generating, using the multi-region style transfer neural network and from the first modified digital image, a second modified digital image that incorporates the second style within the second region while maintaining the first style within the first region. (Liu, [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
Consider Claim 3:
Liu discloses the computer-implemented method of claim 2, wherein generating the first modified digital image from the digital image using the multi-region style transfer neural network comprises modifying the digital image using the multi-region style transfer neural network over a plurality of iterations by modifying parameters of the multi-region style transfer neural network for one or more iterations using a plurality of loss functions. (Liu, [0043], “For each safe phrase 410, a corresponding safe image 422 is generated and paired with to form a safe phrase-image pair 412. These pairs of safe phrases and safe images 412 can be used as training data to train a new diffusion model 424. The safe image-phrase pairs 412 can be inputted into a loss generator 426, which generates and outputs at least a loss value 428. The loss value 428 may include an identity loss and/or a directional loss. The generated loss value 428 is used by a model trainer 104 to train a new diffusion model 424.”)
Consider Claim 4:
Liu discloses the computer-implemented method of claim 3, wherein modifying the parameters of the multi-region style transfer neural network for the one or more iterations using the plurality of loss functions comprises modifying the parameters for the one or more iterations using at least two of a masked directional loss function, a masked patch loss function, a content loss function, an identity loss function, or a relational loss function. (Liu, [0043], “For each safe phrase 410, a corresponding safe image 422 is generated and paired with to form a safe phrase-image pair 412. These pairs of safe phrases and safe images 412 can be used as training data to train a new diffusion model 424. The safe image-phrase pairs 412 can be inputted into a loss generator 426, which generates and outputs at least a loss value 428. The loss value 428 may include an identity loss and/or a directional loss. The generated loss value 428 is used by a model trainer 104 to train a new diffusion model 424.”)
Consider Claim 5:
Liu discloses the computer-implemented method of claim 1, further comprising generating a style-region mapping prompt that includes the single natural language text input and an example style-region mapping that corresponds to an example natural language text input, wherein determining, from the single natural language text input, to modify the first region using the first style and to modify the second region using the second style comprises determining, using a large language model and from the style-region mapping prompt, a style-region mapping that maps the first style to the first region and maps the second style to the second region. (Liu, [0017-0020], [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
Consider Claim 6:
Liu discloses the computer-implemented method of claim 1, further comprising generating, using a segmentation model, a first segmentation mask for the first region of the digital image and a second segmentation mask for the second region of the digital image, wherein modifying the digital image using the multi-region style transfer neural network comprises modifying the digital image using the multi-region style transfer neural network, the first segmentation mask, and the second segmentation mask. (Liu, [0054], “A text-image match gradient calculator 312 receives the input 600 and the processed image 608. The processed image 608 is then back-propagated through the text-image match gradient calculator 312 to calculate a gradient 610 against the input 600. To get feedback from the CLIP model for text-and-image consistency, a plurality of patches of the original generated image are randomly determined and fed into the CLIP model. For each of the plurality of patches, an image embedding is generated based on the processed image 608, and a text embedding is generated based on the region and the input text that are associated with the patch. The gradient 610 is calculated based on a differential between the image embedding and the text embedding.”)
Consider Claim 7:
Liu discloses the computer-implemented method of claim 6, further comprising generating, using a text grounding model, a first bounding box for the first region of the digital image and a second bounding box for the second region of the digital image, wherein generating, using the segmentation model, the first segmentation mask and the second segmentation mask comprises generating, using the segmentation model, the first segmentation mask from the first bounding box and the second segmentation mask from the second bounding box. (Liu, [0032], “At the second stage 332, the final first stage image 330 is inputted into the diffusion model 304, which processes the final first stage image 330 to generate a second stage processed image 334. The second stage processed image 334 outputted by the diffusion model 304 is back-propagated through the text-image match gradient calculator 312 to calculate a second stage gradient 336 against the input text 116. The gradient applicator 326 then applies the second stage gradient 336 to the second stage processed image 334 to generate an updated second stage image 338. The updated second stage image 338 generated by the gradient applicator 326 is inputted back into the diffusion model 304, and the process continues for a second predetermined number of iterations. The number of iterations can vary. In some embodiments, the second predetermined number of iterations may be between 5 to 15 iterations. In further embodiments, the second predetermined number of iterations is 10 iterations. After the second predetermined number of iterations is performed at the second stage 332, a final second stage image 340 generated by the gradient applicator 326 is inputted into the diffusion model 304 at a third stage 342. It will be appreciated that, unlike the first stage 300, the second stage 332 does not include a step for processing an image using the gradient estimator model 308. By neither back-propagating through the diffusion model 304 nor the gradient estimator model 308, the image generation process will be much faster than conventional methods. From the iterations performed during the first stage, the current second stage processed image 334 is at an acceptable level of quality such that the second stage gradient 336 output from the text-image match gradient calculator 312 is adequate to revise the second stage processed image 334 directly.”)
Consider Claim 9:
Liu discloses a system comprising: (Liu, See Abstract.)
one or more memory devices; and one or more processors coupled to the one or more memory devices that cause the system to perform operations comprising: (Liu, [0022], “The processor 202 is a microprocessor that includes one or more of a central processing unit (CPU), a graphical processing unit (GPU), an application specific integrated circuit (ASIC), a system on chip (SOC), a field-programmable gate array (FPGA), a logic circuit, or other suitable type of microprocessor configured to perform the functions recited herein. Volatile memory 206 can include physical devices such as random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), etc., which temporarily stores data only for so long as power is applied during execution of programs. Non-volatile memory 208 can include physical devices that are removable and/or built in, such as optical memory (e.g., CD, DVD, HD-DVD, Blu-Ray Disc, etc.), semiconductor memory (e.g., ROM, EPROM, EEPROM, FLASH memory, etc.), and/or magnetic memory (e.g., hard-disk drive, floppy-disk drive, tape drive, MRAM, etc.), or other mass storage device technology.”)
extracting, using a large language model and from a single natural language text input, a first text segment indicating a first style for modifying a first region of a digital image and a second text segment indicating a second style for modifying a second region of the digital image; (Liu, [0003], [0019], [0039], [0053], [0091-0093], [0073], “At step 1102, an input is received from a user. The input can include a first input text and a second input text. In some implementations, the input includes at least a third input text. The input texts can include phrases that describe objects, scenes, and/or scenarios. The phrases can further include an artistic phrase describing an artistic style in which to render the image. The input can also include information specifying regions.”)
determining, using a segmentation model, a first segmentation mask for the first region of the digital image and a second segmentation mask for the second region; (Liu, [0019], [0074], [0054], “A text-image match gradient calculator 312 receives the input 600 and the processed image 608. The processed image 608 is then back-propagated through the text-image match gradient calculator 312 to calculate a gradient 610 against the input 600. To get feedback from the CLIP model for text-and-image consistency, a plurality of patches of the original generated image are randomly determined and fed into the CLIP model. For each of the plurality of patches, an image embedding is generated based on the processed image 608, and a text embedding is generated based on the region and the input text that are associated with the patch. The gradient 610 is calculated based on a differential between the image embedding and the text embedding.”)
generating, using a multi-region style transfer neural network and from the digital image and the first segmentation mask, a first modified digital image that incorporates the first style within the first region; and generating, using the multi-region style transfer neural network and from the first modified digital image and the second segmentation mask, a second modified digital image that incorporates the second style within the second region. (Liu, [0077], “At step 1116, the initial image is updated with an image generated by applying the gradient to the processed image. Steps 1106 through steps 1116 are performed for a predetermined number of iterations. The predetermined number of iterations can vary. In some implementations, the predetermined number of iterations is between 70 to 100 iterations. At step 1118, a final image is outputted. The final image is the current updated initial image after the predetermined number of iterations has been performed.”)
Consider Claim 10:
Liu discloses the system of claim 9, wherein: generating, using the multi-region style transfer neural network, the first modified digital image from the digital image comprises modifying, using the multi-region style transfer neural network, the digital image over a first set of iterations to generate the first modified digital image; and generating, using the multi-region style transfer neural network, the second modified digital image from the first modified digital image comprises modifying, using the multi-region style transfer neural network, the first modified digital image over a second set of iterations to generate the second modified digital image. (Liu, [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
Consider Claim 11:
Liu discloses the system of claim 10, wherein modifying, using the multi-region style transfer neural network, the digital image over the first set of iterations comprises: generating a modified digital image from the digital image using the multi-region style transfer neural network with a set of parameters; modifying the set of parameters of the multi-region style transfer neural network using the modified digital image and a plurality of loss functions; and generating an additional modified digital image from the modified digital image using the multi-region style transfer neural network with the modified set of parameters. (Liu, [0043], “For each safe phrase 410, a corresponding safe image 422 is generated and paired with to form a safe phrase-image pair 412. These pairs of safe phrases and safe images 412 can be used as training data to train a new diffusion model 424. The safe image-phrase pairs 412 can be inputted into a loss generator 426, which generates and outputs at least a loss value 428. The loss value 428 may include an identity loss and/or a directional loss. The generated loss value 428 is used by a model trainer 104 to train a new diffusion model 424.”)
Consider Claim 12:
Liu discloses the system of claim 11, wherein modifying the set of parameters of the multi-region style transfer neural network using the plurality of loss functions comprises modifying the set of parameters of the multi-region style transfer neural network using a masked directional loss function, a masked patch loss function, a content loss function, an identity loss function, and a relational loss function. (Liu, [0043], “For each safe phrase 410, a corresponding safe image 422 is generated and paired with to form a safe phrase-image pair 412. These pairs of safe phrases and safe images 412 can be used as training data to train a new diffusion model 424. The safe image-phrase pairs 412 can be inputted into a loss generator 426, which generates and outputs at least a loss value 428. The loss value 428 may include an identity loss and/or a directional loss. The generated loss value 428 is used by a model trainer 104 to train a new diffusion model 424.”)
Consider Claim 13:
Liu discloses the system of claim 9, wherein: extracting, using the large language model, the first text segment indicating the first style for modifying the first region of the digital image and the second text segment indicating the second style for modifying the second region of the digital image comprises generating, using the large language model, a style-region mapping that maps the first style to the first region based on the first text segment and maps the second style to the second region based on the second text segment; and the operations further comprise determining to use the first segmentation mask for generating the first modified digital image to incorporate the first style within the first region based on determining that the style-region mapping maps the first style to the first region. (Liu, [0091], “The following paragraphs provide additional support for the claims of the subject application. One aspect provides a computer system for generating an output image corresponding to an input text, the computing system including a processor and memory of a computing device. The processor is configured to execute a program using portions of the memory to receive an input from a user, the input including a first input text and a second input text, provide an initial image, and, for a predetermined number of iterations, define a first region of the initial image associated with the first input text, define a second region of the initial image associated with the second input text, define a plurality of patches of the initial image, each patch associated with at least one of the regions, input the initial image into a diffusion process to generate a processed image, back-propagate the processed image through a text-image match gradient calculator to calculate a gradient against the input from the user, and update the initial image with an image generated by applying the calculated gradient to the processed image. Back-propagating the processed image through the text-image match gradient calculator is performed by, for each of the plurality of patches, generating an image embedding based on the processed image, generating a text embedding based on the region and the input text that are associated with the patch, and calculating a differential between the image embedding and the text embedding. In this aspect, additionally or alternatively, the input further includes a third input text, and the processor is further configured to define a third region of the initial image associated with the third input text. In this aspect, additionally or alternatively, each of the plurality of patches is associated with the region that has a largest intersection with the respective patch. In this aspect, additionally or alternatively, the generated text embedding for each of the plurality of patches is based on a weighted average of sub-text embeddings from the regions, where weights for the weighted average are proportional to an intersected area of the region and the respective patch. In this aspect, additionally or alternatively, the regions are defined based upon the input. In this aspect, additionally or alternatively, the first region is defined based upon the first input text. In this aspect, additionally or alternatively, the diffusion process is a denoising diffusion implicit model. In this aspect, additionally or alternatively, the diffusion process includes a gradient estimator model. In this aspect, additionally or alternatively, the diffusion process includes a diffusion model that has been trained using a curated dataset with safe content. In this aspect, additionally or alternatively, the predetermined number of iterations is between 70 and 100 iterations.”)
Consider Claim 14:
Liu discloses the system of claim 9, wherein extracting, using the large language model and from the single natural language text input, the first text segment indicating the first style for modifying the first region of the digital image comprises extracting, using the large language model and from the single natural language text input, aan additional text segment indicating the first style for modifying the first region. (Liu, [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
Consider Claim 15:
Liu discloses the system of claim 9, wherein: the operations further comprise generating, using a text grounding model, an indication of an association between the first text segment included in the single natural language text input and the first region of the digital image; and determining, using the segmentation model, the first segmentation mask for the first region of the digital image comprises determining, using the segmentation model, the first segmentation mask for the first region of the digital image based on the indication of the association between the first text segment and the first region. (Liu, [0032], “At the second stage 332, the final first stage image 330 is inputted into the diffusion model 304, which processes the final first stage image 330 to generate a second stage processed image 334. The second stage processed image 334 outputted by the diffusion model 304 is back-propagated through the text-image match gradient calculator 312 to calculate a second stage gradient 336 against the input text 116. The gradient applicator 326 then applies the second stage gradient 336 to the second stage processed image 334 to generate an updated second stage image 338. The updated second stage image 338 generated by the gradient applicator 326 is inputted back into the diffusion model 304, and the process continues for a second predetermined number of iterations. The number of iterations can vary. In some embodiments, the second predetermined number of iterations may be between 5 to 15 iterations. In further embodiments, the second predetermined number of iterations is 10 iterations. After the second predetermined number of iterations is performed at the second stage 332, a final second stage image 340 generated by the gradient applicator 326 is inputted into the diffusion model 304 at a third stage 342. It will be appreciated that, unlike the first stage 300, the second stage 332 does not include a step for processing an image using the gradient estimator model 308. By neither back-propagating through the diffusion model 304 nor the gradient estimator model 308, the image generation process will be much faster than conventional methods. From the iterations performed during the first stage, the current second stage processed image 334 is at an acceptable level of quality such that the second stage gradient 336 output from the text-image match gradient calculator 312 is adequate to revise the second stage processed image 334 directly.”)
Consider Claim 16:
Liu discloses the system of claim 15, wherein generating, using the text grounding model, the indication of the association between the first text segment and the first region of the digital image includes generating, using the text grounding model, a bounding box around the first region of the digital image based on the first text segment. (Liu, [0032], “At the second stage 332, the final first stage image 330 is inputted into the diffusion model 304, which processes the final first stage image 330 to generate a second stage processed image 334. The second stage processed image 334 outputted by the diffusion model 304 is back-propagated through the text-image match gradient calculator 312 to calculate a second stage gradient 336 against the input text 116. The gradient applicator 326 then applies the second stage gradient 336 to the second stage processed image 334 to generate an updated second stage image 338. The updated second stage image 338 generated by the gradient applicator 326 is inputted back into the diffusion model 304, and the process continues for a second predetermined number of iterations. The number of iterations can vary. In some embodiments, the second predetermined number of iterations may be between 5 to 15 iterations. In further embodiments, the second predetermined number of iterations is 10 iterations. After the second predetermined number of iterations is performed at the second stage 332, a final second stage image 340 generated by the gradient applicator 326 is inputted into the diffusion model 304 at a third stage 342. It will be appreciated that, unlike the first stage 300, the second stage 332 does not include a step for processing an image using the gradient estimator model 308. By neither back-propagating through the diffusion model 304 nor the gradient estimator model 308, the image generation process will be much faster than conventional methods. From the iterations performed during the first stage, the current second stage processed image 334 is at an acceptable level of quality such that the second stage gradient 336 output from the text-image match gradient calculator 312 is adequate to revise the second stage processed image 334 directly.”)
Consider Claim 17:
Liu discloses a non-transitory computer-readable medium storing instructions thereon that, when executed by at least one processor, cause the at least one processor to perform operations comprising: (Liu, [0064], [0083], “Non-volatile storage device 1206 includes one or more physical devices configured to hold instructions executable by the logic processors to implement the methods and processes described herein. When such methods and processes are implemented, the state of non-volatile storage device 1206 may be transformed—e.g., to hold different data.”)
receiving a single natural language text input for modifying a digital image, the single natural language text input indicating a first style, a second style, a first region of the digital image, and a second region of the digital image; (Liu, [0003], [0019], [0039], [0053], [0091-0093], [0073], “At step 1102, an input is received from a user. The input can include a first input text and a second input text. In some implementations, the input includes at least a third input text. The input texts can include phrases that describe objects, scenes, and/or scenarios. The phrases can further include an artistic phrase describing an artistic style in which to render the image. The input can also include information specifying regions.”)
determining, from the single natural language text input, to modify the first region of the digital image using the first style and to modify the second region of the digital image using the second style; modifying, using a multi-region style transfer neural network, the digital image by incorporating the first style within the first region and incorporating the second style within the second region; and (Liu, [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
providing the modified digital image for display on a graphical user interface of a client device. (Liu, [0077], “At step 1116, the initial image is updated with an image generated by applying the gradient to the processed image. Steps 1106 through steps 1116 are performed for a predetermined number of iterations. The predetermined number of iterations can vary. In some implementations, the predetermined number of iterations is between 70 to 100 iterations. At step 1118, a final image is outputted. The final image is the current updated initial image after the predetermined number of iterations has been performed.”)
Consider Claim 18:
Liu discloses the non-transitory computer-readable medium of claim 17, wherein modifying, using the multi-region style transfer neural network, the digital image by incorporating the first style within the first region and incorporating the second style within the second region comprises modifying, using the multi-region style transfer neural network, the digital image by incorporating the first style within the first region and incorporating the second style within the second region while maintaining an initial style within a third region of the digital image. (Liu, [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
Consider Claim 19:
Liu discloses the non-transitory computer-readable medium of claim 17, wherein receiving the single natural language text input for modifying the digital image comprises receiving a single string of text indicating a plurality of image regions to modify and a plurality of image styles for modifying the plurality of image regions. (Liu, [0073], “At step 1104, an initial image, which may include an image of random noise is provided. At step 1106, a first region of the initial image is defined. The first region is associated with the first input text. At step 1108, a second region of the initial image is defined. The second region is associated with the second input text. The regions can be defined and determined in many different ways. In some implementations, the regions are determined based on information in the input received from the user. For example, the input could specify a region of the image where the content of the input text is to be generated. In some implementations, the regions are determined by applying natural language processing techniques on the input text.”)
Consider Claim 20:
Liu discloses the non-transitory computer-readable medium of claim 17, wherein modifying, using the multi-region style transfer neural network, the digital image by incorporating the first style within the first region and incorporating the second style within the second region comprises modifying the digital image over a plurality of modification iterations by using, for one or more modification iterations, the multi-region style transfer neural network having updated parameters determined using one or more loss functions. (Liu, [0043], “For each safe phrase 410, a corresponding safe image 422 is generated and paired with to form a safe phrase-image pair 412. These pairs of safe phrases and safe images 412 can be used as training data to train a new diffusion model 424. The safe image-phrase pairs 412 can be inputted into a loss generator 426, which generates and outputs at least a loss value 428. The loss value 428 may include an identity loss and/or a directional loss. The generated loss value 428 is used by a model trainer 104 to train a new diffusion model 424.”)
Claim Rejections - 35 USC § 103
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. U.S. Patent Application Publication No. 2024/0153153 A1 as applied to claim 1 above, and further in view of Dundar U.S. Patent Application Publication No. 2019/0244060 A1 hereinafter Dundar.
Consider Claim 8:
Liu discloses the computer-implemented method of claim 1, and while disclosing modifying the digital image using the multi-region style transfer however does not specify that the neural network comprises modifying the digital image using a convolutional neural network.
Dundar however teaches that it was a known technique by those having ordinary skill in the art before the effective filing date of the invention to provide a neural network which comprises modifying the digital image using a convolutional neural network. (Dundar, [0037], “Stylization is formulated as an image reconstruction problem with feature projections. Referring back to FIG. 1A, the encoders 165 and decoder 170 form an auto-encoder comprising an encoder 165 coupled directly to the decoder 170 (without the second encoder and projection modules P.sub.C and P.sub.S) for general image reconstruction. In one embodiment, the encoder 165 implements a pre-trained VGG-19 convolutional neural network (the weights are kept fixed) and the decoder 170 is trained for reconstructing input images provided to the encoder 165. In an embodiment, the decoder 170 is symmetrical to the encoder 165. In an embodiment, the decoder 170 is trained by minimizing the sum of the L.sub.2 reconstruction loss and perceptual loss.”)
It therefore would have been obvious to those having ordinary skill in the art before the effective filing date of the invention to provide convolutional neural network for image modification as this was a known technique as taught by Dundar and would have been utilized for the purpose of training datasets may be automatically generated. Additionally, the performance of neural networks trained using the stylized synthetic training datasets generated from the synthetic images is improved because covariate alignment between the photorealistic images and the stylized synthetic images is improved. The stylization operation not only transfers the style of the real images to the synthetic images, but also more closely aligns the covariate of the synthetic images to the covariate of the real images. Furthermore, the style transfer neural network model 610 does not need to be trained. Therefore, the stylization process can potentially be performed “on the fly,” which means there is no requirement to pre-stylize all the synthetic images, streamlining the process and reducing storage requirements. (Dundar, [0163])
Dundar however teaches that it was a known technique by those having ordinary skill in the art before the effective filing date of the invention to provide wherein generating, using the text grounding model, the indication of the association between the first text segment and the first region of the digital image includes generating, using the text grounding model, a bounding box around the first region of the digital image based on the first text segment. (Dundar, [0065], “FIG. 2C illustrates a photorealistic style image, a photorealistic content image and corresponding style segmentation data and content segmentation data, respectively, in accordance with an embodiment. In an embodiment, the style and/or content segmentation data is provided by semantic label maps. The photorealistic style image is segmented into a first style region and a second style region. The first style region in the style segmentation data identifies (i.e., is labeled as) the road and the second style region in the style segmentation data identifies the landscape. The photorealistic content image is segmented into a first content region and a second content region. The first content region in the content segmentation data identifies the road and corresponds with the first style region. The second content region in the content segmentation data identifies the landscape and corresponds with the second style region. The style and/or content segmentation data may define additional regions, such as the sky.”)
It therefore would have been obvious to those having ordinary skill in the art before the effective filing date of the invention to provide bounding box for image modification as this was a known technique as taught by Dundar and would have been utilized for the purpose of training datasets may be automatically generated. Additionally, the performance of neural networks trained using the stylized synthetic training datasets generated from the synthetic images is improved because covariate alignment between the photorealistic images and the stylized synthetic images is improved. The stylization operation not only transfers the style of the real images to the synthetic images, but also more closely aligns the covariate of the synthetic images to the covariate of the real images. Furthermore, the style transfer neural network model 610 does not need to be trained. Therefore, the stylization process can potentially be performed “on the fly,” which means there is no requirement to pre-stylize all the synthetic images, streamlining the process and reducing storage requirements. (Dundar, [0163])
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Prior art made of record and not relied upon which is still considered pertinent to applicant's disclosure is cited in a current or previous PTO-892. The prior art cited in a current or previous PTO-892 reads upon the applicants claims in part, in whole and/or gives a general reference to the knowledge and skill of persons having ordinary skill in the art before the effective filing date of the invention. Applicant, when responding to this Office action, should consider not only the cited references applied in the rejection but also any additional references made of record.
In the response to this office action, the Examiner respectfully requests support be shown for any new or amended claims. More precisely, indicate support for any newly added language or amendments by specifying page, line numbers, and/or figure(s). This will assist The Office in compact prosecution of this application. The Office has cited particular columns, paragraphs, and/or line numbers in the applied rejection of the claims above for the convenience of the applicant. Citations are representative of the teachings in the art and are applied to the specific limitations within each claim, however other passages and figures may apply. Applicant, in preparing a response, should fully consider the cited reference(s) in its entirety and not only the cited portions as other sections of the reference may expand on the teachings of the cited portion(s).
Applicant Representatives are reminded of CFR 1.4(d)(2)(ii) which states “A patent practitioner (§ 1.32(a)(1) ), signing pursuant to §§ 1.33(b)(1) or 1.33(b)(2), must supply his/her registration number either as part of the S-signature, or immediately below or adjacent to the S-signature. The number (#) character may be used only as part of the S-signature when appearing before a practitioner’s registration number; otherwise the number character may not be used in an S-signature.” When an unsigned or improperly signed amendment is received the amendment will be listed in the contents of the application file, but not entered. The examiner will notify applicant of the status of the application, advising him or her to furnish a duplicate amendment properly signed or to ratify the amendment already filed. In an application not under final rejection, applicant should be given a two month time period in which to ratify the previously filed amendment (37 CFR 1.135(c) ).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. Granting of After Final Interviews: “Interviews merely to restate arguments of record or to discuss new limitations which would require more than nominal reconsideration or new search should be denied.” See MPEP § 713.09.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL J JANSEN II whose telephone number is (571)272-5604. The examiner can normally be reached Normally Available Monday-Friday 9am-4pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Temesghen Ghebretinsae can be reached on 571-272-3017. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Michael J Jansen II/ Primary Examiner, Art Unit 2626