DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: “314a-314c” in paragraph 0082 lines 3 and 5, and “470” in paragraph 0110 lines 1-3 and 7. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 514 in Fig. 18 and 556 in Fig. 19. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to because perhaps “286a” in Fig. 12 should read “288a” to match paragraph 0084 line 9 and Fig. 6. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
In paragraph 0115 line 7, “training of the a relevant model” should read “training of a relevant model”.
In paragraph 0119 line 1, “step 510” should read “step 512”.
In paragraph 0119 line 3, “step 512” should read “step 514”.
Appropriate correction is required.
The use of the terms “WiFi”, “3GPP”, “macOS”, “Microsoft Windows”, “Android”, “Linux”, “Bluetooth”, “Zigbee”, etc. which are trade names or marks used in commerce, have been noted in this application. These terms should be accompanied by the generic terminology; furthermore the terms should be capitalized wherever it appears or, where appropriate, include a proper symbol indicating use in commerce such as ™, SM , or ® following the term.
Although the use of trade names and marks used in commerce (i.e., trademarks, service marks, certification marks, and collective marks) are permissible in patent applications, the proprietary nature of the marks should be respected and every effort made to prevent their use in any manner which might adversely affect their validity as commercial marks.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 7-10, 14-16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Couleaud et al. (US 20250078346 A1), hereinafter Couleaud.
Regarding claim 1, Coudleaud teaches a system (Paragraph 0038 – “image processing system”), comprising:
a processor (Paragraph 0095 – “Control circuitry 804 may be based on any suitable control circuitry such as processing circuitry 806. As referred to herein, control circuitry should be understood to mean circuitry based on one or more microprocessors… control circuitry 804 executes instructions for the image processing system”; Note: the system contains control circuitry, which is equivalent to the processor);
and a non-transitory memory storing instructions that, when executed, cause the processor to (Paragraph 0096 – “The image processing system or application may be implemented as software or a set of executable instructions. The instructions for performing any of the embodiments discussed herein of the image processing system or application may be encoded on non-transitory computer-readable media (e.g., a hard drive, random-access memory on a DRAM integrated circuit, read-only memory on a BLU-RAY disk, etc.). For example, in FIG. 3, the instructions may be stored in storage 808, and executed by control circuitry 804 of a device 800”):
receive a request to generate an interface element package for a selected item (Paragraph 0038-0039 – “the image processing system may receive, as shown in FIG. 1, input of a text prompt 102. Such input may be received at a user interface of a computing device from a user in any suitable form…As a non-limiting illustrative example, as shown in FIG. 2, text prompt 102 may correspond to text prompt 202 of FIG. 2 of ‘a picture of a young woman holding a cat in front of a Victorian era building, 1870, high quality, soft focus, f/18, 60 mm, in the style of Auguste Renoir’”; Note: the prompt is equivalent to the request to generate an interface element package for a selected item, which in this case corresponds to a young woman, a cat, and/or a Victorian era building);
in response to receiving the request: generate a foreground image element (Fig. 2, Paragraph 0040, 0065, 0124 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… the image processing system may perform segmentation 114 of image 110 to obtain images 116, 118, and 120 (e.g., different identified portions, segments, and/or objects of image 210 of FIG. 2), which may respectively correspond to images 216, 218, and 220…based on segmentation 414, corresponding masks (e.g., 222 and 224 of FIG. 2) may be obtained for one or more of images 416, 418 and 420, respectively… a particular mask (e.g., mask 222 of FIG. 2) and/or other suitable information may be used to generate an object image (e.g., woman 236) without any empty regions or holes”; Note: a foreground image element 236 is generated, in response to an input prompt/request. It is shown in Fig. 2; see screenshot below);
generate a contextually appropriate background image (Fig. 2, Paragraph 0040, 0056 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)”; Note: image 234 is a background image that is contextually appropriate based on the input prompt. Screenshot of Fig. 2 shows the background image 234);
PNG
media_image1.png
672
426
media_image1.png
Greyscale
Screenshot of Fig. 2 (taken from Couleaud)
and generate an integrated image by integrating the foreground image element into the contextually appropriate background image at a contextually appropriate position, wherein the foreground image element is unmodified in the integrated image (Fig. 2 and 3, Paragraph 0040, 0061 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image, e.g., composite image 240, 242, 244, or 246”; Note: the multi-layer/composite image is equivalent to the integrated image since it combines the foreground image 236 and background image 234. Fig. 3 shows examples of integrated images, where the foreground element is at a contextually appropriate position (not floating in the sky) and its composition is unmodified).
PNG
media_image2.png
565
448
media_image2.png
Greyscale
Screenshot of Fig. 3 (taken from Couleaud)
In the above embodiment, the system of Couleaud does not directly teach generating at least one textual interface element and generating the interface element package including at least the integrated image and the textual interface element. However, in a different embodiment, Couleaud teaches causing the processor to:
generate at least one textual interface element (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the text elements 718 describing the image layers in the GUI are equivalent to the generated textual interface elements; see screenshot of Fig. 7 below);
and generate the interface element package including at least the integrated image and the textual interface element (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the components of the GUI make-up the interface element package. It includes the multi-layer/integrated image and text elements 718 describing the image layers; see screenshot of Fig. 7 below).
PNG
media_image3.png
336
433
media_image3.png
Greyscale
Screenshot of Fig. 7 (taken from Couleaud)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud, to have obtained the above, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP $2143(A). The prior art included each element recited in claim 1, although not necessarily in a single embodiment, with the only difference being between the claimed element and the prior art being the lack of actual combination of certain elements in a single prior art embodiment, as described above. One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 2, Couleaud teaches the system of claim 1. Couleaud further teaches wherein the foreground image element comprises an isolated foreground image (Fig. 2, Paragraph 0061 – “images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image”; Note: the image 236 is equivalent to the foreground image element, and it is isolated, as shown in Fig. 2; see screenshot of Fig. 2 above).
Regarding claim 3, Couleaud teaches the system of claim 2. Couleaud further teaches wherein the isolated foreground image is generated by an image segmentation model (Paragraph 0043, 0065, 0124 – “Any suitable number or types of techniques may be used to perform such segmentation, such as, for example: machine learning…based on segmentation 414, corresponding masks (e.g., 222 and 224 of FIG. 2) may be obtained for one or more of images 416, 418 and 420, respectively…a particular mask (e.g., mask 222 of FIG. 2) and/or other suitable information may be used to generate an object image (e.g., woman 236) without any empty regions or holes”; Note: the object image 236, which is equivalent to the isolated foreground image, is generated using masks based on image segmentation).
Regarding claim 7, Couleaud teaches the system of claim 1. Couleaud further teaches wherein the textual interface element is descriptive of the integrated image (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the text elements 718 describe the multi-layer/integrated image 714; see screenshot of Fig. 7 above).
Regarding claim 8, Couleaud teaches a computer-implemented method (Paragraph 0006, 0092 – “systems, methods, and apparatuses are disclosed herein for generating a multi-layer image based on text input… FIGS. 8-9 show illustrative devices, systems, servers, and related hardware for generating a multi-layer image, in accordance with some embodiments of this disclosure. FIG. 8 shows generalized embodiments of illustrative computing devices 800 and 801, which may correspond to, e.g., a smart phone; a tablet; a laptop computer; a personal computer; a desktop computer;”), comprising:
receiving a request to generate an interface element package for a selected item (Paragraph 0038-0039 – “the image processing system may receive, as shown in FIG. 1, input of a text prompt 102. Such input may be received at a user interface of a computing device from a user in any suitable form…As a non-limiting illustrative example, as shown in FIG. 2, text prompt 102 may correspond to text prompt 202 of FIG. 2 of ‘a picture of a young woman holding a cat in front of a Victorian era building, 1870, high quality, soft focus, f/18, 60 mm, in the style of Auguste Renoir’”; Note: the prompt is equivalent to the request to generate an interface element package for a selected item, which in this case corresponds to a young woman, a cat, and/or a Victorian era building);
in response to receiving the request: generating a foreground image element (Fig. 2, Paragraph 0040, 0065, 0124 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… the image processing system may perform segmentation 114 of image 110 to obtain images 116, 118, and 120 (e.g., different identified portions, segments, and/or objects of image 210 of FIG. 2), which may respectively correspond to images 216, 218, and 220…based on segmentation 414, corresponding masks (e.g., 222 and 224 of FIG. 2) may be obtained for one or more of images 416, 418 and 420, respectively… a particular mask (e.g., mask 222 of FIG. 2) and/or other suitable information may be used to generate an object image (e.g., woman 236) without any empty regions or holes”; Note: a foreground image element 236 is generated, in response to an input prompt/request. It is shown in Fig. 2; see screenshot above);
generating a contextually appropriate background image (Fig. 2, Paragraph 0040, 0056 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)”; Note: image 234 is a background image that is contextually appropriate based on the input prompt. Screenshot of Fig. 2 shows the background image 234);
and generating an integrated image by integrating the foreground image element into the contextually appropriate background image at a contextually appropriate position, wherein the foreground image element is unmodified in the integrated image (Fig. 2 and 3, Paragraph 0040, 0061 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image, e.g., composite image 240, 242, 244, or 246”; Note: the multi-layer/composite image is equivalent to the integrated image since it combines the foreground image 236 and background image 234. Fig. 3 shows examples of integrated images, where the foreground element is at a contextually appropriate position (not floating in the sky) and its composition is unmodified).
In the above embodiment, the computer-implemented method of Couleaud does not directly teach generating at least one textual interface element and generating the interface element package including at least the integrated image and the textual interface element. However, in a different embodiment, Couleaud teaches
generating at least one textual interface element (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the text elements 718 describing the image layers in the GUI are equivalent to the generated textual interface elements; see screenshot of Fig. 7 above);
and generating the interface element package including at least the integrated image and the textual interface element (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the components of the GUI make-up the interface element package. It includes the multi-layer/integrated image and text elements 718 describing the image layers; see screenshot of Fig. 7 above).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud, to have obtained the above, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP $2143(A). The prior art included each element recited in claim 8, although not necessarily in a single embodiment, with the only difference being between the claimed element and the prior art being the lack of actual combination of certain elements in a single prior art embodiment, as described above. One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 9, Couleaud teaches the computer-implemented method of claim 8. Couleaud further teaches wherein the foreground image element comprises an isolated foreground image (Fig. 2, Paragraph 0061 – “images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image”; Note: the image 236 is equivalent to the foreground image element, and it is isolated, as shown in Fig. 2; see screenshot of Fig. 2 above).
Regarding claim 10, Couleaud teaches the computer-implemented method of claim 9. Couleaud further teaches wherein the isolated foreground image is generated by an image segmentation model (Paragraph 0043, 0065, 0124 – “Any suitable number or types of techniques may be used to perform such segmentation, such as, for example: machine learning…based on segmentation 414, corresponding masks (e.g., 222 and 224 of FIG. 2) may be obtained for one or more of images 416, 418 and 420, respectively…a particular mask (e.g., mask 222 of FIG. 2) and/or other suitable information may be used to generate an object image (e.g., woman 236) without any empty regions or holes”; Note: the object image 236, which is equivalent to the isolated foreground image, is generated using masks based on image segmentation).
Regarding claim 14, Couleaud teaches the computer-implemented method of claim 8. Couleaud further teaches wherein the textual interface element is descriptive of the integrated image (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the text elements 718 describe the multi-layer/integrated image 714; see screenshot of Fig. 7 above).
Regarding claim 15, Couleaud teaches a non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by at least one processor, cause at least one device to perform operations comprising (Paragraph 0095-0096 – “Control circuitry 804 may be based on any suitable control circuitry such as processing circuitry 806. As referred to herein, control circuitry should be understood to mean circuitry based on one or more microprocessors… control circuitry 804 executes instructions for the image processing system…The image processing system or application may be implemented as software or a set of executable instructions. The instructions for performing any of the embodiments discussed herein of the image processing system or application may be encoded on non-transitory computer-readable media (e.g., a hard drive, random-access memory on a DRAM integrated circuit, read-only memory on a BLU-RAY disk, etc.). For example, in FIG. 3, the instructions may be stored in storage 808, and executed by control circuitry 804 of a device 800”):
receiving a request to generate an interface element package for a selected item (Paragraph 0038-0039 – “the image processing system may receive, as shown in FIG. 1, input of a text prompt 102. Such input may be received at a user interface of a computing device from a user in any suitable form…As a non-limiting illustrative example, as shown in FIG. 2, text prompt 102 may correspond to text prompt 202 of FIG. 2 of ‘a picture of a young woman holding a cat in front of a Victorian era building, 1870, high quality, soft focus, f/18, 60 mm, in the style of Auguste Renoir’”; Note: the prompt is equivalent to the request to generate an interface element package for a selected item, which in this case corresponds to a young woman, a cat, and/or a Victorian era building);
in response to receiving the request: generating a foreground image element (Fig. 2, Paragraph 0040, 0065, 0124 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… the image processing system may perform segmentation 114 of image 110 to obtain images 116, 118, and 120 (e.g., different identified portions, segments, and/or objects of image 210 of FIG. 2), which may respectively correspond to images 216, 218, and 220…based on segmentation 414, corresponding masks (e.g., 222 and 224 of FIG. 2) may be obtained for one or more of images 416, 418 and 420, respectively… a particular mask (e.g., mask 222 of FIG. 2) and/or other suitable information may be used to generate an object image (e.g., woman 236) without any empty regions or holes”; Note: a foreground image element 236 is generated, in response to an input prompt/request. It is shown in Fig. 2; see screenshot above);
generating a contextually appropriate background image (Fig. 2, Paragraph 0040, 0056 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)”; Note: image 234 is a background image that is contextually appropriate based on the input prompt. Screenshot of Fig. 2 shows the background image 234);
and generating an integrated image by integrating the foreground image element into the contextually appropriate background image at a contextually appropriate position, wherein the foreground image element is unmodified in the integrated image (Fig. 2 and 3, Paragraph 0040, 0061 – “Based on input prompt 102, model 108 may be configured to output image 110 (e.g., representing an interpretation of text input 202, as determined by model 108), which may correspond to first image 210… images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image, e.g., composite image 240, 242, 244, or 246”; Note: the multi-layer/composite image is equivalent to the integrated image since it combines the foreground image 236 and background image 234. Fig. 3 shows examples of integrated images, where the foreground element is at a contextually appropriate position (not floating in the sky) and its composition is unmodified).
In the above embodiment, the non-transitory computer readable medium of Couleaud does not directly teach generating at least one textual interface element and generating the interface element package including at least the integrated image and the textual interface element. However, in a different embodiment, Couleaud teaches
generating at least one textual interface element (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the text elements 718 describing the image layers in the GUI are equivalent to the generated textual interface elements; see screenshot of Fig. 7 above);
and generating the interface element package including at least the integrated image and the textual interface element (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the components of the GUI make-up the interface element package. It includes the multi-layer/integrated image and text elements 718 describing the image layers; see screenshot of Fig. 7 above).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud, to have obtained the above, and the results of the modification would have been obvious and predictable to one of ordinary skill in the art as of the effective filing date of the claimed invention. See MPEP $2143(A). The prior art included each element recited in claim 15, although not necessarily in a single embodiment, with the only difference being between the claimed element and the prior art being the lack of actual combination of certain elements in a single prior art embodiment, as described above. One of ordinary skill in the art could have combined the elements as claimed by known methods, and in that combination, each element merely performs the same function as it does separately. One of ordinary skill in the art would have also recognized that the results of the combination were predictable as of the effective filing date of the claimed invention.
Regarding claim 16, Couleaud teaches the non-transitory computer readable medium of claim 15. Couleaud further teaches wherein the foreground image element comprises an isolated foreground image generated by an image segmentation model (Paragraph 0043, 0065, 0124 – “Any suitable number or types of techniques may be used to perform such segmentation, such as, for example: machine learning…based on segmentation 414, corresponding masks (e.g., 222 and 224 of FIG. 2) may be obtained for one or more of images 416, 418 and 420, respectively…a particular mask (e.g., mask 222 of FIG. 2) and/or other suitable information may be used to generate an object image (e.g., woman 236) without any empty regions or holes”; Note: the image 236 is equivalent to the foreground image element, and it is isolated, as shown in Fig. 2; see screenshot of Fig. 2 above. The image 236 is generated using masks based on image segmentation).
Regarding claim 20, Couleaud teaches the non-transitory computer readable medium of claim 15. Couleaud further teaches wherein the textual interface element is descriptive of the integrated image (Fig. 7, Paragraph 0087 – “Based on receiving user selection of option 710, the image processing system may, using the techniques described herein, generate for display image 714 at output portion 712 of GUI 700. Image 714 may be a multi-layer image comprising one or more portions that correspond to input 708 (e.g., “A nice watch and a teacup on a desk with a bookshelf in the background”) and a foreground portion of image 714 may include image 704…In some embodiments, as shown at 718, each object or portion of image 714 may be listed”; Note: the text elements 718 describe the multi-layer/integrated image 714; see screenshot of Fig. 7 above).
Claims 4-5, 11-12, and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Couleaud in view of Saharia et al. (Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding), hereinafter Saharia.
Regarding claim 4, Couleaud teaches the system of claim 1. Couleaud further teaches wherein the contextually appropriate background image is generated by an image generation model (Paragraph 0056, 0070, 0073 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)…machine learning model 500 may be configured to receive text 502 as input and output image 508 corresponding to input text 502. In some embodiments, machine learning model 500 may correspond to machine learning model 108 of FIG. 1…machine learning model 500 may be implemented based at least in part on the techniques described in Saharia et al. “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv preprint arXiv:2205.11487 (2022), the contents of which is hereby incorporated by reference herein in its entirety”; Note: the background image, 234, is generated by machine learning model 500, which is implemented based on techniques from Saharia. The techniques from Saharia are done by a diffusion-based model, as discussed below). Couleaud does not directly teach that the image generation model is diffusion-based. However, Saharia teaches a diffusion-based image generation model (Paragraph 5 on Page 3 – “Imagen consists of a text encoder that maps text to a sequence of embeddings and a cascade of conditional diffusion models that map these embeddings to images of increasing resolutions”; Note: Imagen is a diffusion-based image generation model). Since Couleaud already teaches that the machine learning model that generates the background image is implemented by techniques from Saharia, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Saharia to have the image generation model be diffusion-based because diffusion models help create high-quality images.
Regarding claim 5, Couleaud in view of Saharia teaches the system of claim 4. Couleaud further teaches wherein the image generation model receives an image generation prompt describing a contextual background (Fig. 2, Paragraph 0056, 0070, 0073, 0119, 0123 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)…machine learning model 500 may be configured to receive text 502 as input and output image 508 corresponding to input text 502. In some embodiments, machine learning model 500 may correspond to machine learning model 108 of FIG. 1…machine learning model 500 may be implemented based at least in part on the techniques described in Saharia et al. “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv preprint arXiv:2205.11487 (2022), the contents of which is hereby incorporated by reference herein in its entirety”; Note: Fig. 2 shows how prompt 227 describes the background; see screenshot above. The prompt is input to the first trained machine learning model 500, which is the image generation model in this case, and it is based on the Saharia model) generated by a prompt generation model (Paragraph 0119, 0123 – “the control circuitry may generate, using a second trained machine learning model (e.g., machine learning model 126 of FIG. 1, which may correspond to machine learning model 510 of FIG. 5B), a plurality of textual descriptions (e.g., 127, 129, and 131 of FIG. 1 or 227, 229 and 231 of FIG. 2) respectively corresponding to the plurality of objects (e.g., 211, 213, and 215 of FIG. 2)…the control circuitry may generate, using the first trained machine learning model and based on the plurality of textual descriptions (e.g., 227, 229, 231 and/or 228, 230, 232) and the plurality of attributes (e.g., canny maps 223 and/or 225 of FIG. 2), a plurality of images (e.g., images 234, 236, and 238) respectively corresponding to the plurality of textual descriptions”; Note: the text prompt 227 is generated by the second trained machine learning model, which is equivalent to the prompt generation model). Couleaud does not teach that the image generation model is diffusion-based, from the limitation: “wherein the diffusion-based image generation model receives an image generation prompt describing a contextual background generated by a prompt generation model”. However, Saharia teaches wherein the diffusion-based image generation model receives an image generation prompt describing a contextual background (Paragraph 2 on Page 1, Paragraph 1 on Page 3 – “Imagen comprises a frozen T5-XXL [52] encoder to map input text into a sequence of embeddings and a 64×64 image diffusion model, followed by two super-resolution diffusion models for generating 256×256 and 1024×1024 images (see Fig. A.4). All diffusion models are conditioned on the text embedding sequence and use classifier-free guidance”; Note: the diffusion-based image generation model receives an input text (image generation prompt). An example of a prompt and generated image is shown in Fig. 1, where the background is described in the prompt; see screenshot below). Since Couleaud already teaches that the machine learning model that receives the prompt/text is implemented by techniques from Saharia, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Saharia to have the image generation model be diffusion-based because diffusion models help create high-quality images.
PNG
media_image4.png
536
454
media_image4.png
Greyscale
Screenshot of Fig. 1 (taken from Saharia)
Regarding claim 11, Couleaud teaches the computer-implemented method of claim 8. Couleaud further teaches wherein the contextually appropriate background image is generated by an image generation model (Paragraph 0056, 0070, 0073 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)…machine learning model 500 may be configured to receive text 502 as input and output image 508 corresponding to input text 502. In some embodiments, machine learning model 500 may correspond to machine learning model 108 of FIG. 1…machine learning model 500 may be implemented based at least in part on the techniques described in Saharia et al. “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv preprint arXiv:2205.11487 (2022), the contents of which is hereby incorporated by reference herein in its entirety”; Note: the background image, 234, is generated by machine learning model 500, which is implemented based on techniques from Saharia. The techniques from Saharia are done by a diffusion-based model, as discussed below). Couleaud does not directly teach that the image generation model is diffusion-based. However, Saharia teaches a diffusion-based image generation model (Paragraph 5 on Page 3 – “Imagen consists of a text encoder that maps text to a sequence of embeddings and a cascade of conditional diffusion models that map these embeddings to images of increasing resolutions”; Note: Imagen is a diffusion-based image generation model). Since Couleaud already teaches that the machine learning model that generates the background image is implemented by techniques from Saharia, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Saharia to have the image generation model be diffusion-based because diffusion models help create high-quality images.
Regarding claim 12, Couleaud in view of Saharia teaches the computer-implemented method of claim 11. Couleaud further teaches wherein the image generation model receives an image generation prompt describing a contextual background (Fig. 2, Paragraph 0056, 0070, 0073, 0119, 0123 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)…machine learning model 500 may be configured to receive text 502 as input and output image 508 corresponding to input text 502. In some embodiments, machine learning model 500 may correspond to machine learning model 108 of FIG. 1…machine learning model 500 may be implemented based at least in part on the techniques described in Saharia et al. “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv preprint arXiv:2205.11487 (2022), the contents of which is hereby incorporated by reference herein in its entirety”; Note: Fig. 2 shows how prompt 227 describes the background; see screenshot above. The prompt is input to the first trained machine learning model 500, which is the image generation model in this case, and it is based on the Saharia model) generated by a prompt generation model (Paragraph 0119, 0123 – “the control circuitry may generate, using a second trained machine learning model (e.g., machine learning model 126 of FIG. 1, which may correspond to machine learning model 510 of FIG. 5B), a plurality of textual descriptions (e.g., 127, 129, and 131 of FIG. 1 or 227, 229 and 231 of FIG. 2) respectively corresponding to the plurality of objects (e.g., 211, 213, and 215 of FIG. 2)…the control circuitry may generate, using the first trained machine learning model and based on the plurality of textual descriptions (e.g., 227, 229, 231 and/or 228, 230, 232) and the plurality of attributes (e.g., canny maps 223 and/or 225 of FIG. 2), a plurality of images (e.g., images 234, 236, and 238) respectively corresponding to the plurality of textual descriptions”; Note: the text prompt 227 is generated by the second trained machine learning model, which is equivalent to the prompt generation model). Couleaud does not teach that the image generation model is diffusion-based, from the limitation: “wherein the diffusion-based image generation model receives an image generation prompt describing a contextual background generated by a prompt generation model”. However, Saharia teaches wherein the diffusion-based image generation model receives an image generation prompt describing a contextual background (Paragraph 2 on Page 1, Paragraph 1 on Page 3 – “Imagen comprises a frozen T5-XXL [52] encoder to map input text into a sequence of embeddings and a 64×64 image diffusion model, followed by two super-resolution diffusion models for generating 256×256 and 1024×1024 images (see Fig. A.4). All diffusion models are conditioned on the text embedding sequence and use classifier-free guidance”; Note: the diffusion-based image generation model receives an input text (image generation prompt). An example of a prompt and generated image is shown in Fig. 1, where the background is described in the prompt; see screenshot above). Since Couleaud already teaches that the machine learning model that receives the prompt/text is implemented by techniques from Saharia, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Saharia to have the image generation model be diffusion-based because diffusion models help create high-quality images.
Regarding claim 17, Couleaud teaches the non-transitory computer readable medium of claim 15. Couleaud further teaches wherein the contextually appropriate background image is generated by an image generation model (Paragraph 0056, 0070, 0073 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)…machine learning model 500 may be configured to receive text 502 as input and output image 508 corresponding to input text 502. In some embodiments, machine learning model 500 may correspond to machine learning model 108 of FIG. 1…machine learning model 500 may be implemented based at least in part on the techniques described in Saharia et al. “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv preprint arXiv:2205.11487 (2022), the contents of which is hereby incorporated by reference herein in its entirety”; Note: the background image, 234, is generated by machine learning model 500, which is implemented based on techniques from Saharia. The techniques from Saharia are done by a diffusion-based model, as discussed below). Couleaud does not directly teach that the image generation model is diffusion-based. However, Saharia teaches a diffusion-based image generation model (Paragraph 5 on Page 3 – “Imagen consists of a text encoder that maps text to a sequence of embeddings and a cascade of conditional diffusion models that map these embeddings to images of increasing resolutions”; Note: Imagen is a diffusion-based image generation model). Since Couleaud already teaches that the machine learning model that generates the background image is implemented by techniques from Saharia, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Saharia to have the image generation model be diffusion-based because diffusion models help create high-quality images.
Regarding claim 18, Couleaud in view of Saharia teaches the non-transitory computer readable medium of claim 17. Couleaud further teaches wherein the image generation model receives an image generation prompt describing a contextual background (Fig. 2, Paragraph 0056, 0070, 0073, 0119, 0123 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)…machine learning model 500 may be configured to receive text 502 as input and output image 508 corresponding to input text 502. In some embodiments, machine learning model 500 may correspond to machine learning model 108 of FIG. 1…machine learning model 500 may be implemented based at least in part on the techniques described in Saharia et al. “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv preprint arXiv:2205.11487 (2022), the contents of which is hereby incorporated by reference herein in its entirety”; Note: Fig. 2 shows how prompt 227 describes the background; see screenshot above. The prompt is input to the first trained machine learning model 500, which is the image generation model in this case, and it is based on the Saharia model) generated by a prompt generation model (Paragraph 0119, 0123 – “the control circuitry may generate, using a second trained machine learning model (e.g., machine learning model 126 of FIG. 1, which may correspond to machine learning model 510 of FIG. 5B), a plurality of textual descriptions (e.g., 127, 129, and 131 of FIG. 1 or 227, 229 and 231 of FIG. 2) respectively corresponding to the plurality of objects (e.g., 211, 213, and 215 of FIG. 2)…the control circuitry may generate, using the first trained machine learning model and based on the plurality of textual descriptions (e.g., 227, 229, 231 and/or 228, 230, 232) and the plurality of attributes (e.g., canny maps 223 and/or 225 of FIG. 2), a plurality of images (e.g., images 234, 236, and 238) respectively corresponding to the plurality of textual descriptions”; Note: the text prompt 227 is generated by the second trained machine learning model, which is equivalent to the prompt generation model). Couleaud does not teach that the image generation model is diffusion-based, from the limitation: “wherein the diffusion-based image generation model receives an image generation prompt describing a contextual background generated by a prompt generation model”. However, Saharia teaches wherein the diffusion-based image generation model receives an image generation prompt describing a contextual background (Paragraph 2 on Page 1, Paragraph 1 on Page 3 – “Imagen comprises a frozen T5-XXL [52] encoder to map input text into a sequence of embeddings and a 64×64 image diffusion model, followed by two super-resolution diffusion models for generating 256×256 and 1024×1024 images (see Fig. A.4). All diffusion models are conditioned on the text embedding sequence and use classifier-free guidance”; Note: the diffusion-based image generation model receives an input text (image generation prompt). An example of a prompt and generated image is shown in Fig. 1, where the background is described in the prompt; see screenshot above). Since Couleaud already teaches that the machine learning model that receives the prompt/text is implemented by techniques from Saharia, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Saharia to have the image generation model be diffusion-based because diffusion models help create high-quality images.
Claims 6, 13, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Couleaud in view of Reddy et al. (US 20250225609 A1), hereinafter Reddy.
Regarding claim 6, Couleaud teaches the system of claim 1. Couleaud further teaches wherein the instructions cause the processor to generate the contextually appropriate background image by: generating an initial background image (Fig. 2, Paragraph 0056 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)”; Note: image 234 is an initial background image Screenshot of Fig. 2 above shows the background image 234); generating a composite image by overlaying the foreground image element on the initial background image (Fig. 2, Paragraph 0061 – “images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image, e.g., composite image 240, 242, 244, or 246, which may comprise suitable transparency for such a multi-layer image”; Note: Fig. 2 shows how the foreground image 236 is overlaid on top of the initial background image 234). Couleaud does not teach generating the contextually appropriate background image by regenerating the initial background image in view of a position of the foreground image element in the composite image. However, Reddy teaches generating the contextually appropriate background image by regenerating the initial background image in view of a position of the foreground image element in the composite image (Paragraph 0058 – “In some examples, the foreground object 126 positioned on the re-dimensioned background 130 does not appear cohesive with the remainder of the re-dimensioned background 130. For instance, the foreground object 126 has different lighting, contrast, shadows, brightness or other factors from the re-dimensioned background 130 that contribute to an unaesthetic appearance of the re-dimensioned digital image 118 overall. To correct this, the blending module 214 adjusts visual properties of the foreground object 126, the re-dimensioned background 130, or both using a machine learning model trained on adjusting properties of foreground objects and backgrounds of digital images to generate visually cohesive images”; Note: the background is regenerated to match the foreground image based on the foreground element position (shadows, etc. depend on position of the object)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Reddy to regenerate the initial background image in view of a position of the foreground image element in the composite image for the benefit of generating “visually cohesive images” (Reddy: Paragraph 0058), wherein the background and foreground match, and together appear more realistic and aesthetic.
Regarding claim 13, Couleaud teaches the computer-implemented method of claim 8. Couleaud further teaches wherein generating the contextually appropriate background image comprises: generating an initial background image (Fig. 2, Paragraph 0056 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)”; Note: image 234 is an initial background image Screenshot of Fig. 2 above shows the background image 234); generating a composite image by overlaying the foreground image element on the initial background image (Fig. 2, Paragraph 0061 – “images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image, e.g., composite image 240, 242, 244, or 246, which may comprise suitable transparency for such a multi-layer image”; Note: Fig. 2 shows how the foreground image 236 is overlaid on top of the initial background image 234). Couleaud does not teach generating the contextually appropriate background image by regenerating the initial background image in view of a position of the foreground image element in the composite image. However, Reddy teaches generating the contextually appropriate background image by regenerating the initial background image in view of a position of the foreground image element in the composite image (Paragraph 0058 – “In some examples, the foreground object 126 positioned on the re-dimensioned background 130 does not appear cohesive with the remainder of the re-dimensioned background 130. For instance, the foreground object 126 has different lighting, contrast, shadows, brightness or other factors from the re-dimensioned background 130 that contribute to an unaesthetic appearance of the re-dimensioned digital image 118 overall. To correct this, the blending module 214 adjusts visual properties of the foreground object 126, the re-dimensioned background 130, or both using a machine learning model trained on adjusting properties of foreground objects and backgrounds of digital images to generate visually cohesive images”; Note: the background is regenerated to match the foreground image based on the foreground element position (shadows, etc. depend on position of the object)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Reddy to regenerate the initial background image in view of a position of the foreground image element in the composite image for the benefit of generating “visually cohesive images” (Reddy: Paragraph 0058), wherein the background and foreground match, and together appear more realistic and aesthetic.
Regarding claim 19, Couleaud teaches the non-transitory computer readable medium of claim 15. Couleaud further teaches wherein generating the contextually appropriate background image comprises: generating an initial background image (Fig. 2, Paragraph 0056 – “textual prompts 128, 130, and 132 (and/or textual prompts 127, 129 and 131) may be input to trained machine learning model 108, and based on such input, trained machine learning model 108 may be configured to generate and output a second plurality of images 134, 136, and 138 (which may correspond to images 234, 236, and 238, respectively, of FIG. 2)”; Note: image 234 is an initial background image Screenshot of Fig. 2 above shows the background image 234); generating a composite image by overlaying the foreground image element on the initial background image (Fig. 2, Paragraph 0061 – “images 234, 236 and 238 may correspond to respective layers of a multi-layer image, and may be composited together to form a multi-layer image, e.g., composite image 240, 242, 244, or 246, which may comprise suitable transparency for such a multi-layer image”; Note: Fig. 2 shows how the foreground image 236 is overlaid on top of the initial background image 234). Couleaud does not teach generating the contextually appropriate background image by regenerating the initial background image in view of a position of the foreground image element in the composite image. However, Reddy teaches generating the contextually appropriate background image by regenerating the initial background image in view of a position of the foreground image element in the composite image (Paragraph 0058 – “In some examples, the foreground object 126 positioned on the re-dimensioned background 130 does not appear cohesive with the remainder of the re-dimensioned background 130. For instance, the foreground object 126 has different lighting, contrast, shadows, brightness or other factors from the re-dimensioned background 130 that contribute to an unaesthetic appearance of the re-dimensioned digital image 118 overall. To correct this, the blending module 214 adjusts visual properties of the foreground object 126, the re-dimensioned background 130, or both using a machine learning model trained on adjusting properties of foreground objects and backgrounds of digital images to generate visually cohesive images”; Note: the background is regenerated to match the foreground image based on the foreground element position (shadows, etc. depend on position of the object)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Couleaud to incorporate the teachings of Reddy to regenerate the initial background image in view of a position of the foreground image element in the composite image for the benefit of generating “visually cohesive images” (Reddy: Paragraph 0058), wherein the background and foreground match, and together appear more realistic and aesthetic.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Han et al. (US 12579703 B1) teaches a method of generating a new background for an image based on a prompt. Clever et al. (US 20250061583 A1) teaches a method of generating synthetic images with realistic augmentations based on foreground image data and input data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE HAU MA whose telephone number is (571)272-2187. The examiner can normally be reached M-Th 7-5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at (571) 270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHELLE HAU MA/ Examiner, Art Unit 2617
/KING Y POON/Supervisory Patent Examiner, Art Unit 2617