DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to because:
FIG. 2E has reference numeral 207 but in paragraph [0086] at line 9 is reference numeral 206 which is not in FIG. 2E, changing reference numeral 206 to 207 at line 9 would resolve this drawing issue;
FIG. 2F has reference numeral 265 but in paragraph [0087] 265 is not present but in paragraph [0087] at line 15 is reference numeral 264 which is not in FIG. 2F, changing reference numeral 264 to 265 at line 15 would resolve this drawing issue; and
paragraph [0091] at lines 3 and 8 refers to RoPE positional embeddings 238, however, FIG. 2G does not have RoPE positional embeddings 238 refer to FIG. 2C which does have RoPE positional embeddings 238.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities: In paragraph [0033] on page 7 line 2 “an graphic” to a graphic.
Appropriate correction is required.
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
CLAIM INTERPRETATION
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
Claims 1-20 have been interpreted under 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) to not invoke 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) claim interpretation.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 11-15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
These claims are unclear and ambiguous as to the actor of each of the claimed steps. This ambiguity fails to put one of ordinary skill in the art on notice the metes and bounds of the method claims. Refer to MPEP 2173.02 Determining Whether Claim Language is Definite [R-01.2024]; MPEP 2173.05(g) Functional Limitations [R-07.2022]; MPEP 2173.06 Practice Compact Prosecution [R-07.2022] II. PRIOR ART REJECTION OF CLAIM REJECTED AS INDEFINITE; and MPEP 2143.03 All Claim Limitations Must Be Considered [R-01.2024] I. INDEFINITE LIMITATIONS MUST BE CONSIDERED II. LIMITATIONS WHICH DO NOT FIND SUPPORT IN THE ORIGINAL SPECIFICATION MUST BE CONSIDERED .
Data processing system claim 1 sets forth an actor for the claimed steps as a “processor alone or in combination with other processors”. Non-transitory computer readable medium claim 16 sets forth an actor for the claimed functions as a programmable device. Paragraph [0003] of Applicant’s specification describes “An example method implemented in a data processing system includes”, paragraph [0002] describes “An example data processing system according to the disclosure includes a processor and a machine-readable medium storing executable instructions.”, paragraph [0004] describes “An example non-transitory computer readable medium data processing system according to the disclosure on which are stored instructions that, when executed, cause a programmable device to perform functions of”, and paragraph [0131] describes “In some examples, at least some of the operations of a method may be performed by one or more processors or processor-implemented modules. Moreover, the one or more processors may also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS). For example, at least some of the operations may be performed by, and/or among, multiple computers (as examples of machines including processors), with these operations being accessible via a network (for example, the Internet) and/or via one or more software interfaces (for example, an application program interface (API)). The performance of certain of the operations may be distributed among the processors, not only residing within a single machine, but deployed across several machines. Processors or processor-implemented modules may be in a single geographic location (for example, within a home or office environment, or a server farm), or may be distributed across multiple geographic locations.”. Also refer to paragraphs [0127]-[0130]. Method claims 11-15 could be amended to claim an actor for each of the steps such as “one or more processors” or “processor-implemented modules”.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 2, 4, 10, 11, 12, 14, 16, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over the Pu et al. article cited on both IDS statements filed on 03/11/2025 and 03/14/2025, hereinafter Pu, in view of Zhang et al., US Patent Application Publication No. 2023/0325996, hereinafter Zhang.
The Pu et al. article has six authors in common with the inventors of this application, refer to MPEP 2153.01(a) Grace Period Inventor-Originated Disclosure Exception [R-01.2024] which states “If, however, the application names fewer joint inventors than a publication (e.g., the application names as joint inventors A and B, and the publication names as authors A, B and C), it would not be readily apparent from the publication that it is an inventor-originated disclosure and the publication would be treated as prior art under AIA 35 U.S.C. 102(a)(1) unless there is evidence of record that an exception under AIA 35 U.S.C. 102(b)(1) applies.”. Currently no evidence is present. Also refer the following ways to overcome the Pu article:
2155 Use of Affidavits or Declarations Under 37 CFR 1.130 To Overcome Prior Art Rejections [R-07.2022]
2155.01 Showing That the Disclosure Was Made by the Inventor or a Joint Inventor [R-07.2022]
A detailed analysis of the claims follows.
Claim 1:
1. A data processing system comprising:
a processor (Pu: For example refer to the computer processor implementable code located on pages 11-12 of the Supplementary Material, inherently a processor is conveyed by Pu.), and
a machine-readable storage medium storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform operations (Pu: For example refer to the computer processor implementable code located on pages 11-12 of the Supplementary Material conveyed to the reader of Pu on a machine-readable storage medium, eg printed on paper, displayed on a screen, etc.):
receiving a text prompt to create a multilayer graphic design (Pu: Figure 1 which has text prompt: “GlobalPrompt: A stark top-down view of a dessert plate hold a tart slice with whipped cream, caramel drizzle, and a silver spoon.”.);
predicting, by a first generative model based on the text prompt, a layout including a plurality of anonymous regions each defined by a bounding box without content or region-wise prompt annotations (Pu: Figure 1, “The Anonymous Region Layout Planner predicts a set of anonymous bounding boxes based on the user-provided text prompt.”, refer to page 2 second column, and section 3.3.);
concurrently generating, by a diffusion transformer, multilayer image latents of a global reference image, a background layer, and a plurality of transparent foreground layers using a Gaussian noise conditioned on the layout and the text prompt, each of the transparent foreground layers corresponding to one of the anonymous regions (Pu: “The Multi-layer Transparent Autoencoder encodes and decodes a variable number of transparent layers at different resolutions using a sequence of latent visual tokens.”, refer to page 2 second column, section 3.1, Figure 4(a), and section 5.);
decoding, by a vision transformer, the multilayer image latents into the global reference image, the background layer, and the plurality of transparent foreground layers as the multilayer graphic design (Pu: “The Anonymous Region Transformer concurrently generates a global reference image, a background image, and multiple cropped transparent foreground layers from Gaussian noise conditioned on the anonymous region layout.”, refer to page 2 second column, section 3.2, Figure 4(b), and section 5.);
composing an output based on the background layer and the plurality of transparent foreground layers (Pu: Figure 2, Tables 7-15 and Figure 6 of the Supplementary Material.);
providing the output to a client device (Pu: silent.); and
causing a user interface of the client device to display the output (Pu: silent.).
Pu is silent regarding the claimed:
providing the output to a client device; and
causing a user interface of the client device to display the output.
Zhang describes compositing by inserting foreground objects into a background image by the use of auto-composite comprising at least one of a scale prediction model, a harmonization model, or a shadow generation model. Zhang describes a server providing an output based on composition of the background layer and the plurality of transparent foreground layers to a client device and causing a user interface of the client device to display the output, refer to FIGs. 1, 2A, 12 paragraphs [0052]-[0055], [0059], [0071], and [0177]-[0182].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to perform processor implemented operations providing an output based on composition of the background layer and the plurality of transparent foreground layers to a client device and causing a user interface of the client device to display the output causing a user interface of the client device to display the output because this will offer to the user flexible composite features for efficiently generating realistic composite images based on a reduced set of user interactions, refer to paragraph [0003] of Zhang.
Claim 2:
2. The data processing system of claim I, wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
encoding relative position information associated with the multilayer image latents based on the layout as multilayer rotary position embeddings (Pu: Rotary Position Embedding (RoPE), refer to page 5 first column.),
wherein the multilayer image latents are generated with the multilayer rotary position embeddings (Pu: Rotary Position Embedding (RoPE), refer to page 5 first column.).
Claim 4:
4. The data processing system of claim 1, wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
generating text embeddings based on the text prompt (Pu: page 5 column 1 “MMDiT is an improved variant of DiT framework [15] that uses two different sets of model weights to process text tokens and image tokens separately.”.); and
providing the text embeddings to the diffusion transformer as text tokens for generating the multilayer image latents (Pu: page 4 column 2 to page 5 column 1 “We choose the latest multimodal diffusion
transformer (MMDiT), e.g., FLUX.1[dev] [32], to build our variable multi-layer image generation model, ART.
MMDiT is an improved variant of DiT framework [15] that uses two different sets of model weights to process text tokens and image tokens separately.”.).
Claim 10:
10. The data processing system of claim 1, wherein
the first generative model is a large language model (Pu: section 3.3 “We propose an anonymous region layout planner, which predicts a set of bounding boxes based on the text input. This planner is implemented by fine-tuning an LLM model
on our layout dataset, specifically using the pre-trained LLaMa-3.1-8B [14].”.), and
the diffusion transformer is a multimodal diffusion transformer (Pu: section 3. “Our approach enables diffusion transformer based models to jointly generate images with multiple transparent layers conditioned on an anonymous region layout
provided by the user or predicted by an LLM. The entire framework consists of three key components: the Multi-layer Transparent Autoencoder (Section 3.1), which jointly encodes and decodes multi-layer images and their corresponding latent representations;”.).
Claims 11, 12, and 14:
Claims 11, 12, and 14 are method claim versions of data processing system claims 1, 2, and 4 and method claims 11, 12, and 14 are rejected for the same reasons given for data processing system claims 1, 2, and 4.
Claims 16, 17, and 19:
Claims 16, 17, and 19 are non-transitory computer readable medium claim versions of data processing system claims 1, 2, and 4 and non-transitory computer readable medium claims 16, 17, and 19 are rejected for the same reasons given for data processing system claim 1, 2, and 4.
Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Leroux, US Patent Application Publication No. 2026/0099968, describes compositing by inserting foreground objects into a background image by the use of text-to-image models that provide prompts to the user during the compositing process.
Yifan Pu et al., ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation, Published: 31 Dec 2024, Last Modified: 15 Oct 2025, Computer Vision Foundation CVF OpenReview.net,
pp. 7952-7962, corresponds to Pu et al. cited by on the IDS statements filed on 03/11/2025 and 03/14/2025 and was published on December 31, 2024 and modified on October 15, 2025.
Allowable Subject Matter
Claims 3, 5-9, 18, and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 13 and 15 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Claim 3:
The prior art of record fails to teach or suggest in the context of parent claim 1 the furthering limitations present in claim 3:
3. The data processing system of claim 1, wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
generating timestep embeddings; and
providing the timestep embeddings to the diffusion transformer as temporal context when generating the multilayer image latents (emphasis added).
Claims 5 and 6:
The prior art of record fails to teach or suggest in the context of parent claim 1 the furthering limitations present in claim 5:
5. The data processing system of claim 1, wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
retrieving training data including a training layout, a training global reference image, a training background layer, and a plurality of training transparent foreground layers;
converting the training background layer into a gray background layer;
converting the plurality of training transparent foreground layers into a plurality of training foreground layers padded with a gray-background;
jointly generating, by a variational autoencoder (VAE) encoder, multilayer image latents of the training global reference image, the gray background layer, and the plurality of training foreground layers;
applying a ceiling-aligned tight crop on the image latents of each of the plurality of training foreground layers;
flattening and concatenating the image latents of the training global reference image, the training background layer, and the cropped image latents of the plurality of training foreground layers, into one sequence of image latents; and
feeding the sequence of image latents into a pre-trained vision transformer to be trained into the vision transformer (emphasis added).
Claim 7:
The prior art of record fails to teach or suggest in the context of parent claim 1 the furthering limitations present in claim 7:
7. The data processing system of claim 1, wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
providing the plurality of transparent foreground layers to the client device; and
causing the user interface of the client device to display the plurality of transparent foreground layers (emphasis added).
Claim 13:
The prior art of record fails to teach or suggest in the context of parent claim 11 the furthering limitations present in claim 13:
13. The method of claim 11, further comprising:
generating timestep embeddings; and
providing the timestep embeddings to the diffusion transformer as temporal context when generating the multilayer image latents (emphasis added).
Claim 15:
The prior art of record fails to teach or suggest in the context of parent claim 1 the furthering limitations present in claim 15:
15. The method of claim 11, further comprising:
retrieving training data including a training layout, a training global reference image, a training background layer, and a plurality of training transparent foreground layers;
converting the training background layer into a gray background layer;
converting the plurality of training transparent foreground layers into a plurality of training foreground layers padded with a gray-background;
jointly generating, by a variational autoencoder (VAE) encoder, multilayer image latents of the training global reference image, the gray background layer, and the plurality of training foreground layers based on Gaussian noise conditioned on the training layout;
applying a ceiling-aligned tight crop on the image latents of each of the plurality of training foreground layers;
flattening and concatenating the image latents of the training global reference image, the training background layer, and the cropped image latents of the plurality of training foreground layers, into one sequence of image tokens; and
feeding the sequence of image tokens into a pre-trained vision transformer to be trained into the vision transformer (emphasis added).
Claim 18:
The prior art of record fails to teach or suggest in the context of parent claim 16 the furthering limitations present in claim 18:
18. The non-transitory computer readable medium of claim 16, wherein the instructions when executed, further cause the programmable device to perform:
generating timestep embeddings; and
providing the timestep embeddings to the diffusion transformer as temporal context when generating the multilayer image latents (emphasis added).
Claim 20:
The prior art of record fails to teach or suggest in the context of parent claim 16 the furthering limitation present in claim 20:
20. The non-transitory computer readable medium of claim 16, wherein the instructions when executed, further cause the programmable device to perform:
retrieving training data including a training layout, a training global reference image, a training background layer, and a plurality of training transparent foreground layers;
converting the training background layer into a gray background layer;
converting the plurality of training transparent foreground layers into a plurality of training foreground layers padded with a gray-background;
jointly generating, by a variational autoencoder (VAE) encoder, multilayer image latents of the training global reference image, the gray background layer, and the plurality of training foreground layers based on Gaussian noise conditioned on the training layout;
applying a ceiling-aligned tight crop on the image latents of each of the plurality of training foreground layers;
flattening and concatenating the image latents of the training global reference image, the training background layer, and the cropped image latents of the plurality of training foreground layers, into one sequence of image tokens; and
feeding the sequence of image tokens into a pre-trained vision transformer to be trained into the vision transformer (emphasis added).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JEFFERY A BRIER whose telephone number is (571)272-7656. The examiner can normally be reached on Mon-Fri from 8:30am-3:00pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao M Wu, can be reached at telephone number 571-272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from Patent Center. Status information for published applications may be obtained from Patent Center. Status information for unpublished applications is available through Patent Center for authorized users only. Should you have questions about access to Patent Center, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form.
JEFFERY A. BRIER
Primary Examiner
Art Unit 2613
/JEFFERY A BRIER/Primary Examiner, Art Unit 2613