Prosecution Insights
Last updated: August 17, 2026
Application No. 19/003,906

Diffusion Models for Multi-Garment Virtual Try-On or Editing

Non-Final OA §103
Filed
Dec 27, 2024
Priority
Dec 29, 2023 — provisional 63/616,294
Examiner
BASHIR, ADEEL
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
89%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
39 granted / 44 resolved
+28.6% vs TC avg
Minimal +4% lift
Without
With
+4.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
13 currently pending
Career history
53
Total Applications
across all art units

Statute-Specific Performance

§101
4.2%
-35.8% vs TC avg
§103
88.4%
+48.4% vs TC avg
§102
5.3%
-34.7% vs TC avg
§112
2.1%
-37.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 44 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Priority Acknowledgment is made of applicant’s priority claim, for U.S. Application No. 19/003,906, to a U.S. Provisional Application filed on 12/29/2023. Status of Claims Claims 1–20 are pending in the application.Claims 1-7, 9-18, 20 are rejected. Claims 8, 19 are objected to. Claim 20 is objected to under 37 C.F.R. § 1.75 as being a substantial duplicate of claim 9. Allowable Subject Matter Claims 8, 19 are objected to as being dependent upon a rejected base claim(s), but would be allowable if rewritten in independent form including all of the limitations of the base claim(s) and any intervening claim(s). Overview of Grounds of Rejection Ground of Rejection Claim(s) Statute(s) Reference(s) Ground 1 1, 2, 6, 12, 13, 17 § 103 Lee et al. (US20220318892A1) and Morelli et al. (NPL) Ground 2 3, 14 § 103 Lee et al. (US20220318892A1) and Morelli et al. (NPL), further in view of Wu et al. (NPL) Ground 3 4, 15 § 103 Lee et al. (US20220318892A1), Morelli et al. (NPL), and Wu et al. (NPL), further in view of Chia et al. (NPL) Ground 4 5, 16 § 103 Lee et al. (US20220318892A1) and Morelli et al. (NPL), further in view of Polanía et al. (NPL) Ground 5 7, 18 § 103 Lee et al. (US20220318892A1) and Morelli et al. (NPL), further in view of Gal et al. (NPL) Ground 6 9, 20 § 103 Lee et al. (US20220318892A1) and Morelli et al. (NPL), further in view of Li et al. (NPL) Ground 7 10 § 103 Lee et al. (US20220318892A1) and Morelli et al. (NPL), further in view of Gu et al. (NPL) Ground 8 11 § 103 Morelli et al. (NPL) and Gu et al. (NPL) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. (Please see the cited paragraphs, sections, pages, or surrounding text in the references for the paraphrased content.) Ground of Rejection 1 Claims 1, 2, 6, 12, 13, 17 are rejected under 35 U.S.C. § 103 as being unpatentable over Lee et al. (US20220318892A1) in view of Morelli et al. (NPL). As per Claim 1, Lee teaches the following portion of Claim 1, which recites:“A computer-implemented method for multi-garment try-on, the method comprising:” Lee et al. teaches a virtual try-on method “executed by at least one processor of a computing device” and directed to “naturally fitting a plurality of clothes on a model object” using deep learning. Lee et al. (US20220318892A1), ¶¶ [0015], [0017]. Thus, Lee et al. teaches the claimed computer-implemented method for multi-garment try-on. Lee teaches the following portion of Claim 1, which recites: “obtaining, by a computing system comprising one or more computing devices, an input set comprising a person image that depicts a person, a first garment image that depicts a first garment, and a second garment image that depicts a second garment;” Lee et al. teaches a system including “one or more computing devices 100 (100-1, 100-2)” and receiving “a top garment image C1, a bottom garment image C2, and a model image PR.” Lee et al. (US20220318892A1), ¶¶ [0045], [0096]. The model image PR corresponds to the claimed person image, while the top garment image C1 and bottom garment image C2 correspond to the claimed first and second garment images. Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), they collectively teach all of the limitation(s). Lee and Morelli teach the following portion of Claim 1, which recites:“processing, by the computing system, the input set with a machine-learned denoising diffusion model to generate, as an output of the machine-learned denoising diffusion model, a synthetic image that depicts the person wearing the first garment and the second garment;” Lee et al. teaches that when “a first clothes image representing first clothes, a second clothes image representing second clothes, and a model image representing a model object are input,” the system provides “a virtual fitting image representing a model object fitting the first clothes and the second clothes.” Lee et al. (US20220318892A1), ¶ [0067]. Lee et al. does not teach using a denoising diffusion model for this generation. Morelli et al. teaches Stable Diffusion comprising “a text time-conditional U-Net denoising model,” whose “training of the denoising network” is performed by minimizing a loss function. Morelli et al. (NPL), Section 3.1, page 3. Morelli et al. further teaches conditioning the “Stable Diffusion denoising network” to “obtain the final image” depicting the model wearing the input garment. Morelli et al. (NPL), Section 3.2, page 4. Thus, Lee et al. teaches processing the three-image input to produce the multi-garment synthetic image, while Morelli et al. teaches performing virtual try-on generation using a machine-learned denoising diffusion model. Lee eaches the following portion of Claim 1, which recites: “and providing, by the computing system, the synthetic image as an output.” Lee et al. teaches generating the virtual fitting image and “output[ting] the generated virtual fitting image through the display 105.” Lee et al. (US20220318892A1), ¶ [0068]. This teaches providing the synthetic image as an output. Before the effective filing date of the claimed invention, a person of ordinary skill in the art would have been motivated to employ Morelli et al.’s trained latent-diffusion denoising network as the image-generation mechanism in Lee et al.’s multi-garment virtual try-on system. Lee et al. recognizes that conventional virtual fitting images may be blurred, distorted, and artificial in appearance, while Morelli et al. teaches that diffusion models improve image realism and preserve garment texture and model characteristics. Lee et al. (US20220318892A1), ¶ [0010]; Morelli et al. (NPL), Section 1, page 2. The combination would have predictably enhanced Lee et al.’s multi-garment output using a known diffusion-based virtual try-on technique. PNG media_image1.png 9 307 media_image1.png Greyscale As per Claim 2, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), they collectively teach all of the limitation(s). Morelli teaches Claim 2, which recites: “The computer-implemented method of claim 1, wherein the denoising diffusion model comprises a single-stage denoising diffusion model.” Morelli et al. (NPL) teaches that Stable Diffusion “consists of . . . a text time-conditional U-Net denoising model 𝜖θ.” Figure 2 depicts one “Denoising UNet 𝜖θ” receiving the conditioning inputs and producing, through decoder D, the final try-on image. Morelli et al. (NPL), Section 3.1, page 3; Figure 2 and Section 3.2, page 4. Thus, Morelli et al. teaches a single denoising diffusion stage, rather than a cascade of multiple diffusion models. The term “single-stage” is not used verbatim, but the one-denoising-UNet architecture shown in Figure 2 provides a reasonably strong mapping. The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim. PNG media_image1.png 9 307 media_image1.png Greyscale As per Claim 6, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), they collectively teach all of the limitation(s). Morelli teaches Claim 6, which recites:“The computer-implemented method of claim 1, wherein the machine-learned denoising diffusion model comprises a person encoder configured to generate a person encoding from the person image, a U-Net encoder, and a U-Net decoder.” Morelli et al. (NPL) teaches that the target-model image I is processed by an encoder E that “compresses an image I into a lower-dimensional latent space,” thereby generating the claimed person encoding. Morelli further teaches a “text time-conditional U-Net denoising model”; Figure 2 depicts its contracting and expanding paths, corresponding respectively to the claimed U-Net encoder and U-Net decoder. Morelli et al. (NPL), Section 3.1, page 3; Figure 2, page 4. The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 12 does not include any additional limitations that would significantly distinguish it from claim 1. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 13 does not include any additional limitations that would significantly distinguish it from claim 2. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 17 does not include any additional limitations that would significantly distinguish it from claim 6. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Ground of Rejection 2 Claims 3, 14 are rejected under 35 U.S.C. § 103 as being unpatentable over Lee et al. (US20220318892A1) in view of Morelli et al. (NPL), and further in view of Wu et al. (NPL). As per Claim 3, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL) and Wu et (NPL), they collectively teach all of the limitation(s). Wu teaches Claim 3, which recites:“The computer-implemented method of claim 1, wherein the input set further comprises a textual layout description.” Wu et al. (NPL) teaches a text input describing the relative layout of multiple objects, such as “A red car is to the left of a black mailbox,” and states that “[t]he relative spatial positions of the object should match the description in D.” Wu further teaches that “the layout predictor is a transformer that takes the text description D as the input” and predicts the pixel region for each object. Wu et al. (NPL), Section 3.1 and Figure 2, page 3; Section 3.3, page 4. Thus, Wu et al. provides a strong mapping for an input set that further includes a textual layout description specifying the relative placement of depicted items. Applied to the multi-garment system of Claim 1, the text may similarly specify the spatial arrangement or layering of the garments. Before the effective filing date of the claimed invention, a POSITA would have been motivated to add Wu et al.’s textual layout description to the virtual try-on system of Lee et al. and Morelli et al. to improve control over the relative placement of multiple garments in the generated image, thereby predictably enhancing spatial fidelity and reducing garment misplacement. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 14 does not include any additional limitations that would significantly distinguish it from claim 3. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Ground of Rejection 3 Claims 4, 15 are rejected under 35 U.S.C. § 103 as being unpatentable over Lee et al. (US20220318892A1) in view of Morelli et al. (NPL), further in view of Wu et al. (NPL), and still further in view of Chia et al. (NPL). As per Claim 4, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), Wu et (NPL), and Chia et al. (NPL), they collectively teach all of the limitation(s). Wu and Chia teach Claim 4, which recites:“The computer-implemented method of claim 3, further comprising: processing, by the computing system, the textual layout description with a text embedding model to generate a text embedding, wherein the text embedding model has been finetuned on training data comprising clothing descriptions.” Wu et al. (NPL) provides the textual layout description of Claim 3. Chia et al. (NPL) teaches FashionCLIP comprising a text encoder and states that training produces, for example, a “textual embedding for the string ‘red long dress.’” Chia et al. further teaches fine-tuning a pre-trained CLIP model using Farfetch fashion-product data containing textual “highlight” terms and “short description[s]”, with a training set of 700,000 products. Chia et al. (NPL), Sections 4.1 and 4.3, pages 4-5; Appendix A, page 10. Thus, Chia et al. robustly teaches processing text with a text embedding model to generate a text embedding, where the model is finetuned on training data comprising clothing descriptions. Before the effective filing date, a POSITA would have been motivated to use Chia et al.’s fashion-finetuned text encoder to process Wu et al.’s textual layout description, thereby improving recognition of clothing-specific terms and producing more accurate text conditioning for the virtual try-on model, with predictable improvements in garment placement and fidelity. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 15 does not include any additional limitations that would significantly distinguish it from claim 4. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Ground of Rejection 4 Claims 5, 16 are rejected under 35 U.S.C. § 103 as being unpatentable over Lee et al. (US20220318892A1) in view of Morelli et al. (NPL), and further in view of Polanía et al. (NPL). As per Claim 5, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), and Polanía, they collectively teach all of the limitation(s). Morelli and Polanía teach Claim 5, which recites:“The computer-implemented method of claim 1, wherein the machine-learned denoising diffusion model comprises a first garment encoder configured to generate a first garment embedding from the first garment image and a second garment encoder configured to generate a second garment embedding form the second garment image.” Morelli et al. (NPL) teaches encoding an input garment using a “CLIP visual encoder V E ” and mapping the extracted visual features into pseudo-word token embeddings that condition the Stable Diffusion denoising network. Morelli et al. (NPL), Section 3.2 and Figure 2, page 4. Polanía et al. (NPL) teaches that a “siamese sub-network maps a pair of apparel images to a pair of features” and comprises “two branches with shared weights.” Figure 1 shows separate garment-image inputs processed through separate truncated CNN branches, producing the “embeddings generated by the siamese sub-network.” Polanía et al. (NPL), Section 2 and Figure 1, page 2. Thus, applying Morelli et al.’s garment-encoding path to Polanía et al.’s two apparel-image branches teaches a first garment encoder and first garment embedding and a second garment encoder and second garment embedding. Before the effective filing date, a POSITA would have been motivated to incorporate Polanía et al.’s two-branch apparel-image encoding architecture into Morelli et al.’s diffusion-based virtual try-on model so that each garment image is independently encoded into a corresponding garment embedding, thereby improving multi-garment conditioning and predictably preserving the features of each garment. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 16 does not include any additional limitations that would significantly distinguish it from claim 5. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Ground of Rejection 5 Claims 7, 18 are rejected under 35 U.S.C. § 103 as being unpatentable over Lee et al. (US20220318892A1) in view of Morelli et al. (NPL), and further in view of Gal et al. (NPL). As per Claim 7, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), and Gal et (NPL), they collectively teach all of the limitation(s). Gal teaches Claim 7, which recites:“The computer-implemented method of claim 6, wherein only the person encoding has been finetuned.” Gal et al. (NPL) teaches learning a concept-specific embedding v * from images depicting the target concept, including varied poses, by “direct optimization” while “keeping both c θ and ϵ θ fixed.” Gal et al. (NPL), Section 3, page 5. When the target concept is the person of Claim 6, v * constitutes the person encoding, and only that encoding is fine-tuned while the text encoder and denoising network remain fixed. Before the effective filing date, a POSITA would have been motivated to apply Gal et al.’s textual-inversion technique to fine-tune only the person encoding while keeping the remaining diffusion model fixed, thereby personalizing the generated person while preserving the model’s learned capabilities and predictably reducing retraining cost and catastrophic forgetting. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 18 does not include any additional limitations that would significantly distinguish it from claim 7. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Ground of Rejection 6 Claims 9, 20 are rejected under 35 U.S.C. § 103 as being unpatentable over Lee et al. (US20220318892A1) in view of Morelli et al. (NPL), and further in view of Li et al. (NPL). As per Claim 9, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), and Li et al. (NPL), they collectively teach all of the limitation(s). Li teaches Claim 9, which recites:“The computer-implemented method of claim 1, wherein the input set further comprises first garment pose data, second garment pose data, and person pose data.” Lee et al. teaches an input set containing a “top garment image C1, a bottom garment image C2, and a model image PR.” Lee et al. (US20220318892A1), ¶ [0096]. Li et al. (NPL) teaches that “human pose and garment keypoints are extracted from source images” and represents them as a pose graph and a garment graph. Li et al. (NPL), Abstract, page 1; Section 3.1 and Figure 3, page 3. Applying the garment-keypoint extraction to each of Lee et al.’s two garment images provides the claimed first garment pose data and second garment pose data, while the extracted human-pose keypoints provide the claimed person pose data. Before the effective filing date, a POSITA would have been motivated to apply Li et al.’s pose- and garment-keypoint extraction to each of Lee et al.’s two garment inputs and person image to improve garment alignment and predictably reduce warping distortion in multi-garment virtual try-on. PNG media_image1.png 9 307 media_image1.png Greyscale Claim 20 does not include any additional limitations that would significantly distinguish it from claim 9. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above. PNG media_image1.png 9 307 media_image1.png Greyscale Ground of Rejection 7 Claim 10 is rejected under 35 U.S.C. § 103 as being unpatentable over Lee et al. (US20220318892A1) in view of Morelli et al. (NPL), and further in view of Gu et al. (NPL). As per Claim 10, Lee alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Morelli et al. (NPL), and Gu et al. (NPL), they collectively teach all of the limitation(s). Gu teaches Claim 10, which recites:“The computer-implemented method of claim 1, wherein the machine-learned denoising diffusion model has been progressively trained on increasing image resolutions.” Gu et al. (NPL) teaches training a diffusion model using a “normal denoising objective” and a “progressive training schedule from lower to higher resolutions.” Gu further states that training is divided into phases that “progressively add higher resolution into the training objective,” with example target resolutions of 64, 256, and 1024. Gu et al. (NPL), Abstract, page 1; Section 3.3, pages 4-5; Appendix B, page 18. Before the effective filing date, a POSITA would have been motivated to apply Gu et al.’s progressive lower-to-higher resolution training to the diffusion model of Morelli et al. to improve training efficiency and high-resolution image quality with predictable results. PNG media_image1.png 9 307 media_image1.png Greyscale Ground of Rejection 8 Claim 11 is rejected under 35 U.S.C. § 103 as being unpatentable over Morelli et al. (NPL) in view of Gu et al. (NPL). As per Claim 11, Morelli teaches the following portion of Claim 11, which recites: “A computer system configured to train a denoising diffusion model to perform virtual try-on by performing operations, the operations comprising:” Morelli et al. teaches training a virtual try-on architecture based on “Latent Diffusion Models” using a “text time-conditional U-Net denoising model” and states that the virtual try-on pipeline is trained for “200k iterations.” Morelli et al. (NPL), Sections 3, 3.1, and 4.2, pages 3 and 6. Morelli teaches the following portion of Claim 11, which recites: “performing a plurality of training iterations, each training iteration comprising:” Morelli et al. teaches training the diffusion virtual try-on pipeline for “200k iterations.” Morelli et al. (NPL), Section 4.2, page 6. Morelli teaches the following portion of Claim 11, which recites: “obtaining an image pair, the image pair comprises a target image of a person wearing a garment and a garment image of the garment;” Morelli et al. teaches datasets containing “image pairs of in-shop garments and model images” and explains that, in the paired setting, “the in-shop garment is the same as the model is wearing.” Morelli et al. (NPL), Section 4.1, pages 5-6. Morelli teaches the following portion of Claim 11, which recites: “creating a garment-agnostic image of the person based on the target image and the garment image;” Morelli et al. creates a masked model image I M by masking the target garment region, with the mask defined to “fully encompass the target garment.” The resulting masked model image removes the garment information while retaining the person information. Morelli et al. (NPL), Sections 3.1-3.2, pages 3-4. Morelli teaches the following portion of Claim 11, which recites: “processing the garment image and the garment-agnostic image of the person with the denoising diffusion model to generate a synthetic image that depicts the person wearing the garment;” Morelli et al. processes the encoded masked model image E I M and encoded warped garment E C W with the “denoising network” to obtain the final image I ~ , “where the model in I is wearing the garment in C .” Morelli et al. (NPL), Section 3.2 and Figure 2, pages 4-5. Morelli alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Gu et al. (NPL), they collectively teach all of the limitation(s). Morelli and Gu teach the following portion of Claim 11, which recites: “and modifying one or more values of one or more parameters of the denoising diffusion model based on a loss function that compares the synthetic image to the target image;” Gu et al. teaches training a denoising network x θ that maps a noisy input to its “clean version x ” using the loss ∥ x θ z t t - x ∥ 2 2 , thereby modifying the model parameters based on a comparison between the predicted clean image and the target image. Gu et al. (NPL), Section 2, page 2. Gu teach the following portion of Claim 11, which recites: “wherein the plurality of training iterations are performed over at least two training stages, wherein a first training stage is performed on images having a first resolution, and wherein a second, subsequent training stage is performed on images having a second resolution that is larger than the first resolution.” Gu et al. teaches dividing training into multiple phases that “progressively add higher resolution into the training objective” and provides successive target resolutions of 64, 256, and 1024. Gu et al. (NPL), Section 3.3, pages 4-5; Appendix B, page 18 Before the effective filing date, a POSITA would have been motivated to apply Gu et al.’s progressive multi-resolution training and clean-image prediction loss to Morelli et al.’s diffusion-based virtual try-on model to improve convergence and high-resolution try-on fidelity with predictable results. PNG media_image1.png 9 307 media_image1.png Greyscale Conclusion The prior art made of record and relied upon in this action is as follows: Patent Literature: Lee et al. (US20220318892A1) — “Method and System for Clothing Virtual Try-On Service Based on Deep Learning” Non-Patent Literature (NPL): Morelli et al. — “LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On,” 2023-08-03. Available at: [https://arxiv.org/pdf/2305.13501] Chia et al. — “Contrastive Language and Vision Learning of General Fashion Concepts,” 2023-04-18 attached arXiv version; peer-reviewed version published in 2022. Available at: [https://arxiv.org/pdf/2204.03972] Wu et al. — “Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image Synthesis,” 2023-04-07. Available at: [https://arxiv.org/pdf/2304.03869] Polanía et al. — “Learning Fashion Compatibility Across Apparel Categories for Outfit Recommendation,” 2019-05-01. Available at: [https://arxiv.org/pdf/1905.03703] Gal et al. — “An Image is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion,” 2022-08-02. Available at: [https://arxiv.org/pdf/2208.01618] Li et al. — “Virtual Try-On with Pose-Garment Keypoints Guided Inpainting,” October 2023. Available at: [https://openaccess.thecvf.com/content/ICCV2023/papers/Li_Virtual_Try-On_with_Pose-Garment_Keypoints_Guided_Inpainting_ICCV_2023_paper.pdf]. For date: [https://ieeexplore.ieee.org/abstract/document/10377825] Gu et al. — “Matryoshka Diffusion Models,” 2023-10-23. Available at: [https://arxiv.org/pdf/2310.15111v1] Note: A PDF copy of each NPL reference is attached with this Office Action. URLs are included for applicant convenience. If a link becomes unavailable in the future, the citation information may be used to locate the reference or access archived versions via the Wayback Machine. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure and is listed as follows: Patent Literature: (none) Non-Patent Literature (NPL): Baldrati et al. — “Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image Editing,” 2023-08-23. Available at: [https://arxiv.org/pdf/2304.02051] Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADEEL BASHIR whose telephone number is (571) 270-0440. The examiner can normally be reached Monday-Thursday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached on (571) 276-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ADEEL BASHIR/ Examiner, Art Unit 2616 /DANIEL F HAJNIK/Supervisory Patent Examiner, Art Unit 2616
Read full office action

Prosecution Timeline

Dec 27, 2024
Application Filed
Jul 16, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700165
REAL-TIME NEURAL APPEARANCE MODELS
2y 6m to grant Granted Aug 04, 2026
Patent 12694629
GEOSPATIAL CREATOR PLATFORM
2y 2m to grant Granted Jul 28, 2026
Patent 12682566
ENHANCING ELEVATION MODELS WITH LANDCOVER FEATURE DATA
2y 6m to grant Granted Jul 14, 2026
Patent 12664746
AUGMENTED REALITY VISUALIZATION OF DETECTED DEFECTS
2y 4m to grant Granted Jun 23, 2026
Patent 12664351
AI-BASED SHAPE-ADAPTIVE CONSISTENT VISUAL EFFECT GENERATION
2y 2m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
89%
Grant Probability
93%
With Interview (+4.2%)
2y 2m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 44 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month