DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 1-3, 5, 9, & 17-18 are objected to because of the following informalities: inconsistent naming of the second loss function (Lcp). Examiner understands the presence and purpose of the first, second, and third loss functions, however, the second loss function (Lcp) is inconsistently referred to as the second loss function (Lcp) and second loss function. The Examiner is requesting unity of naming if the significance of (Lcp) does not impact the claimed limitations. Additionally, the third loss function Lcs is interpreted solely as a third loss function, with the term Lcs being treated solely for naming.
Claim 10 is objected to because of the following informalities: undefined acronym. The Examiner would like to note that a person having ordinary skill in the art would likely understand CLIP and DINO scores, however without the written meaning of both acronyms, the reader cannot be certain.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1-18 are rejected under 35 U.S.C. 103 as being unpatentable over Nataniel Ruiz et al. (“DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation”, herein after “Ruiz”) in view of Mehmet Akçakaya et al. (Pat. Pub. US-20260195598-A1, herein after “Mehmet”. Examiner notes that US-20260195598-A1 is granted an earlier filing date from provisional US 63/743,541, filed on 2025-01-09, of which the cited case is supported by).
In regard to claims 1 & 18, Ruiz teaches [a] computer-implemented method for fine-tuning a pre-trained diffusion-based generative model to synthesize subject-specific images, comprising:
obtaining a pre-trained diffusion model “we propose a technique to represent a given subject with rare token identifiers and fine-tune a pre-trained, diffusion-based text-to-image framework” (Ruiz, Page 22501, Col. 1), a set of subject images, and a set of prior images generated by the pre-trained diffusion model “With just a few images (typically 3-5) of a subject (left), DreamBooth—our AI-powered photo booth—can generate a myriad of images of the subject in different contexts (right)” (Ruiz, Page 22500, Fig. 1); and
PNG
media_image1.png
371
990
media_image1.png
Greyscale
outputting a fine-tuned diffusion model configured to generate subject-specific images while retaining generalization capabilities of the pre-trained diffusion model “Figure 3 illustrates the model fine-tuning with the class-generated samples and prior-preservation loss” (Ruiz, Page 22504, Col. 1) where the fine-tuned model is the product which is capable of generating subject-specific images and utilizes the pre-trained model’s capabilities.
PNG
media_image2.png
552
472
media_image2.png
Greyscale
Ruiz, Fig. 3.
Ruiz does not explicitly teach fine-tuning the pre-trained diffusion model by optimizing a first loss function (Ls) configured to predict diffusion noise for the subject images during a reverse diffusion process to preserve an identity of a subject in the subject images subject to a noise consistency regularization that optimizes a second loss function (Lcp) to minimize a discrepancy between prediction of the diffusion noise of the pre-trained and fine-tuned diffusion models for the prior images.
Mehmet teaches fine-tuning the pre-trained diffusion model by optimizing a first loss function (Ls) configured to predict diffusion noise for the subject images “parameters are injected into a fixed, pretrained diffusion sampling pipeline composed of score model prediction (SMP), data fidelity (DF), and DDIM updates. The diffusion model remains frozen, while the auxiliary network parameters are optimized using a supervised reconstruction loss computed between the final reconstruction x0 and the fully sampled reference xref.” (Mehmet, ¶ [0024]) where a pretrained diffusion model composed of model prediction is taught during a reverse diffusion process to preserve an identity of a subject in the subject images subject to a noise consistency regularization “The diffusion model then learns a reverse diffusion process, where a neural network is trained to gradually remove noise and reconstruct the original image” (Mehmet, ¶ [0053]) that optimizes a second loss function (Lcp) to minimize a discrepancy between prediction of the diffusion noise of the pre-trained and fine-tuned diffusion models for the prior images “This also alleviates the need for backpropagation across the score function network, leading to further savings in computational time. The fine-tuning is performed using a physics-inspired loss function that evaluates the consistency of the final estimate and the measurements” (Mehmet, ¶ [0079]) where another loss function is taught to evaluate the consistency of the final estimate and the measurements.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images taught by Ruiz with the use of multiple loss functions to accurately assess the image noise taught by Mehmet to regulate the noise of the subject-specific image as it is put through the diffusion process. The suggestion/motivation to do so would have been to identify when primitives are connected, and when they are crossing so that they may be separated properly.
In regard to claim 18, claim 1 is substantially similar to claim 18, hence the rejection analysis for claim 1 is also applied to claim 18. Ruiz in view of Mehmet teach the additional limitations of [a] non-transitory computer-readable storage medium embodied thereon a program executable by a processor for performing a method for fine-tuning a pre-trained diffusion-based generative model to synthesize subject-specific images “memory 1028 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 1022 to control the one or more data acquisition systems 1024, and/or receive data from the one or more data acquisition systems 1024” (Mehmet, ¶ [0158]), the method comprising:
obtaining a pre-trained diffusion model, a set of subject images, and a set of prior images generated by the pre-trained diffusion model (Ruiz, Page 22501, Col. 1 & 22500, Fig. 1);
fine-tuning the pre-trained diffusion model by optimizing a first loss function (Ls) configured to predict diffusion noise for the subject images during a reverse diffusion process to preserve an identity of a subject in the subject images subject to a noise consistency regularization that optimizes a second loss function (Lcp) to minimize a discrepancy between prediction of the diffusion noise of the pre-trained and fine-tuned diffusion models for the prior images (Mehmet, ¶ [0024], [0053], & [0079]); and
outputting a fine-tuned diffusion model configured to generate subject-specific images while retaining generalization capabilities of the pre-trained diffusion model (Ruiz, Page 22504, Col. 1).
In regard to claim 2, Ruiz in view of Mehmet teach [t]he method of claim 1, wherein computing the second loss function comprises:
generating a latent representation of the prior images “diffusion models define a generative process by introducing a sequence of latent variables that interpolate between structured data and pure noise” (Mehmet, ¶ [0119]) where latent variables are applied;
processing the latent representation through both the pre-trained diffusion model and the fine-tuned diffusion model “Rather than modeling data directly, these models construct a tractable family of intermediate distributions that allow sampling from a complex target distribution through iterative refinement” (Mehmet, ¶ [0119]) where the models iteratively refine the image using the latent values; and
computing the second loss function Lcp to minimize the discrepancy between the noise predictions of the pre-trained and fine-tuned diffusion models for each of the prior images “This also alleviates the need for backpropagation across the score function network, leading to further savings in computational time. The fine-tuning is performed using a physics-inspired loss function that evaluates the consistency of the final estimate and the measurements” (Mehmet, ¶ [0079]) where another loss function is taught to evaluate the consistency of the final estimate and the measurements.
In regard to claim 4, Ruiz in view of Mehmet teach [t]he method of claim 1, wherein the prior images are selected from a class similar to the subject images “We fine-tune the text-to-image model with the input images and text prompts containing a unique identifier followed by the class name of the subject (e.g., “A [V] dog”)” (Ruiz, Page 22501, Col. 1) where the model is encouraged to generate instances similar to the given class.
In regard to claim 5, Ruiz in view of Mehmet teach [t]he method of claim 1, wherein the first loss function LS minimizes a difference between a ground-truth noise added during a forward diffusion process and a predicted noise during the reverse diffusion process “A supervised loss is computed by comparing the output image to the fully sampled reference or ground truth image, and gradients are backpropagated exclusively through the auxiliary network, while all parameters of the pretrained diffusion model remain frozen” (Mehmet, ¶ [0091]) where the supervised loss compares the output, which is read as the reverse diffusion process with the ground truth which has the forward model applied “All measurements were obtained through applying the forward model to a ground truth image” (Mehmet, ¶ [0103]).
In regard to claim 7, Ruiz in view of Mehmet teach [t]he method of claim 1, wherein the fine-tuned diffusion model is configured to generate images in response to textual prompts describing the subject “text-to-image diffusion models … as guided by simple and intuitive text prompts” (Ruiz, Page 22501, Col. 1) where the diffusion model can create images based on a text prompt provided.
In regard to claim 10, Ruiz in view of Mehmet teach [t]he method of claim 1, further configured to evaluate the quality of synthesized images using one or more evaluation metrics, including CLIP scores and DINO scores “One important aspect to evaluate is subject fidelity: the preservation of subject details in generated images. For this, we compute two metrics: CLIP-I and DINO[10] … The second important aspect to evaluate is prompt fidelity, measured as the average cosine similarity between prompt and image CLIP embeddings. We denote this as CLIP-T” (Ruiz, Page 22505, Col. 1) where CLIP and DINO scores are used for evaluation.
PNG
media_image3.png
329
481
media_image3.png
Greyscale
Ruiz, Tables 1 & 2, depicting CLIP and DINO scores
In regard to claim 17, Ruiz teaches [a] system for fine-tuning a diffusion-based generative model to synthesize subject-specific images “we propose a technique to represent a given subject with rare token identifiers and fine-tune a pre-trained, diffusion-based text-to-image framework” (Ruiz, Page 22501, Col. 1), and “There are various approaches to control generative models, where some of them might prove to be viable directions for subject-driven prompt-guided image synthesis” (Ruiz, Page 22502, Col. 1), comprising: a memory for storing instructions; and a processor configured to execute the instructions to:
obtain a pre-trained diffusion model “we propose a technique to represent a given subject with rare token identifiers and fine-tune a pre-trained, diffusion-based text-to-image framework” (Ruiz, Page 22501, Col. 1), a set of subject images and a set of prior images “With just a few images (typically 3-5) of a subject (left), DreamBooth—our AI-powered photo booth—can generate a myriad of images of the subject in different contexts (right)” (Ruiz, Page 22500, Fig. 1) where the images of a subject are read as subject images and the generated images are read as prior images; and
output a fine-tuned diffusion model configured to generate subject-specific images with improved fidelity and diversity “Figure 3 illustrates the model fine-tuning with the class-generated samples and prior-preservation loss” (Ruiz, Page 22504, Col. 1) where the fine-tuned model is the product which is capable of generating subject-specific images and utilizes the pre-trained model’s capabilities.
Ruiz fails to explicitly teach
fine-tune the pre-trained diffusion model by
optimizing a first loss function Ls to predict diffusion noise for the subject images during reverse diffusion, and preserving a subject's identity;
optimizing a second loss function Lcp to enforce consistency between the noise predictions of the pre-trained and fine-tuned diffusion models for the prior images; and
optimizing a third loss function Lcs to enforce consistency between noise predictions for original and perturbed latent representations of the subject images; and
Mehmet teaches fine-tune the pre-trained diffusion model by
optimizing a first loss function Ls to predict diffusion noise for the subject images during reverse diffusion, and preserving a subject's identity “parameters are injected into a fixed, pretrained diffusion sampling pipeline composed of score model prediction (SMP), data fidelity (DF), and DDIM updates. The diffusion model remains frozen, while the auxiliary network parameters are optimized using a supervised reconstruction loss computed between the final reconstruction x0 and the fully sampled reference xref.” (Mehmet, ¶ [0024]) where a pretrained diffusion model composed of model prediction is taught and “[t]he diffusion model then learns a reverse diffusion process, where a neural network is trained to gradually remove noise and reconstruct the original image” (Mehmet, ¶ [0053]);
optimizing a second loss function Lcp to enforce consistency between the noise predictions of the pre-trained and fine-tuned diffusion models for the prior images “This also alleviates the need for backpropagation across the score function network, leading to further savings in computational time. The fine-tuning is performed using a physics-inspired loss function that evaluates the consistency of the final estimate and the measurements” (Mehmet, ¶ [0079]) where another loss function is taught to evaluate the consistency of the final estimate and the measurements.
optimizing a third loss function Lcs to enforce consistency between noise predictions for original and perturbed latent representations of the subject images “A supervised loss is computed by comparing the output image to the fully sampled reference or ground truth image, and gradients are backpropagated exclusively through the auxiliary network, while all parameters of the pretrained diffusion model remain frozen” (Mehmet, ¶ [0136]) where the supervised loss compares the output image, which is read as the perturbed image with the fully sampled image, which is read as the original image.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images taught by Ruiz with the use of multiple loss functions to accurately assess the image noise taught by Mehmet to regulate the noise of the subject-specific image as it is put through the diffusion process. The suggestion/motivation to do so would have been to identify when primitives are connected, and when they are crossing so that they may be separated properly.
Claims 3 & 9 is rejected under 35 U.S.C. 103 as being unpatentable over Ruiz in view of Mehmet and Tay Wi et al. (Pat. Pub. WO-2024201980-A1, herein after “Wi”).
In regard to claim 3, Ruiz in view of Mehmet teach [t]he method of claim 1, further comprising:
optimizing a third loss function Lcs to minimize the discrepancy between the noise predictions for the original and perturbed latent representations of the subject images, thereby improving diversity in the generated subject-specific images “A supervised loss is computed by comparing the output image to the fully sampled reference or ground truth image, and gradients are backpropagated exclusively through the auxiliary network, while all parameters of the pretrained diffusion model remain frozen” (Mehmet, ¶ [0136]) where the supervised loss compares the output image, which is read as the perturbed image with the fully sampled image, which is read as the original image.
Ruiz in view of Mehmet fail to teach applying multiplicative Gaussian noise to latent representations of the subject images to create perturbed latent representations
processing both original and perturbed latent representations through the fine-tuned diffusion model.
Wi teaches applying multiplicative Gaussian noise to latent representations of the subject images to create perturbed latent representations “The latent representation converted by the image encoder 401 is then repeatedly added with noise (e.g., Gaussian noise) at multiple time steps in a diffusion process 402 to generate a noise image 403” (Wi, Page 7) where gaussian noise is repeatedly added to the latent representation to generate a noise image;
processing both original and perturbed latent representations through the fine-tuned diffusion model “The image encoder 401 converts the image 41 (pixel image) into a latent representation (latent image embedding representation) in a latent space. This embeds a high-dimensional image into a low-dimensional latent space, reducing the load on computational processing. The latent representation converted by the image encoder 401 is then repeatedly added with noise (e.g., Gaussian noise) at multiple time steps in a diffusion process 402” (Wi, Page 7), additionally “the noise added in the diffusion process 402 is gradually removed from the noise image 403” (Wi, Page 7) where the images are processed through the diffusion model.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the latent representation taught by Wi to decrease computation space. The suggestion/motivation to do so would have been to reduce the work strain during performance.
In regard to claim 9, Ruiz in view of Mehmet and Wi teach [t]he method of claim 3, wherein the third loss function Lcs is scaled by a configurable hyperparameter to balance diversity and identity preservation “… dynamic automated hyperparameter tuning during the inference phase to enhance the reconstruction quality of solving linear noisy inverse problems using diffusion models …” (Mehmet, [0118]) where hyperparameter tuning can be used in the diffusion model to enhance the reconstruction quality.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Ruiz in view of Mehmet, Wi, and Zizheng Pan et al. (Pat. Pub. US-20250103968-A1, herein after “Pan”. Examiner notes that US-20250103968-A1 is granted an earlier filing date from provisional application No. US 63/540,598, filed on 2023-09-26, of which the cited case is supported by).
In regard to claim 6, Ruiz in view of Mehmet and Wi teach [t]he method of claim 3.
Ruiz in view of Mehmet and Wi fail to explicitly teach wherein the multiplicative Gaussian noise applied to the latent representations is defined by Em~ N (1, σ2,I) where σ2 is a variance parameter.
Pan teaches wherein the multiplicative Gaussian noise applied to the latent representations is defined by Em~ N (1, σ2,I) where σ2 is a variance parameter “denote the distribution of noised samples by injecting σ2-variance Gaussian noise” (Pan, ¶ [0037]) additionally, the function provided in ¶ [0039-0040] provides a similar distribution. Examiner would also like to note that no absolute variable definition is given for the latent representation definition beyond what a person having ordinary skill in the art would know.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the multiplicative Gaussian noise using a variance parameter taught by Pan to regulate the noise of the subject-specific image as it is put through the diffusion process. The suggestion/motivation to do so would have been to apply or remove noise in a specific manner based on the desired image results.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Ruiz in view of Mehmet and Xin-yu Yang et al. (Pat. Pub. CN-118052705-A, herein after “Yang”).
In regard to claim 8, Ruiz in view of Mehmet teach [t]he method of claim 1.
Ruiz in view of Mehmet fail to teach wherein the fine-tuning is performed using a low-rank adaptation (LoRA) method to update only a subset of parameters of the pre-trained diffusion model or additional parameters added to the pre-trained diffusion model.
Yang teaches wherein the fine-tuning is performed using a low-rank adaptation (LoRA) method to update only a subset of parameters of the pre-trained diffusion model or additional parameters added to the pre-trained diffusion model “according to the expression of the potential vector and the potential space, using the diffusion layer to remove the noise in the expression of the potential space, obtaining the mask image after removing the noise, at the same time, using the low rank adaptive model Lora, adjusting the weight of the deep learning layer and the variable self-encoder” (Yang, Page 5) where the low rank adaptive model is used during the noise reduction, or fine-tuning, to update parameters.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the low-rank adaptation to update parameters taught by Yang to adjust the weight of the models learning. The suggestion/motivation to do so would have been to adjust a portion of the model based on the desired parameters.
Claims 11-13 & 16 are rejected under 35 U.S.C. 103 as being unpatentable over Ruiz in view of Mehmet and Duffy Brian et al. (Pat. Pub. KR-20150123872-A, herein after “Brian”).
In regard to claim 11, Ruiz in view of Mehmet teach [t]he method of claim 1.
Ruiz in view of Mehmet fail to teach wherein the subject-specific image synthesis is applied to visual inspection and quality control in factory automation, further comprising:
generating diverse representations of components under varying operational conditions to train machine vision systems for defect detection; and
enhancing the accuracy and robustness of defect detection systems by simulating potential variations in component features or manufacturing defects.
Brian teaches wherein the subject-specific image synthesis is applied to visual inspection and quality control in factory automation “The embodiments described herein are applicable to inspection, metrology, and test applications (e.g., electrical, image based, etc.) that can benefit from a large number of actual, written, simulated, (Or time-intensive) image acquisition sources in combination with a wide variety of image acquisition systems” (Brian, Page 2) where the image inspection is capable of being implemented in a factory “virtual and physical systems can interact with each other according to standard factory automation protocols” (Brian, Page 7), further comprising:
generating diverse representations of components under varying operational conditions to train machine vision systems for defect detection “The functions of the full inspection system that can be performed by the virtual system (s) as a proxy for the full inspection system include, but are not limited to, defect detection, defect classification, image defect source analysis …” (Brian, Page 9) where various conditions of objects are inspected and used for defect detection; and
enhancing the accuracy and robustness of defect detection systems by simulating potential variations in component features or manufacturing defects “virtual systems can be configured to perform modeling of manufacturing processes performed on a sample to simulate that the sample "looks" at some point in the process. Virtual systems can also be configured to perform modeling of the actual process performed by one or more of the actual systems to simulate that the output … "appears" to the sample on which the actual process is performed” (Brian, Page 11) where modeling can be performed on the various manufacture products.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the defect detection taught by Brian to gather clear images of a production line for detecting product quality. The suggestion/motivation to do so would have been to quickly run quality control on a production line. The examiner would like to note that the computer-implemented method claimed in claim 1 being used in factory automation is intended use.
In regard to claim 12, Ruiz in view of Mehmet teach [t]he method of claim 1.
Ruiz in view of Mehmet fail to teach wherein the subject-specific image synthesis is utilized for robotics and machine vision, the method further comprising:
generating synthetic images of tools, components, and factory environments to train robotic systems for object recognition and manipulation; and
allowing robotic systems to adapt to dynamic factory setups and modular production lines without requiring extensive reprogramming.
Brian teaches wherein the subject-specific image synthesis is utilized for robotics and machine vision “image composition may be performed using any image data generated by any of the real or virtual systems described herein” (Brian, Page 16) where machine vision is used and “multi-layer image synthesis, images can be formed by combining images obtained in different steps in a computationally determined process flow” (Brian, Page 16) where image synthesis is used, the method further comprising:
generating synthetic images of tools, components, and factory environments to train robotic systems for object recognition and manipulation “The post-processing analysis may include such operations as sampling, filtering, data or image "synthesis" for subsequent operations” (Brian, Page 13) where image synthesis and object recognition can be used for defect detection; and
allowing robotic systems to adapt to dynamic factory setups and modular production lines without requiring extensive reprogramming “one or more processes performed on the sample (s) by actual systems are production control related processes such as inspection, review, metrology, testing, and the like” (Brian, Page 3), additionally “virtual systems can be configured for adaptive sampling of points of interest and regions of interest in a particular imaging or modeling mode. Virtual systems can also be used for adaptive sampling with automated defect classification (ADC) to create a self-learning optical-electron beam classification engine for high defect volumes, which can be used for engineering experiments under factory automation” (Brian, Page 14) teaches adaptive point sampling for defect detection in a factory automation setting.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the object recognition taught by Brian to gather clear images of a production line for detecting products. The suggestion/motivation to do so would have been to quickly determine the objects on the production line.
In regard to claim 13, Ruiz in view of Mehmet teach [t]he method of claim 1.
Ruiz in view of Mehmet fail to teach wherein the subject-specific image synthesis is applied to product prototyping and design, the method further comprising:
creating diverse visual prototypes of machinery or components to facilitate rapid evaluation and iteration of design configurations; and
customizing factory layouts and setups by generating tailored visualizations of workflows and production processes.
Brian teaches wherein the subject-specific image synthesis is applied to product prototyping and design, the method further comprising:
creating diverse visual prototypes of machinery or components to facilitate rapid evaluation and iteration of design configurations “The virtual system (s) may also be configured to obtain a modeled representation of the modeled image based on one or more of the images or other output of the images or real systems generated by one of the actual systems described herein” (Brian, Page 11) where the image may be configured to produce a model, which is read as a prototype of a component; and
customizing factory layouts and setups by generating tailored visualizations of workflows and production processes “the real system may generate a modeled representation and the virtual system (s) may obtain a modeled representation from the real system. In a similar manner, virtual systems can acquire modeled images based on different outputs of the actual system (s) by performing modeling or receiving such modeled images from the actual system that created them” (Brian, Page 11-12) where modeled representations of the system can be generated in a number of ways.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the visualization and modeling capabilities taught by Brian to create and modify a visualization of an object. The suggestion/motivation to do so would have been to allow for a user to visualize the object or product and perceive any defects.
In regard to claim 16, Ruiz in view of Mehmet teach [t]he method of claim 1.
Ruiz in view of Mehmet fail to explicitly teach wherein the subject-specific image synthesis supports real-time factory monitoring, the method further comprising:
integrating synthesized images into monitoring systems to provide real-time visualization of machinery states; and
enabling anomaly detection through comparison of real-time data with diverse synthetic representations of normal and faulty operating conditions.
Brian teaches wherein the subject-specific image synthesis supports real-time factory monitoring “Previous data may come from any of the sources connected to the virtual systems” (Brian, Page 13-14) where the data is collected in real-time by the system and is used in image synthesis, the method further comprising:
integrating synthesized images into monitoring systems to provide real-time visualization of machinery states “Previous data can also be generated in simulated real-time (e.g., as data is acquired, the recipe can be changed based on that data)” (Brian, Page 14) where the data is from synthesized images and may be gathered as the system is used; and
enabling anomaly detection through comparison of real-time data with diverse synthetic representations of normal and faulty operating conditions “For data analysis recipes, recipes can include inspection, defect source analysis …” (Brian, Page 15) where the data includes defect source analysis which is read as anomaly detection by comparing the real-time data with normal conditions.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the real-time anomaly detection taught by Brian to visualize an object and detect faults of an object. The suggestion/motivation to do so would have been to quickly detect and verify the quality of the produced object.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Ruiz in view of Mehmet, Brian, and Charles Cella et al. (Pat. Pub. JP-2023524250-A, herein after “Cella”).
In regard to claim 14, Ruiz in view of Mehmet teach [t]he method of claim 1.
Ruiz in view of Mehmet fail to teach wherein the subject-specific image synthesis is applied to predictive maintenance, the method further comprising:
generating imagery of machinery operating under different conditions, including normal states and early signs of wear or failure; and
simulating rare failure scenarios to train maintenance systems for proactive identification and resolution of potential issues.
Brian teaches wherein the subject-specific image synthesis is applied to predictive maintenance “virtual machines may be configured for a predictive factory resource scheduling and may include a programmable state machine based resource allocation / management structure. In this way, the system can be configured for state machine resource management with the predictive features linked to the factory automation system” (Brian, Page 12) where the system may be used in predictive resource scheduling, which is read as maintenance under broadest reasonable interpretation, since the maintenance is not specified in this limitation, the method further comprising:
generating imagery of machinery operating under different conditions, including normal states and early signs of wear or failure “The post-processing analysis may include such operations as sampling, filtering, data or image "synthesis" for subsequent operations, systematic pattern failure, and any other post-processing analysis that may be performed by the actual systems described herein” (Brian, Page 13) where the post-processing analysis is capable of image synthesis for systematic pattern failure.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the visualization and modeling capabilities of the object at different conditions taught by Brian to create and modify a visualization of an object. The suggestion/motivation to do so would have been to allow for a user to visualize the object or product and perceive any defects.
Ruiz in view of Mehmet and Brian fail to teach simulating rare failure scenarios to train maintenance systems for proactive identification and resolution of potential issues.
Cella teaches simulating rare failure scenarios to train maintenance systems for proactive identification and resolution of potential issues “the digital twin 60136 may generate a large set of realistic accident scenarios and then reliably simulate the vehicle's 60104 response in such scenarios. In embodiments, the digital twin 60136 may display the trajectory taken by the vehicle 60104 in the event of a brake failure and the effect on the occupants or other vehicles” (Cella, Page 77) where a failure simulation is taught which can improve maintenance in the future.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions and modeling capabilities to accurately assess the image noise and product quality taught by Ruiz, Mehmet and Brian with the failure simulation taught by Cella to improve future quality by testing critical areas. The suggestion/motivation to do so would have been to reinforce quality by preparing for failed outcomes.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Ruiz in view of Mehmet and Erol Baris et al. (Pat. Pub. WO-2024035397-A1, herein after “Baris”).
In regard to claim 15, Ruiz in view of Mehmet teach [t]he method of claim 1.
Ruiz in view of Mehmet fail to explicitly teach wherein the subject-specific image synthesis is applied to training and simulation environments in factory automation, the method further comprising:
creating high-fidelity virtual factory models with detailed images of machinery, workflows, and production lines for operator training; and
enabling the simulation of complex workflows through digital twins to optimize production processes and support continuous improvement.
Baris teaches wherein the subject-specific image synthesis is applied to training and simulation environments in factory automation “the simulation engine 204 may be executed to render synthetic images of a part on the production line acquired by one or more cameras” (Baris, ¶ [0023]), the method further comprising:
creating high-fidelity virtual factory models with detailed images of machinery, workflows, and production lines for operator training “the optimization of the camera configuration can be generalized for different types of parts. In this case, the above-described methodology may be implemented by using 3D models of different parts having different nominal geometries at each iteration including simulation, surface coverage measurement … and optimization” (Baris, ¶ [0055]) where factory models are included in the simulation; and
enabling the simulation of complex workflows through digital twins to optimize production processes and support continuous improvement “A landmark may be representative of all points on a given planar surface. The landmarks may be specified, for example, in the part information used for creating the digital twin 202” (Baris, ¶ [0037]) where digital twins are used to specify landmarks.
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of fine-tuning a diffusion model for generating subject-specific images with the use of multiple loss functions to accurately assess the image noise taught by Ruiz and Mehmet with the visualization and modeling capabilities taught by Baris to create and modify a visualization of an object. The suggestion/motivation to do so would have been to allow for a user to visualize the object or product and perceive any defects.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Nicholas Kolkin et al. (Pat Pub. US-20240135610-A1) discloses a diffusion model for image processing including the difference between multiple output styles (abstract).
Jun-Yan Zhu et al. "Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks," 2017 discloses translation of image styles between multiple photos.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAIDEN ALEXANDER USSERY whose telephone number is (571)272-1192. The examiner can normally be reached Monday - Friday* 7:30AM - 5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached at (571) 272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/C.A.U./Examiner, Art Unit 2611