Prosecution Insights
Last updated: August 17, 2026
Application No. 18/892,031

IMAGE OBJECT MASK GENERATION

Non-Final OA §102§103
Filed
Sep 20, 2024
Examiner
RIVERA-MARTINEZ, GUILLERMO M
Art Unit
2677
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
8m
Est. Remaining
81%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
398 granted / 511 resolved
+15.9% vs TC avg
Minimal +3% lift
Without
With
+3.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
31 currently pending
Career history
542
Total Applications
across all art units

Statute-Specific Performance

§101
6.1%
-33.9% vs TC avg
§103
44.9%
+4.9% vs TC avg
§102
23.2%
-16.8% vs TC avg
§112
23.6%
-16.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 511 resolved cases

Office Action

§102 §103
DETAILED ACTION This Office action is in response to the Application filed on September 20, 2024. An action on the merits follows. Claims 1-20 are pending on the application. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 3, 5-6, 8, 13, and 15-20 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Zhang et al. (US PG Publication No. 2024/0169541 A1), hereafter referred to as Zhang, Applicant cited prior art furnished via IDS. Regarding claim 1, Zhang discloses a device (Par. [0004-7]: systems and methods for using a machine learning to perform instance segmentation on an image… generating a segmentation mask for the input image… using a diffusion model… a mask network configured to generate an instance mask for an input image; Par. [0008]: FIG. 1 is an illustrative depiction of users interacting with an image processing system, including a mask network) comprising: a memory configured to store image data (Par. [0007]: the apparatus and system include… a memory; Par. [0036-40]: submit data (e.g., image(s)) for analysis… a computer network configured to provide on-demand availability of computer system resources, such as data storage… FIG. 2 shows a block diagram of an example of an image generator… an image generator 200 receives an original image… the image generator 200 can include a computer system 280 including… computer memory 220; Par. [0124-127]: FIG. 13 shows an example of a computer system… the computing device includes… memory subsystem 1320… memory subsystem 1320 includes one or more memory devices; a memory configured to store image data (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models include one or more memory devices configured to perform data storage, including submitted and received input data for analysis, such as images (i.e. a memory configured to store image data), as indicated above), for example); and one or more processors coupled to the memory and configured to (Par. [0007]: the apparatus and system include… a processor; a memory containing instructions executable by the processor; Par. [0039-40]: FIG. 2 shows a block diagram of an example of an image generator according to aspects of the present disclosure… an image generator 200 receives an original image… the image generator 200 can include a computer system 280 including one or more processors 210, computer memory 220; Par. [0124-127]: FIG. 13 shows an example of a computer system according to aspects of the present disclosure… the computing device includes processor(s) 1310, memory subsystem 1320… The computer system 1300 can be configured to perform the operations described above and illustrated in FIG. 1-12… computing device 1300 includes one or more processors 1310… a processor 1310 is configured to operate a memory array using a memory controller… a processor is configured to execute computer-readable instructions stored in a memory to perform various functions… memory subsystem 1320 includes one or more memory devices… memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor to perform various functions described herein; and one or more processors coupled to the memory and configured to (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models include a computing device which includes one or more processors and one or more memory devices containing computer-readable software instructions that, when executed, cause a processor to perform various functions (i.e. and one or more processors coupled to the memory and configured to), as indicated above), for example): obtain a first group of feature sets from a first sampling iteration of multiple sampling iterations associated with a diffusion model, wherein the multiple sampling iterations are configured to generate a latent representation of a first image (Par. [0040-62]: image generator 200 can include… a diffusion model 260… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors… noise component 250 generates an iterative noise map for each of the set of masks with successively reduced noise to produce the output image… diffusion model 260 iteratively produces a set of output images… FIG. 3 shows a block diagram of an example of a guided diffusion model 300… The guided latent diffusion model 300 depicted in FIG. 3 is an example of, or includes aspects of, the corresponding diffusion model element 260 described with reference to FIG. 2… Diffusion models are a class of generative neural networks which can be trained to generate new data with features similar to features found in training data… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301… as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels… Next, a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… an output image 390 is created from each of the various noise levels. The output image 390 can be compared to the original image 301 to train the reverse diffusion process 330… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution and an initial number of channels, and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415. The intermediate features 415 are then down-sampled using a down-sampling layer 420… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435. The up-sampled features 435 can be combined with intermediate features 415 having a same resolution and number of channels via a skip connection 440. These inputs are processed using a final neural network layer 445 to produce output features 450… FIG. 5 shows a diffusion process… As described above with reference to FIG. 3, a diffusion model 300 can include both a forward diffusion process 505 for adding noise to an image (or features in a latent space) and a reverse diffusion process 510 for denoising the images (or features) to obtain a denoised image… the forward diffusion process 505 is used during training to generate images with successively greater noise, and a neural network is trained to perform the reverse diffusion process 510 (i.e., to successively remove the noise)… The neural network may be trained to perform the reverse process. During the reverse diffusion process 510… At each step t−1, the reverse diffusion process 510 takes xt, such as first intermediate image 520, and t as input. Here, t represents a step in the sequence of transitions associated with different noise levels. The reverse diffusion process 510 outputs xt-1, such as second intermediate image 525 iteratively until xT is reverted back to x0, the original image 530; Par. [0068-77]: provide an initial image 510 to an image processing system 130… The image processing system 130 can be configured to identify the objects 512, 515 in the image 510… image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… FIG. 7 is an illustrative depiction of a diagram of an image processing system to perform amodal instance segmentation using… a diffusion model…image processing system 130 can include a mask network having an image encoder 720 that can identify multiple objects of the same class as distinct individual objects or instances 512, 515 in the image 510. The image encoder 720 can generate features… The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130…The image encoder 720 can generate one or more feature maps 740; obtain a first group of feature sets from a first sampling iteration of multiple sampling iterations associated with a diffusion model, wherein the multiple sampling iterations are configured to generate a latent representation of a first image (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), for example, in which the intermediate features are down-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using a down-sampling layer and up-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using an up-sampling process to obtain up-sampled features (i.e. obtain a first, second, third… Nth group of feature sets from a first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model), as indicated above), for example); and generate, based on the first group of feature sets, first mask data that indicates a first mask associated with a first object of the first image (Par. [0004-7]: systems and methods for using a machine learning to perform instance segmentation on an image… generating a segmentation mask for the input image… using a diffusion model… a mask network configured to generate an instance mask for an input image; Par. [0025]: a neural network model, referred to as a mask model, is provided that can identify the visible parts of objects in an image; Par. [0040-43]: image generator 200 can include… a mask component 240, a noise component 250, and a diffusion model 260… first machine learning model generates an output mask that is then used as input for a second machine learning model, e.g., diffusion model 260. An example of the diffusion model 260 is provided with reference to FIG. 3… the mask component 240 can generate a mask, which is a set of binary values for a region of interest in the image… mask component 240 identifies a mask indicating the region of the original image. A mask can be used to specify a region of an image that contains an object of interest based identified through object detection. Different masks can identify different regions and objects in an image; Par. [0068-77]: image processing system 130 can be configured to identify the objects 512, 515 in the image 510, generate masks for the objects… identify a plurality of object instance 512, 515 in the image 510, where the plurality of object instances may be determined through instance segmentation of the image… a mask network is used to provide an instance mask prediction… based on the features generated by an image encoder… the image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… an instance mask can be generated for each of the object instances 512, 515, where the instance masks specify the visible portions of each of the objects 512, 515. The instance masks can be generated by a trained mask network, that can distinguish separate objects, where the instance masks can be generated based on the output of the instance segmentation… segmentation mask can be predicted and generated by a diffusion model of the mask network. The segmentation mask may indicate the visible region and the occluded region of the object… a diffusion model can generate a complete image 590 of the occluded object 515, and provide the amodal segmentation mask… FIG. 7 is an illustrative depiction of a diagram of an image processing system to perform amodal instance segmentation using a mask network having a segmentation network and a diffusion model… For example, the image processing system depicted in FIG. 7 may be an example of a mask network including the mask component 240 of FIG. 2… user 110 can submit the initial image 510… to an image processing system 130. The image processing system 130 can include a mask network having an image encoder 720 that can identify multiple objects of the same class as distinct individual objects or instances 512, 515 in the image 510. The image encoder 720 can generate features that include information… The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130… the image can be divided into fragments, for example, patches of a predetermined size. The image encoder 720 can generate one or more feature maps 740, which may be semantic rich feature maps(s). Feature maps generated by a convolution operation can represent the spatial arrangement of the activations in a convolutional layer; Par. [0082-97]: the image processing system 130 can include a transformer decoder 730 that can be configured to generate masks, tokens, labels, and confidence scores. Tokens can represent a patch of an image, where a token can be created for each small patch of the image. The image patch can be mapped to a feature vector referred to as the token. An output from the image encoder 720, including multi-level features, can be input to the transformer decoder 730, along with a set 732 of input tokens 734… the input token set 732 can have one input token 734 generated for each of the individual objects or instances 512, 515 in the image 510…. the output token set 736 can have one output token 738 generated for each of the individual objects or instances 512, 515 in the image 510… the neural network layers 750 can include a first head configured to generate an instance mask… the first kernel 770 and the second kernel 775 can be applied to the feature maps 740, where the first kernel 770 applied to the feature maps 740 can generate a positive mask for the visible region 780 of the occluded object 515… The first kernel 770 can generate the positive mask through convolution applied to the feature map(s) 740… The diffusion model can determine the shape and size of the occluded object from instance masks and the occlusion mask based on the different regions… FIG. 8 is an illustrative depiction of a method of mask determination… instance masks 810, 820 can be generated by a mask network; generate, based on the first group of feature sets, first mask data that indicates a first mask associated with a first object of the first image (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), as indicated above, for example, including generating different masks identifying different regions and objects in an image, respectively, by using kernels that generate masks through convolution applied to the generated features (i.e. generate, based on the first, second, third… Nth group of feature sets, first, second, third... Nth mask data that indicates a first, second, third... Nth mask associated with a first, second, third... Nth object of the first, second, third... Nth image), as indicated above, for example). Regarding claim 3, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), wherein a first feature set of the first group of feature sets has a first resolution, and wherein a second feature set of the first group of feature sets has a second resolution that is distinct from the first resolution (Par. [0040-58]: image generator 200 can include… a diffusion model 260… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors… noise component 250 generates an iterative noise map for each of the set of masks with successively reduced noise to produce the output image… diffusion model 260 iteratively produces a set of output images… FIG. 3 shows a block diagram of an example of a guided diffusion model 300… The guided latent diffusion model 300 depicted in FIG. 3 is an example of, or includes aspects of, the corresponding diffusion model element 260 described with reference to FIG. 2… Diffusion models are a class of generative neural networks which can be trained to generate new data with features similar to features found in training data… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301… as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels… Next, a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… an output image 390 is created from each of the various noise levels. The output image 390 can be compared to the original image 301 to train the reverse diffusion process 330… The U-Net 400 depicted in FIG. 4 is an example of, or includes aspects of, the architecture used within the reverse diffusion process described with reference to FIG. 3… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution… and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415. The intermediate features 415 are then down-sampled using a down-sampling layer 420 such that down-sampled features 425 features have a resolution less than the initial resolution and a number of channels greater than the initial number of channels… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435. The up-sampled features 435 can be combined with intermediate features 415 having a same resolution and number of channels via a skip connection 440. These inputs are processed using a final neural network layer 445 to produce output features 450; wherein a first feature set of the first group of feature sets has a first resolution, and wherein a second feature set of the first group of feature sets has a second resolution that is distinct from the first resolution (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images, for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features having an initial resolution (i.e. group of feature sets has a first resolution) and processes the input features to produce intermediate features that are down-sampled, such that down-sampled features have a resolution less than the initial resolution (i.e. group of feature sets has a second resolution that is distinct from the first resolution), including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. wherein a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets has a first resolution, and wherein a first, second, third… Nth feature set of the first, second, third… Nth group of feature sets has a second resolution that is distinct from the first resolution), as indicated above), for example). Regarding claim 5, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), wherein: the one or more processors are configured to scale one or more feature sets of the first group of feature sets to generate input feature sets, each of the input feature sets having a same resolution; and the first mask data is based on the input feature sets (Par. [0040-58]: image generator 200 can include… a diffusion model 260… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors… noise component 250 generates an iterative noise map for each of the set of masks with successively reduced noise to produce the output image… diffusion model 260 iteratively produces a set of output images… FIG. 3 shows a block diagram of an example of a guided diffusion model 300… The guided latent diffusion model 300 depicted in FIG. 3 is an example of, or includes aspects of, the corresponding diffusion model element 260 described with reference to FIG. 2… Diffusion models are a class of generative neural networks which can be trained to generate new data with features similar to features found in training data… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301… as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels… Next, a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… an output image 390 is created from each of the various noise levels. The output image 390 can be compared to the original image 301 to train the reverse diffusion process 330… The U-Net 400 depicted in FIG. 4 is an example of, or includes aspects of, the architecture used within the reverse diffusion process described with reference to FIG. 3… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution… and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415. The intermediate features 415 are then down-sampled using a down-sampling layer 420 such that down-sampled features 425 features have a resolution less than the initial resolution and a number of channels greater than the initial number of channels… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435. The up-sampled features 435 can be combined with intermediate features 415 having a same resolution and number of channels via a skip connection 440. These inputs are processed using a final neural network layer 445 to produce output features 450… the output features 450 have the same resolution as the initial resolution; scale one or more feature sets of the first group of feature sets to generate input feature sets, each of the input feature sets having a same resolution; and the first mask data is based on the input feature sets (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images, for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features having an initial resolution (i.e. group of feature sets has a first resolution) and processes the input features to produce intermediate features that are down-sampled (i.e. scaled), such that down-sampled features have a resolution less than the initial resolution (i.e. scale one or more feature sets of the first group of feature sets to generate input feature sets), including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. and the first, second, third… Nth mask data is based on the input feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. scale one or more feature sets of the first, second, third… Nth group of feature sets to generate input feature sets), for example, in which the down-sampled features are up-sampled to obtain up-sampled features that are combined with intermediate features having a same resolution and the output features have the same resolution as the initial resolution (i.e. each of the input feature sets having a same resolution), for example, and these inputs are processed using a final neural network layer to produce output features, as indicated above), for example). Regarding claim 6, claim 5 is incorporated and Zhang discloses the device (Par. [0004-7]), wherein: the one or more processors are configured to aggregate the input feature sets to generate an aggregated feature set (Par. [0055-59]: obtain guidance features 360 in guidance space 370. The guidance features 360 can be combined with the noisy images 320 at one or more layers of the reverse diffusion process 330 to ensure that the output image 390 includes content described by the text prompt 340. For example, guidance features 360 can be combined with the noisy features using a cross-attention block within the reverse diffusion process 330… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution and an initial number of channels, and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435. The up-sampled features 435 can be combined with intermediate features 415 having a same resolution and number of channels via a skip connection 440. These inputs are processed using a final neural network layer 445 to produce output features 450. In some cases, the output features 450 have the same resolution as the initial resolution… In some cases, U-Net 400 takes additional input features to produce conditionally generated output. For example, the additional input features could include a vector representation of an input prompt. The additional input features can be combined with the intermediate features 415 within the neural network at one or more layers. For example, a cross-attention module can be used to combine the additional input features and the intermediate features 415; aggregate the input feature sets to generate an aggregated feature set (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images, for example, including up-sampled features that are combined (i.e. aggregated) with intermediate features and additional input features that combined with the intermediate features (i.e. aggregate the input feature sets to generate an aggregated feature set), as indicated above), for example); and the first mask data is based on the aggregated feature set (Par. [0029]: diffusion model can recreate an unseen, occluded portion of an occluded object by generating modal segmentation masks for the visible portions of the occluding and occluded objects, and an occlusion mask that combines the modal segmentation masks of each object; Par. [0068-73]: image processing system 130 can be configured to identify the objects 512, 515 in the image 510, generate masks for the objects, generate an amodal segmentation mask for an occluded portion of an occluded object 515, and generate an image patch to fill in the unseen, occluded portion of the occluded object… identify a plurality of object instance 512, 515 in the image 510, where the plurality of object instances may be determined through instance segmentation of the image… a mask network is used to provide an instance mask prediction and a rough occlusion prediction based on the features generated by an image encoder in the first step. The outputs of the first step are then provided to the diffusion model to perform amodal mask completion … an amodal segmentation mask can be generated by the mask network, where the amodal segmentation mask combines the inferred occluded region with the visible region of the occluded object 515. The amodal segmentation mask can be predicted and generated by a diffusion model of the mask network; Par. [0073]: an amodal segmentation mask can be generated by the mask network, where the amodal segmentation mask combines the inferred occluded region with the visible region of the occluded object 515. The amodal segmentation mask can be predicted and generated by a diffusion model of the mask network; and the first mask data is based on the aggregated feature set (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images, for example, for example, including up-sampled features that are combined (i.e. aggregated) with intermediate features and additional input features that are combined with the intermediate features (i.e. aggregate the input feature sets to generate an aggregated feature set), for example, including a segmentation mask that is predicted and generated by a diffusion model of the mask network which combines the segmentation masks of each object (i.e. and the first mask data is based on the aggregated feature set), as indicated above), for example). Regarding claim 8, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), wherein: the one or more processors are configured to obtain a second group of feature sets from a second sampling iteration of the multiple sampling iterations (Par. [0040-62]: image generator 200 can include… a diffusion model 260… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors… noise component 250 generates an iterative noise map for each of the set of masks with successively reduced noise to produce the output image… diffusion model 260 iteratively produces a set of output images… FIG. 3 shows a block diagram of an example of a guided diffusion model 300… The guided latent diffusion model 300 depicted in FIG. 3 is an example of, or includes aspects of, the corresponding diffusion model element 260 described with reference to FIG. 2… Diffusion models are a class of generative neural networks which can be trained to generate new data with features similar to features found in training data… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301… as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels… Next, a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… an output image 390 is created from each of the various noise levels. The output image 390 can be compared to the original image 301 to train the reverse diffusion process 330… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution and an initial number of channels, and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415. The intermediate features 415 are then down-sampled using a down-sampling layer 420… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435. The up-sampled features 435 can be combined with intermediate features 415 having a same resolution and number of channels via a skip connection 440. These inputs are processed using a final neural network layer 445 to produce output features 450… FIG. 5 shows a diffusion process… As described above with reference to FIG. 3, a diffusion model 300 can include both a forward diffusion process 505 for adding noise to an image (or features in a latent space) and a reverse diffusion process 510 for denoising the images (or features) to obtain a denoised image… the forward diffusion process 505 is used during training to generate images with successively greater noise, and a neural network is trained to perform the reverse diffusion process 510 (i.e., to successively remove the noise)… The neural network may be trained to perform the reverse process. During the reverse diffusion process 510… At each step t−1, the reverse diffusion process 510 takes xt, such as first intermediate image 520, and t as input. Here, t represents a step in the sequence of transitions associated with different noise levels. The reverse diffusion process 510 outputs xt-1, such as second intermediate image 525 iteratively until xT is reverted back to x0, the original image 530; Par. [0068-77]: provide an initial image 510 to an image processing system 130… The image processing system 130 can be configured to identify the objects 512, 515 in the image 510… image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… FIG. 7 is an illustrative depiction of a diagram of an image processing system to perform amodal instance segmentation using… a diffusion model…image processing system 130 can include a mask network having an image encoder 720 that can identify multiple objects of the same class as distinct individual objects or instances 512, 515 in the image 510. The image encoder 720 can generate features… The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130…The image encoder 720 can generate one or more feature maps 740; obtain a second group of feature sets from a second sampling iteration of the multiple sampling iterations (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), for example, in which the intermediate features are down-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using a down-sampling layer and up-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using an up-sampling process to obtain up-sampled features (i.e. obtain a first, second, third… Nth group of feature sets from a first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model), as indicated above), for example); and the first mask data is further based on the second group of feature sets (Par. [0004-7]: systems and methods for using a machine learning to perform instance segmentation on an image… generating a segmentation mask for the input image… using a diffusion model… a mask network configured to generate an instance mask for an input image; Par. [0025]: a neural network model, referred to as a mask model, is provided that can identify the visible parts of objects in an image; Par. [0040-43]: image generator 200 can include… a mask component 240, a noise component 250, and a diffusion model 260… first machine learning model generates an output mask that is then used as input for a second machine learning model, e.g., diffusion model 260. An example of the diffusion model 260 is provided with reference to FIG. 3… the mask component 240 can generate a mask, which is a set of binary values for a region of interest in the image… mask component 240 identifies a mask indicating the region of the original image. A mask can be used to specify a region of an image that contains an object of interest based identified through object detection. Different masks can identify different regions and objects in an image; Par. [0068-77]: image processing system 130 can be configured to identify the objects 512, 515 in the image 510, generate masks for the objects… identify a plurality of object instance 512, 515 in the image 510, where the plurality of object instances may be determined through instance segmentation of the image… a mask network is used to provide an instance mask prediction… based on the features generated by an image encoder… the image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… an instance mask can be generated for each of the object instances 512, 515, where the instance masks specify the visible portions of each of the objects 512, 515. The instance masks can be generated by a trained mask network, that can distinguish separate objects, where the instance masks can be generated based on the output of the instance segmentation… segmentation mask can be predicted and generated by a diffusion model of the mask network. The segmentation mask may indicate the visible region and the occluded region of the object… a diffusion model can generate a complete image 590 of the occluded object 515, and provide the amodal segmentation mask… FIG. 7 is an illustrative depiction of a diagram of an image processing system to perform amodal instance segmentation using a mask network having a segmentation network and a diffusion model… For example, the image processing system depicted in FIG. 7 may be an example of a mask network including the mask component 240 of FIG. 2… user 110 can submit the initial image 510… to an image processing system 130. The image processing system 130 can include a mask network having an image encoder 720 that can identify multiple objects of the same class as distinct individual objects or instances 512, 515 in the image 510. The image encoder 720 can generate features that include information… The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130… the image can be divided into fragments, for example, patches of a predetermined size. The image encoder 720 can generate one or more feature maps 740, which may be semantic rich feature maps(s). Feature maps generated by a convolution operation can represent the spatial arrangement of the activations in a convolutional layer; Par. [0082-97]: the image processing system 130 can include a transformer decoder 730 that can be configured to generate masks, tokens, labels, and confidence scores. Tokens can represent a patch of an image, where a token can be created for each small patch of the image. The image patch can be mapped to a feature vector referred to as the token. An output from the image encoder 720, including multi-level features, can be input to the transformer decoder 730, along with a set 732 of input tokens 734… the input token set 732 can have one input token 734 generated for each of the individual objects or instances 512, 515 in the image 510…. the output token set 736 can have one output token 738 generated for each of the individual objects or instances 512, 515 in the image 510… the neural network layers 750 can include a first head configured to generate an instance mask… the first kernel 770 and the second kernel 775 can be applied to the feature maps 740, where the first kernel 770 applied to the feature maps 740 can generate a positive mask for the visible region 780 of the occluded object 515… The first kernel 770 can generate the positive mask through convolution applied to the feature map(s) 740… The diffusion model can determine the shape and size of the occluded object from instance masks and the occlusion mask based on the different regions… FIG. 8 is an illustrative depiction of a method of mask determination… instance masks 810, 820 can be generated by a mask network; and the first mask data is further based on the second group of feature sets (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), as indicated above, for example, including generating different masks identifying different regions and objects in an image, respectively, by using kernels that generate masks through convolution applied to the generated features (i.e. the first, second, third… Nth mask data is further based on the first, second, third… Nth group of feature sets), as indicated above, for example). Regarding claim 13, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), wherein the one or more processors are configured to generate, based on a group of feature sets from at least one sampling iteration of second sampling iterations associated with the diffusion model, second mask data that indicates a second mask associated with a second object of a second image (Par. [0004-7]: systems and methods for using a machine learning to perform instance segmentation on an image… generating a segmentation mask for the input image… using a diffusion model… a mask network configured to generate an instance mask for an input image; Par. [0025]: a neural network model, referred to as a mask model, is provided that can identify the visible parts of objects in an image; Par. [0040-43]: image generator 200 can include… a mask component 240, a noise component 250, and a diffusion model 260… first machine learning model generates an output mask that is then used as input for a second machine learning model, e.g., diffusion model 260. An example of the diffusion model 260 is provided with reference to FIG. 3… the mask component 240 can generate a mask, which is a set of binary values for a region of interest in the image… mask component 240 identifies a mask indicating the region of the original image. A mask can be used to specify a region of an image that contains an object of interest based identified through object detection. Different masks can identify different regions and objects in an image; Par. [0068-77]: image processing system 130 can be configured to identify the objects 512, 515 in the image 510, generate masks for the objects… identify a plurality of object instance 512, 515 in the image 510, where the plurality of object instances may be determined through instance segmentation of the image… a mask network is used to provide an instance mask prediction… based on the features generated by an image encoder… the image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… an instance mask can be generated for each of the object instances 512, 515, where the instance masks specify the visible portions of each of the objects 512, 515. The instance masks can be generated by a trained mask network, that can distinguish separate objects, where the instance masks can be generated based on the output of the instance segmentation… segmentation mask can be predicted and generated by a diffusion model of the mask network. The segmentation mask may indicate the visible region and the occluded region of the object… a diffusion model can generate a complete image 590 of the occluded object 515, and provide the amodal segmentation mask… FIG. 7 is an illustrative depiction of a diagram of an image processing system to perform amodal instance segmentation using a mask network having a segmentation network and a diffusion model… For example, the image processing system depicted in FIG. 7 may be an example of a mask network including the mask component 240 of FIG. 2… user 110 can submit the initial image 510… to an image processing system 130. The image processing system 130 can include a mask network having an image encoder 720 that can identify multiple objects of the same class as distinct individual objects or instances 512, 515 in the image 510. The image encoder 720 can generate features that include information… The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130… the image can be divided into fragments, for example, patches of a predetermined size. The image encoder 720 can generate one or more feature maps 740, which may be semantic rich feature maps(s). Feature maps generated by a convolution operation can represent the spatial arrangement of the activations in a convolutional layer; Par. [0082-97]: the image processing system 130 can include a transformer decoder 730 that can be configured to generate masks, tokens, labels, and confidence scores. Tokens can represent a patch of an image, where a token can be created for each small patch of the image. The image patch can be mapped to a feature vector referred to as the token. An output from the image encoder 720, including multi-level features, can be input to the transformer decoder 730, along with a set 732 of input tokens 734… the input token set 732 can have one input token 734 generated for each of the individual objects or instances 512, 515 in the image 510…. the output token set 736 can have one output token 738 generated for each of the individual objects or instances 512, 515 in the image 510… the neural network layers 750 can include a first head configured to generate an instance mask… the first kernel 770 and the second kernel 775 can be applied to the feature maps 740, where the first kernel 770 applied to the feature maps 740 can generate a positive mask for the visible region 780 of the occluded object 515… The first kernel 770 can generate the positive mask through convolution applied to the feature map(s) 740… The diffusion model can determine the shape and size of the occluded object from instance masks and the occlusion mask based on the different regions… FIG. 8 is an illustrative depiction of a method of mask determination… instance masks 810, 820 can be generated by a mask network; generate, based on a group of feature sets from at least one sampling iteration of second sampling iterations associated with the diffusion model, second mask data that indicates a second mask associated with a second object of a second image (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), as indicated above, for example, including generating different masks identifying different regions and objects in an image, respectively, by using kernels that generate masks through convolution applied to the generated features (i.e. generate, based on a group of feature sets from at least one sampling iteration of first, second, third… Nth sampling iterations associated with the diffusion model, first, second, third… Nth mask data that indicates a first, second, third… Nth mask associated with a first, second, third… Nth object of a first, second, third… Nth image), as indicated above, for example), wherein the second sampling iterations are configured to generate a latent representation of the second image (Par. [0040-62]: image generator 200 can include… a diffusion model 260… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors… noise component 250 generates an iterative noise map for each of the set of masks with successively reduced noise to produce the output image… diffusion model 260 iteratively produces a set of output images… FIG. 3 shows a block diagram of an example of a guided diffusion model 300… The guided latent diffusion model 300 depicted in FIG. 3 is an example of, or includes aspects of, the corresponding diffusion model element 260 described with reference to FIG. 2… Diffusion models are a class of generative neural networks which can be trained to generate new data with features similar to features found in training data… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301… as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels… Next, a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… an output image 390 is created from each of the various noise levels. The output image 390 can be compared to the original image 301 to train the reverse diffusion process 330… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution and an initial number of channels, and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415. The intermediate features 415 are then down-sampled using a down-sampling layer 420… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435. The up-sampled features 435 can be combined with intermediate features 415 having a same resolution and number of channels via a skip connection 440. These inputs are processed using a final neural network layer 445 to produce output features 450… FIG. 5 shows a diffusion process… As described above with reference to FIG. 3, a diffusion model 300 can include both a forward diffusion process 505 for adding noise to an image (or features in a latent space) and a reverse diffusion process 510 for denoising the images (or features) to obtain a denoised image… the forward diffusion process 505 is used during training to generate images with successively greater noise, and a neural network is trained to perform the reverse diffusion process 510 (i.e., to successively remove the noise)… The neural network may be trained to perform the reverse process. During the reverse diffusion process 510… At each step t−1, the reverse diffusion process 510 takes xt, such as first intermediate image 520, and t as input. Here, t represents a step in the sequence of transitions associated with different noise levels. The reverse diffusion process 510 outputs xt-1, such as second intermediate image 525 iteratively until xT is reverted back to x0, the original image 530; Par. [0068-77]: provide an initial image 510 to an image processing system 130… The image processing system 130 can be configured to identify the objects 512, 515 in the image 510… image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… FIG. 7 is an illustrative depiction of a diagram of an image processing system to perform amodal instance segmentation using… a diffusion model…image processing system 130 can include a mask network having an image encoder 720 that can identify multiple objects of the same class as distinct individual objects or instances 512, 515 in the image 510. The image encoder 720 can generate features… The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130…The image encoder 720 can generate one or more feature maps 740; wherein the second sampling iterations are configured to generate a latent representation of the second image (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), for example, in which the intermediate features are down-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using a down-sampling layer and up-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using an up-sampling process to obtain up-sampled features (i.e. wherein the first, second, third… Nth sampling iterations are configured to generate a latent representation of the first, second, third… Nth image), as indicated above), for example). Regarding claim 15, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), further comprising: an input device coupled to the one or more processors, wherein: the one or more processors are configured to receive, from the input device, an input (Par. [0040]: image generator 200 can include a computer system 280 including one or more processors 210, computer memory 220, an object detector 230, a mask component 240, a noise component 250, and a diffusion model 260. The computer system 280 of the image generator 200 can be operatively coupled to a display device 290 (e.g., computer screen) for presenting prompts and images to a user 110, and operatively coupled to input devices to receive input from the user; Par. [0129]: I/O interface 1340 is controlled by an I/O controller to manage input and output signals for computing device 1340) that indicates an object type of the first object (Par. [0040-59]: mage generator 200 can include a computer system 280 including one or more processors 210, computer memory 220, an object detector 230, a mask component 240, a noise component 250, and a diffusion model 260. The computer system 280 of the image generator 200 can be operatively coupled to a display device 290 (e.g., computer screen) for presenting prompts and images to a user 110, and operatively coupled to input devices to receive input from the user… reverse diffusion process 330 can also be guided based on a text prompt or description 340, or another guidance prompt, such as an image, a layout, a segmentation map, etc. The text prompt 340 can be encoded using a text encoder 350 (e.g., a multimodal encoder) to obtain guidance features 360 in guidance space 370. The guidance features 360 can be combined with the noisy images 320 at one or more layers of the reverse diffusion process 330 to ensure that the output image 390 includes content described by the text prompt 340… U-Net 400 takes additional input features to produce conditionally generated output. For example, the additional input features could include a vector representation of an input prompt; Par. [0107]: diffusion model can also receive the class of the occluded object as a text prompt or description that can guide the generation); and the diffusion model is configured to generate, based on the object type of the first object, the latent representation of the first image including the first object (Par. [0040-59]: mage generator 200 can include a computer system 280 including one or more processors 210, computer memory 220, an object detector 230, a mask component 240, a noise component 250, and a diffusion model 260. The computer system 280 of the image generator 200 can be operatively coupled to a display device 290 (e.g., computer screen) for presenting prompts and images to a user 110, and operatively coupled to input devices to receive input from the user… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors… the conditioning information can be a text prompt (e.g., TP, “ ”, and AP are text prompts)… guided latent diffusion model 300 depicted in FIG. 3 is an example of, or includes aspects of, the corresponding diffusion model element 260… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… The reverse diffusion process 330 can also be guided based on a text prompt or description 340, or another guidance prompt, such as an image, a layout, a segmentation map, etc. The text prompt 340 can be encoded using a text encoder 350 (e.g., a multimodal encoder) to obtain guidance features 360 in guidance space 370. The guidance features 360 can be combined with the noisy images 320 at one or more layers of the reverse diffusion process 330 to ensure that the output image 390 includes content described by the text prompt 340… U-Net 400 takes additional input features to produce conditionally generated output. For example, the additional input features could include a vector representation of an input prompt) Regarding claim 16, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), wherein the one or more processors are configured to: generate an input latent representation based on an encoded image and noise data (Par. [0052-53]: Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301 in a pixel space 305 as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels; Par. [0070]: image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515); use the diffusion model to process the input latent representation to generate the latent representation of the first image (Par. [0052-60]: Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301 in a pixel space 305 as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405… and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415… a diffusion model 300 can include both a forward diffusion process 505 for adding noise to an image (or features in a latent space) and a reverse diffusion process 510 for denoising the images (or features) to obtain a denoised image); use a mask decoder to generate the first mask data based on the first group of feature sets (Par. [0069]: a mask network is used to provide an instance mask prediction… based on the features generated by an image encoder in the first step; Par. [0082]: image processing system 130 can include a transformer decoder 730 that can be configured to generate masks… An output from the image encoder 720, including multi-level features, can be input to the transformer decoder 730); and update one or more parameters of the mask decoder based on a comparison of the first mask data and training mask data, the training mask data indicating a mask associated with a representation of the first object in the encoded image (Par. [0052]: Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion); Par. [0070]: image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515; Par. [0115-120]: method of training a mask network, according to aspects of the present disclosure. One or more machine learning models including a mask network and a diffusion model can be trained through supervised learning. The mask training method 1200 can involve adjusting parameters of transformers, encoders, decoders, deep neural networks, and diffusion models based on error scores between ground truth masks and predicted segmentation masks… a mask network can generate an instance mask for the training image… to train the mask network, a first head of the network can generate an instance mask for each object and a second head can generate an occlusion mask for the union of objects. The generated instance mask for each object can be compared to the ground truth instance masks, and an error calculated. The generated occlusion mask can be compared to the ground truth occlusion mask, and an error calculated. The errors for the generated instance masks and generated occlusion mask can be back-propagated through the mask network to update the parameters of the mask network… a predicted segmentation mask can be generated for the training image by the mask network… the predicted amodal segmentation mask can be compared to the ground truth segmentation mask, and an error score calculated… the parameters of the mask network can be updated, for example, through back propagation based on the comparison, to more closely match the expected position and area of the masks and amodal segmentation mask. Updating the parameters of the diffusion model of the mask network based on the comparison can train the diffusion model to generate more accurate segmentation masks). Regarding claim 17, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), further comprising a modem coupled to the one or more processors (Par. [0126-129]: computing device 1300 includes one or more processors 1310… a processor 1310 includes special purpose components for modem processing… I/O interface 1340 is controlled by an I/O controller to manage input and output signals for computing device 1340… the I/O controller represents or interacts with a user interface component, including… a modem), the modem configured to transmit the latent representation of the first image and the first mask data (Par. [0025]: a neural network model, referred to as a mask model, is provided that can identify the visible parts of objects in an image; Par. [0040-43]: image generator 200 can include… a mask component 240, a noise component 250, and a diffusion model 260… first machine learning model generates an output mask that is then used as input for a second machine learning model, e.g., diffusion model 260. An example of the diffusion model 260 is provided with reference to FIG. 3… the mask component 240 can generate a mask, which is a set of binary values for a region of interest in the image… mask component 240 identifies a mask indicating the region of the original image. A mask can be used to specify a region of an image that contains an object of interest based identified through object detection. Different masks can identify different regions and objects in an image; Par. [0126-129]: computing device 1300 includes one or more processors 1310… a processor 1310 includes special purpose components for modem processing… transmission processing… communication interface 1350 operates at a boundary between communicating entities (such as computing device 1300, one or more user devices, a cloud, and one or more databases) and channel (e.g., bus) 1330 and can record and process communications. … communication interface 1350 is provided to enable a processing system coupled to a transceiver (e.g., a transmitter and/or a receiver) … the I/O controller represents or interacts with a user interface component, including… a modem). Regarding claim 18, is a corresponding method claim rejected as applied to the apparatus claim 1 above. Regarding claim 19, claim 18 is incorporated and Zhang discloses the device (Par. [0004-7]), further comprising using the diffusion model to process an input latent representation of noise data to generate the latent representation of the first image (Par. [0044-48]: noise component 250 can generate Gaussian noise for the diffusion model 260… noise component 250 generates a noise map based on the original image and the mask, where the amodal segmentation mask is generated based on the noise map… noise component 250 generates an iterative noise map for each of the set of masks with successively reduced noise to produce the output image… the noise component 250 is an element of the diffusion model 260; Par. [0052-53]: Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301 in a pixel space 305 as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels), the noise data sampled from a noise distribution (Par. [0044-46]: the noise component 250 can generate Gaussian noise for the diffusion model 260. Gaussian noise is signal noise that has a probability density function (PDF) equal to that of the normal distribution, also referred to as a Gaussian distribution… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors). Regarding claim 20, Zhang discloses a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to (Par. [0004-7]: systems and methods for using a machine learning to perform instance segmentation on an image… method, apparatus, non-transitory computer readable medium, and system for amodal segmentation mask generation… method, apparatus, non-transitory computer readable medium, and system include identifying a training image and a ground truth segmentation mask for the training image… the apparatus and system include… a processor; a memory containing instructions executable by the processor; Par. [0039-40]: FIG. 2 shows a block diagram of an example of an image generator according to aspects of the present disclosure… an image generator 200 receives an original image… the image generator 200 can include a computer system 280 including one or more processors 210, computer memory 220; Par. [0124-134]: FIG. 13 shows an example of a computer system according to aspects of the present disclosure… the computing device includes processor(s) 1310, memory subsystem 1320… The computer system 1300 can be configured to perform the operations described above and illustrated in FIG. 1-12… computing device 1300 includes one or more processors 1310… a processor 1310 is configured to operate a memory array using a memory controller… a processor is configured to execute computer-readable instructions stored in a memory to perform various functions… memory subsystem 1320 includes one or more memory devices… memory is used to store computer-readable, computer-executable software including instructions that, when executed, cause a processor to perform various functions described herein… Computer-readable media includes both non-transitory computer storage media and communication media including any medium that facilitates transfer of code or data. A non-transitory storage medium may be any available medium that can be accessed by a computer. For example, non-transitory computer-readable media can comprise random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disk (CD) or other optical disk storage, magnetic disk storage, or any other non-transitory medium for carrying or storing data or code; a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models include a computing device which includes one or more processors and a non-transitory computer readable medium containing computer-readable software instructions that, when executed, cause a processor to perform various functions (i.e. a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to), as indicated above), for example). The steps of the computer-readable software program further recited in claim 20 correspond to claim 1 when executed and are rejected as applied to the apparatus claim 1 above. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as applied to claim 1 above, in view of Saharia et al. (PG Publication No. 2023/0103638 A1), hereafter referred to as Saharia. Regarding claim 2, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), but fails to teach the following as further recited in claim 2. However, Saharia teaches wherein the first sampling iteration corresponds to a final sampling iteration of the multiple sampling iterations (Par. [0003-4]: image processing system that can process a noisy image to generate a denoised version of the noisy image… iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image; updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data; and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image… providing, by the computing device, the trained multi-task diffusion model; Par. [0042-72]: iterative refinement process enables the image processing system described herein to generate higher quality outputs… Diffusion-based models also may be used for image generation… diffusion models convert samples from a standard Gaussian distribution into samples from an empirical data distribution through an iterative denoising process… multi-task diffusion model 120 generates a forward diffusion process 160 by iteratively adding noise to the denoised version. After generating the forward diffusion process 160, the multi-task diffusion model 120 learns a reverse diffusion process 170 that can be applied to denoise an image… multi-task diffusion model 120 is a neural network… at a first iteration, a current noisy estimate of the denoised version of the noisy image may be initialized to generate an initial estimate of the noisy image. Also, for example, noise data may be sampled from a predetermined noise distribution. The method involves iteratively generating the forward diffusion process 160 by predicting, at each iteration in a sequence of iterations (e.g., T iterations), and based on a current noisy estimate 140 of the denoised version of the noisy image, noise data 130 to predict a next noisy estimate 150 of the denoised version of the noisy image… In a subsequent iteration, the next noisy estimate 150 is re-initialized as current noisy estimate 140, and provided as input to the multi-task diffusion model 120, which then predicts updated noise data 130… The iterative process may continue until a desired next noisy estimate 150 is achieved… After the multi-task diffusion model 120 generates the forward diffusion process 160 based on the iterative process outlined above, the multi-task diffusion model 120 learns the reverse diffusion process 170 by inverting the forward diffusion process 160. Accordingly, a trained multi-task diffusion model 120 can be configured to predict the denoised version of the noisy image… the sampling process can start at pure Gaussian noise, followed by T steps of iterative refinement… each iteration of the reverse process may be computed… This process may be iterated repeatedly to produce the final predicted denoised image, ŷ0… The training of the neural network may include applying a forward Gaussian diffusion process that adds Gaussian noise to a corresponding target version of each of the plurality of pairs of images to enable iterative denoising of the input image, wherein the iterative denoising is based on a reverse Markov chain associated with the forward Gaussian diffusion process; Par. [0153-158]: training, based on the training data, a multi-task diffusion model to perform a plurality of image-to-image translation tasks, wherein the training comprises: iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image… the method involves providing, by the computing device, the trained multi-task diffusion model… embodiments involve sampling, at a first iteration of the sequence of iterations, an initial noise data from a predefined noise distribution… each iteration in the sequence of iterations is associated with a respective noise level parameter, and the predicting of the noise data at each iteration is based on the noise level parameter associated with the iteration… for each iteration in the sequence of iterations, the updating of the current noisy estimate to the next noisy estimate is performed by combining the predicted noise data with the current estimate in accordance with the noise level parameter associated with the iteration… for each iteration prior to the final iteration in the sequence of iterations, the updating of the current estimate includes sampling additional noise data from a predefined noise distribution; wherein the first sampling iteration corresponds to a final sampling iteration of the multiple sampling iterations (e.g. image processing system that processes noisy images to generate denoised versions of the noisy images by using diffusion models to convert samples from a standard Gaussian distribution into samples from an empirical data distribution through an iterative denoising process, for example, include an iterative denoising process that is iterated repeatedly (i.e. first, second, third… Nth sampling iteration) to produce a final predicted denoised image (i.e. first, second, third… Nth sampling iteration corresponds to a final sampling iteration of the multiple sampling iterations), for example, by sampling, at a first, second, third… Nth iteration of the sequence of iterations, initial noise data from a predefined noise distribution followed by T steps of iterative refinement, for example, including iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of a denoised version of a noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image, as indicated above), for example). Zhang and Saharia are considered to be analogous art because they pertain to image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, including a guided latent diffusion model that iteratively produces a set of output images (as disclosed by Zhang) with wherein the first sampling iteration corresponds to a final sampling iteration of the multiple sampling iterations (as taught by Saharia, Abstract, Par. [0003-4, 42-72, 153-158]) to generate a sharp image, to predict the denoised version of the noisy image, to recognize whether an input image has image degradations and to apply a trained neural network to remove the image degradations in the input image, and to improve images by removing image degradations, thereby enhancing their actual and/or perceived quality (Saharia, Abstract, Par. [0003-4, 41-72, 153-158]). Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as applied to claim 1 above, in view of Pakhomov et al. (PG Publication No. 2024/0135514 A1), hereafter referred to as Pakhomov. Regarding claim 4, claim 1 is incorporated and Zhang discloses the device (Par. [0004-7]), wherein: the diffusion model includes a downsampling stage; and each feature set of the first group of feature sets corresponds to each downsampling stage of the diffusion model (Par. [0054-59]: a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… guidance features 360 can be combined with the noisy images 320 at one or more layers of the reverse diffusion process 330 to ensure that the output image 390 includes content described… FIG. 4 shows an example of a U-Net 400 according to aspects of the present disclosure. The U-Net 400 depicted in FIG. 4 is an example of, or includes aspects of, the architecture used within the reverse diffusion process described with reference to FIG. 3… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution and an initial number of channels, and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415. The intermediate features 415 are then down-sampled using a down-sampling layer 420 such that down-sampled features 425 features have a resolution less than the initial resolution… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435… U-Net 400 takes additional input features to produce conditionally generated output. For example, the additional input features could include a vector representation of an input prompt. The additional input features can be combined with the intermediate features 415 within the neural network at one or more layers; wherein: the diffusion model includes a downsampling stage; and each feature set of the first group of feature sets corresponds to each downsampling stage of the diffusion model (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), for example, in which the intermediate features are down-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using a down-sampling layer (i.e. diffusion model includes a downsampling stage and each feature set of the first group of feature sets corresponds to each downsampling stage of the diffusion model) and up-sampled (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations) using an up-sampling process to obtain up-sampled features, as indicated above), for example), but fails to teach the following as further recited in claim 4. However, Pakhomov teaches multiple downsampling stages; and a respective downsampling stage of the multiple downsampling stages (Par. [0127-128]: object detection machine learning model 308 includes lower neural network layers and higher neural network layers. In general, the lower neural network layers collectively form the encoder 302 and the higher neural network layers collectively form the detection heads 304 (e.g., decoder)… the encoder 302 includes convolutional layers that encodes a digital image into feature vectors, which are outputted from the encoder 302 and provided as input to the detection heads 304… the detection heads 304 comprise fully connected layers that analyze the feature vectors and output the detected objects (potentially with approximate boundaries around the objects)… the encoder 302, in one or more implementations, comprises convolutional layers that generate a feature vector in the form of a feature map. To detect objects within the digital image 316, the object detection machine learning model 308 processes the feature map utilizing a convolutional layer in the form of a small network that is slid across small windows of the feature map. The object detection machine learning model 308 further maps each sliding window to a lower-dimensional feature… the object detection machine learning model 308 processes this feature using two separate detection heads that are fully connected layers; Par. [0597] FIG. 46B illustrates an example architecture of a diffusion neural network implemented by the scene-based image editing system 106 for multi-layered scene completion in accordance with one or more embodiments… the diffusion neural network 4602 comprises a modified U-Net architecture as shown. Specifically, the diffusion neural network 4602 comprises upsampling and downsampling layers and skip connections; multiple downsampling stages (i.e. layers, levels, etc.); and a respective downsampling stage of the multiple downsampling stages (e.g. object detection machine learning model includes convolutional layers that generate a feature vector in the form of a feature map, for example, including a diffusion neural network for multi-layered scene completion that comprises a modified U-Net architecture as shown in Fig. 46B, for example, in which the diffusion neural network comprises upsampling and downsampling layers (i.e. multiple downsampling stages and a respective downsampling stage of the multiple downsampling stages) and skip connections), for example). Zhang and Pakhomov are considered to be analogous art because they pertain to image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, including a guided latent diffusion model that iteratively produces a set of output images (as disclosed by Zhang) with multiple downsampling stages; and a respective downsampling stage of the multiple downsampling stages (as taught by Pakhomov, Abstract, Par. [0002-3, 127-128, 161, 597]) to provide a variety of image-related tasks, such as object identification, classification, segmentation, composition, style transfer, image inpainting, etc., to generate more accurate inpainted digital images, to accurately predict attributes, to offer a robust set of objects that can be completed for the modification of digital images, and to reduce user interactions that would typically be required under conventional systems for completing objects and portions of the background occluded by those objects by implementing multi-layered scene completion (Pakhomov, Abstract, Par. [0127-128, 315, 597, 615]). Claim 7, 9-12, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, as applied to claim 1 above, in view of KULAL et al. (PG Publication No. 2024/0169701 A1), hereafter referred to as KULAL, Applicant cited prior art furnished via IDS. Regarding claim 7, claim 6 is incorporated and Zhang discloses the device (Par. [0004-7]), but fails to disclose the following as further erecited in claim 7. However, KULAL teaches wherein the one or more processors are configured to concatenate the input feature sets to generate the aggregated feature set (Par. [0004-7]: Systems, methods, and software are described herein for using a diffusion mode… an apparatus for inserting an object into a background includes: one or more processors and one or more memories including instructions executable by the one or more processors to: obtain an object image depicting the object and a background image including a region for inserting the object; encode, using an image encoder, the object image to obtain an encoded object; encode, using a condition encoder, the background image to obtain an encoded background; and generate, using a diffusion model, a modified image based on the encoded object and the encoded background, wherein the modified image depicts the object within the region; Par. [0062]: method for training the diffusion model according to aspects of the present disclosure… object image 122 may be output to a multi-modal encoder 148 of the model to generate conditioning features to inpaint the image… The latent features 142, the re-scaled mask 146, and a noise map 144 may be concatenated together for output to a De-noising network 152 (e.g., a time-conditional U-Net), where the conditioning features are passed to the network 152 via cross-attention to generate predicted latent features 156. The noise map 144 may be obtained by noising (e.g., adding noise) the features 154 in FIG. 5A. The mask 117 and the masked background image 119 may be concatenated as they are spatially aligned with the final output; Par. [0091-99]: method of FIG. 8A further includes encoding the masked background image to generate first features of a first dimension (S804). For example, the encoder 265 of FIG. 2 could generate the first feature…The method of FIG. 8A further includes concatenating noise to the first features to generate second features of a second dimension (S805). The noise may be one of a plurality of different noise maps. The second dimension is higher than the first dimension since rather than change values of the first features, additional channels of noise are concatenated after the first features… method of FIG. 9 further includes generating first guidance features from the selected object (S903). For example, the encoder 285 may generate the first guidance features… method of FIG. 9 further includes adding to noise to an encoding of the masked background image 119 to generate second guidance features (S904). The noise added to the encoding may be one of a plurality of different noise maps, where each different noise map can contribute to generating a different final modified image. For example, the encoder 265 of FIG. 2 may encode the masked background image 119 and concatenate the noise to the encoded mask background image; wherein the one or more processors are configured to concatenate the input feature sets to generate the aggregated feature set (e.g. systems, methods, and software for using a diffusion mode include concatenating noise to first features to generate second features of a second dimension, and the noise includes one of a plurality of different noise maps, for example, and including additional channels of noise which are concatenated after the first features (i.e. concatenate the input feature sets to generate the aggregated feature set), as indicated above), for example). Zhang and KULAL are considered to be analogous art because they pertain to image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, including a guided latent diffusion model that iteratively produces a set of output images (as disclosed by Zhang) with wherein the one or more processors are configured to concatenate the input feature sets to generate the aggregated feature set (as taught by KULAL, Abstract, Par. [0004-7, 62, 91-99]) to insert, a background, and potentially other objects, and all of the pixels of an object image except for the object to insert or part of the object to insert are removed from the preliminary object image to generate the object image, to provide a mask input using a client user interface to crop out all or part of the object to insert, to divide the preliminary object image into one or more objects, to selects one of the objects resulting from the segmentation, and to generate an object image from the selected one object, to generate a modified image from received data using a previously trained Diffusion model (KULAL, Abstract, Par. [0002-7, 21-30, 62, 91-99]). Regarding claim 9, claim 1 is incorporated and the combination of Zhang and KULAL, as a whole, teaches the device (Zhang, Par. [0004-7]), wherein the one or more processors are configured to: obtain a background image (KULAL, Par. [0005]: method of inserting an object into a background includes: obtaining a background image… encoding the background image to obtain an encoded background… and generating a modified image based on the encoded background using a diffusion model; Par. [0042-43]: FIG. 2 illustrates the Diffusion model (e.g., 200) according to an exemplary embodiment of the disclosure… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… the Diffusion model 200 is a latent diffusion model since noise is added to image features generated by an encoder. The Diffusion model 200 iteratively adds noise to data input during a forward process and then learns to recover the data by denoising the data during a reverse process. For example, during training, the Diffusion model 200 takes a background image 118 in a pixel space 210 as input and applies a forward diffusion process 230 to gradually add noise to the background image 118 to obtain noisy images 235 at various noise level; Par. [0104]: generate a masked background image 119 from a background image 118, the segmentation component 825 is used to generate an object image 122 from a preliminary object image, and the diffusion component 815 generates a modified image 124 using the Diffusion model 200 based on the object image 122, and the masked background image 119); and generate, based on the first image and the first mask data, an output image that includes a representation of the first object and at least a portion of the background image (KULAL, Par. [0031-47]: FIG. 1 shows an example of the client-server environment, where a user uses a graphical user interface 112 of a client device 110 to create a modified image 124 from a masked background image 119 generated from masking out a region of a background image, and an object image 122… the user interface 112 enables the user to select the background image from a list of available images or use a camera 115 to capture the background image… mark a region of the background image for inserting the object, and generate the masked background image 119 by masking out the region from the background image… the server interface 114 outputs the masked background image 119… the server 130 forwards the received data (e.g., masked background image 119 and the object image 122 when present) to an image generator 134. The image generator 134 generates a modified image 124 from the received data using a previously trained Diffusion model… client device 110 includes one or more processors, and one or more computer-readable media. The computer-readable media may include computer-readable instructions executable by the one or more processors. The instructions may correspond to one or more applications, such as software to manage the graphical user interface 112, software to output the data (e.g., masked background image 119, and object image 122 when present)… computing device includes one or more processors, and one or more computer-readable media. The computer-readable media may include computer-readable instructions executable by the one or more processors. The instructions may correspond to one or more applications, such as software to interface with the client device 110 for receiving the data (e.g., the masked image background 119, and the object image 122 when present) and outputting the modified image 124… FIG. 2 illustrates the Diffusion model (e.g., 200)… Diffusion models are a class of generative artificial neural network that can be trained to generate new data with features similar to features found in training data… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… the Diffusion model 200 is a latent diffusion model since noise is added to image features generated by an encoder. The Diffusion model 200 iteratively adds noise to data input during a forward process and then learns to recover the data by denoising the data during a reverse process. For example, during training, the Diffusion model 200 takes a background image 118 in a pixel space 210 as input and applies a forward diffusion process 230 to gradually add noise to the background image 118 to obtain noisy images 235 at various noise levels. During training, a mixture of masks may be used to hide all or part of objects within the background image 118. The background image 118 may include various objects… For fixed time steps T, the forward diffusion process 230 gradually adds noise and approximates samples at t=T as uniform Gaussian noise… Next, a reverse diffusion process 240 (e.g., a U-Net ANN) gradually removes the noise from the noisy images 235 at the various noise levels to obtain the modified image 124. The modified image 124 can be compared to the background image 118 to train the reverse diffusion process 240; Par. [0104]: generate a masked background image 119 from a background image 118, the segmentation component 825 is used to generate an object image 122 from a preliminary object image, and the diffusion component 815 generates a modified image 124 using the Diffusion model 200 based on the object image 122, and the masked background image 119; obtain a background image and generate, based on the first image and the first mask data, an output image that includes a representation of the first object and at least a portion of the background image (e.g. systems, methods, and software for using a diffusion mode include obtaining a background image (i.e. obtain a background image) to generate a masked background image from masking out a region (i.e. at least a portion of) of the background image, for example, and forwards (i.e. transmits, sends, etc.) received data, including the masked background image and an object image (i.e. based on the first image and the first mask data) when present, to an image generator which generates a modified image (i.e. a representation of the first object and at least a portion of the background image) from the received data using a previously trained Diffusion model (i.e. obtain a background image and generate, based on the first image and the first mask data, an output image that includes a representation of the first object and at least a portion of the background image), as indicated above), for example). The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 7. Regarding claim 10, claim 9 is incorporated and the combination of Zhang and KULAL, as a whole, teaches the device (Zhang, Par. [0004-7]), further comprising a camera coupled to the one or more processors (KULAL, Par. [0113-115]: computing device 1000 includes bus 1010 that directly or indirectly couples the following devices: memory 1012, one or more processors 1014… computing device 1000 may be equipped with depth cameras, such as, stereoscopic camera systems, infrared camera systems, RGB camera systems, and combinations of these), wherein the camera is configured to generate the background image (KULAL, Par. [0033]: user interface 112 enables the user to select the background image from a list of available images or use a camera 115 to capture the background image). The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 7. Regarding claim 11, claim 9 is incorporated and Zhang discloses the device (Par. [0004-7]), further comprising a display device coupled to the one or more processors, wherein the display device is configured to display the output image (Par. [0040]: image generator 200 can include a computer system 280 including one or more processors 210, computer memory 220, an object detector 230, a mask component 240, a noise component 250, and a diffusion model 260. The computer system 280 of the image generator 200 can be operatively coupled to a display device 290 (e.g., computer screen) for presenting prompts and images to a user 110; Par. [0130]: user interface component(s) 1360 enable a user to interact with computing device 1300… user interface component(s) 1360 include… an external display device such as a display screen). Regarding claim 12, claim 11 is incorporated and Zhang discloses the device (Par. [0004-7]), further comprising a speaker coupled to the one or more processors, wherein the speaker is configured to, concurrently with the output image being displayed at the display device, output audio associated with the first object (Par. [0040-42]: image generator 200 can include a computer system 280 including one or more processors 210, computer memory 220, an object detector 230, a mask component 240, a noise component 250, and a diffusion model 260. The computer system 280 of the image generator 200 can be operatively coupled to a display device 290 (e.g., computer screen) for presenting prompts and images to a user 110… Image segmentation partitions a digital image into multiple image objects or regions and locates the objects, regions, and boundaries within the image; Par. [0130]: user interface component(s) 1360 enable a user to interact with computing device 1300… user interface component(s) 1360 include an audio device, such as an external speaker system, an external display device such as a display screen). Regarding claim 14, claim 13 is incorporated and the combination of Zhang and KULAL, as a whole, teaches the device (Zhang, Par. [0004-7]), wherein: the one or more processors are configured to generate an output image including a representation of the first object, a representation of the second object (Zhang, Par. [0040-62]: image generator 200 can include… a diffusion model 260… Diffusion models are a class of generative models that convert Gaussian noise into images from a learned data distribution using an iterative denoising process. Diffusion models are also latent variable models with latent vectors… noise component 250 generates an iterative noise map for each of the set of masks with successively reduced noise to produce the output image… diffusion model 260 iteratively produces a set of output images… FIG. 3 shows a block diagram of an example of a guided diffusion model 300… The guided latent diffusion model 300 depicted in FIG. 3 is an example of, or includes aspects of, the corresponding diffusion model element 260 described with reference to FIG. 2… Diffusion models are a class of generative neural networks which can be trained to generate new data with features similar to features found in training data… Diffusion models may also be characterized by whether the noise is added to the image itself, or to image features generated by an encoder (i.e., latent diffusion)… Diffusion models work by iteratively adding noise to the data during a forward process and then learning to recover the data by denoising the data during a reverse process. For example, during training, guided latent diffusion model 300 may take an original image 301… as input and apply forward diffusion process 310 to gradually add noise to the original image 305 to obtain noisy images 320 at various noise levels… Next, a reverse diffusion process 330 (e.g., a U-Net Artificial Neural Network (ANN)) gradually removes the noise from the noisy images 320 at the various noise levels to obtain an output image 390… an output image 390 is created from each of the various noise levels. The output image 390 can be compared to the original image 301 to train the reverse diffusion process 330… diffusion models are based on a neural network architecture known as a U-Net. The U-Net 400 takes input features 405 having an initial resolution and an initial number of channels, and processes the input features 405 using an initial neural network layer 410 (e.g., a convolutional network layer) to produce intermediate features 415. The intermediate features 415 are then down-sampled using a down-sampling layer 420… This process is repeated multiple times, and then the process is reversed. That is, the down-sampled features 425 are up-sampled using up-sampling process 430 to obtain up-sampled features 435. The up-sampled features 435 can be combined with intermediate features 415 having a same resolution and number of channels via a skip connection 440. These inputs are processed using a final neural network layer 445 to produce output features 450… FIG. 5 shows a diffusion process… As described above with reference to FIG. 3, a diffusion model 300 can include both a forward diffusion process 505 for adding noise to an image (or features in a latent space) and a reverse diffusion process 510 for denoising the images (or features) to obtain a denoised image… the forward diffusion process 505 is used during training to generate images with successively greater noise, and a neural network is trained to perform the reverse diffusion process 510 (i.e., to successively remove the noise)… The neural network may be trained to perform the reverse process. During the reverse diffusion process 510… At each step t−1, the reverse diffusion process 510 takes xt, such as first intermediate image 520, and t as input. Here, t represents a step in the sequence of transitions associated with different noise levels. The reverse diffusion process 510 outputs xt-1, such as second intermediate image 525 iteratively until xT is reverted back to x0, the original image 530; generate an output image including a representation of the first object, a representation of the second object (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. generate an output image including a representation of the first, second, third… Nth object), as indicated above), for example), and at least a portion of a background image (KULAL, Par. [0005]: method of inserting an object into a background includes: obtaining a background image… encoding the background image to obtain an encoded background… and generating a modified image based on the encoded background using a diffusion model; Par. [0043]: the Diffusion model 200 takes a background image 118 in a pixel space 210 as input and applies a forward diffusion process 230 to gradually add noise to the background image 118 to obtain noisy images 235 at various noise levels; Par. [0104]: generate a masked background image 119 from a background image 118, the segmentation component 825 is used to generate an object image 122 from a preliminary object image, and the diffusion component 815 generates a modified image 124 using the Diffusion model 200 based on the object image 122, and the masked background image 119); the representation of the first object is based on the first image and the first mask data; and the representation of the second object is based on the second image and the second mask data (Par. [0004-7]: systems and methods for using a machine learning to perform instance segmentation on an image… generating a segmentation mask for the input image… using a diffusion model… a mask network configured to generate an instance mask for an input image; Par. [0025]: a neural network model, referred to as a mask model, is provided that can identify the visible parts of objects in an image; Par. [0040-43]: image generator 200 can include… a mask component 240, a noise component 250, and a diffusion model 260… first machine learning model generates an output mask that is then used as input for a second machine learning model, e.g., diffusion model 260. An example of the diffusion model 260 is provided with reference to FIG. 3… the mask component 240 can generate a mask, which is a set of binary values for a region of interest in the image… mask component 240 identifies a mask indicating the region of the original image. A mask can be used to specify a region of an image that contains an object of interest based identified through object detection. Different masks can identify different regions and objects in an image; Par. [0068-77]: image processing system 130 can be configured to identify the objects 512, 515 in the image 510, generate masks for the objects… identify a plurality of object instance 512, 515 in the image 510, where the plurality of object instances may be determined through instance segmentation of the image… a mask network is used to provide an instance mask prediction… based on the features generated by an image encoder… the image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515… an instance mask can be generated for each of the object instances 512, 515, where the instance masks specify the visible portions of each of the objects 512, 515. The instance masks can be generated by a trained mask network, that can distinguish separate objects, where the instance masks can be generated based on the output of the instance segmentation… segmentation mask can be predicted and generated by a diffusion model of the mask network. The segmentation mask may indicate the visible region and the occluded region of the object… a diffusion model can generate a complete image 590 of the occluded object 515, and provide the amodal segmentation mask… FIG. 7 is an illustrative depiction of a diagram of an image processing system to perform amodal instance segmentation using a mask network having a segmentation network and a diffusion model… For example, the image processing system depicted in FIG. 7 may be an example of a mask network including the mask component 240 of FIG. 2… user 110 can submit the initial image 510… to an image processing system 130. The image processing system 130 can include a mask network having an image encoder 720 that can identify multiple objects of the same class as distinct individual objects or instances 512, 515 in the image 510. The image encoder 720 can generate features that include information… The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130… the image can be divided into fragments, for example, patches of a predetermined size. The image encoder 720 can generate one or more feature maps 740, which may be semantic rich feature maps(s). Feature maps generated by a convolution operation can represent the spatial arrangement of the activations in a convolutional layer; Par. [0082-97]: the image processing system 130 can include a transformer decoder 730 that can be configured to generate masks, tokens, labels, and confidence scores. Tokens can represent a patch of an image, where a token can be created for each small patch of the image. The image patch can be mapped to a feature vector referred to as the token. An output from the image encoder 720, including multi-level features, can be input to the transformer decoder 730, along with a set 732 of input tokens 734… the input token set 732 can have one input token 734 generated for each of the individual objects or instances 512, 515 in the image 510…. the output token set 736 can have one output token 738 generated for each of the individual objects or instances 512, 515 in the image 510… the neural network layers 750 can include a first head configured to generate an instance mask… the first kernel 770 and the second kernel 775 can be applied to the feature maps 740, where the first kernel 770 applied to the feature maps 740 can generate a positive mask for the visible region 780 of the occluded object 515… The first kernel 770 can generate the positive mask through convolution applied to the feature map(s) 740… The diffusion model can determine the shape and size of the occluded object from instance masks and the occlusion mask based on the different regions… FIG. 8 is an illustrative depiction of a method of mask determination… instance masks 810, 820 can be generated by a mask network; the representation of the first object is based on the first image and the first mask data; and the representation of the second object is based on the second image and the second mask data (e.g. systems and methods for performing instance segmentation and generating segmentation masks for input images using diffusion models, including latent variable models with latent vectors, for example, include a guided latent diffusion model that iteratively (i.e. repeatedly, recurrently, etc.) produces (i.e. generates, creates, constructs, etc.) a set (i.e. first, second, third… Nth) of output images (i.e. first, second, third… Nth sampling iteration of multiple sampling iterations associated with a diffusion model configured to generate a latent representation of a first, second, third… Nth image), for example, including a neural network architecture known as a U-Net that takes input (i.e. first, second, third… Nth sampling) features, including image features generated by a mask network having an image encoder (i.e. latent diffusion) that generates one or more (i.e. first, second, third… Nth) feature maps (i.e. a first, second, third… Nth group of feature sets), for example, by iteratively processing the input features to produce intermediate features and output features (i.e. a first, second, third… Nth feature set of a first, second, third… Nth group of feature sets), as indicated above, for example, including generating different masks identifying different regions and objects in an image, respectively, by using kernels that generate masks through convolution applied to the generated features (i.e. the representation of the first object is based on the first image and the first mask data and the representation of the second object is based on the second image and the second mask data), as indicated above, for example). The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 7. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to GUILLERMO M RIVERA-MARTINEZ whose telephone number is (571) 272-4979. The examiner can normally be reached on 9 am to 5 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached on 571-270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GUILLERMO M RIVERA-MARTINEZ/ Primary Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Sep 20, 2024
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694563
SYSTEM AND METHOD FOR USING DYNAMIC OBJECTS TO ESTIMATE CAMERA POSE
2y 6m to grant Granted Jul 28, 2026
Patent 12682441
CHARACTERIZATION SYSTEM AND METHOD IMPLEMENTING IMAGE ENHANCEMENT FOR IMPROVED DEFECT DETECTION
4y 4m to grant Granted Jul 14, 2026
Patent 12651392
STATIONARY MULTI-SOURCE AI-POWERED REAL-TIME TOMOGRAPHY (SMART)
2y 9m to grant Granted Jun 09, 2026
Patent 12648816
CALCULATING RANGE OF MOTION
3y 1m to grant Granted Jun 09, 2026
Patent 12639785
DEEP LEARNING ROBUSTNESS AGAINST DISPLAY FIELD OF VIEW VARIATIONS
3y 9m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
81%
With Interview (+3.3%)
2y 6m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 511 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month