Prosecution Insights
Last updated: October 02, 2026
Application No. 18/673,841

Semi-Generative Artificial Intelligence

Final Rejection §103
Filed
May 24, 2024
Examiner
CASCAIS, JUSTIN PHILIP
Art Unit
2669
Tech Center
2600 — Communications
Assignee
The Boeing Company
OA Round
2 (Final)
75%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
48 granted / 64 resolved
+13.0% vs TC avg
Moderate +14% lift
Without
With
+13.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
18 currently pending
Career history
74
Total Applications
across all art units

Statute-Specific Performance

§101
9.6%
-30.4% vs TC avg
§103
61.2%
+21.2% vs TC avg
§102
13.2%
-26.8% vs TC avg
§112
11.2%
-28.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 64 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Amendment Applicant submitted amendments on 8/18/2026. The Examiner acknowledges the amendment and has reviewed the claims accordingly. Information Disclosure Statement The IDS(s) dated 10/13/2025 that has been previously considered remains placed in the application file. Overview Claims 1-20 are pending in this application and have been considered below. Claims 1-20 are rejected. Applicant Arguments In regards to Argument 1, Applicant states that the Office relied on Ma’s same target denoising U-Net cross-attention teaching both for “feeding the embeddings into dual diffusion implicit bridges” and for target reverse diffusion in accordance with the metadata (See Remarks, page 8). In regards to Argument 2, Applicant states that the Su-Ma combination would route metadata only into the final target reverse diffusion stage and not into the separate DDIB mapping between the first and second Gaussian distributions (See Remarks, page 8-9). In regards to Argument 3, Applicant states that neither Su nor Ma provides a reason to place embedded constraints into the DDIB mapping and that doing so would result from hindsight based on Applicant’s disclosure (See Remarks, page 9-10). Examiner’s Response In response to Argument 1, it has been considered but is moot in view of new ground(s) of rejection based on the amendments. However, in response to the Applicant alleging that the rejection of the claim language previously presented was based on an unreasonable construction, the Examiner respectfully disagrees. Su defines a DDIB as a two step process comprising source model encoding followed by target model decoding and describes the source to latent and latent to target processes as two concatenated Schrödinger bridges (Su, Abstract, pp. 4-5, Algorithm 1, Proposition 3.2). The instant specification in ¶25 and ¶45 similarly describes the DDIB as a “back-to-back concatenation of source to latent space and from latent space to target Schrödinger bridge” and states that metadata constrains the target model Schrödinger bridge. The former claim required feeding the embedding into the DDIB and performing target reverse diffusion in accordance with the metadata, but it did not require the embedding to constrain the inter-Gaussian mapping independently of that reverse diffusion. Under the previously presented language, Ma’s target model conditioning teaches both overlapping requirements because the target bridge was part of the composite DDIB. Applicants’ amendment now adds a different relationship. The present rejection therefore does not rely on Ma’s target U-Net passage to meet the newly added independent mapping constraint. A new reference, Zhang, has been introduced which in at least pp. 3553-3554, Eqs. 5-7, and Algorithm 2 discloses the added relationship. After reviewing the amendments, the Examiner interprets that Su and Ma in view of Zhang teaches on the amended claims that were presented. The details of the rejection are listed below. In response to Argument 2, it has been considered but is moot in view of new ground(s) of rejection based on the amendments. A new reference, Zhang, has been introduced which in pp. 3553-3554, Eqs. 5-7, and Algorithm 2 discloses the relationship added by the amendment. Zhang introduces the conditioned derived mean shift into a Gaussian diffusion trajectory, identifies E(c) as a shift predictor that maps the condition into latent space, and defines the reverse process starting prior as p(x_T) = N(s_T,Σ). Zhang’s Algorithm 2 first calculates s_T = k_T ·E(c) and samples x_T from N(s_T,Σ), and then enters the reverse loop. Zhang also states that the condition does not need to be fed into the noise prediction network because it has already been encoded into the condition dependent trajectory. Su supplies the source model Gaussian latent and two model DDIB handoff. Applying Zhang’s embedding derived shift at Su’s handoff establishes the condition dependent target Gaussian prior before the target reverse ODE. Therefore, the embedding constrains the inter-Gaussian mapping as an operation distinct from the subsequent target reverse diffusion. After reviewing the amendments, the Examiner interprets that Su and Ma in view of Zhang teaches on the amended claims that were presented. The details of the rejection are listed below. In response to Argument 3, it has been considered but is moot in view of new ground(s) of rejection based on the amendments. However, in response to the Applicant alleging that the combination rationale in the prior rejection was derived from the Applicant’s disclosure, the Examiner respectfully disagrees. The claim language then pending did not require the metadata embedding to constrain a distinct inter-Gaussian mapping independently of target reverse diffusion. The prior rejection proposed conditioning Su’s target DDIB stage with Ma’s metadata-derived text embedding. The reason for that modification was supplied by the prior art. Su teaches classifier guidance that steers sampling toward selected classes and reports target images that retain source content, including pose, while accounting for target domain characteristics (Su, p. 8, § 4.4). Ma teaches that text prompts guide diffusion models to produce images adhering to specified content and that CAD derived visual and text prompts provide control over object geometry, pose, viewpoint, distance, background, and color (Ma, pp. 1-2, 4-6, §§ 3.1-3.3). Applying Ma’s known conditioning to Su’s target DDIB stage would have predictably controlled the content and characteristics of Su’s target image. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-5, 7-12, 14-18, and 20 is/are rejected under 35 U.S.C. 103 as obvious over Su et al (Su, X., Song, J., Meng, C., & Ermon, S. (2022). Dual diffusion implicit bridges for image-to-image translation. arXiv preprint arXiv:2203.08382., hereafter referred to as Su) and Ma et al (Ma, W., Liu, Q., Wang, J., Wang, A., Liu, Y., Kortylewski, A., & Yuille, A. (2023). Adding 3d geometry control to diffusion models. arXiv preprint arXiv:2306.08103., hereafter referred to as Ma), further in view of Zhang et al (Zhang, Z., Zhao, Z., Yu, J., & Tian, Q. (2023, June). Shiftddpms: Exploring conditional diffusion models by shifting diffusion trajectories. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 3, pp. 3552-3560)., hereafter referred to as Zhang). Claim 1 Regarding Claim 1, Su teaches A computer-implemented method for semi-generative artificial intelligence modelling between domains, the method comprising: using a number of processors to perform: receiving a source image of an object in a first domain (Su in p. 2, Figure 1 discloses “Given a source image x(s)”; p.4, Algorithm 1 discloses receiving “data sample from source domain x(s) ∼ ps(x)”); diffusing the source image through a source diffusion model to generate a first Gaussian distribution in the first domain (Su in pp. 3-4, §§2.1-3, Algorithm 1 discloses applying the source-domain probability-flow ODE from t=0 to t=1 to obtain the source model’s latent code; pp. 3-4, §§2.2 discloses that an SGM’s endpoint marginal p1 is the standard Gaussian prior; p.14 discloses, Appendix B.1, discloses modeling the diffusion conditionals as Gaussians and samples ɛt from N(0,I). Under BRI, the source model’s endpoint marginal is the first, source model associated Gaussian distribution); feeding the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain (Su in pp. 3-5, Algorithm 1, Proposition 3.2 discloses DDIBs having a source to latent stage and a latent to target stage, with the source latent supplied as the initial condition to the target stage. Su maps a source domain image through the source model to source latent x(l), supplies x(l) as the target model’s initial latent at t=1, and maps that latent through the target model to the target domain. Su identifies the endpoint marginal of each domain specific diffusion model as a Gaussian prior); sampling from the first Gaussian distribution (Su in p. 4, Algorithm 1 discloses beginning with source sample x(s) ∼ ps(x) and deterministically transports it through the source probability-flow ODE to latent x(l); pp. 3-4, §§2.1-2.2 discloses that the probability-flow ODE carries the same marginals as the diffusion process and that the endpoint marginal is the standard Gaussian prior. x(l) is a sample having the source model’s Gaussian endpoint law); mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain (Su in p. 4, §3, Algorithm 1 discloses supplying the source latent x(l) as the target latent code at t=1 for the independently trained target domain model), reverse diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata (Su in pp. 2/4, Figure 1, §3, Algorithm 1 discloses solving the target model from t=1 to t=0 to construct target-domain image x(t) while preserving source content). Su does not explicitly teach all of generating embeddings from metadata, wherein the metadata provides constraints for image reconstruction; feeding the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain; mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain, wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model; and reverse diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata. However, Ma teaches generating embeddings from metadata, wherein the metadata provides constraints for image reconstruction (Ma in p. 4, §3.1, Eq. 2 discloses preprocessing text prompt T with pretrained text encoder θ to produce representation θ(T), which is used as key/value information in U-Net cross-attention; pp. 5-6, Figure 2, §§3.2-3.3, forms the prompt from CAD class names, tags or keywords, and generated descriptions that constrain object identity, background, color, shape, pose, viewpoint, and distance. θ(T) is the embedding; the separate CAD-derived edge map is additional visual conditioning); feeding the embeddings into dual diffusion implicit bridges (Ma in p. 4, §3.1 discloses feeding θ(T) into successive target-denoising U-Net cross-attention layers); and reverse diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata (Ma in pp. 4-6, §§3.1-3.4, Figure 2, Algorithm 1 discloses reverse denoising conditioned on θ(T), optionally with a CAD-derived edge prompt, and decoding z0 into a target image constrained by that information). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Su by conditioning the target DDIB stage with metadata-derived text embeddings that is taught by Ma, since both reference are analogous art in the field of diffusion-based computer vision image generation and translation; thus, one of ordinary skilled in the art would be motivated to combine the references since Su’s source-latent-target DDIB process with Ma’s fixed cross-attention condition yields the predictable result of metadata-conditioned target reverse diffusion, thereby improving control over target object attributes, background, color, viewpoint, and geometry. Su in view of Ma does not explicitly teach all of feeding the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain; and mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain, wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model. However, Zhang teaches feeding the embeddings into dual diffusion implicit bridges (Zhang in p. 3558 discloses feeding pre-trained sentence embeddings into a shift predictor that generates a shift to guide a diffusion trajectory) that map between the first Gaussian distribution and a second Gaussian distribution in a second domain (Zhang in pp. 3553-3554, Eqs. 5-7 discloses mapping a condition into latent space using E(c) and shifting a Gaussian distribution to the condition dependent Gaussian distribution N(s_T, Σ)); mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain (Zhang in p. 3554, Algorithm 2 discloses calculating s_T = k_T · E(c) and sampling x_T ~ N(s_T, Σ) before beginning the reverse loop), wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model (Zhang in pp. 3553-3554, Eqs. 5-7, Algorithm 2 discloses deriving shifts s_T from condition embedding E(c) and using the shift to establish and sample the condition dependent Gaussian prior before reverse diffusion). Zhang is analogous art because Zhang, like Su and Ma, concerns diffusion model image generation, and Zhang is pertinent to the problem of incorporating condition information into a Gaussian diffusion trajectory before reverse denoising. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the latent handoff between the source diffusion model and target diffusion model of Su in view of Ma with the condition embedding derived Gaussian shift of Zhang because Zhang teaches that shifting a diffusion trajectory using a condition derived embedding establishes a condition dependent Gaussian prior before reverse diffusion and improves latent space utilization and model learning capacity (Zhang, pp. 3553-3554, 3556-3558), and one of ordinary skill in the art would have recognized that incorporating this feature into the DDIB method of Su would map Su’s source model Gaussian latent to a condition dependent target model Gaussian prior before target reverse diffusion, with a reasonable expectation of success because Su implements its DDIB probability flow ODE using DDIM, Zhang adapts ShiftDDPMs to ShiftDDIM, and both use compatible Gaussian latent representations. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 2 Regarding Claim 2, Su and Ma, further in view of Zhang teaches The method of claim 1, wherein the first domain comprises three-dimensional CAD (computer assisted drawing) image data (Ma in pp. 2/5-6, Figures 1-2, §3.2 discloses CAD models rendered from selected viewpoints and distances into two-dimensional CAD-derived sketches and compact edge-map images carrying the CAD object’s 3D shape and pose). Claim 3 Regarding Claim 3, Su and Ma, further in view of Zhang teaches The method of claim 1, wherein the second domain comprises real world image data (Su in p. 7, §4.3 discloses an image domain containing “real photos taken via a camera or a satellite” and translation in either direction). Claim 4 Regarding Claim 4, Su and Ma, further in view of Zhang teaches The method of claim 3, wherein the target image comprises a photorealistic image (Ma in pp. 1/5, Abstract, Introduction, Figure 2 discloses generating “photo-realistic images” using encoded text and CAD-derived visual conditions). Claim 5 Regarding Claim 5, Su and Ma, further in view of Zhang teaches The method of claim 1, wherein the metadata comprises at least one of: clustering of data; prompt embedding; segmentation mask; two-dimensional drawing information; text description of the target object; audio description of the target object; graph representation of the target object; material; or background (Ma in p. 4, §3.1 discloses generating prompt embedding θ(T); pp. 5-6, §§3.2-3.3, Figure 2 discloses target-object descriptions and background descriptions). Claim 7 Regarding Claim 7, Su and Ma, further in view of Zhang teaches The method of claim 1, wherein the source diffusion model and target diffusion model comprise Schrödinger bridges (Su in pp. 1/4-5, Abstract, §§2.2-3, Proposition 3.2 discloses characterizing its source-to-Gaussian and Gaussian-to-target PF-ODE stages as two concatenated special Schrödinger Bridges). Claim 8 Regarding Claim 8, Su teaches A system for semi-generative artificial intelligence modelling between domains, the system comprising: a storage device that stores program instructions (Su in Abstract discloses “image translation with DDIBs relies on two diffusion models trained independently on each domain, and is a two-step process: DDIBs first obtain latent encodings for source images with the source diffusion model, and then decode such encodings using the target model to construct target images”. The computer-implemented image processing would predictably be executed on conventional reliable hardware); one or more processors operably connected to the storage device (Su in Abstract discloses “image translation with DDIBs relies on two diffusion models trained independently on each domain, and is a two-step process: DDIBs first obtain latent encodings for source images with the source diffusion model, and then decode such encodings using the target model to construct target images”. The computer-implemented image processing would predictably be executed on conventional reliable hardware) and configured to execute the program instructions to cause the system to: receive a source image of an object in a first domain (Su in p. 2, Figure 1 discloses “Given a source image x(s)”; p.4, Algorithm 1 discloses receiving “data sample from source domain x(s) ∼ ps(x)”); diffuse the source image through a source diffusion model to generate a first Gaussian distribution in the first domain (Su in pp. 3-4, §§2.1-3, Algorithm 1 discloses applying the source-domain probability-flow ODE from t=0 to t=1 to obtain the source model’s latent code; pp. 3-4, §§2.2 discloses that an SGM’s endpoint marginal p1 is the standard Gaussian prior; p.14 discloses, Appendix B.1, discloses modeling the diffusion conditionals as Gaussians and samples ɛt from N(0,I). Under BRI, the source model’s endpoint marginal is the first, source model associated Gaussian distribution); feed the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain (Su in pp. 3-5, Algorithm 1, Proposition 3.2 discloses DDIBs having a source to latent stage and a latent to target stage, with the source latent supplied as the initial condition to the target stage. Su maps a source domain image through the source model to source latent x(l), supplies x(l) as the target model’s initial latent at t=1, and maps that latent through the target model to the target domain. Su identifies the endpoint marginal of each domain specific diffusion model as a Gaussian prior); sample from the first Gaussian distribution (Su in p. 4, Algorithm 1 discloses beginning with source sample x(s) ∼ ps(x) and deterministically transports it through the source probability-flow ODE to latent x(l); pp. 3-4, §§2.1-2.2 discloses that the probability-flow ODE carries the same marginals as the diffusion process and that the endpoint marginal is the standard Gaussian prior. x(l) is a sample having the source model’s Gaussian endpoint law); through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain (Su in p. 4, §3, Algorithm 1 discloses supplying the source latent x(l) as the target latent code at t=1 for the independently trained target domain model), reverse diffuse the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata (Su in pp. 2/4, Figure 1, §3, Algorithm 1 discloses solving the target model from t=1 to t=0 to construct target-domain image x(t) while preserving source content). Su does not explicitly teach all of generate embeddings from metadata, wherein the metadata provides constraints for image reconstruction; feed the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain; map, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain, wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model; and reverse diffuse the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata. However, Ma teaches generate embeddings from metadata, wherein the metadata provides constraints for image reconstruction (Ma in p. 4, §3.1, Eq. 2 discloses preprocessing text prompt T with pretrained text encoder θ to produce representation θ(T), which is used as key/value information in U-Net cross-attention; pp. 5-6, Figure 2, §§3.2-3.3, forms the prompt from CAD class names, tags or keywords, and generated descriptions that constrain object identity, background, color, shape, pose, viewpoint, and distance. θ(T) is the embedding; the separate CAD-derived edge map is additional visual conditioning); feed the embeddings into dual diffusion implicit bridges (Ma in p. 4, §3.1 discloses feeding θ(T) into successive target-denoising U-Net cross-attention layers); and reverse diffuse the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata (Ma in pp. 4-6, §§3.1-3.4, Figure 2, Algorithm 1 discloses reverse denoising conditioned on θ(T), optionally with a CAD-derived edge prompt, and decoding z0 into a target image constrained by that information). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Su by conditioning the target DDIB stage with metadata-derived text embeddings that is taught by Ma, since both reference are analogous art in the field of diffusion-based computer vision image generation and translation; thus, one of ordinary skilled in the art would be motivated to combine the references since Su’s source-latent-target DDIB process with Ma’s fixed cross-attention condition yields the predictable result of metadata-conditioned target reverse diffusion, thereby improving control over target object attributes, background, color, viewpoint, and geometry. Su in view of Ma does not explicitly teach all of feed the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain; and map, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain, wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model. However, Zhang teaches feed the embeddings into dual diffusion implicit bridges (Zhang in p. 3558 discloses feeding pre-trained sentence embeddings into a shift predictor that generates a shift to guide a diffusion trajectory) that map between the first Gaussian distribution and a second Gaussian distribution in a second domain (Zhang in pp. 3553-3554, Eqs. 5-7 discloses mapping a condition into latent space using E(c) and shifting a Gaussian distribution to the condition dependent Gaussian distribution N(s_T, Σ)); map, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain (Zhang in p. 3554, Algorithm 2 discloses calculating s_T = k_T · E(c) and sampling x_T ~ N(s_T, Σ) before beginning the reverse loop), wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model (Zhang in pp. 3553-3554, Eqs. 5-7, Algorithm 2 discloses deriving shifts s_T from condition embedding E(c) and using the shift to establish and sample the condition dependent Gaussian prior before reverse diffusion). Zhang is analogous art because Zhang, like Su and Ma, concerns diffusion model image generation, and Zhang is pertinent to the problem of incorporating condition information into a Gaussian diffusion trajectory before reverse denoising. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the latent handoff between the source diffusion model and target diffusion model of Su in view of Ma with the condition embedding derived Gaussian shift of Zhang because Zhang teaches that shifting a diffusion trajectory using a condition derived embedding establishes a condition dependent Gaussian prior before reverse diffusion and improves latent space utilization and model learning capacity (Zhang, pp. 3553-3554, 3556-3558), and one of ordinary skill in the art would have recognized that incorporating this feature into the DDIB method of Su would map Su’s source model Gaussian latent to a condition dependent target model Gaussian prior before target reverse diffusion, with a reasonable expectation of success because Su implements its DDIB probability flow ODE using DDIM, Zhang adapts ShiftDDPMs to ShiftDDIM, and both use compatible Gaussian latent representations. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 9 Regarding Claim 9, Su and Ma, further in view of Zhang teaches The system of claim 8, wherein the first domain comprises three-dimensional CAD (computer assisted drawing) image data (Ma in pp. 2/5-6, Figures 1-2, §3.2 discloses CAD models rendered from selected viewpoints and distances into two-dimensional CAD-derived sketches and compact edge-map images carrying the CAD object’s 3D shape and pose). Claim 10 Regarding Claim 10, Su and Ma, further in view of Zhang teaches The system of claim 8, wherein the second domain comprises real world image data (Su in p. 7, §4.3 discloses an image domain containing “real photos taken via a camera or a satellite” and translation in either direction). Claim 11 Regarding Claim 11, Su and Ma, further in view of Zhang teaches The system of claim 10, wherein the target image comprises a photorealistic image (Ma in pp. 1/5, Abstract, Introduction, Figure 2 discloses generating “photo-realistic images” using encoded text and CAD-derived visual conditions). Claim 12 Regarding Claim 12, Su and Ma, further in view of Zhang teaches The system of claim 8, wherein the metadata comprises at least one of: clustering of data; prompt embedding; segmentation mask; two-dimensional drawing information; text description of the target object; audio description of the target object; graph representation of the target object; material; or background (Ma in p. 4, §3.1 discloses generating prompt embedding θ(T); pp. 5-6, §§3.2-3.3, Figure 2 discloses target-object descriptions and background descriptions). Claim 14 Regarding Claim 14, Su and Ma, further in view of Zhang teaches The system of claim 8, wherein the source diffusion model and target diffusion model comprise Schrödinger bridges (Su in pp. 1/4-5, Abstract, §§2.2-3, Proposition 3.2 discloses characterizing its source-to-Gaussian and Gaussian-to-target PF-ODE stages as two concatenated special Schrödinger Bridges). Claim 15 Regarding Claim 15, Su teaches A computer program product for semi-generative artificial intelligence modelling between domains, the computer program product comprising: a computer-readable storage medium having program instructions embodied thereon to perform the steps of: receiving a source image of an object in a first domain (Su in p. 2, Figure 1 discloses “Given a source image x(s)”; p.4, Algorithm 1 discloses receiving “data sample from source domain x(s) ∼ ps(x)”); diffusing the source image through a source diffusion model to generate a first Gaussian distribution in the first domain (Su in pp. 3-4, §§2.1-3, Algorithm 1 discloses applying the source-domain probability-flow ODE from t=0 to t=1 to obtain the source model’s latent code; pp. 3-4, §§2.2 discloses that an SGM’s endpoint marginal p1 is the standard Gaussian prior; p.14 discloses, Appendix B.1, discloses modeling the diffusion conditionals as Gaussians and samples ɛt from N(0,I). Under BRI, the source model’s endpoint marginal is the first, source model associated Gaussian distribution); feeding the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain (Su in pp. 3-5, Algorithm 1, Proposition 3.2 discloses DDIBs having a source to latent stage and a latent to target stage, with the source latent supplied as the initial condition to the target stage. Su maps a source domain image through the source model to source latent x(l), supplies x(l) as the target model’s initial latent at t=1, and maps that latent through the target model to the target domain. Su identifies the endpoint marginal of each domain specific diffusion model as a Gaussian prior); sampling from the first Gaussian distribution (Su in p. 4, Algorithm 1 discloses beginning with source sample x(s) ∼ ps(x) and deterministically transports it through the source probability-flow ODE to latent x(l); pp. 3-4, §§2.1-2.2 discloses that the probability-flow ODE carries the same marginals as the diffusion process and that the endpoint marginal is the standard Gaussian prior. x(l) is a sample having the source model’s Gaussian endpoint law); mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain (Su in p. 4, §3, Algorithm 1 discloses supplying the source latent x(l) as the target latent code at t=1 for the independently trained target domain model), reverse diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata (Su in pp. 2/4, Figure 1, §3, Algorithm 1 discloses solving the target model from t=1 to t=0 to construct target-domain image x(t) while preserving source content). Su does not explicitly teach all of generating embeddings from metadata, wherein the metadata provides constraints for image reconstruction; feeding the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain; mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain, wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model; and reverse diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata. However, Ma teaches generating embeddings from metadata, wherein the metadata provides constraints for image reconstruction (Ma in p. 4, §3.1, Eq. 2 discloses preprocessing text prompt T with pretrained text encoder θ to produce representation θ(T), which is used as key/value information in U-Net cross-attention; pp. 5-6, Figure 2, §§3.2-3.3, forms the prompt from CAD class names, tags or keywords, and generated descriptions that constrain object identity, background, color, shape, pose, viewpoint, and distance. θ(T) is the embedding; the separate CAD-derived edge map is additional visual conditioning); feeding the embeddings into dual diffusion implicit bridges (Ma in p. 4, §3.1 discloses feeding θ(T) into successive target-denoising U-Net cross-attention layers); and reverse diffusing the second Gaussian distribution through a target diffusion model to generate a target image of the object in the second domain in accordance with the metadata (Ma in pp. 4-6, §§3.1-3.4, Figure 2, Algorithm 1 discloses reverse denoising conditioned on θ(T), optionally with a CAD-derived edge prompt, and decoding z0 into a target image constrained by that information). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Su by conditioning the target DDIB stage with metadata-derived text embeddings that is taught by Ma, since both reference are analogous art in the field of diffusion-based computer vision image generation and translation; thus, one of ordinary skilled in the art would be motivated to combine the references since Su’s source-latent-target DDIB process with Ma’s fixed cross-attention condition yields the predictable result of metadata-conditioned target reverse diffusion, thereby improving control over target object attributes, background, color, viewpoint, and geometry. Su in view of Ma does not explicitly teach all of feeding the embeddings into dual diffusion implicit bridges that map between the first Gaussian distribution and a second Gaussian distribution in a second domain; and mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain, wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model. However, Zhang teaches feeding the embeddings into dual diffusion implicit bridges (Zhang in p. 3558 discloses feeding pre-trained sentence embeddings into a shift predictor that generates a shift to guide a diffusion trajectory) that map between the first Gaussian distribution and a second Gaussian distribution in a second domain (Zhang in pp. 3553-3554, Eqs. 5-7 discloses mapping a condition into latent space using E(c) and shifting a Gaussian distribution to the condition dependent Gaussian distribution N(s_T, Σ)); mapping, through the dual diffusion implicit bridges, the sample from the first Gaussian distribution to the second Gaussian distribution in the second domain (Zhang in p. 3554, Algorithm 2 discloses calculating s_T = k_T · E(c) and sampling x_T ~ N(s_T, Σ) before beginning the reverse loop), wherein the mapping is constrained by the embeddings independently of reverse diffusion performed by a target diffusion model (Zhang in pp. 3553-3554, Eqs. 5-7, Algorithm 2 discloses deriving shifts s_T from condition embedding E(c) and using the shift to establish and sample the condition dependent Gaussian prior before reverse diffusion). Zhang is analogous art because Zhang, like Su and Ma, concerns diffusion model image generation, and Zhang is pertinent to the problem of incorporating condition information into a Gaussian diffusion trajectory before reverse denoising. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the latent handoff between the source diffusion model and target diffusion model of Su in view of Ma with the condition embedding derived Gaussian shift of Zhang because Zhang teaches that shifting a diffusion trajectory using a condition derived embedding establishes a condition dependent Gaussian prior before reverse diffusion and improves latent space utilization and model learning capacity (Zhang, pp. 3553-3554, 3556-3558), and one of ordinary skill in the art would have recognized that incorporating this feature into the DDIB method of Su would map Su’s source model Gaussian latent to a condition dependent target model Gaussian prior before target reverse diffusion, with a reasonable expectation of success because Su implements its DDIB probability flow ODE using DDIM, Zhang adapts ShiftDDPMs to ShiftDDIM, and both use compatible Gaussian latent representations. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 16 Regarding Claim 16, Su and Ma, further in view of Zhang teaches The computer program product of claim 15, wherein the first domain comprises three-dimensional CAD (computer assisted drawing) image data (Ma in pp. 2/5-6, Figures 1-2, §3.2 discloses CAD models rendered from selected viewpoints and distances into two-dimensional CAD-derived sketches and compact edge-map images carrying the CAD object’s 3D shape and pose). Claim 17 Regarding Claim 17, Su and Ma, further in view of Zhang teaches The computer program product of claim 15, wherein the second domain comprises real world image data (Su in p. 7, §4.3 discloses an image domain containing “real photos taken via a camera or a satellite” and translation in either direction). Claim 18 Regarding Claim 18, Su and Ma, further in view of Zhang teaches The computer program product of claim 15, wherein the metadata comprises at least one of: clustering of data; prompt embedding; segmentation mask; two-dimensional drawing information; text description of the target object; audio description of the target object; graph representation of the target object; material; or background (Ma in p. 4, §3.1 discloses generating prompt embedding θ(T); pp. 5-6, §§3.2-3.3, Figure 2 discloses target-object descriptions and background descriptions). Claim 20 Regarding Claim 20, Su and Ma, further in view of Zhang teaches The computer program product of claim 15, wherein the source diffusion model and target diffusion model comprise Schrödinger bridges (Su in pp. 1/4-5, Abstract, §§2.2-3, Proposition 3.2 discloses characterizing its source-to-Gaussian and Gaussian-to-target PF-ODE stages as two concatenated special Schrödinger Bridges). Claim(s) 6, 13, and 19 is/are rejected under 35 U.S.C. 103 as obvious over Su et al (Su, X., Song, J., Meng, C., & Ermon, S. (2022). Dual diffusion implicit bridges for image-to-image translation. arXiv preprint arXiv:2203.08382., hereafter referred to as Su), Ma et al (Ma, W., Liu, Q., Wang, J., Wang, A., Liu, Y., Kortylewski, A., & Yuille, A. (2023). Adding 3d geometry control to diffusion models. arXiv preprint arXiv:2306.08103., hereafter referred to as Ma), and Zhang et al (Zhang, Z., Zhao, Z., Yu, J., & Tian, Q. (2023, June). Shiftddpms: Exploring conditional diffusion models by shifting diffusion trajectories. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 3, pp. 3552-3560)., hereafter referred to as Zhang), further in view of Jin et al (Jin, L., Li, Z., & Tang, J. (2020). Deep semantic multimodal hashing network for scalable image-text and video-text retrievals. IEEE Transactions on Neural Networks and Learning Systems, 34(4), 1838-1851., hereafter referred to as Jin). Claim 6 Regarding Claim 6, Su and Ma, further in view of Zhang teaches The method of claim 1. Su and Ma, further in view of Zhang does not explicitly teach all of wherein the metadata comprises information provided by an artificial intelligence design parser that cross-references text and specifications against images and video. However, Jin teaches wherein the metadata comprises information provided by an artificial intelligence design parser that cross-references text and specifications against images and video (The Examiner interprets under the broadest reasonable interpretation that ‘information provided by an artificial intelligence design parser that cross-references text and specifications against images and video' is given its BRI consistent with the spec ¶22 as metadata produced by a machine-implemented component that extract information from text or specification sources and associates that information with image or video data. The recitations limit the metadata by its source and does not require any particular parser architecture, training procedure, or accuracy. Jin in p. 1, Abstract discloses a deep semantic multimodal hashing network for both image-text and video-text retrieval, using 2-D CNN processing for image information and 3-D CNN processing for spatial and temporal video information while preserving intermodality similarity; pp. 3-4, §III.A-C discloses image/video modality X and text modality Y, using deep modality-specific networks to map new samples from those modalities into a common Hamming space, and measuring similarity for cross-modal text-visual pairs. Under BRI, the deep network is a machine-implemented parser that obtains information from text sources and associates that information with image or video data; the resulting associated textual/semantic information is parser-provided metadata). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Su and Ma, further in view of Zhang by incorporating a cross-modal network that is taught by Jin, since both reference are analogous art in the field of learned computer vision image processing; thus, one of ordinary skilled in the art would be motivated to combine the references since Su and Ma, further in view of Zhang’s image translation and metadata conditioning framework with Jin’s text-image/video common space association yields the predictable result of providing semantically matched parser-derived metadata to the target diffusion stage, thereby reducing mismatched reconstruction constraints and improving consistency between the source object and the metadata conditioned target image. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 13 Regarding Claim 13, Su and Ma, further in view of Zhang teaches The system of claim 8. Su and Ma, further in view of Zhang does not explicitly teach all of wherein the metadata comprises information provided by an artificial intelligence design parser that cross-references text and specifications against images and video. However, Jin teaches wherein the metadata comprises information provided by an artificial intelligence design parser that cross-references text and specifications against images and video (The Examiner interprets under the broadest reasonable interpretation that ‘information provided by an artificial intelligence design parser that cross-references text and specifications against images and video' is given its BRI consistent with the spec ¶22 as metadata produced by a machine-implemented component that extract information from text or specification sources and associates that information with image or video data. The recitations limit the metadata by its source and do not require any particular parser architecture, training procedure, or accuracy. Jin in p. 1, Abstract discloses a deep semantic multimodal hashing network for both image-text and video-text retrieval, using 2-D CNN processing for image information and 3-D CNN processing for spatial and temporal video information while preserving intermodality similarity; pp. 3-4, §III.A-C discloses image/video modality X and text modality Y, using deep modality-specific networks to map new samples from those modalities into a common Hamming space, and measuring similarity for cross-modal text-visual pairs. Under BRI, the deep network is a machine-implemented parser that obtains information from text sources and associates that information with image or video data; the resulting associated textual/semantic information is parser-provided metadata). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Su and Ma, further in view of Zhang by incorporating a cross-modal network that is taught by Jin, since both reference are analogous art in the field of learned computer vision image processing; thus, one of ordinary skilled in the art would be motivated to combine the references since Su and Ma, further in view of Zhang’s image translation and metadata conditioning framework with Jin’s text-image/video common space association yields the predictable result of providing semantically matched parser-derived metadata to the target diffusion stage, thereby reducing mismatched reconstruction constraints and improving consistency between the source object and the metadata conditioned target image. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Claim 19 Regarding Claim 19, Su and Ma, further in view of Zhang teaches The computer program product of claim 15. Su and Ma, further in view of Zhang does not explicitly teach all of wherein the metadata comprises information provided by an artificial intelligence design parser that cross-references text and specifications against images and video. However, Jin teaches wherein the metadata comprises information provided by an artificial intelligence design parser that cross-references text and specifications against images and video (The Examiner interprets under the broadest reasonable interpretation that ‘information provided by an artificial intelligence design parser that cross-references text and specifications against images and video' is given its BRI consistent with the spec ¶22 as metadata produced by a machine-implemented component that extract information from text or specification sources and associates that information with image or video data. The recitations limit the metadata by its source and do not require any particular parser architecture, training procedure, or accuracy. Jin in p. 1, Abstract discloses a deep semantic multimodal hashing network for both image-text and video-text retrieval, using 2-D CNN processing for image information and 3-D CNN processing for spatial and temporal video information while preserving intermodality similarity; pp. 3-4, §III.A-C discloses image/video modality X and text modality Y, using deep modality-specific networks to map new samples from those modalities into a common Hamming space, and measuring similarity for cross-modal text-visual pairs. Under BRI, the deep network is a machine-implemented parser that obtains information from text sources and associates that information with image or video data; the resulting associated textual/semantic information is parser-provided metadata). Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Su and Ma, further in view of Zhang by incorporating a cross-modal network that is taught by Jin, since both reference are analogous art in the field of learned computer vision image processing; thus, one of ordinary skilled in the art would be motivated to combine the references since Su and Ma, further in view of Zhang’s image translation and metadata conditioning framework with Jin’s text-image/video common space association yields the predictable result of providing semantically matched parser-derived metadata to the target diffusion stage, thereby reducing mismatched reconstruction constraints and improving consistency between the source object and the metadata conditioned target image. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN P CASCAIS whose telephone number is (703)756-5576. The examiner can normally be reached Monday-Friday 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mr. O’Neal Mistry can be reached on (313) 446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.P.C./Examiner, Art Unit 2674 /ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674 Date: 9/2/2026
Read full office action

Prosecution Timeline

May 24, 2024
Application Filed
Aug 05, 2026
Non-Final Rejection mailed — §103
Aug 18, 2026
Response Filed
Sep 10, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749203
SYSTEMS AND METHODS FOR ANNOTATING IMAGE SEQUENCES WITH LANDMARKS
3y 10m to grant Granted Sep 29, 2026
Patent 12748805
CONTENT-AWARE ARTIFICIAL INTELLIGENCE GENERATED FRAMES FOR DIGITAL IMAGES
2y 9m to grant Granted Sep 29, 2026
Patent 12743764
APPARATUS AND METHOD FOR INSPECTING ASSEMBLY HOLE OF VEHICLE
3y 7m to grant Granted Sep 22, 2026
Patent 12743806
POSITION ESTIMATING DEVICE, FORKLIFT, POSITION ESTIMATING METHOD, AND NON-TRANSITORY COMPUTER READABLE STORAGE MEDIUM STORING PROGRAM
2y 11m to grant Granted Sep 22, 2026
Patent 12737873
INSPECTION APPARATUS AND STORAGE MEDIUM STORING COMPUTER PROGRAM
2y 10m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
75%
Grant Probability
89%
With Interview (+13.7%)
2y 10m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 64 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month