Prosecution Insights
Last updated: October 01, 2026
Application No. 18/481,719

IMAGE AND DEPTH MAP GENERATION USING A CONDITIONAL MACHINE LEARNING

Final Rejection §103
Filed
Oct 05, 2023
Examiner
WU, YANNA
Art Unit
2615
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
4 (Final)
81%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
369 granted / 456 resolved
+18.9% vs TC avg
Strong +34% interview lift
Without
With
+34.4%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
23 currently pending
Career history
474
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
69.7%
+29.7% vs TC avg
§102
6.8%
-33.2% vs TC avg
§112
8.1%
-31.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 456 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This is in response to applicant’s amendment/response filed on 07/27/2026, which has been entered and made of record. Claim 1, 3, 4, 6, 15-16, 18, 21, 23-24, 26 are amended. Claims 2, 17, 22 are canceled. Claims 1, 3-7, 15-16, 18-19 and 21, 23-27 are pending in the application. Response to Arguments Applicant arguments regarding claim rejections under 103 are considered, but are moot in view of new ground of rejections. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1,3-4, 6-7, 15-16, 18-19, 21, 23-24, 26-27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Skypnyk et al. (US 2024/0355064 A1) in view of Lim et al. (US 2009/0190852 A1) and further in view of Couleaud et al. (US 2025/0078347 A1). Regarding claim 1, Skypnyk teaches: A method for image generation, comprising: obtaining a prompt and a noise input;([0147], “At operation 502, the interaction system 100 (alone or in combination with one or more elements of the personal AI agent 302) identifies a prompt of a user indicating a user's intent.” [0163], “The stable diffusion model generates populated image templates (such as images with color and depth information) by learning to reverse the process of adding noise to the input image template 612.”) generating, by a diffusion model, an image and a depth map corresponding to the image by denoising the noise input based on the prompt.( [0162], “In the stable diffusion model, the user prompt 610 is converted into an embedding using a pre-trained language model, such as a transformer-based architecture (e.g., GPT or BERT). This embedding encodes the semantic information of the user prompt 610 and is used to condition the image/populated image template 612 generation process.” [0163], “The stable diffusion model is then conditioned on both the image template 612 and the text embedding by incorporating the text embedding into the stable diffusion model architecture or by conditioning the stable diffusion model's latent space on the textual information. The stable diffusion model generates populated image templates (such as images with color and depth information) by learning to reverse the process of adding noise to the input image template 612.”) However, Skypnyk does not, but Lim teaches: identifying an image occlusion area and a depth occlusion area based on the image and the depth map, respectively; and generating, …a modified image and a modified depth map by inpainting the image occlusion area and the depth occlusion area, respectively, wherein the modified image and the modified depth map correspond to a modified view of the image. ([0020], “In a case where a depth image and a color image after a change in viewpoint are to be inpainted by using a depth image and a color image obtained from a viewpoint, if a portion that is not shown before the viewpoint change is shown after the viewpoint change, depth information and color information on the portion are to be inpainted from the depth image and the color image before the viewpoint change. When a region of an object in an image is hidden by an object having a depth value lower than that of the former, the region may not be shown from a predetermined viewpoint. This region is referred to as an occlusion region. A size of the occlusion region is determined according to a degree of change in viewpoint and a difference between the depth values of the objects. By combining an image inpainted by inpainting a depth value and a color value of the occlusion region with a depth image and a color image seen from the viewpoint after the viewpoint change, an image after the viewpoint change can be obtained. More specifically, if the depth image and the color image after the viewpoint change are three-dimensionally modeled by using the depth image and the color image from the original viewpoint, depth information and color information on the occlusion region that is not shown from the original viewpoint do not exist. Therefore, for the perfect inpainting, the depth information and the color information are needed. Therefore, according to an embodiment of the present invention, by inpainting the depth information and the color information on the occlusion region and combining the inpainted information with the depth image and the color image after the viewpoint change which can be acquired from the depth image and the color image from the original viewpoint, a 3D image after the viewpoint change can be obtained. Hereinafter, an embodiment of an apparatus for inpainting the depth value and the color value of the occlusion region is described.”) Skypnyk teaches generating an output color image and depth image based on prompt. Lim teaches how to generate better quality color image and depth image for a different viewpoint based on existing color image and depth image. it would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Skypnyk with the specific teachings of Lim to generate different viewpoint images based on exiting images. The benefit would be to provide a easy way to generate images with different viewpoints. However, Skypnyk in view of Lim does not, but Couleaud teaches: Generating image, by the diffusion model, by inpainting image occlusion area ([0066], “n some embodiments, as an alternative to the regeneration of such images (e.g., 216, 218, and 220 of FIG. 2) as images (e.g., 234, 236, and 238 of FIG. 2) based on a text prompt (e.g., one or more of text prompts 227-232 of FIG. 2), the image processing system may be configured to fill such holes or empty regions (e.g., 219, 221 in FIG. 2). As an example, the image processing system may perform completion (e.g., interpolation or extrapolation of image content) or inpainting 422 of such holes or empty regions in images 416, 418, and 420 to obtain updated images 434, 436, and 438. In some embodiments, such inpainting may be performed using one or more of the techniques described in Zheng et al., “Image Inpainting with Cascaded Modulation GAN and Object-Aware Training,” Computer Vision—ECCV 2022: 17th European Conference, Tel Aviv, Israel, Oct. 23-27, 2022, Proceedings, Part XVI, the contents of which is hereby incorporated by reference herein in its entirety.”) Skypnyk in view of Lim teaches handling occlusion in image generation, Couleaud teaches a specific method of handling occlusion. It would have been obvious for ordinary skills in the art to have combined the teachings of Skypnyk in view of Lim with the specific teachings of Couleaud to produce high quality output images. Regarding claim 3, Skrypnyk in view of Lim and Couleaud teaches: The method of claim 2, wherein identifying the occlusion area and the depth occlusion area comprises: computing a camera view of the image and the depth map; and shifting the camera view to obtain the modified view. (Lim [0020], “In a case where a depth image and a color image after a change in viewpoint are to be inpainted by using a depth image and a color image obtained from a viewpoint, if a portion that is not shown before the viewpoint change is shown after the viewpoint change, depth information and color information on the portion are to be inpainted from the depth image and the color image before the viewpoint change. When a region of an object in an image is hidden by an object having a depth value lower than that of the former, the region may not be shown from a predetermined viewpoint. This region is referred to as an occlusion region. A size of the occlusion region is determined according to a degree of change in viewpoint and a difference between the depth values of the objects. By combining an image inpainted by inpainting a depth value and a color value of the occlusion region with a depth image and a color image seen from the viewpoint after the viewpoint change, an image after the viewpoint change can be obtained. More specifically, if the depth image and the color image after the viewpoint change are three-dimensionally modeled by using the depth image and the color image from the original viewpoint, depth information and color information on the occlusion region that is not shown from the original viewpoint do not exist. Therefore, for the perfect inpainting, the depth information and the color information are needed. Therefore, according to an embodiment of the present invention, by inpainting the depth information and the color information on the occlusion region and combining the inpainted information with the depth image and the color image after the viewpoint change which can be acquired from the depth image and the color image from the original viewpoint, a 3D image after the viewpoint change can be obtained. Hereinafter, an embodiment of an apparatus for inpainting the depth value and the color value of the occlusion region is described.” The combination rationale of claim 1 is incorporated here.) Regarding claim 4, Skrypnyk in view of Lim and Couleaud teaches: The method of claim 1, further comprising: generating, by the diffusion model, an additional modified image based on the modified image. (Skypnyk [0191], “As such, the interaction system 100 applies the modified 3D mesh generated from a stable diffusion model to a live camera feed and create an augmented reality experience. This allows users to view and interact with the mesh in real-time, creating immersive and engaging experiences that combine the virtual and real worlds.” [0178]-[0191] teaches the details of generating modified images using a diffusion model.) Regarding claim 5, Skrypnyk in view of Lim and Couleaud teaches: The method of claim 1, further comprising: generating a video file based on the image and the modified image.( Skypnyk [0207], “FIG. 14 illustrates the content augmentation 1402 applied to the user when the user's head is centered, according to some examples. The generated augmentation content 1402 is applied to the user's face in the live camera feed when their head is centered and aligned. The camera feed shows the live video stream with the user's face centered, similar to FIG. 12, but with the applied augmentation. The user's face remains aligned and framed within the camera feed.” [0243], “The image display driver 1820 commands and controls the image display of optical assembly 1818. The image display driver 1820 may deliver image data directly to the image display of optical assembly 1818 for presentation or may convert the image data into a signal or data format suitable for delivery to the image display device. For example, the image data may be video data formatted according to compression formats, such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, or the like, and still image data may be formatted according to compression formats such as Portable Network Group (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF) or exchangeable image file format (EXIF) or the like.”) Regarding claim 7, Skrypnyk in view of Lim and Couleaud teaches: The method of claim 1, wherein: the prompt comprises a text prompt.( Skrypnyk [0147], “The user prompt includes a textual description or keyword(s) provided by the user, which defines the desired characteristics or subject of the generated image.”) Regarding claim 15, Skrypnyk in view of Lim and Couleaud teaches: A system for image generation, comprising: a memory component; and a processing device coupled to the memory component, the processing device configured to perform operations comprising: (Skrypnyk FIG. 19) the rest of claim 15 recites similar limitations of claim 1, thus are rejected accordingly. Regarding claim 16, Skrypnyk in view of Lim and Couleaud teaches: The system of claim 15, the system further comprising: an occlusion component (FIG. 19, processor) configured to generate the image occlusion area and the depth occlusion area. (Lim ([0020], “In a case where a depth image and a color image after a change in viewpoint are to be inpainted by using a depth image and a color image obtained from a viewpoint, if a portion that is not shown before the viewpoint change is shown after the viewpoint change, depth information and color information on the portion are to be inpainted from the depth image and the color image before the viewpoint change. When a region of an object in an image is hidden by an object having a depth value lower than that of the former, the region may not be shown from a predetermined viewpoint. This region is referred to as an occlusion region. A size of the occlusion region is determined according to a degree of change in viewpoint and a difference between the depth values of the objects. By combining an image inpainted by inpainting a depth value and a color value of the occlusion region with a depth image and a color image seen from the viewpoint after the viewpoint change, an image after the viewpoint change can be obtained. More specifically, if the depth image and the color image after the viewpoint change are three-dimensionally modeled by using the depth image and the color image from the original viewpoint, depth information and color information on the occlusion region that is not shown from the original viewpoint do not exist. Therefore, for the perfect inpainting, the depth information and the color information are needed. Therefore, according to an embodiment of the present invention, by inpainting the depth information and the color information on the occlusion region and combining the inpainted information with the depth image and the color image after the viewpoint change which can be acquired from the depth image and the color image from the original viewpoint, a 3D image after the viewpoint change can be obtained. Hereinafter, an embodiment of an apparatus for inpainting the depth value and the color value of the occlusion region is described.” The combination rationale of claim 15 is incorporated here.) Regarding claim 18, Skrypnyk in view of Lim and Couleaud teaches: The system of claim 15, the system further comprising: an encoder configured to generate a guidance embedding based on the prompt, wherein the image and the depth map are generated based on the guidance embedding (Skrypnyk [0162], “In the stable diffusion model, the user prompt 610 is converted into an embedding using a pre-trained language model, such as a transformer-based architecture (e.g., GPT or BERT). This embedding encodes the semantic information of the user prompt 610 and is used to condition the image/populated image template 612 generation process.” [0163], “The stable diffusion model is then conditioned on both the image template 612 and the text embedding by incorporating the text embedding into the stable diffusion model architecture or by conditioning the stable diffusion model's latent space on the textual information. The stable diffusion model generates populated image templates (such as images with color and depth information) by learning to reverse the process of adding noise to the input image template 612.”) Regarding claim 19, Skrypnyk in view of Lim and Couleaud teaches: The system of claim 15, the system further comprising: a training component configured to train the diffusion model. (Skrypnyk [0161], “the interaction system 100 trains a stable diffusion model to receive image templates and prompts and output populated image template that form the characteristics of the template.”) Regarding claim 21, Skrypnyk in view of Lim and Couleaud teaches: A non-transitory computer readable medium storing code for image generation, the code comprising instructions that, when executed by at least one processor, cause the at least one processor to perform operations (Skrypnyk FIG. 19) The rest of claims 21 recites similar limitations of claim 1, thus are rejected accordingly. claims 23, 24, 26, 27 recites similar limitations of claim 3, 4, 6,7 thus are rejected accordingly. Claim(s) 5, 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Skrypnyk in view of Lim and Couleaud and in view of Liao et al. (US 2021/0374904 A1). Regarding claim 5, Skrypnyk in view of Lim and Couleaud teaches: The method of claim 4, wherein generating the additional modified image comprises: However, Skrypnyk in view of Lim and Couleaud does not teach, but Liao teaches: averaging pixel information of the image and the modified image to obtain average pixel information, wherein the additional modified image is based on the average pixel information. ([0036], “In one embodiment, for each target pixel within the target inpainting region of the first image frame, based on the corresponding depth map, the method may further map the target pixel within the target inpainting region of the first image frame to a candidate pixel in a second image frame included in the image frames. The method may further determine a candidate color to fill the target pixel. The method may further perform Poisson image editing on the first image frame to achieve color consistency between inside and outside of the target inpainting region of the first image frame. The method may further use video fusion inpainting to inpaint occluded areas within the target inpainting region. For each pixel in the target inpainting region of the first image frame, the method may trace the pixel into neighboring frames and replacing an original color of the pixel with an average of colors sampled from the neighboring frames.”) Skrypnyk in view of Lim and Couleaud teaches inpainting occluded regions. Liao teaches a specific method of inpainting them. It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to have combined the teachings of Skrypnyk in view of Lim and Couleaud with the specific teachings of Liao to effectively and accurately inpainting occlusion areas. Claim 25 recites similar limitations of claim 5, thus is rejected accordingly. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to YANNA WU whose telephone number is (571)270-0725. The examiner can normally be reached Monday-Thursday 8:00-5:30 ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YANNA WU/Primary Examiner, Art Unit 2615
Read full office action

Prosecution Timeline

Show 8 earlier events
Apr 06, 2026
Request for Continued Examination
Apr 07, 2026
Response after Non-Final Action
Apr 29, 2026
Non-Final Rejection mailed — §103
Jul 15, 2026
Interview Requested
Jul 22, 2026
Examiner Interview Summary
Jul 22, 2026
Applicant Interview (Telephonic)
Jul 27, 2026
Response Filed
Sep 02, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743561
Generating Technical Drawings From Building Information Models
2y 0m to grant Granted Sep 22, 2026
Patent 12724582
SYSTEMS, APPARATUSES, AND METHODS FOR REAL-TIME COLLABORATION WITH A GRAPHICAL RENDERING PROGRAM
3y 0m to grant Granted Sep 01, 2026
Patent 12725352
INFORMATION PROCESSING DEVICE AND INFORMATION PROCESSING METHOD
2y 1m to grant Granted Sep 01, 2026
Patent 12718475
SITE MODEL UPDATING METHOD AND SYSTEM
3y 2m to grant Granted Aug 25, 2026
Patent 12711690
SPECULATIVE EXECUTION OF HIT AND INTERSECTION SHADERS ON PROGRAMMABLE RAY TRACING ARCHITECTURES
1y 10m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
81%
Grant Probability
99%
With Interview (+34.4%)
2y 2m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 456 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month