Prosecution Insights
Last updated: August 18, 2026
Application No. 18/975,445

DEPTH MAP GENERATION METHOD USING PREVIOUS FRAME INFORMATION FOR FAST DEPTH ESTIMATION

Non-Final OA §103§112
Filed
Dec 10, 2024
Priority
Dec 28, 2023 — RE 10-2023-0193784
Examiner
LEE, JONATHAN S
Art Unit
Tech Center
Assignee
Korea Electronics Technology Institute
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
6m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
507 granted / 599 resolved
+24.6% vs TC avg
Moderate +9% lift
Without
With
+9.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
21 currently pending
Career history
611
Total Applications
across all art units

Statute-Specific Performance

§101
4.2%
-35.8% vs TC avg
§103
47.2%
+7.2% vs TC avg
§102
26.3%
-13.7% vs TC avg
§112
12.1%
-27.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 599 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “extraction unit” and generation unit” in claim 10. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-4 and 6-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (TemporalStereo: Efficient Spatial-Temporal Stereo Matching Network, 1-5 October 2023, IEEE/RSJ International Conference on Intelligent Robots and Systems, Pages 9528-9535), hereinafter “Zhang”, in view of Stack Overflow (What is the difference between disparity and depth?, 6 April 2020, Stack Overflow, computer vision - What is difference between disparity and depth? - Stack Overflow). Claim 1 is met by the combination of Zhang and Stack Overflow, wherein Zhang discloses: A depth map generation method (See the Abstract.) comprising: a step of extracting a feature map of a 1/n scale resolution on a stereo image of a current frame (See page 9529, section III.A, 1st paragraph: “Given a single stereo pair, a backbone extracts multi-scale features at 1/4, 1/8, and 1/16 of the original resolution.” Also see the first line on page 9530: “Then, the resulting feature maps are processed by two 2D convolutions with kernels of 3, obtaining Fs l , Fs r.”); and a first generation step of generating a [disparity] map of a 1/n scale resolution of the current frame by using a [disparity] map of a 1/n scale resolution on a stereo image of a previous frame, and the extracted feature map (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities, costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).” Then see Fig. 2, where in temporal mode, cached past candidates (including past disparities) are used to generate the next disparity map (that is at the same 1/n scale before the depicted upsampling operation).). Zhang does not explicitly disclose the use of depth maps, instead relying on the inverse depth (i.e., disparity maps). However, Stack Overflow discloses (as part of the answer given by Harshit Kumar) in a similar field of art: “The depth (the actual z location of 3d point) can be calculated by using the disparity of the corresponding point e.g. in simple cases, as follows: PNG media_image1.png 46 428 media_image1.png Greyscale where baseline is the distance b/w the cameras. By getting the depth of every pixel, you get the depth map/image.” Modifying the system and method of Zhang by adding the capability to calculate and use depth maps, as taught by Stack Overflow, would yield the expected and predictable result of better visualization and physical interpretation (over disparity) to humans viewing the maps. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang and Stack Overflow in this way. Claim 2 is met by the combination of Zhang and Stack Overflow, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 1, wherein the first generation step comprises: And Zhang (as modified by Stack Overflow) further discloses: a step of warping the depth map of the 1/n scale resolution of the previous frame (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities [or as modified by Stack Overflow, past depth values], costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).”); and a second generation step of generating the depth map of the 1/n scale resolution of the current frame by using the warped depth map and the extracted feature map (See Fig. 2, where in temporal mode, cached past candidates (including aligned/warped past disparities/depth values) are used to generate the next disparity/depth map (that is at the same 1/n scale before the depicted upsampling operation).). Claim 3 is met by the combination of Zhang and Stack Overflow, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 2, wherein the step of warping comprises: And Zhang (as modified by Stack Overflow) further discloses: a step of calculating an optical flow between the previous frame and the current frame; and a step of warping the depth map of the previous frame by using the calculated optical flow (See page 9531, right column, 1st partial paragraph: “Instead, in this work, we suppose that the camera is calibrated and the pose is given–since pose can be provided by external IMU sensors or estimated from the stereo pairs to build the rigid motion field needed to align the candidates.” The examiner asserts that this is a form of optical flow.). Claim 4 is met by the combination of Zhang and Stack Overflow, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 2, wherein the step of warping comprises: And Zhang (as modified by Stack Overflow) further discloses: The depth map generation method of claim 2, wherein the second generation step comprises generating the depth map of the current frame from the warped depth map and the extracted feature map, by using a neural network that is trained to generate a depth map of a current frame from a feature map and a depth map of a previous frame (See the network architecture in Fig. 2.). Claim 6 is met by the combination of Zhang and Stack Overflow, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 1, further comprising a step of And Zhang (as modified by Stack Overflow) further discloses: up-scaling the depth map generated at the second generation step to an original scale resolution (See page 9530, Fig. 2, an upsampled disparity map (depth map as modified by Stack Overflow) is generated.). Claim 7 is met by the combination of Zhang and Stack Overflow, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 6, wherein the step of up-scaling comprises And Zhang further discloses: up-scaling the depth map generated at the second generation step to the original scale resolution by using a neural network that is trained to generate a depth map of an original scale resolution from a depth map of a 1/n scale resolution (See the network in Fig. 2.). Claim 8 is met by the combination of Zhang and Stack Overflow, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 1, wherein And Zhang further discloses: n is a single value (See page 9529, section III.A, 1st full paragraph: “Given a single stereo pair, a backbone extracts multi-scale features at 1/4, 1/8, and 1/16 of the original resolution. Then, three stages predict disparity maps starting from these features.” For each stage in Fig. 2, n has a single value.). Claim 9 is met by the combination of Zhang and Stack Overflow, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 1, wherein the step of extracting comprises And Zhang further discloses: extracting a feature map of a 1/n scale resolution on a left-eye image and a feature map of a 1/n scale resolution on a right-eye image (See page 9529, section III.A., 2nd full paragraph: “In each stage s ∈ {1,2,3}, a decoder processes the current features, together with those from the previous stage Fs−1 l Fs−1 l , Fs−1 r , Fs−1 r if s > 1. In particular, are bilinearly upsampled by a factor 2 and concatenated with left and right features from the backbone.”). Claim 10 is met by the combination of Zhang and Stack Overflow for the reasons given in the treatment of claim 1. Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (TemporalStereo: Efficient Spatial-Temporal Stereo Matching Network, 1-5 October 2023, IEEE/RSJ International Conference on Intelligent Robots and Systems, Pages 9528-9535), in view of Stack Overflow (What is the difference between disparity and depth?, 6 April 2020, Stack Overflow, computer vision - What is difference between disparity and depth? - Stack Overflow), in view of Lipson et al. (RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching, 2021, International Conference on 3D Vision, Pages 218-227), hereinafter “Lipson”. Claim 5 is met by the combination of Zhang, Stack Overflow, and Lipson, wherein The combination of Zhang and Stack Overflow discloses: The depth map generation method of claim 2, wherein, when the current frame is a first frame, the step of warping is not performed, and (See page 9530, Fig. 2 caption: “In single-pair mode, the model predicts the disparity map in a coarse-to-fine manner. If past pairs are available, the same model switches to temporal mode”. The current frame is a first frame and no warping is performed in the case of single-pair mode.) The combination of Zhang and Stack Overflow does not disclose the following; however, Lipson discloses: the second generation step comprises generating the depth map of the 1/n scale resolution of the current frame by using a depth map which is filled with 0 and the extracted feature map (See page 219, Fig. 1 and its caption: “’Context’ image features (white) and an initial hidden state are also extracted from the context encoder. The disparity field is initialized to zero. Every iteration, the GRU(s) (green) use the current disparity estimate to sample from the correlation pyramid. The resulting correlation features, initial image features and current hidden state(s) are used by the GRU(s) to produce a new hidden state and an update to the disparity.”). Zhang, Stack Overflow, and Lipson together disclose the limitations of claim 5. Lipson is directed to a similar field of art (video stereo matching). Therefore, Zhang, Stack Overflow, and Lipson are combinable. Modifying the system and method of Zhang and Stack Overflow by adding the capability of “generating the depth map of the 1/n scale resolution of the current frame by using a depth map which is filled with 0 and the extracted feature map”, as disclosed by Lipson (as modified by Stack Overflow to generate depth maps), would yield the expected and predictable result of an improved initial depth map estimate that accounts for high-uncertainty motion regions by zeroing. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang, Stack Overflow, and Lipson in this way. Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (TemporalStereo: Efficient Spatial-Temporal Stereo Matching Network, 1-5 October 2023, IEEE/RSJ International Conference on Intelligent Robots and Systems, Pages 9528-9535), in view of Stack Overflow (What is the difference between disparity and depth?, 6 April 2020, Stack Overflow, computer vision - What is difference between disparity and depth? - Stack Overflow), in view of Xu et al. (Multiscale Attention Fusion for Depth Map Super-Resolution Generative Adversarial Networks, 23 May 2023, Entropy, Vol. 25, No. 836, Pages 1-15), hereinafter “Xu”. Claim 11 is met by the combination of Zhang, Stack Overflow, and Xu, wherein Zhang discloses: A depth map generation method (See the Abstract.) comprising: a step of warping a [disparity] map of a 1/n scale resolution on a stereo image of a previous frame (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities [or as modified by Stack Overflow, past depth values], costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).”); a step of generating a [disparity] map of a 1/n scale resolution of a current frame by using the warped [disparity] map and a feature map of a 1/n scale resolution which is extracted from a stereo image of a current frame (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities, costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).” Then see Fig. 2, where in temporal mode, cached past candidates (including past disparities) are used to generate the next disparity map (that is at the same 1/n scale before the depicted upsampling operation).); and Zhang does not explicitly disclose the use of depth maps, instead relying on the inverse depth (i.e., disparity maps). However, Stack Overflow discloses (as part of the answer given by Harshit Kumar) in a similar field of art: “The depth (the actual z location of 3d point) can be calculated by using the disparity of the corresponding point e.g. in simple cases, as follows: PNG media_image1.png 46 428 media_image1.png Greyscale where baseline is the distance b/w the cameras. By getting the depth of every pixel, you get the depth map/image.” Modifying the system and method of Zhang by adding the capability to calculate and use depth maps, as taught by Stack Overflow, would yield the expected and predictable result of better visualization and physical interpretation (over disparity) to humans viewing the maps. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang and Stack Overflow in this way. The combination of Zhang and Stack Overflow does not explicitly disclose the following; however, Xu discloses: a step of up-scaling the generated depth map to an n scale resolution (See in Fig. 1, a network that generates a high-resolution depth map from a low-resolution depth map. In Table 2 on page 10, upsampling factors of 4x, 8x, and 16x are used, which match the reduced resolutions used in Zhang (i.e., 1/4, 1/8, and 1/16).). Zhang, Stack Overflow, and Xu together disclose the limitations of claim 11. Xu is directed to a similar field of art (a network for stereo matching of video sequences). Therefore, Zhang, Stack Overflow, and Xu are combinable. Modifying the system and method of Zhang and Stack Overflow by adding the capability of “up-scaling the generated depth map to an n scale resolution”, as taught by Xu, would yield the expected and predictable result of sharpening object boundaries for improved 3D reconstruction and autonomous vehicle navigation applications. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang, Stack Overflow, and Xu in this way. Contact Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN S LEE whose telephone number is (571)272-1981. The examiner can normally be reached 11:30 AM - 7:30 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at (571)270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Jonathan S Lee/Primary Examiner, Art Unit 2677
Read full office action

Prosecution Timeline

Dec 10, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705752
SYSTEMS AND METHODS FOR CONTINUOUS ADAPTATION OF SEMANTIC IMAGE SEGMENTATION MODEL
3y 6m to grant Granted Aug 11, 2026
Patent 12682438
Systems and Methods for Automated Clinical Image Quality Assessment
3y 9m to grant Granted Jul 14, 2026
Patent 12657857
ILLUMINATION ADJUSTING DEVICE, ILLUMINATION ADJUSTING METHOD, AND PRODUCT RECOGNITION SYSTEM
2y 7m to grant Granted Jun 16, 2026
Patent 12657687
AN APPARATUS AND METHOD FOR GENERATING RANDOM IMAGES OF A HERBAL FORMULATION
2y 3m to grant Granted Jun 16, 2026
Patent 12639851
SYSTEM AND METHOD FOR LEARNING TO SYNTHESIZE HAND-OBJECT INTERACTION SCENE
2y 7m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
94%
With Interview (+9.3%)
2y 3m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 599 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month