Prosecution Insights
Last updated: October 02, 2026
Application No. 19/002,798

PANORAMIC DEPTH MAP GENERATION METHOD, MODEL TRAINING METHOD, ELECTRONIC DEVICE, AND UNMANNED VEHICLE

Non-Final OA §102§103
Filed
Dec 27, 2024
Priority
Nov 28, 2024 — continuation of PCTCN2024135410
Examiner
PERLMAN, DAVID S
Art Unit
2673
Tech Center
2600 — Communications
Assignee
Arashi Vision Inc.
OA Round
1 (Non-Final)
81%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
445 granted / 550 resolved
+18.9% vs TC avg
Moderate +13% lift
Without
With
+12.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
11 currently pending
Career history
556
Total Applications
across all art units

Statute-Specific Performance

§101
9.6%
-30.4% vs TC avg
§103
55.5%
+15.5% vs TC avg
§102
19.9%
-20.1% vs TC avg
§112
11.9%
-28.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 550 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 12/27/2024 and 08/20/2025 have been considered by the examiner. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 7-8, 12-14, and 16-17 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chen et al. (“Unsupervised OmniMVS: Efficient Omnidirectional Depth Inference via Establishing Pseudo-Stereo Supervision”). Regarding claim 1, Chen discloses, a method for generating a target panoramic depth map, comprising: grouping target images involving different orientations in a target scene to form at least two image groups, (See Chen Fig. 3, “Spherical sweeping & Panorama feature”, where there are four feature volumes and two groups each containing two feature volumes, where each feature volume comes from a fisheye image. As shown in Fig. 2, there are four fisheye cameras each at a different orientation.) the target images in each of the at least two image groups covering a panoramic field of view of the target scene; (See Chen p. 4 left col 2nd para, “2) Light Cost Volume: According to Eq. (5), we can obtain four panoramic feature maps with the size of CxNxH^fxW^f by projecting four fisheye features to N depth planes, as shown in Fig. 3.”) performing feature combination on target image feature volumes of each of the image groups to obtain at least two target panoramic feature volumes, each of the target image feature volumes representing three-dimensional stereoscopic features of one of the target images; (See Chen p. 4 left col 3rd para, “First, we establish the four spherical feature volumes from the fisheye features and crop their FoV to 180°. Then, two 360° spherical feature volumes can be obtained by stitching the four 180° feature volumes.”) performing correlation processing on every two of the at least two target panoramic feature volumes to obtain at least one target correlation volume; (See Chen p. 4 left col 3rd para, “The two processed panoramic feature volumes are more discriminative and can be denoted as F1omni, F2omni. Our cost volume is calculated as the following formulation: V = Sum, i=1…2 (Fiomni-Fbar)^2 / 2 (6), where Fbar is the average volume of F1omni and F2omni.”) and performing panoramic depth estimation based on an initial depth map and the at least one target correlation volume to obtain the target panoramic depth map for the target scene. (See Chen p. 4 left col, 4th para, “3) Disparity regression: in the last stage, we leverage the 3D codec to process the cost volume and get a single channel volume V* with the size of N x H x W. The the disparity map can be obtained by a softargmin as D(Ө,φ) = Sum [D=Dmin…Dmax] D x softmax(V*(Ө,φ,D), where Dmax and Dmin are the maximum and minimum values of the disparity. Depth map is Ddepth = 1/D(Ө,φ).” Where the cost volume is considered to be a correlation volume, and the output of D(Ө,φ) for D=Dmin is considered to be the initial disparity/depth map.) Regarding claim 7, Chen discloses, the method according to claim 1, wherein the grouping target images involving different orientations in the target scene to form the at least two image groups comprises: grouping every two target images facing back-to-back among the target images involving different orientations in a target scene into one image group to obtain the at least two image groups, wherein the two target images facing back-to-back represent that the orientations of the two target images are opposite. (See Chen p. 4 left col 3rd para, “First, we establish the four spherical feature volumes from the fisheye features and crop their FoV to 180°. Then, two 360° spherical feature volumes can be obtained by stitching the four 180° feature volumes.” As shown in Fig. 2, the fisheye images are from four orientations where the cameras are facing back-to-back. To obtain a 360-degree feature volume from two 180 feature volumes, the feature volumes must come from back-to-back fisheye images.) Regarding claim 8, Chen discloses the method according to claim 1, wherein the performing correlation processing on every two of the at least two target panoramic feature volumes to obtain at least one target correlation volume comprises: performing an inner product calculation on the every two of the at least two target panoramic feature volumes to obtain the at least one target correlation volume. (See Chen p. 4 left col 3rd para, “The two processed panoramic feature volumes are more discriminative and can be denoted as F1omni, F2omni. Our cost volume is calculated as the following formulation: V = Sum, i=1…2 (Fiomni-Fbar)^2 / 2 (6), where Fbar is the average volume of F1omni and F2omni.” The squared value in the equation involves matrix multiplication of Fiomni by Fbar. Matrix multiplication uses dot products to multiply column vectors with row vectors.) Regarding claim 12, Chen discloses, the method according to claim 1, further comprising: performing feature extraction on each of the target images to obtain a target feature map respectively; (See Chen Fig. 3, “feature extraction”, where the four fisheye images are input to a neural network that includes a FDAM (Frequency Attention Module) in order to output four feature volumes. Further see Chen p. 4 left col 3rd para, “First, we establish the four spherical feature volumes from the fisheye features.”) and performing spherical scanning processing on the target feature map to obtain a target image feature volume for each of the target images. (See Chen Fig. 3, caption, “Spherical sweeping & Panorama Feature: Complete the projection of input features to spherical features and the stitching of panoramic features.”) Regarding claim 13, Chen discloses the method according to claim 1, wherein the target images involving different orientations in the target scene include target images in at least four orientations. (See Chen Fig. 2, where the fisheye images are from four different fisheye camera orientations.) Regarding claim 14, Chen discloses, the method according to claim 1, wherein target images involving different orientations in the target scene are acquired by fisheye lenses at different orientations in the target scene. (See Chen Fig. 2, where the fisheye images are from four different fisheye camera orientations.) Regarding claim 16, Chen discloses, a method for training a depth estimation model, comprising: grouping sample images involving different orientations in a sample scene based on training samples to form at least two sample image groups, (See Chen Fig. 3, “Spherical sweeping & Panorama feature”, where there are four feature volumes and two groups each containing two feature volumes, where each feature volume comes from a fisheye image. As shown in Fig. 2, there are four fisheye cameras each at a different orientation.” Further see Chen p. 5 left col 2nd para, “Dataset: OmniMVS [24] provides three synthetic datasets, in which depth labels, extrinsic parameters, intrinsic parameters, and four fisheye images (H1 = 768; W1 = 800) are provided. … The OmniThings contains 10,240 diverse scenes, and it is the largest and richest dataset. We select 4000 images as the training set and 1000 images as the test set.”) the sample images in each of the at least two sample image groups covering a panoramic field of view of the sample scene; (See Chen p. 4 left col 2nd para, “2) Light Cost Volume: According to Eq. (5), we can obtain four panoramic feature maps with the size of CxNxH^fxW^f by projecting four fisheye features to N depth planes, as shown in Fig. 3.”) performing feature combination on sample image feature volumes of each of the sample image groups to obtain at least two sample panoramic feature volumes, each of the sample image feature volumes representing three-dimensional features of one of the sample images; (See Chen p. 4 left col 3rd para, “First, we establish the four spherical feature volumes from the fisheye features and crop their FoV to 180°. Then, two 360° spherical feature volumes can be obtained by stitching the four 180° feature volumes.”) performing correlation processing on every two of the at least two sample panoramic feature volumes to obtain at least one sample correlation volume; (See Chen p. 4 left col 3rd para, “The two processed panoramic feature volumes are more discriminative and can be denoted as F1omni, F2omni. Our cost volume is calculated as the following formulation: V = Sum, i=1…2 (Fiomni-Fbar)^2 / 2 (6), where Fbar is the average volume of F1omni and F2omni.”) performing panoramic depth estimation based on an initial depth map and the at least one sample correlation volume to obtain a predicted panoramic depth map for the sample scene; (See Chen p. 4 left col, 4th para, “3) Disparity regression: in the last stage, we leverage the 3D codec to process the cost volume and get a single channel volume V* with the size of N x H x W. The the disparity map can be obtained by a softargmin as D(Ө,φ) = Sum [D=Dmin…Dmax] D x softmax(V*(Ө,φ,D), where Dmax and Dmin are the maximum and minimum values of the disparity. Depth map is Ddepth = 1/D(Ө,φ).” Where the cost volume is considered to be a correlation volume, and the output of D(Ө,φ) for D=Dmin is considered to be the initial disparity/depth map.) determining loss information of the depth estimation model according to the predicted panoramic depth map; (See Chen p. 4 right col, Section D, “The total optimization goal is defined as follows (β1 = 1, β2 = 2, β3 = 1): Ltotal = β1Lp + β2L2 + B3Lg,, (9) where Lp, Ls, and Lg are the photometric loss, smoothness loss and gradient loss.”) adjusting network parameters of the depth estimation model iteratively according to the loss information until the loss information satisfies an iteration stop condition, and determining the network parameters obtained when the loss information satisfies the iteration stop condition as the trained depth estimation model. (See Chen p. 5 left col Section B, “We use PyTorch to implement the proposed end-to-end network. During training, the number of spheres N is set to 32, and the number of cost volume channels C is set to 32. The minimum distance is set to 0.55 meters, and the maximum distance is set to 105 meters. The spherical depth is selected according to the 32 equal divisions of the inverse depth, and the disparity map resolution is set to 640_320. Our network is trained using an Adam optimizer [11] with an exponentially decaying learning rate initialized to 1x10^-4 for the 30 epochs. In the inference, it takes 0.33s for our method to accomplish an iteration, which is 2x faster than the 0.67s of OmniMVS.”) Regarding claim 17, Chen discloses, an electronic device, comprising: at least one processor; and at least one memory for storing at least one program, wherein, the at least one processor, when executing the at least one program, is configured to: (The image processing of Chen inherently uses a processor and memory for storing at least one program.) group target images involving different orientations in a target scene to form at least two image groups, (See Chen Fig. 3, “Spherical sweeping & Panorama feature”, where there are four feature volumes and two groups each containing two feature volumes, where each feature volume comes from a fisheye image. As shown in Fig. 2, there are four fisheye cameras each at a different orientation.) the target images in each of the at least wo image groups covering a panoramic field of view of the target scene; (See Chen p. 4 left col 2nd para, “2) Light Cost Volume: According to Eq. (5), we can obtain four panoramic feature maps with the size of CxNxH^fxW^f by projecting four fisheye features to N depth planes, as shown in Fig. 3.”) perform feature combination on target image feature volumes of each of the image groups to obtain at least two target panoramic feature volumes, each of the target image feature volumes representing three-dimensional stereoscopic features of one of the target images; (See Chen p. 4 left col 3rd para, “First, we establish the four spherical feature volumes from the fisheye features and crop their FoV to 180°. Then, two 360° spherical feature volumes can be obtained by stitching the four 180° feature volumes.”) perform correlation processing on every two of the at least two target panoramic feature volumes to obtain at least one target correlation volume; (See Chen p. 4 left col 3rd para, “The two processed panoramic feature volumes are more discriminative and can be denoted as F1omni, F2omni. Our cost volume is calculated as the following formulation: V = Sum, i=1…2 (Fiomni-Fbar)^2 / 2 (6), where Fbar is the average volume of F1omni and F2omni.”) and perform panoramic depth estimation based on an initial depth map and the at least one target correlation volume to obtain the target panoramic depth map for the target scene. (See Chen p. 4 left col, 4th para, “3) Disparity regression: in the last stage, we leverage the 3D codec to process the cost volume and get a single channel volume V* with the size of N x H x W. The the disparity map can be obtained by a softargmin as D(Ө,φ) = Sum [D=Dmin…Dmax] D x softmax(V*(Ө,φ,D), where Dmax and Dmin are the maximum and minimum values of the disparity. Depth map is Ddepth = 1/D(Ө,φ).” Where the cost volume is considered to be a correlation volume, and the output of D(Ө,φ) for D=Dmin is considered to be the initial disparity/depth map.) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 15 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (“Unsupervised OmniMVS: Efficient Omnidirectional Depth Inference via Establishing Pseudo-Stereo Supervision”) in view of Zhang et al. (US Pub. No. 2020/0201361 A1). Regarding claim 15, Chen discloses the method according to claim 1, but he fails to disclose, further comprising: determining obstacles according to the panoramic depth map; and controlling an unmanned vehicle to perform obstacle avoidance processing according to the obstacles. However, Zhang discloses, further comprising: determining obstacles according to the panoramic depth map; (See Zhang ¶39, “In some embodiments, during the process of generating the obstacle map, the obstacle map may be generated based on multiple frames of the depth data.”) and controlling an unmanned vehicle to perform obstacle avoidance processing according to the obstacles. (See Zhang ¶41, “In some embodiments, after generating the obstacle map, the UAV have already known the location distribution situation of the obstacles in the flight space. The UAV may determine whether to trigger an obstacle avoidance operation based on the generated obstacle map and the location of the UAV.”) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include the obstacle avoidance for a UAV using a depth data map as suggested by Zhang to Chen’s panorama depth map using known engineering techniques, with a reasonable expectation of success. The motivation for doing so is because knowing the exact 3D coordinates of a hazard lets the vehicle plot a smooth, collision free route around the object rather than just stopping blindly. Regarding claim 18, Chen and Zhang disclose, the electronic device according to claim 17, wherein the at least one processor is further configured to: determine obstacles according to the target panoramic depth map; and control an unmanned vehicle to perform obstacle avoidance processing according to the obstacles. (See the rejection of claim 15 as it is equally applicable for claim 18 as well.) Regarding claim 19, Chen and Zhang disclose, an unmanned vehicle comprising the electronic device according to claim 17. (See Zhang ¶70, “The present disclosure provides a UAV control device (i.e., a device for controlling a UAV). FIG. 10 is a schematic illustration of a structure of the UAV control device.”) Regarding claim 20, Chen and Zhang disclose the unmanned vehicle of claim 18, wherein the unmanned vehicle comprises an unmanned aerial vehicle or an unmanned robot. (See Zhang ¶70, “The present disclosure provides a UAV control device (i.e., a device for controlling a UAV). FIG. 10 is a schematic illustration of a structure of the UAV control device.”) Allowable Subject Matter Claims 2-6 and 9-11 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding claim 2, the method according to claim 1, wherein each of the target image feature volumes has a corresponding target weight, and the performing feature combination on target image feature volumes of each of the image groups to obtain at least two target panoramic feature volumes comprises: performing weighted summation of target weight of each of the target image feature volumes in each of the image groups to obtain the at least two target panoramic feature volumes. (The disclosed prior art of record fails to disclose the limitations of this claim.) Regarding claim 9, the method according to claim 1, wherein the performing panoramic depth estimation based on the initial depth map and the at least one target correlation volume to obtain the target panoramic depth map for the target scene comprises: performing the panoramic depth estimation based on the initial depth map, a preset target context feature volume and the at least one target correlation volume to obtain the target panoramic depth map for the target scene, wherein the target context feature volume is determined based on at least one of the target panoramic feature volumes. (The disclosed prior art of record fails to disclose the limitations of this claim.) Regarding claims 3-6 and 10-11, these claims are objected to, since they depend on objected to claims 2 and 9 respectively. Conclusion Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure. Calligaro et al. (US Pub. No. 2024/0305739 A1) method is described for forming a panoramic image (I); the method comprising: receiving (B1) a plurality of images (I1, I2 . . . Ir) of an environment from a corresponding plurality of cameras (C1, C2 . . . Cr) at a given time instant (t), receiving (B1) data on the position of the points relative to the physical objects of said environment shot by said plurality of cameras by at least one depth sensor (L1, L2 . . . Lf) at said given time instant, processing (B2) data obtained from the at least one depth sensor to construct the distance of all the objects contained inside said environment by a virtual camera (C), obtaining a three-dimensional map of the positions of the points relative to the physical objects of said environment. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID PERLMAN whose telephone number is (571) 270-1417. The examiner can normally be reached on Monday - Friday; 10:00am -6:30pm. Examiner interviews are available via telephone and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is (571) 273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at (866) 217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call (800) 786-9199 (IN USA OR CANADA) or (571) 272-1000. /DAVID PERLMAN/Primary Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Dec 27, 2024
Application Filed
Sep 17, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749281
METHOD AND APPARATUS FOR DETECTING CARGO IN CONTAINER IMAGE USING CONTAINER WALL BACKGROUND REMOVAL
3y 4m to grant Granted Sep 29, 2026
Patent 12738006
IMAGE PROCESSING APPARATUS, CONTROL METHOD THEREFOR, IMAGE CAPTURING APPARATUS, AND STORAGE MEDIUM
3y 8m to grant Granted Sep 15, 2026
Patent 12738367
BLOCKCHAIN-BASED AND HUMAN CHARACTERISTICS INTELLIGENCE RECOGNITION FOR APPOINTMENT VISUALIZATION ELDERLY CARE SYSTEM
3y 4m to grant Granted Sep 15, 2026
Patent 12738019
IMAGE MATCHING USING CHROMA-ENHANCED OR ADAPTABLE SIZE FEATURES
2y 4m to grant Granted Sep 15, 2026
Patent 12738056
APPARATUS FOR RECOGNIZING ACTIVITY IN SPORTS VIDEO USING CROSS GRANULARITY ACCUMULATION MODULE AND METHOD THEREOF
1y 10m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
81%
Grant Probability
94%
With Interview (+12.6%)
2y 6m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 550 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month