DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “extraction unit” and generation unit” in claim 10.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-4 and 6-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (TemporalStereo: Efficient Spatial-Temporal Stereo Matching Network, 1-5 October 2023, IEEE/RSJ International Conference on Intelligent Robots and Systems, Pages 9528-9535), hereinafter “Zhang”, in view of Stack Overflow (What is the difference between disparity and depth?, 6 April 2020, Stack Overflow, computer vision - What is difference between disparity and depth? - Stack Overflow).
Claim 1 is met by the combination of Zhang and Stack Overflow, wherein
Zhang discloses:
A depth map generation method (See the Abstract.) comprising:
a step of extracting a feature map of a 1/n scale resolution on a stereo image of a current frame (See page 9529, section III.A, 1st paragraph: “Given a single stereo pair, a backbone extracts multi-scale features at 1/4, 1/8, and 1/16 of the original resolution.” Also see the first line on page 9530: “Then, the resulting feature maps are processed by two 2D convolutions with kernels of 3, obtaining Fs l , Fs r.”); and
a first generation step of generating a [disparity] map of a 1/n scale resolution of the current frame by using a [disparity] map of a 1/n scale resolution on a stereo image of a previous frame, and the extracted feature map (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities, costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).” Then see Fig. 2, where in temporal mode, cached past candidates (including past disparities) are used to generate the next disparity map (that is at the same 1/n scale before the depicted upsampling operation).).
Zhang does not explicitly disclose the use of depth maps, instead relying on the inverse depth (i.e., disparity maps). However, Stack Overflow discloses (as part of the answer given by Harshit Kumar) in a similar field of art: “The depth (the actual z location of 3d point) can be calculated by using the disparity
of the corresponding point e.g. in simple cases, as follows:
PNG
media_image1.png
46
428
media_image1.png
Greyscale
where baseline is the distance b/w the cameras. By getting the depth of every pixel, you get the depth map/image.” Modifying the system and method of Zhang by adding the capability to calculate and use depth maps, as taught by Stack Overflow, would yield the expected and predictable result of better visualization and physical interpretation (over disparity) to humans viewing the maps. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang and Stack Overflow in this way.
Claim 2 is met by the combination of Zhang and Stack Overflow, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 1, wherein the first generation step comprises:
And Zhang (as modified by Stack Overflow) further discloses:
a step of warping the depth map of the 1/n scale resolution of the previous frame (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities [or as modified by Stack Overflow, past depth values], costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).”); and a second generation step of generating the depth map of the 1/n scale resolution of the current frame by using the warped depth map and the extracted feature map (See Fig. 2, where in temporal mode, cached past candidates (including aligned/warped past disparities/depth values) are used to generate the next disparity/depth map (that is at the same 1/n scale before the depicted upsampling operation).).
Claim 3 is met by the combination of Zhang and Stack Overflow, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 2, wherein the step of warping comprises:
And Zhang (as modified by Stack Overflow) further discloses:
a step of calculating an optical flow between the previous frame and the current frame; and a step of warping the depth map of the previous frame by using the calculated optical flow (See page 9531, right column, 1st partial paragraph: “Instead, in this work, we suppose that the camera is calibrated and the pose is given–since pose can be provided by external IMU sensors or estimated from the stereo pairs to build the rigid motion field needed to align the candidates.” The examiner asserts that this is a form of optical flow.).
Claim 4 is met by the combination of Zhang and Stack Overflow, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 2, wherein the step of warping comprises:
And Zhang (as modified by Stack Overflow) further discloses:
The depth map generation method of claim 2, wherein the second generation step comprises generating the depth map of the current frame from the warped depth map and the extracted feature map, by using a neural network that is trained to generate a depth map of a current frame from a feature map and a depth map of a previous frame (See the network architecture in Fig. 2.).
Claim 6 is met by the combination of Zhang and Stack Overflow, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 1, further comprising a step of
And Zhang (as modified by Stack Overflow) further discloses:
up-scaling the depth map generated at the second generation step to an original scale resolution (See page 9530, Fig. 2, an upsampled disparity map (depth map as modified by Stack Overflow) is generated.).
Claim 7 is met by the combination of Zhang and Stack Overflow, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 6, wherein the step of up-scaling comprises
And Zhang further discloses:
up-scaling the depth map generated at the second generation step to the original scale resolution by using a neural network that is trained to generate a depth map of an original scale resolution from a depth map of a 1/n scale resolution (See the network in Fig. 2.).
Claim 8 is met by the combination of Zhang and Stack Overflow, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 1, wherein
And Zhang further discloses:
n is a single value (See page 9529, section III.A, 1st full paragraph: “Given a single stereo pair, a backbone extracts multi-scale features at 1/4, 1/8, and 1/16 of the original resolution. Then, three stages predict disparity maps starting from these features.” For each stage in Fig. 2, n has a single value.).
Claim 9 is met by the combination of Zhang and Stack Overflow, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 1, wherein the step of extracting comprises
And Zhang further discloses:
extracting a feature map of a 1/n scale resolution on a left-eye image and a feature map of a 1/n scale resolution on a right-eye image (See page 9529, section III.A., 2nd full paragraph: “In each stage s ∈ {1,2,3}, a decoder processes the current features, together with those from the previous stage Fs−1 l Fs−1 l , Fs−1 r , Fs−1 r if s > 1. In particular, are bilinearly upsampled by a factor 2 and concatenated with left and right features from the backbone.”).
Claim 10 is met by the combination of Zhang and Stack Overflow for the reasons given in the treatment of claim 1.
Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (TemporalStereo: Efficient Spatial-Temporal Stereo Matching Network, 1-5 October 2023, IEEE/RSJ International Conference on Intelligent Robots and Systems, Pages 9528-9535), in view of Stack Overflow (What is the difference between disparity and depth?, 6 April 2020, Stack Overflow, computer vision - What is difference between disparity and depth? - Stack Overflow), in view of Lipson et al. (RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching, 2021, International Conference on 3D Vision, Pages 218-227), hereinafter “Lipson”.
Claim 5 is met by the combination of Zhang, Stack Overflow, and Lipson, wherein
The combination of Zhang and Stack Overflow discloses:
The depth map generation method of claim 2, wherein, when the current frame is a first frame, the step of warping is not performed, and (See page 9530, Fig. 2 caption: “In single-pair mode, the model predicts the disparity map in a coarse-to-fine manner. If past pairs are available, the same model switches to temporal mode”. The current frame is a first frame and no warping is performed in the case of single-pair mode.)
The combination of Zhang and Stack Overflow does not disclose the following; however, Lipson discloses:
the second generation step comprises generating the depth map of the 1/n scale resolution of the current frame by using a depth map which is filled with 0 and the extracted feature map (See page 219, Fig. 1 and its caption: “’Context’ image features (white) and an initial hidden state are also extracted from the context encoder. The disparity field is initialized to zero. Every iteration, the GRU(s) (green) use the current disparity estimate to sample from the correlation pyramid. The resulting correlation features, initial image features and current hidden state(s) are used by the GRU(s) to produce a new hidden state and an update to the disparity.”).
Zhang, Stack Overflow, and Lipson together disclose the limitations of claim 5. Lipson is directed to a similar field of art (video stereo matching). Therefore, Zhang, Stack Overflow, and Lipson are combinable. Modifying the system and method of Zhang and Stack Overflow by adding the capability of “generating the depth map of the 1/n scale resolution of the current frame by using a depth map which is filled with 0 and the extracted feature map”, as disclosed by Lipson (as modified by Stack Overflow to generate depth maps), would yield the expected and predictable result of an improved initial depth map estimate that accounts for high-uncertainty motion regions by zeroing. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang, Stack Overflow, and Lipson in this way.
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (TemporalStereo: Efficient Spatial-Temporal Stereo Matching Network, 1-5 October 2023, IEEE/RSJ International Conference on Intelligent Robots and Systems, Pages 9528-9535), in view of Stack Overflow (What is the difference between disparity and depth?, 6 April 2020, Stack Overflow, computer vision - What is difference between disparity and depth? - Stack Overflow), in view of Xu et al. (Multiscale Attention Fusion for Depth Map Super-Resolution Generative Adversarial Networks, 23 May 2023, Entropy, Vol. 25, No. 836, Pages 1-15), hereinafter “Xu”.
Claim 11 is met by the combination of Zhang, Stack Overflow, and Xu, wherein
Zhang discloses:
A depth map generation method (See the Abstract.) comprising:
a step of warping a [disparity] map of a 1/n scale resolution on a stereo image of a previous frame (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities [or as modified by Stack Overflow, past depth values], costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).”);
a step of generating a [disparity] map of a 1/n scale resolution of a current frame by using the warped [disparity] map and a feature map of a 1/n scale resolution which is extracted from a stereo image of a current frame (See page 9529, left column, 1st partial paragraph: “4) additionally, when multiple pairs are available, our network can easily switch to temporal mode, in which past disparities, costs and cached features could also be aligned to the current reference frame and used to boost current estimates with a negligible runtime increase (4ms as shown in Tab. VI).” Then see Fig. 2, where in temporal mode, cached past candidates (including past disparities) are used to generate the next disparity map (that is at the same 1/n scale before the depicted upsampling operation).); and
Zhang does not explicitly disclose the use of depth maps, instead relying on the inverse depth (i.e., disparity maps). However, Stack Overflow discloses (as part of the answer given by Harshit Kumar) in a similar field of art: “The depth (the actual z location of 3d point) can be calculated by using the disparity
of the corresponding point e.g. in simple cases, as follows:
PNG
media_image1.png
46
428
media_image1.png
Greyscale
where baseline is the distance b/w the cameras. By getting the depth of every pixel, you get the depth map/image.” Modifying the system and method of Zhang by adding the capability to calculate and use depth maps, as taught by Stack Overflow, would yield the expected and predictable result of better visualization and physical interpretation (over disparity) to humans viewing the maps. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang and Stack Overflow in this way.
The combination of Zhang and Stack Overflow does not explicitly disclose the following; however, Xu discloses:
a step of up-scaling the generated depth map to an n scale resolution (See in Fig. 1, a network that generates a high-resolution depth map from a low-resolution depth map. In Table 2 on page 10, upsampling factors of 4x, 8x, and 16x are used, which match the reduced resolutions used in Zhang (i.e., 1/4, 1/8, and 1/16).).
Zhang, Stack Overflow, and Xu together disclose the limitations of claim 11. Xu is directed to a similar field of art (a network for stereo matching of video sequences). Therefore, Zhang, Stack Overflow, and Xu are combinable. Modifying the system and method of Zhang and Stack Overflow by adding the capability of “up-scaling the generated depth map to an n scale resolution”, as taught by Xu, would yield the expected and predictable result of sharpening object boundaries for improved 3D reconstruction and autonomous vehicle navigation applications. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang, Stack Overflow, and Xu in this way.
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN S LEE whose telephone number is (571)272-1981. The examiner can normally be reached 11:30 AM - 7:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached at (571)270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Jonathan S Lee/Primary Examiner, Art Unit 2677