DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. JP2024-054469, filed on 03/28/2024.
Status of Claims
Claims 1-11 are currently pending in this application.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/31/2025 has been considered by the examiner.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
“a feature generation unit” in claims 1-3, 8 and 10-11
“a map generation unit” in claims 1, 4, 8 and 10-11
“a correction unit” in claims 4 and 5
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
Claims 1-3, 8 and 10-11: “a feature generation unit” corresponds to elements 301R and 301L in Figure 3 “The model 300 includes two feature generation units 301R and 301L and a map generation unit 302.” (Paragraph [0028]).
Claims 1, 4, 8 and 10-11: “a map generation unit” corresponds to elements 302 in Figure 3 “The model 300 includes two feature generation units 301R and 301L and a map generation unit 302.” (Paragraph [0028]).
Claims 4 and 5: “a correction unit” corresponds to element 701 in Figure 7 “The model 700 is different from the model 300 in further including a correction unit 701 provided downstream of the map generation unit 302.” (Paragraph [0046]).
Claims 6-7 and 9 are similarly interpreted for their dependency from claims 1 and 8.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-3, 8 and 10-11 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li et al. (Li, Zhaoshuo, et al. "Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers." 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2021.) (hereinafter, “Li”).
Regarding claim 1, Li discloses a learning apparatus for performing machine learning, the learning apparatus configured to (Page 6178 left column first paragraph “We take advantage of the recent Transformer architecture [36] proposed for language processing and recent advances in feature matching [31], and present a new end-to-end-trained stereo depth estimation network named STereo TRansformer (STTR).”; See Figure 1, emphasized below):
Figure 1
PNG
media_image1.png
431
892
media_image1.png
Greyscale
acquire teaching data including input data and ground truth data (ground truth disparity dgt,I on Page 6182 equate to ground truth data), the input data including a first image and a second image (left and right pair of images on Page 6179 equate to the first and second image) (Page 6179 right column Section 3 “In the following sections, we denote the height and width of the rectified left and right pair of images as Ih and Iw. We denote the channel dimension of feature descriptors as C.”; Page 6181 right column Subsection 3.4 continuing to page 6182 left column first paragraph “We adopt the Relative Response loss Lrr proposed in [23] on the assignment matrix T for both sets of matched pixels M and sets of unmatched pixels U due to occlusion. The goal of the network is to maximize the attention on the true target location. Since disparity is subpixel, we use linear interpolation between the nearest integer pixels to find the matching probability t*. Specifically, for the i-th pixel in the left image with ground truth disparity dgt,i,
PNG
media_image2.png
140
706
media_image2.png
Greyscale
);
generate output data representing a disparity (disparity estimate on Page 6178 equate to output data) between the first image and the second image by inputting the input data (stereo images on Page 6178 equate to the input data) to a model (Page 6178 Figure 1 Caption “STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”); and
update a parameter of the model to reduce a loss (Page 6181 right column Subsection 3.4 continuing to page 6182 left column first paragraph “We adopt the Relative Response loss Lrr proposed in [23] on the assignment matrix T for both sets of matched pixels M and sets of unmatched pixels U due to occlusion.) obtained by inputting the output data (Ld1,r and Ld1,f on Page 6182 equate to the output data) and the ground truth data (ground truth disparity dgt,I on Page 6182 equate to ground truth data) to a loss function (total loss on Page 6182 equation 13 equates to loss function) (Page 6181 right column Subsection 3.4 continuing to page 6182 left column first paragraph “We adopt the Relative Response loss Lrr proposed in [23] on the assignment matrix T for both sets of matched pixels M and sets of unmatched pixels U due to occlusion. The goal of the network is to maximize the attention on the true target location. Since disparity is subpixel, we use linear interpolation between the nearest integer pixels to find the matching probability t*. Specifically, for the i-th pixel in the left image with ground truth disparity dgt,i,
PNG
media_image2.png
140
706
media_image2.png
Greyscale
…We use smooth L1 loss [13] on both raw and final disparities, denoted as Ld1,r and Ld1,f. The final occlusion map is supervised via a binary-entropy loss Lbe,f. The total loss is the summation:
PNG
media_image3.png
78
644
media_image3.png
Greyscale
where w are the loss weights.”), wherein the model (See Figure 1, emphasized below) includes:
PNG
media_image4.png
454
856
media_image4.png
Greyscale
a feature generation unit (feature extractor in Figure 1 equates to feature generation unit) configured to generate a first feature based on the first image and generate a second feature based on the second image (Page 6178 Figure 1 Caption “(a) STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate.”; See Figure 1, emphasized below);
PNG
media_image5.png
431
892
media_image5.png
Greyscale
and a map generation unit configured to generate a disparity map (See Figure 1 emphasized below) of the disparity between the first image and the second image based on the first feature and the second feature (Page 6178 Figure 1 caption “A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality. (b-f) Inference of STTR trained only on synthetic Scene Flow dataset. Top row shows the left images. Bottom row shows predicted disparities. The color map used to visualize disparity is relative to the image width and is shown on the right.”; See Figure 1, emphasized below),
PNG
media_image6.png
555
1573
media_image6.png
Greyscale
the map generation unit includes a cross-attention layer configured to receive an input based on the first feature and an input based on the second feature (Page 6178 Figure 1 Caption “The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”; See Figure 1, emphasized below
PNG
media_image7.png
445
892
media_image7.png
Greyscale
), and
the disparity map is based on an output from the cross-attention layer (See Figure 1, emphasized below).
PNG
media_image8.png
445
892
media_image8.png
Greyscale
Regarding claim 2, which claim 1 is incorporated, Li discloses wherein the feature generation unit includes a self-attention layer (Page 6179 right column subsection Section 3.2 “Self-attention computes attention between pixels along the epipolar line in the same image”),
an input to the self-attention layer is based on the first image (Page 6180 left column “For each attention head h, a set of linear projections are used to compute the query vectors Qh, key vectors Kh and value vectors Vh using feature descriptors eI as input…For self-attention, the Qh, Kh, Vh are computed from the same image.”), and
the first feature is based on an output of the self-attention layer (Page 6180 left column first paragraph “The output value vector VO can be computed as:
PNG
media_image9.png
68
630
media_image9.png
Greyscale
where WO ∈ RCe×Ce and bO ∈ RCe. The output value vector VO is then added to the original feature descriptors to form a residual connection”).
Regarding claim 3, which claim 2 is incorporated, Li discloses wherein the feature generation unit includes a path that bypasses the self-attention layer (Page 6180 left column first paragraph “The output value vector VO is then added to the original feature descriptors to form a residual connection” See Figure 2, emphasized below).
PNG
media_image10.png
446
734
media_image10.png
Greyscale
Regarding claim 8, Li discloses an estimation apparatus for performing disparity estimation, the estimation apparatus configured to (Page 6178 left column first paragraph “We take advantage of the recent Transformer architecture [36] proposed for language processing and recent advances in feature matching [31], and present a new end-to-end-trained stereo depth estimation network named STereo TRansformer (STTR).”):
acquire input data including a first image and a second image (See Figure 1, emphasized below)
Figure 1
PNG
media_image1.png
431
892
media_image1.png
Greyscale
; and
estimate a disparity between the first image and the second image by inputting the input data (stereo images on Page 6178 equate to the input data) to a model, wherein the model includes (Page 6178 Figure 1 Caption “STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”; See Figure 1, emphasized below):
PNG
media_image4.png
454
856
media_image4.png
Greyscale
a feature generation unit (feature extractor in Figure 1 equates to feature generation unit) configured to generate a first feature based on the first image and generate a second feature based on the second image (Page 6178 Figure 1 Caption “(a) STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate.”; See Figure 1, emphasized below);
PNG
media_image5.png
431
892
media_image5.png
Greyscale
and a map generation unit configured to generate a disparity map (predicted disparities in Figure 1 equates to disparity map) of the disparity between the first image and the second image based on the first feature and the second feature (Page 6178 Figure 1 caption “A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality. (b-f) Inference of STTR trained only on synthetic Scene Flow dataset. Top row shows the left images. Bottom row shows predicted disparities. The color map used to visualize disparity is relative to the image width and is shown on the right.”; See Figure 1, emphasized below),
PNG
media_image11.png
555
1573
media_image11.png
Greyscale
the map generation unit includes a cross-attention layer configured to receive an input based on the first feature and an input based on the second feature (Page 6178 Figure 1 Caption “The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”; See Figure 1, emphasized below
PNG
media_image7.png
445
892
media_image7.png
Greyscale
), and
the disparity map is based on an output from the cross-attention layer (See Figure 1, emphasized below).
PNG
media_image8.png
445
892
media_image8.png
Greyscale
Regarding claim 10, Li discloses a method for performing machine learning (Page 6178 left column first paragraph “We take advantage of the recent Transformer architecture [36] proposed for language processing and recent advances in feature matching [31], and present a new end-to-end-trained stereo depth estimation network named STereo TRansformer (STTR).”; See Figure 1, emphasized below):
Figure 1
PNG
media_image1.png
431
892
media_image1.png
Greyscale
, the method comprising:
acquiring teaching data including input data and ground truth data (ground truth disparity dgt,I on Page 6182 equate to ground truth data), the input data including a first image and a second image (left and right pair of images on Page 6179 equate to the first and second image) (Page 6179 right column Section 3 “In the following sections, we denote the height and width of the rectified left and right pair of images as Ih and Iw. We denote the channel dimension of feature descriptors as C.”; Page 6181 right column Subsection 3.4 continuing to page 6182 left column first paragraph “We adopt the Relative Response loss Lrr proposed in [23] on the assignment matrix T for both sets of matched pixels M and sets of unmatched pixels U due to occlusion. The goal of the network is to maximize the attention on the true target location. Since disparity is subpixel, we use linear interpolation between the nearest integer pixels to find the matching probability t*. Specifically, for the i-th pixel in the left image with ground truth disparity dgt,i,
PNG
media_image2.png
140
706
media_image2.png
Greyscale
);
generating output data representing a disparity (disparity estimate on Page 6178 equate to output data) between the first image and the second image by inputting the input data (stereo images on Page 6178 equate to the input data) to a model (Page 6178 Figure 1 Caption “STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”); and
updating a parameter of the model to reduce a loss (Page 6181 right column Subsection 3.4 continuing to page 6182 left column first paragraph “We adopt the Relative Response loss Lrr proposed in [23] on the assignment matrix T for both sets of matched pixels M and sets of unmatched pixels U due to occlusion.) obtained by inputting the output data (Ld1,r and Ld1,f on Page 6182 equate to the output data) and the ground truth data (ground truth disparity dgt,I on Page 6182 equate to ground truth data) to a loss function (total loss on Page 6182 equation 13 equates to loss function) Page 6181 right column Subsection 3.4 continuing to page 6182 left column first paragraph “We adopt the Relative Response loss Lrr proposed in [23] on the assignment matrix T for both sets of matched pixels M and sets of unmatched pixels U due to occlusion. The goal of the network is to maximize the attention on the true target location. Since disparity is subpixel, we use linear interpolation between the nearest integer pixels to find the matching probability t*. Specifically, for the i-th pixel in the left image with ground truth disparity dgt,i,
PNG
media_image2.png
140
706
media_image2.png
Greyscale
…We use smooth L1 loss [13] on both raw and final disparities, denoted as Ld1,r and Ld1,f. The final occlusion map is supervised via a binary-entropy loss Lbe,f. The total loss is the summation:
PNG
media_image3.png
78
644
media_image3.png
Greyscale
where w are the loss weights.”), wherein the model includes (See Figure 1, emphasized below):
PNG
media_image4.png
454
856
media_image4.png
Greyscale
a feature generation unit (feature extractor in Figure 1 equates to feature generation unit) configured to generate a first feature based on the first image and generate a second feature based on the second image (Page 6178 Figure 1 Caption “(a) STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate.”; See Figure 1, emphasized below);
PNG
media_image5.png
431
892
media_image5.png
Greyscale
and a map generation unit configured to generate a disparity map (See Figure 1) of the disparity between the first image and the second image based on the first feature and the second feature (Page 6178 Figure 1 caption “A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality. (b-f) Inference of STTR trained only on synthetic Scene Flow dataset. Top row shows the left images. Bottom row shows predicted disparities. The color map used to visualize disparity is relative to the image width and is shown on the right.”; See Figure 1, emphasized below),
PNG
media_image12.png
555
1573
media_image12.png
Greyscale
the map generation unit includes a cross-attention layer configured to receive an input based on the first feature and an input based on the second feature (Page 6178 Figure 1 Caption “The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”; See Figure 1, emphasized below
PNG
media_image7.png
445
892
media_image7.png
Greyscale
), and
the disparity map is based on an output from the cross-attention layer (See Figure 1, emphasized below).
PNG
media_image8.png
445
892
media_image8.png
Greyscale
Regarding claim 11, Li discloses a method for disparity estimation, the method comprising: acquiring input data including a first image and a second image (See Figure 1, emphasized below)
PNG
media_image1.png
431
892
media_image1.png
Greyscale
; and
estimating a disparity between the first image and the second image by inputting the input data (stereo images on Page 6178 equate to the input data) to a model (Page 6178 Figure 1 Caption “STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”), wherein the model (See Figure 1, emphasized below) includes:
PNG
media_image4.png
454
856
media_image4.png
Greyscale
a feature generation unit (feature extractor in Figure 1 equates to feature generation unit) configured to generate a first feature based on the first image and generate a second feature based on the second image (Page 6178 Figure 1 Caption “(a) STTR estimates disparity by first extracting features from stereo images using a shared feature extractor. The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate.”; See Figure 1, emphasized below);
PNG
media_image5.png
431
892
media_image5.png
Greyscale
and a map generation unit configured to generate a disparity map of the disparity between the first image and the second image based on the first feature and the second feature (Page 6178 Figure 1 caption “A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality. (b-f) Inference of STTR trained only on synthetic Scene Flow dataset. Top row shows the left images. Bottom row shows predicted disparities. The color map used to visualize disparity is relative to the image width and is shown on the right.”; See Figure 1, emphasized below),
PNG
media_image13.png
555
1573
media_image13.png
Greyscale
the map generation unit includes a cross-attention layer configured to receive an input based on the first feature and an input based on the second feature (Page 6178 Figure 1 Caption “The extracted feature descriptors are then used by a Transformer for dense self- and cross-attention computation, yielding a raw disparity estimate. A context adjustment layer further refines the disparity with information across epipolar lines conditioned on the left image for cross epipolar line optimality.”; See Figure 1, emphasized below)
PNG
media_image7.png
445
892
media_image7.png
Greyscale
, and
the disparity map is based on an output from the cross-attention layer (See Figure 1, emphasized below).
PNG
media_image8.png
445
892
media_image8.png
Greyscale
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 4, 6-7, and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (Li, Zhaoshuo, et al. "Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers." 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2021.) (hereinafter, “Li”) as applied to claims 1 and 8 above; in view of Lee et al. (US 2019/0080464 A1) (hereinafter, “Lee”).
Regarding claim 4, which claim 1 is incorporated, Li fails to teach wherein the input data includes time-series data of image pairs of the first image and the second image, the model further includes a correction unit configured to correct the disparity map generated by the map generation unit, and the correction unit corrects, based on the disparity map generated by the map generation unit for the image pair at a first time point, the disparity map generated by the map generation unit for the image pair at a second time point after the first time point.
Lee teaches wherein the input data includes time-series data of image pairs of the first image and the second image (current frame and a previous frame in Paragraph [0054] equate to the first and second image) (Paragraph [0054] “a stereo matching apparatus 100 acquires distance information from a current frame using a previous frame.”; Paragraph [0125] “the I/O interface 1005 respectively receives a left image and a right image for consecutive frames generated by the stereo camera.”,
the model further includes a correction unit (stereo matching apparatus 100 in Paragraph [0056] equates to a correction unit) configured to correct the disparity map generated by the map generation unit (Paragraph [0056] “The stereo matching apparatus 100 compares the left image 111 and the right image 113 captured from the same object at the respective different viewpoints of the camera 101 and the camera 103. The stereo matching apparatus 100 matches the corresponding two images and acquires a difference between the two images based on a matching result. This difference may be represented as a change or variance in position of the same objects between the left image 111 and the right image 113”), and
the correction unit corrects, based on the disparity map generated by the map generation unit for the image pair at a first time point (information of the previous frame in Paragraph [0063] equates to image pair at a first time point), the disparity map generated by the map generation unit for the image pair at a second time point (disparity map of the current frame in Paragraph [0063] equates to disparity map…for the image pair at a second time point) after the first time point (Paragraph [0063] “When calculating a disparity, the stereo matching apparatus 100 may perform spatial-temporal stereo matching based on different pixels of different times or frames. If a change of a target and a camera motion, e.g., of the stereo camera, are determined insignificant between a previous frame and a current frame, the stereo matching apparatus 100 may further generate a disparity map of the current frame based on information of the previous frame, which may result in a final disparity map that is more accurate than if a disparity map is generated solely on information from the current frame.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include wherein the input data includes time-series data of image pairs of the first image and the second image, the model further includes a correction unit configured to correct the disparity map generated by the map generation unit, and the correction unit corrects, based on the disparity map generated by the map generation unit for the image pair at a first time point, the disparity map generated by the map generation unit for the image pair at a second time point after the first time point taught by Lee’s reference. The motivation for doing so would have been to reduce inaccurate information from being applied to the distance information of the current frame as suggested by Lee (see Lee, Paragraph [0054]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Lee with Li to obtain the invention specified in claim 4.
Regarding claim 6, which claim 1 is incorporated, Li fails to teach wherein the first image and the second image are two images captured by a stereo camera of a mobile body.
Lee teaches wherein the first image and the second image are two images captured by a stereo camera of a mobile body (Paragraph [0055] “The stereo matching apparatus 100 may be representative of being, or included in, corresponding computing apparatuses or systems in a variety of fields, for example, a display device, a robot, a vehicle, and included in the example provision of various technological functionalities in three-dimensional (3D) object restoration”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include wherein the first image and the second image are two images captured by a stereo camera of a mobile body taught by Lee’s reference. The motivation for doing so would have been to acquire distance information of an external object and restore the external appearance of the object using stereo matching as suggested by Lee (see Lee, Paragraph [0055]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Lee with Li to obtain the invention specified in claim 6.
Regarding claim 7, which claim 1 is incorporated, Li fails to teach a non-transitory computer-readable storage medium storing a program for causing a computer to function as the learning apparatus according to claim 1.
Lee teaches a non-transitory computer-readable storage medium storing a program for causing a computer to function as the learning apparatus according to claim 1 (Paragraph [0132] “The instructions or software to control computing hardware, for example, one or more processors or computers, to implement the hardware components and perform the methods as described above, and any associated data, data files, and data structures, may be recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include a non-transitory computer-readable storage medium storing a program for causing a computer to function as the learning apparatus according to claim 1 taught by Lee’s reference. The motivation for doing so would have been to provide instructions to acquire distance information of an external object and restore the external appearance of the object using stereo matching as suggested by Lee (see Lee, Paragraph [0055]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Lee with Li to obtain the invention specified in claim 7.
Regarding claim 9, which claim 8 is incorporated, Li fails to teach a non-transitory computer-readable storage medium storing a program for causing a computer to function as the estimation apparatus according to claim 8.
Lee teaches a non-transitory computer-readable storage medium storing a program for causing a computer to function as the estimation apparatus according to claim 8 (Paragraph [0132] “The instructions or software to control computing hardware, for example, one or more processors or computers, to implement the hardware components and perform the methods as described above, and any associated data, data files, and data structures, may be recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li’s reference to include non-transitory computer-readable storage medium storing a program for causing a computer to function as the estimation apparatus according to claim 8 taught by Lee’s reference. The motivation for doing so would have been to provide instructions to acquire distance information of an external object and restore the external appearance of the object using stereo matching as suggested by Lee (see Lee, Paragraph [0055]).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Lee with Li to obtain the invention specified in claim 9.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (Li, Zhaoshuo, et al. "Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers." 2021 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2021.) (hereinafter, “Li”) in view of Lee et al. (US 2019/0080464 A1) (hereinafter, (“Lee”) as applied to claim 4 above; and further in view of Lipson et al. (Lipson, Lahav, Zachary Teed, and Jia Deng. "Raft-stereo: Multilevel recurrent field transforms for stereo matching." 2021 International conference on 3D vision (3DV). IEEE, 2021.) (hereinafter, “Lipson”).
Regarding claim 5, which claim 4 is incorporated, Li and Lee both fail to teach wherein the correction unit is configured by a convolutional gated recurrent unit (ConvGRU).
Lipson teaches wherein the correction unit is configured by a convolutional gated recurrent unit (ConvGRU) (Page 218 left column Introduction “In the standard setup, two frames—a left frame and a right frame—are provided as input. The task is to estimate a pixelwise displacement map between the input images.”; Page 221 right column Figure 3 Caption “Multilevel GRU. We use a 3-level convolutional GRU which acts on feature maps at 1/32, 1/16, and 1/8 the input image resolution. Information is passed between GRUs at adjacent resolutions using upsampling and downsampling operations. The GRU at the highest resolution (red) performs lookups from the correlation pyramid and updates the disparity estimate.”).
Therefore, it would have been obvious to one of ordinary skill of the art before the effective filing date to modify Li in view of Lee to include wherein the correction unit is configured by a convolutional gated recurrent unit (ConvGRU) taught by Lipson’s reference. The motivation for doing so would have been to improve the ability of the update operator to propagate information across the image as suggested by Lipson (see Lipson, Page 218 right column paragraph 4).
Further, one skilled in the art could have combined the elements described above by known methods with no change to the respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Lipson with Li and Lee to obtain the invention specified in claim 5.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Jeong et al. (US 2024/0404093 A1) discloses a system for generating disparity information by combining a first and second disparity information to generate a refined disparity map.
Csordás et al. (US 10,380,753 B1) discloses a method for generating a displacement map of a first input dataset and a second input dataset by using a feature extractor for processing the first input dataset and the second input dataset so to generate a feature map hierarchy.
Ye et al. (US 2023/0122373 A1) discloses a method for training a model by obtaining sample images and generating two kinds of training signals from them: sample depth images and sample residual maps.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to UROOJ FATIMA whose telephone number is (571)272-2096. The examiner can normally be reached M-F 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/UROOJ FATIMA/Examiner, Art Unit 2676
/Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676