DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 6, 10, and 12-13 objected to because of the following informalities:
Claim 6 should not have a second sentence. The examiner suggests changing part of the claim to “… long short-term memory, wherein c(t) denotes an input value…”.
Claim 6’s first instance of “a specific element” should be “a specific element x of c(t)”.
Claim 6’s “previous point” should be “previous time”.
Claim 6’s “select(x)” should be “selectt(x)”.
Claim 9 should not have a second sentence. The examiner suggests changing part of the claim to “… Gumbel noise is added, wherein G denotes the Gumbel noise…”.
Claim 10 should have a period at the end of the claim.
Claim 10’s “current time” (line 20) should be “current time of the long short-term memory”.
Claim 12’s “Loss(selection)” variable should be defined.
Claim 12 should not have a second sentence. The examiner suggests changing part of the claim to “temperature parameter of Gumbel-softmax, wherein seq denotes a length…”.
Claim 13 should not have a second sentence. The examiner suggests changing part of the claim to “camera pose as in [Equation 6] below, wherein seq denotes a length…”.
Claim 13’s “Loss(pose)” variable should be defined.
Appropriate correction is required. The examiner additionally suggests checking for other additional objections, as all may not have been listed.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 6, 9, and 13 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claim 6 recites the variable “a”. This variable is neither defined in the original specification nor in the original claims, and thus calls into question whether the inventor (or joint inventor) had possession of the claimed invention.
Claims 7-9 are rejected for being dependent on claim 6.
Claim 9 recites the variable “g(t)” or “g(i)” (it is unclear whether the letter is an “i” or a “t”). This variable is neither defined in the original specification nor in the original claims, and thus calls into question whether the inventor (or joint inventor) had possession of the claimed invention.
Claim 13 recites the variables “ r’ “ and “ v’ “. These variables each have an arrow over them. These variables are neither defined in the original specification nor in the original claims, and thus call into question whether the inventor (or joint inventor) had possession of the claimed invention.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 5-10, 13, and 16-17 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 5 recites “the pattern”. There is insufficient antecedent basis for this term.
Claim 6 recites an element “K(i)” in the second equation. K(i) is not defined. As per the specification, K(i) is interpreted to be a convolution kernel.
Claims 7-9 are rejected on account of their dependence on claim 6.
Claim 9 recites “g(t)” or “g(i)” (it is unclear if the subscript is an “i ” or a “t”). It is not defined.
Claim 10 recites “the output result” (line 18). There is insufficient antecedent basis for this term.
Claim 10 recites “at the previous time”. There is insufficient antecedent basis for this term.
Claim 10 recites “the long short-term memory” (lines 17-18). There is insufficient antecedent basis for this term.
Claim 10 recites “the previous time of the tensor”. There is insufficient antecedent basis for this term.
Claim 13 recites “ r’ “. This variable has an arrow over it. This variable is not defined.
Claim 13 recites “ v’ “. This variable has an arrow over it. This variable is not defined.
Claim 16 recites “the output value” (lines 19-20 and twice in line 22). It is unclear whether this “output value” is reciting back to “an output value corresponding to a specific element” or “an output value at a previous time of a long short-term memory” (both instances in claim 15). There is insufficient antecedent basis for this term.
Claim 17 recites “the output value” three times. It is unclear whether this “output value” is reciting back to “an output value corresponding to a specific element” or “an output value at a previous time of a long short-term memory” (both instances in claim 15). There is insufficient antecedent basis for this term.
The examiner additionally suggests checking for other additional antecedent basis issues as well as other terms that are not defined, as all may not have been listed.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim does not fall within at least one of the four categories of patent eligible subject matter because “A computer program stored in a computer-readable medium…” under the broadest reasonable interpretation of the claim covers forms of non-transitory tangible media and transitory propagating signals per se in view of the ordinary and customary meaning of computer readable media (CRM). Examiner suggests amending the claim to “A non-transitory computer-readable medium storing therein a program, that when executed by a processor…” commensurate with the disclosure of the published specification and in concordance with the Kappos memo on Subject Matter Eligibility of Computer Readable Media, February 23, 2010, 1351 OG 212.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 10-14, and 19 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Yang et al., "Efficient Deep Visual and Inertial Odometry with Adaptive Visual Modality Selection", arXiv:2205.06187v2, hereinafter referred to as Yang.
Regarding claim 1, Yang teaches:
A method of visual inertial odometry performed in an apparatus for visual inertial odometry (pg. 5, Fig. 2), the method comprising:
PNG
media_image1.png
644
838
media_image1.png
Greyscale
(a) acquiring image data from a camera sensor (pg. 5, Fig. 2, “Images” at t and t+1) and acquiring inertial data from an inertial measurement unit (pg. 5, Fig. 2, ”IMU” from t to t+1);
(b) performing a convolution operation on the image data to extract image feature information (pg. 5. Fig. 2, Images are input into “Visual Encoder (2D CNN)”) and performing the convolution operation on the inertial data to extract inertial feature information (pg. 5, Fig. 2, IMU data is input into “Inertial Encoder (1D CNN)”);
(c) determining whether to use the image feature information (pg. 5, Fig. 2 description, “the policy network takes the inertial features and the previous hidden state vector to decide whether to use the visual modality or not”); and
(d) estimating a camera pose based on the used image feature information and the inertial feature information determined to be used according to a result of determining whether to use the image feature information (pg. 5, Fig. 2 description, “Once the policy network decides to use the visual modality, the current image is passed through the visual encoder, and the corresponding visual features are fed to the pose regression LSTM together with the inertial features for pose estimation”).
Regarding claim 10, the rejection of claim 1 is incorporated herein. Yang teaches the method of claim 1, and further teaches:
wherein the (d) includes: forming a tensor by connecting the used image feature information and inertial feature information along a dimension axis (pg. 5, Fig. 2, visual and inertial info are concatenated into a tensor);
acquiring an output result at a current time of the long short-term memory based on the output result at the previous time of the tensor (“h(t-1)”) and the long short-term memory (pg. 5, Fig. 2, the “output result” is the output of the middle LSTM block, where the h(t-1) is input into the middle LSTM block);
and representing the output result at the current time as a vector for each axis through a fully connected layer (pg. 6, Section 3.1, “The RNN employes fully connected layers as the last step for the final 6-DoF pose regression as in: [Equation 2], where ht−1 and ht are the hidden latent vectors of the RNN at time t and t−1. ˆϕt and ˆvt denotes the estimated rotational vector and translational vector, respectively.”).
PNG
media_image2.png
248
733
media_image2.png
Greyscale
Regarding claim 11, the rejection of claim 1 is incorporated herein. Yang teaches the method of claim 1, and further teaches:
wherein a loss function (pg. 8, Eq. 10, “L”) is defined as a sum of a loss function for determining whether to use the image feature information (pg. 8, Eq. 10, “L(eff)”) and a loss function for estimating the camera pose (pg. 8, Eq. 10, “L(pose)”).
PNG
media_image3.png
145
739
media_image3.png
Greyscale
Regarding claim 12, the rejection of claim 11 is incorporated herein. Yang teaches the method of claim 11, and further teaches:
wherein the loss function (pg. 8, “L(eff)”) for determining whether to use the image feature information is as in [Equation 5] below where a penalty term is applied as a temperature parameter (pg. 7, “where τ is a temperature parameter”) of Gumbel-softmax (pg. 8, Eq. 9).
PNG
media_image4.png
109
265
media_image4.png
Greyscale
Here, seq denotes a length of a sequence (pg. 8, “T is the sequence length of training”), λ denotes the penalty term (pg. 8, “we apply an additional penalty factor λ”), and gt denotes a result value of applying the Gumbel-softmax to an output value (pg. 7, “Sampled dt that follows a Bernoulli distribution”).
PNG
media_image5.png
94
533
media_image5.png
Greyscale
Regarding claim 13, the rejection of claim 11 is incorporated herein. Yang teaches the method of claim 11, and further teaches:
wherein the loss function for estimating the camera pose corresponds to a square loss of a Euclidean norm of a translation vector and a rotation vector representing the camera pose as in [Equation 6] below (pg. 7, Eq. 8).
PNG
media_image6.png
104
419
media_image6.png
Greyscale
Here, seq denotes a length of a sequence, alpha=100 (pg. 8, “We set α = 100”), v denotes the translation vector and r denotes the rotation vector (pg. 7, “vt and ϕt denote the ground-truth translational and rotational vectors”).
PNG
media_image7.png
97
596
media_image7.png
Greyscale
Regarding claim 14, the rejection of claim 1 is applied, mutatis mutandis, to claim 14.
Regarding claim 19, the rejection of claim 1 is applied, mutatis mutandis, to claim 19.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2-3 and 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang as applied to claim 1 above.
Regarding claim 2, the rejection of claim 1 is incorporated herein. Yang teaches the method of claim 1, and further teaches:
performing a two-dimensional convolution operation on the image data acquired from the camera sensor (pg. 5, Fig. 2, “Visual Encoder (2D CNN)”); and
extracting image feature information corresponding to a texture, an edge, or a color pattern through the two-dimensional convolution operation (description of Fig. 2, “the current image is passed through the visual encoder, and the corresponding visual features are fed to the pose regression LSTM together with the inertial features for pose estimation”).
Although Yang does not explicitly teach extracting feature information that is a “texture, edge, or color pattern”, it would have been obvious to one of ordinary skill in the art to do so, as feature extraction as taught by Yang is a common and well-known task for 2D CNNs, especially considering Yang’s inputs are images and the outputs are “visual features”.
Regarding claim 3, the rejection of claim 2 is incorporated herein. Yang teaches the method of claim 2, and further teaches:
prior to the (b), training to estimate an optical flow for extracting the image feature information (pg. 8, Section 4.1, “Implementation Details”, “For the visual encoder, we adopt the FlowNet-S network [8] (except for the last layer) pretrained on the FlyingChairs dataset [8] for optical flow estimation.”).
Regarding claim 5, the rejection of claim 1 is incorporated herein. Yang teaches the method of claim 1, and further teaches:
performing a one-dimensional convolution operation on the inertial data acquired from the inertial measurement unit (Pg 5, Fig. 2, “Inertial Encoder (1D CNN)”);
extracting the inertial feature information corresponding to the pattern or change over time through the one-dimensional convolution operation (“The inertial encoder contains three 1D convolutional layers and a fully connected layer to generate the inertial feature of size 256,” pgs. 8-9, “Implementation Details”).
Although Yang does not explicitly teach extracting inertial feature information, it would have been obvious to one of ordinary skill in the art to do so, as inertial feature extraction as taught by Yang is a common and well-known task for 1D CNNs, especially considering Yang’s inputs are IMU values from t to t+1 (Fig. 2) and the outputs are “inertial features”.
Claim(s) 4 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang as applied to claim 1 and 14 above, and further in view of Fischer et al., "FlowNet: Learning Optical Flow with Convolutional Networks", arXiv:1504.06852v2, hereinafter referred to as Fischer.
Regarding claim 4, the rejection of claim 2 is incorporated herein. Yang teaches the method of claim 2, but is not relied upon to teach the following limitations. Fischer, however, further teaches:
wherein the (b) includes:
calculating an amount of change in the image data (“Our networks predict optical flow at up to 10 image pairs per second,” Section 1 “Introduction”); and
updating a weight to minimize a calculation error for the camera pose depending on the amount of change (Section 5.1 “Network and Training Details”, “As training loss we use the endpoint error (EPE), which is the standard error measure for optical flow estimation. It is the Euclidean distance between the predicted flow vector and the ground truth, averaged over all pixels.”).
Fischer is considered to be analogous to the claimed invention because they are both in the field of calculating optical flow using a CNN. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Fischer into Yang for the benefit of more precise visual frame tracking.
Regarding claim 18, the rejection of claim 4 applies, mutatis mutandis, to claim 18.
Claim(s) 6-9 and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yang as applied to claims 1 and 14 above, and further in view of Panda et al., "AdaMML:Adaptive Multi-Modal Learning for Efficient Video Recognition", arXiv:2105.05165v2, hereinafter referred to as Panda.
Regarding claim 6, the rejection of claim 1 is incorporated herein. Yang teaches the method of claim 1, but is not relied upon to teach the following limitations. Panda, however, further teaches:
wherein the (c) includes calculating an output value (Eq. 1, “h(t)”) corresponding to a specific element through [Equation 1] below based on the inertial feature information (Eq. 1, “f(t)”) and an output result (Eq. 1, “h(t-1)”) at a previous time of a long short-term memory (Section 3.2. “Learning Adaptive Multi-Modal Policy”, “Multi-Modal Policy Network”, Eq. 1).
PNG
media_image8.png
50
431
media_image8.png
Greyscale
PNG
media_image9.png
130
302
media_image9.png
Greyscale
Here, ct denotes an input value (“f(t)”, “h(t-1)”, “o(t-1)”), ht-1 denotes the output result at the previous point of the long short-term memory (“h(t-1)”), et denotes the inertial feature information (“f(t)”), and select(x) denotes an output value corresponding to a specific element x of ct (“h(t)”).
Panda is considered to be analogous to the claimed invention because they are both in the field of selectively choosing certain types of feature information. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Panda into Yang for the benefit of better decision-making for determining when to use visual information.
Regarding claim 7, the rejection of claim 6 is incorporated herein. Yang in view of Panda teaches the method of claim 6, and Panda further teaches:
wherein the (c) includes:
applying Gumbel-softmax to the output value to calculate a probability value (3.2, Learning Adaptive Multi-Modal Policy, “Training using Gumbel-Softmax Sampling”, “we first generate the logits zk ∈ R2 (i.e, output scores of policy network for modality k) from hidden states ht by a fully-connected layer zk = FCk(ht,θFCk) for each modality and then use the Gumbel-Max trick [22] to draw discrete samples from a categorical distribution,”);
and multiplying the probability value by the output value to change the output value to 0 or 1 (using the argmax equation, Eq. 2).
PNG
media_image10.png
93
573
media_image10.png
Greyscale
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Panda into Yang for the benefit of better decision-making for determining when to use visual information.
Regarding claim 8, the rejection of claim 7 is incorporated herein. Yang in view of Panda teach the method of claim 7, and Yang further teaches:
wherein the calculating of the probability value includes:
generating Gumbel noise to apply the Gumbel-softmax to the output value (pg. 7, Eq. 7);
adding the Gumbel noise to the output value (pg. 7, Eq. 7, Gumbel noise “g(k)” is added to output value “log(p(k))”);
and applying a softmax function that adjusts continuity of an output distribution through a temperature parameter to the output value to which the Gumbel noise is added (pg. 7, Eq. 7, the temperature parameter τ divides the Gumbel noise “g(k)” and the output value “log(p(k))”).
PNG
media_image11.png
94
613
media_image11.png
Greyscale
Regarding claim 9, the rejection of claim 8 is incorporated herein. Yang teaches the method of claim 8, and further teaches:
wherein the calculating of the probability value includes: generating the Gumbel noise as in [Equation 2] below (pg. 7, “where gk = −log(−logUk) is a standard Gumbel distribution with a random variable Uk sampled from a uniform distribution U(0,1)”);
PNG
media_image12.png
74
154
media_image12.png
Greyscale
adding the Gumbel noise to the output value as in [Equation 3] below (pg. 7, Eq. 6, g(k) is added to log(p(k))); and
PNG
media_image13.png
68
606
media_image13.png
Greyscale
PNG
media_image14.png
84
203
media_image14.png
Greyscale
applying the softmax function as in [Equation 4] below based on the output value to which the Gumbel noise is added (pg. 7, Eq. 7).
PNG
media_image15.png
90
623
media_image15.png
Greyscale
PNG
media_image16.png
112
256
media_image16.png
Greyscale
Here, G denotes the Gumbel noise, U denotes a random number sampled from a uniform distribution between 0 and 1, logits denotes the output value, (logitsgumbel)i denotes the output value to which the Gumbel noise is added for an i-th class, K denotes the number of classes to be classified, and τ denotes the temperature parameter.
Regarding claims 15-17, the rejection of claims 6-8 applies, mutatis mutandis, to claims 15-17.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Han et al., "DeepVIO: Self-supervised Deep Learning of Monocular Visual Inertial Odometry using 3D Geometric Constraints", arXiv:1906.11435v2, teaches a self-supervised network for visual inertial odometry.
Liu et al., "ATVIO: Attention Guided Visual-Inertial Odometry", ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), teaches a method for a deep framework for visual inertial odometry.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RACHEL A OMETZ whose telephone number is (571)272-2535. The examiner can normally be reached 8:30am-5:30pm ET Monday-Thursday, 7:30am-3:30pm ET every other Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at 571-272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Rachel Anne Ometz/Examiner, Art Unit 2668 8/12/26
Rachel.ometz@uspto.gov
/VU LE/Supervisory Patent Examiner, Art Unit 2668