Prosecution Insights
Last updated: October 01, 2026
Application No. 18/583,185

LOW-FOOTPRINT MODEL APPLICABLE TO OPTICAL FLOW ESTIMATION AND STEREO MATCHING

Non-Final OA §101§102§103§112
Filed
Feb 21, 2024
Examiner
FITCH, GRANT FREDERICK
Art Unit
Tech Center
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
5 currently pending
Career history
5
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is in response to submission of application on 07/14/2025. Claims 1-20 are presented for examination. Information Disclosure Statement The information disclosure statements (IDS) submitted on 02/27/2024 & 05/15/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Drawings The Drawings filed on 02/21/2024 & 07/14/2025 are acceptable for examination purposes. Specification The Specification filed on 02/21/2024 & 07/14/2025 are acceptable for examination purposes. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 1, 17, & 20 recite a machine learning model that “incorporates a softmax with norm folding mechanism.” The term “norm folding mechanism” does not appear to have an established meaning in the art in the context recited. Although the specification describes “norm folding” [¶0072-0082], it describes the inverse accumulation sum associated with softmax normalization as being applied at a later stage of the attention computation rather than immediately to the softmax numerator values. For example, the inverse accumulation sum may be applied to third data V prior to a second matrix multiplication or to the output of the second matrix multiplication. Thus, the disclosed “norm folding” does not appear to identify a particular type of softmax mechanism itself, but instead describes relocating the normalization operation associated with the softmax computation into subsequent operations of an attention mechanism. Accordingly, it is unclear what structure or operations are required for a machine learning model to incorporate a softmax with norm folding mechanism, particularly where the claims do not recite the attention architecture or subsequent operations into which the normalization operation is folded. Further the term ‘norm folding’ does not have an established meaning in the art in as used within the claims, however there is a recognized term in the art known as batch-normalization folding, or batch norm folding, which is substantially different from the instant applications use of the term. Claims 2-16 depend from indefinite claim 1 and are therefore also rejected as being indefinite. Claims 18 & 19 depend from indefinite claim 17 and are therefore also rejected as being indefinite. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims will follow the 2019 Revised Patent Subject Matter Eligibility Guidelines (“2019 PEG”). Step 1: Is the claim directed at one of the four statutory categories? Claims 1-16 are directed to A device comprising: a memory … and one or more processors, which falls within the statutory category of a machine. Claims 17-19 are directed to A method, which falls within the statutory category of a process. Claims 20 are directed to A non-transitory computer-readable medium, which falls within the statutory category of an article of manufacture. Therefore, claims 1-20 are directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter. (MPEP § 2106.03) Independent Claims Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, independent claims 1, 17, & 20 recites an abstract idea in the form of a mathematical concept. Mathematical concepts are defined as mathematical relationships, mathematical formulas or equations, or mathematical calculations (MPEP § 2106.04(a)(2)(I)). The following limitations of claims 1, 17, & 20 are mathematical concepts: […] a softmax with norm folding mechanism [The softmax function is a known mathematical formula, and integrating a ‘norm folding mechanism’ as defined in the specification [¶0074-0075] where the normalization step happens at a later stage is merely further defining this formula. Therefore this limitation is considered an abstract idea in the form of a mathematical concept (MPEP § 2106.04(a)(2)(I))] Therefore, the independent claims recite a judicial exception Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No. The judicial exception recited in the above discussed claims is not integrated into a practical application. store input data [Obtaining and storing data represents insignificant extra-solution activity of data gathering. (MPEP § 2106.05(g))] process the input data using a machine learning (ML) model that incorporates […] [This is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity. (MPEP § 2106.05(f))] Therefore, under MPEP § 2106.04(d), the additional elements of the claims do not integrate the judicial exception into a practical application. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No. The claims do not recite additional elements that are sufficient for the claims to amount to significantly more than the judicial exception. Additional elements that are considered an insignificant extra solution activity are re-evaluated under MPEP § 2106.05(d) and have been determined to merely specify the activity of data gathering and storing, this is considered WURC (Well-Understood, Routine, Conventional activity) and has been recognized by the courts as such (MPEP §2106.05(d)(II)). Further, additional elements that do no more than mere instructions to apply an exception using a generic class of computer algorithms do not constitute significantly more than a judicial exception under MPEP § 2106.05(f). Therefore, the additional elements identified in the Step 2A Prong Two analysis do not constitute significantly more than a judicial exception. Dependent Claims The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claim 2: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the input data includes a first image and a second image, and wherein the ML model corresponds to a depth from stereo or optical flow architecture. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely limits the judicial exception to a particular computing environment, namely that of image processing, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 3: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the ML model includes a language model, a vision model, or a multimodal model [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely limits the judicial exception to a particular computing environment, namely within the fields of language, vision, or multi-modal models, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claims 4 & 18: Step 2A, Prong 1: wherein the softmax with norm folding mechanism is included in a streamable attention mechanism [These further limitations merely further define the mathematical concept recited in the parent claim and are therefore considered to be part of the mathematical concept of the parent claim. This claim does not recite any non-abstract additional elements for the purposes of Step 2A Prong Two and Step 2B analysis] Step 2A, Prong 2 and Step 2B: There are no additional elements recited, as such the claim does not provide a practical application and is not considered to be significantly more. Claim 5 & 19: Step 2A, Prong 1: Generate a softmax input stream based on a first matrix multiplication operation of a particular row of first data and a corresponding row of second data; [Generating the softmax input as defined, using a matrix multiplication operation on particular rows of data, is a mathematical concept as defined by mathematical relationships, formulas, equations, or calculations (MPEP § 2106.04(a)(2)(I))] Apply an exponentiation operation of the softmax with norm folding mechanism to the softmax input stream to generate a stream of softmax numerator values; [Generating a stream of numerator values as defined by an exponentiation operation is a mathematical concept as defined by mathematical relationships, formulas, equations, or calculations (MPEP § 2106.04(a)(2)(I))] input the stream of softmax numerator values as a first input to a second matrix multiplication operation; and generate an accumulation sum of the softmax numerator values. [Generating the accumulation sum as defined by second matrix multiplication applied to previous mathematical results is a mathematical concept as defined by mathematical relationships, formulas, equations, or calculations (MPEP § 2106.04(a)(2)(I))] Step 2A, Prong 2 and Step 2B: There are no additional elements recited, as such the claim does not provide a practical application and is not considered to be significantly more. Claim 6: Step 2A, Prong 1: wherein the streamable attention mechanism is configured to perform a norm operation to apply the accumulation sum to third data to generate a second input to the second matrix multiplication operation. [Performing a normalization operation and a matrix multiplication is a mathematical concept as defined by mathematical relationships, formulas, equations, or calculations (MPEP § 2106.04(a)(2)(I))] Step 2A, Prong 2 and Step 2B: There are no additional elements recited, as such the claim does not provide a practical application and is not considered to be significantly more. Claim 7: Step 2A, Prong 1: wherein the streamable attention mechanism is configured to perform a norm operation to apply the accumulation sum to an output of the second matrix multiplication operation. [Performing a normalization operation is a mathematical concept as defined by mathematical relationships, formulas, equations, or calculations (MPEP § 2106.04(a)(2)(I))] Step 2A, Prong 2 and Step 2B: There are no additional elements recited, as such the claim does not provide a practical application and is not considered to be significantly more. Claim 8: Step 2A, Prong 1: a second input to the second matrix multiplication operation corresponds to a displacement matrix of coordinates. [These further limitations merely further define the mathematical concept recited in the parent claim and are therefore considered to be part of the mathematical concept of the parent claim. This claim does not recite any non-abstract additional elements for the purposes of Step 2A Prong Two and Step 2B analysis] Step 2A, Prong 2 and Step 2B: the first data corresponds to features associated with a first image; the second data corresponds to features associated with a second image; [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely limits the judicial exception to a particular computing environment, namely within the field of image processing, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 9: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the ML model generates probabilistic geometry measures without generating a cost volume data structure. [This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity (MPEP § 2106.05(f)) and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 10: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the ML model is configured to perform regression for depth from stereo or optical flow geometric coordinates using just-in-time computations. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely limits the judicial exception to a particular computing environment, namely image processing using depth from stereo or optical flow, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 11: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: Further comprising an image sensor configured to generate image data corresponding to the input data. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely limits the judicial exception to a particular computing environment, namely the field of image processing, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 12: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: Further comprising a modem coupled to the one or more processors, the modem configured to receive image data corresponding to the input data from a second device. [This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity. (MPEP § 2106.05(f)) This limitation merely defines the input data as being received through a modem, a generic class of computer equipment performing its ordinary purpose, and as such does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 13: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2: wherein the one or more processors are integrated in a headset device that includes a display, [This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity. (MPEP § 2106.05(f))] and wherein the headset device is configured, when worn by a user, to display an output image based on an output of the ML model. [This additional element represents the intended use as well as merely outputting data to a display and is considered insignificant extra solution activity (MPEP § 2106.05(g)] Step 2B: Additional elements that are considered an insignificant extra solution activity are re-evaluated under MPEP § 2106.05(d) and have been determined to merely output data, this is considered WURC (Well-Understood, Routine, Conventional activity) and has been recognized by the courts as such (MPEP §2106.05(d)(II)). Further, additional elements that do no more than mere instructions to apply an exception using a generic class of computer algorithms do not constitute significantly more than a judicial exception under MPEP § 2106.05(f). Therefore, the additional elements identified in the Step 2A Prong Two analysis do not constitute significantly more than a judicial exception. Claim 14: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the one or more processors are integrated in at least one of a mobile phone, a tablet computer device, a wearable electronic device, or a camera device. [This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity. (MPEP § 2106.05(f))] Claim 15: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the one or more processors are integrated in a vehicle, the vehicle further including one or more cameras configured to capture image data corresponding to the input data. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely limits the judicial exception to a particular computing environment, namely image processing, and therefore does not integrate the judicial exception into a practical application or amount to significantly more than the abstract idea itself.] Claim 16: Step 2A, Prong 1: There are no additional abstract idea limitations. Step 2A, Prong 2 and Step 2B: wherein the one or more processors are included in an integrated circuit.[This additional element is mere recitation that a judicial exception is to be performed using generic computer equipment running general class of computer algorithms in their ordinary capacity. (MPEP § 2106.05(f))] The following references are relied upon for the art rejection set forth below: Dao, Tri Flashattention-2: Faster Attention with Better Parallelism and Work Partitioning, July 2023, hereinafter Dao Xu et al. GMFlow: Learning Optical Flow via Global Matching, July 2022, hereinafter Xu Kim et al. US Pub. No. US 2018/0130221 A1, May 2018, Hereinafter Kim Lee, Kwangyong US Pub. No. US 2021/0142127 A1, May 2021, hereinafter Lee Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 3-5, 7, 17-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Dao. Regarding Claim 1, Dao discloses A device comprising: a memory configured to store input data; and one or more processors configured to [Dao: 2.1] “the A100 GPU has 40-80GB of high bandwidth memory (HBM)…and 192KB of on-chip SRAM per each of 108 streaming multiprocessors … Each kernel loads inputs from HBM to registers and SRAM, computes, then writes outputs to HBM”. Dao discloses a GPU having memory configured to store input data and executing the FlashAttention-2 algorithm, which constitutes the disclosed device under Broadest Reasonable interpretation (BRI). process the input data using a machine learning (ML) model [Dao: Abstract] “Scaling Transformers to longer sequence lengths has been a major problem … in language modeling and high-resolution image understanding” [Dao: §4] “We evaluate the impact of using FlashAttention-2 to train Transformer models.” Dao applies the FlashAttention-2 attention algorithm in Transofmer models, thereby disclosing processing input data using a machine learning model, which constitutes processing the input data using a machine learning model under BRI. that incorporates a softmax with norm folding mechanism. [Dao: §2.3.1] “online softmax [11,13] can split the attention computation into blocks, and rescale the output of each block to finally get the right result” [Dao: §3.1.1] “We do not have to rescale both terms of the output update by diag(𝓁(2))-1: … We can instead maintain an “un-scaled” version of O(2) and keep around the statistics 𝓁(2): … Only at the every end of the loop do we scale the final Õ(last) by diag(𝓁(last))-1 to get the right output.” [Dao Algorithm 1 lines 9-12] Dao computes and accumulates n unscaled attention output during the loop and subsequently applies the accumulated normalization term to the output after completion of the loop. Dao delays application of the normalization associated with the softmax computation until a later stage of the attention computation, corresponding to the claimed softmax with norm folding mechanism claimed under BRI. Regarding Claim 3, Dao discloses the device according to claim 1, wherein the ML model includes a language model, a vision model, or a multi-modal model. [Dao: Abstract] “… when used end-to-end to train GPT-style models, FlashAttention-2 reaches training speed…” [Dao: §1] “there have been several language models with much longer context than before: GPT-4 [12] with context length 32k“ Dao discloses using a GPT (Generative Pre-trained Transformer) style model provides an example of the GPT-4 being a type of language model, corresponding to the claimed language model under BRI. Regarding Claim 4, Dao discloses the device according to claim 1, wherein the softmax with norm folding mechanism is included in a streamable attention mechanism. [Dao: §2.3.1] “FlashAttention applies the classical technique of tiling to reduce memory IOs, by (1) loading blocks of inputs from HBM to SRAM, (2) computing attention with respect to that block, and then (3) updating the output without writing the large intermediate matrices S and P to HBM. As the softmax couples entire rows or blocks of row, online softmax [11,13] can split the attention computation into blocks, and rescale the output of each block to finally get the right result.” Dao teaches the blockwise computation of the attention operation using online softmax, the output is updated without writing the intermediate attention matrices in HBM. Thus, Dao’s attention computation proceeds blockwise while avoiding storing of the large intermediate attention matrices, corresponding to the claimed streamable attention mechanism under BRI. Regarding Claim 5, Dao discloses the device according to claim 4, wherein the streamable attention mechanism is configured to: generate a softmax input stream based on a first matrix multiplication operation of a particular row of first data and a corresponding row of second data; [Dao: Algorithm 1 Line 8] “On chip, compute S i ( j ) = Q i K j T ∈ R B r × B c .” Dao divides first data Q and second data K into blocks and computes the attention-score block S i ( j ) by matrix multiplication of Q i and K j T . The matrix multiplication performs row-wise dot products between the respective data, generating the values that are then supplied to the online softmax computation (Algorithm 1, line 9), corresponding to the claimed softmax input stream based on the first matrix multiplication under BRI. apply an exponentiation operation of the softmax with norm folding mechanism to the softmax input stream to generate a stream of softmax numerator values; [Dao: §2.3.1] “Online softmax instead computes ‘local’ softmax with respect to each block and rescale to get the right output at the end.” Dao explains that FlashAttention computes attention blockwise using online softmax, this is implemented by computing [Dao: Algorithm 1 Line 9] “   P ~ i j = exp ⁡ S i j - m i j ”. During the blockwise processing of the softmax input values S i j , Dao applies an exponentiation operation to generate the corresponding values P ~ i J . These exponentiated values are then used to compute the softmax normalization and attention output, corresponding to the claimed stream of softmax numerator values generated by an exponentiation operation under BRI. input the stream of softmax numerator values as a first input to a second matrix multiplication operation; [Dao: Algorithm 1 Line 10] “On chip, compute O i j = d i a g e m i j - 1 - m i j - 1 O i j - 1 + P ~ i j V j .” Dao uses the previously generated exponentiated softmax values P ~ i j as an input to the multiplication with V j to update the attention output O i j , corresponding to the claimed stream of softmax numerator values as a first input to a second matrix multiplication operation under BRI. and generate an accumulation sum of the softmax numerator values. [Dao: Algorithm 1, line 9] “ l i j = e m i j - 1 - m i j   l i j - 1 + r o w s u m ( P ~ i j ) ” Dao generates and updates l i j using the row sum of the exponentiated softmax numerator values P ~ i j , accumulating the softmax numerator values during the blockwise computation, corresponding to the claimed generation of an accumulation sum of the softmax numerator values under BRI. Regarding Claim 7, Dao discloses the device according to claim 5, wherein the streamable attention mechanism is configured to perform a norm operation to apply the accumulation sum to an output of the second matrix multiplication operation. [Dao: § 3.1.1] “We can instead maintain an “un-scaled” version of O 2 and keep around the statistics l 2 : O ~ 2 = d i a g l 1 - 1 O 1 + e S 2 - m 2 V 2 .   Only at the every end of the loop do we scale the final O ~ l a s t by d i a g ( l l a s t ) - 1 to get the right output.” Dao maintains the unscaled attention output during the blockwise computation and, after the accumulation is complete, applies the inverse of the accumulated softmax numerator sum l to the unscaled output. As shown in Algorithm 1, the unscaled output is generated by the second matrix multiplication operation at line 10, and the normalization is applied at line 12, corresponding to the claimed performing a norm operation to apply the accumulation sum to an output of the second matrix multiplication operation under BRI. Claims 17, 18, & 19 are substantially similar in scope and spirit as claims 1, 4, & 5, respectively, differing only in reciting A method comprising the recited operations – operations that Dao discloses as being performed by a computing device [Dao: § 3.1.1] “We describe the full FlashAttention-2 forward pass in Algorithm 1.” Therefore the rejection of claims 1, 4, & 5 are applied accordingly. Claim 20 is substantially similar in scope and spirit as claim 1, differing only in reciting A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the recited operations – a computer-implemented form that Dao discloses as software implemented and executed on processing hardware [Dao: Abstract] “We empirically validate that when used end-to-end to train GPT-style models, FlashAttention-2 reaches training speed of up to 225 TFLOPs/s per A100 GPU”; Therefore the rejection of claim 1 is applied accordingly. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 2, 8, & 9 are rejected under 35 U.S.C. 103 as being unpatentable over Dao in view of Xu. Regarding Claim 2, Dao discloses the device according to claim 1, however Dao does not expressly disclose, but in combination with Xu does teach: wherein the input data includes a first image and a second image, [Xu: Figure 2] “We first extract 8x downsampled dense features from two input video frames with a weight-sharing convolutional network” Xu discloses using two consecutive video frames as the input data, corresponding to the claimed input data including a first image and a second image under BRI. and wherein the ML model corresponds to a depth from stereo or optical flow architecture. [Xu: Abstract] “we propose a GMFlow framework, which consists of three main components: a customized Transformer… a new paradigm for accurate and efficient optical flow estimation.” Xu discloses that GMFlow is an optical flow architecture using a Transformer model, corresponding to the claimed ML model corresponding to an optical flow architecture under BRI. Dao and Xu are analogous art because they are from the same field of endeavor, specifically machine-learning models using Transformer-based attention mechanisms for processing input data. It would have been obvious to a person having ordinary skill in the art (PHOSITA), before the effective filing date of the claimed invention, to combine the FlashAttention-2 attention processing of Dao with the optical-flow architecture of Xu in order to improve the speed and efficiency of the optical flow estimation. The suggestion/motivation for doing so would have been provided by Xu itself, which teaches [Xu: Introduction] “such a large number of sequential refinements introduce linearly increasing inference time, which makes it hard for speed optimization and hinders its deployment in real-world applications.” Xu further teaches that [Xu: §2] “our new framework streamlines the optical flow pipeline and estimates large displacements with both high accuracy and efficiency, achieved by a reformulation of the optical flow problem and a strong Transformer.” A PHOSITA would have applied Dao’s efficient Transformer-based attention processing Xu’s optical-flow architecture to reduce the computational burden and inference time associated with optical-flow estimation, combining prior-art elements according to known methods to yield the predictable result of faster and more efficient optical-flow processing (MPEP § 2143.01(A)). Regarding Claim 8, Dao discloses the device according to claim 5, Xu further discloses wherein: the first data corresponds to features associated with a first image; the second data corresponds to features associated with a second image; [Xu § 3.1] “Given two consecutive video frames I 1 and I 2 , we first extract downsampled dense features F 1 ,   F 2 ∈ R H × W × D with a weight-sharing convolutional network, where H ,   W ,   a n d   D denote height, width and feature dimension, respectively … each element in the correlation matrix C represents the correlation value between coordinates p i = i ,   j   i n   F 1   a n d   p 2 = k ,   l   i n   F 2 ” Xu’s correlation matrix uses the features obtained from two consecutive video frames, corresponding to the claimed first and second data corresponding to features associated with a first and second image under BRI. and a second input to the second matrix multiplication operation corresponds to a displacement matrix of coordinates. [Xu: § 3.1] “we normalize the last two dimensions of C with the softmax operation, which gives us a matching distribution M = s o f t m a x C … The, the correspondence G ^ can be obtained by taking a weighted average of the 2D coordinates of pixel grid G … with the matching distribution M :   G ^ = M G   ∈ R H × W × 2 .   Finally, the optical flow V can be obtained by computing the difference between the corresponding pixel coordinates V =   G ^ - G ”. Xu teaches multiplying the softmax matching distribution M by G , which represents the 2-D pixel coordinates, to obtain corresponding coordinates G ^ , the difference between corresponding and original pixel coordinates provides optical-flow displacement V , corresponding to the displacement matrix of coordinates as claimed under BRI. The rationale to combine Dao and Xu is the same as for claim 2 as above. Regarding Claim 9, Dao discloses the device according to claim 1, Xu further discloses wherein the ML model generates probabilistic geometry measures without generating a cost volume data structure. [Xu: §1] “Inspired by sparse matching, we propose to completely remove the additional convolutional layers operating on a predefined local cost volume, and reformulate optical flow…” [Xu: § 3.1] “we normalize the last two dimensions of C with the softmax operation which gives us a matching distribution M”. Xu removes the conventional processing associated with a predefined local cost volume and instead generates a softmax-normalized matching distribution from the correlation matrix. The matching distribution represents correspondence probabilities between locations in the first and second images, corresponding to the claimed probabilistic geometry measures without generating a cost volume data structure under BRI. The rationale to combine Dao and Kim is the same as for claim 2 as above. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Dao in view of Xu, in further view of Kim. Regarding Claim 10, Dao discloses the device according to claim 1, Dao does not expressly disclose but Xu does teach : wherein the ML model is configured to perform regression for depth from stereo or optical flow geometric coordinates [Xu: §2] “For flow prediction at each stage, their pipeline is conceptual similar, i.e., regressing optical flow from a local cost volume with convolutions.” Xu further teaches determining corresponding pixel coordinates and obtaining the optical flow by computing the difference between the corresponding pixel coordinates, corresponding to regression for optical flow geometric coordinates under BRI. The rationale to combine Dao and Xu is the same as above. The combination of Dao and Xu does not expressly disclose, but Kim does teach: using just-in-time computations. [Kim ¶0056] “A processor may sequentially store the image data in an order from a first address to a last address of the (N+1)th line buffer” [Kim: ¶0057] “the processor may use one line buffer to store the image data and use N remaining line buffers to read the received image data. The process may perform an operation of storing image data and an operation of reading image data, simultaneously” Kim’s use of line buffers to simultaneously store and read received image data corresponds to processing image data as it becomes available using just-in-time computations under BRI. Dao, Xu and Kim are analogous art. Dao and Xu are analogous art as set forth above for claim 2. Kim is analogous art to both because Kim is reasonably pertinent to the same problem with which Dao and Xu are concerned, namely performing dense image-correspondence computation within a constrained memory and hardware budget (Kim: ¶[0008]…"another aspect also provides a method and system for improving a performance of a stereo matching system by increasing a matching accuracy while using a minimized amount of hardware resources"). It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to perform the optical-flow regression of Dao combined with Xu using the line-buffered, compute-as-received image data handling of Kim, in order to reduce the memory and hardware resources that the combined pipeline requires The suggestion/motivation for doing so would have been provided by Kim itself, which states the memory-reduction benefit its line-buffer scheme obtains, Kim teaching [Kim: ¶0068] "The stereo matching system may minimize a memory usage for storing the image data using a horizontally long rectangular window with a horizontal length of 2N×(N/2-1) instead of an N×N square window in a process of the area-based stereo matching", and by Dao, which identifies memory traffic as the constraint its own attention design exists to relieve, Dao teaching [Dao: Abstract] "The attention layer is the main bottleneck in scaling to longer sequences, as its runtime and memory increase quadratically in the sequence length". A PHOSITA addressing that shared memory constraint would have had reason to feed the combined attention pipeline from Kim's line buffers, and a reasonable expectation of success, because Kim's scheme delivers image data to the computation in the same row-wise order in which Dao's blockwise attention consumes it. This is the applicable rationale under MPEP § 2143.01(A), combining prior-art elements according to known methods to yield predictable results. Claims 11, 14, & 16 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Dao in view of Kim. Regarding Claim 11, Dao discloses the device according to claim 1, Dao does not expressly disclose, but Kim does teach: further comprising an image sensor configured to generate image data corresponding to the input data. [Kim: ¶0038] “The preprocessor 111 may receive the left image and the right image acquired using the stereo camera 100. The preprocessor 111 may remove noise from the received left and right images…” Kim teaches a stereo camera that acquires the left and right images received by the preprocessor, corresponding to the claimed image sensor configured to generate image data corresponding to the input data under BRI. Dao and Kim are analogous art because Kim is reasonably pertinent to the problem addressed by Dao of efficiently implementing computational processing while reducing memory and hardware-resource requirements. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to combine the attention processing of Dao with the image processing hardware of Kim in order to improve processing performance while reducing the memory and hardware requirements for processing image data. The suggestion/motivation for doing so would have been provided by Kim itself, which teaches [Kim: ¶0008] “a method and system for improving a performance of a stereo matching system by increasing a matching accuracy while using a minimized amount of hardware resources.” A PHOSITA would have been motivated to implement Dao’s attention processing using Kim’s image processing hardware to obtain the improved performance on reduced hardware taught by Kim, combining prior-art elements according to known methods to yield predictable results (MPEP § 2143.01(A)). Regarding Claim 14, Dao discloses the device according to claim 1, Kim further discloses wherein the one or more processors are integrated in at least one of a mobile phone, a tablet computer device, a wearable electronic device, or a camera device. [Kim: ¶0045] “The preprocessor 111 may include a memory configured to store the left image and the right image received from the stereo camera 100” [Kim: ¶0096] “The processing device described herein may be implemented using hardware components, software components, and/or a combination thereof. For example, the processing device and the component described herein may be implemented using one or more general-purpose or special purpose computers, such as, for example, a processor, … , a microprocessor, or any other device capable of responding to and executing instructions in a defined manner... The processing device also may access, store, manipulate, process, and create data in response to execution of the software.” Kim teaches a stereo matching system having processing components that operate on images received from a stereo camera, and further teaches that the components of its disclosed embodiments may be implemented using a processor, corresponding to the claimed one or more processors integrated in a camera device under BRI. The rationale to combine Dao and Kim is the same as for claim 11 as above. Regarding Claim 16, Dao discloses the device according to claim 1, Kim further discloses wherein the one or more processors are included in an integrated circuit [Kim: ¶0096] “The components described in the exemplary embodiments of the present invention may be achieved by hardware components including at least one DSP (Digital Signal Processor), a processor, a controller, an ASIC (Application Specific Integrated Circuit), a programmable logic element such as an FPGA (Field Programmable Gate Array), other electronic devices…” Kim teaches implementing the processing components of its stereo-matching system using integrated circuit hardware, including an ASIC or FPGA, corresponding to the claimed one or more processors comprising an integrated circuit under BRI. The rationale to combine Dao and Kim is the same as for claim 11 as above. Claims 12, 13, & 15 are rejected under 35 U.S.C. 103 as being unpatentable over Dao in view of Lee. Regarding Claim 12, Dao discloses the device according to claim 1, Dao does not expressly disclose, but Lee does teach: further comprising a modem coupled to the one or more processors, the modem configured to receive image data corresponding to the input data from a second device. [Lee: ¶0139-0140] “The communication unit 110 may also be referred to as a communication modem… the input unit 120 may include a camera 121 for image signal input” [Lee: ¶0172] “The processor 180 may receive the image data through a camera 121 or may receive the image data photographed by an external device through a communication unit 110.” Lee teaches receiving image data photographed by an external device through a communication unit 110, which Lee expressly identifies as a communication modem, corresponding to the claimed modem configured to receive image data corresponding to the input data from a second device under BRI. Dao and Lee are analogous art because they are from the same field of endeavor, specifically artificial-intelligence systems designed for processing input data using machine-learning models. It would have been obvious to a PHOSITA, before the effective filing date of the claimed invention, to incorporate the efficient Transformer-based attention architecture of Dao into the AI system of Lee in order to improve the performance and training efficiency of Lee’s machine-learning models for image processing and recognition. The suggestion/motivation for doing so would have been provided by both Dao and Lee themselves. Dao teaches [Dao: Abstract] “Scaling Transformers to longer sequence lengths has been a major problem… promising to improve performance in language modeling and high-resolution image understanding” noting that “The attention layer is the main bottleneck in scaling to longer sequences, as its runtime and memory increase quadratically in the sequence length.” Dao further teaches [Dao: Abstract] “FlashAttention-2, with better work partitioning to address these issues… yield around 2x speedup compared to FlashAttention.” Dao’s FlashAttention-2 offers improved performance and training for machine-learning models directed towards image understanding and processing. Lee similarly recognizes [Lee: ¶0003] “in order to learn the image recognition model used to recognize the object included in the image data, a lot of training data is required” and further teaches that [Lee: ¶0004] “if there is a model that is capable of recognizing the object in the image data by only using little training data, the performance of the image recognition function may be more improved.” A PHOSITA seeking to improve performance and training efficiency of Lee’s image-processing and recognition system would therefore have been motivated to incorporate Dao’s FlashAttention-2 architecture to obtain the predictable benefit of improved performance and training efficiency for image understanding, combining prior-art elements according to known methods to yield predictable results (MPEP § 2143.01(A)) Regarding Claim 13, Dao discloses the device according to claim 1, however Dao does not expressly disclose, but Lee teaches: wherein the one or more processors are integrated in a headset device that includes a display, and wherein the headset device is configured, when worn by a user, to display an output image based on an output of the ML model. [Lee: ¶0113-115] “The XR device 100c, to which the AI technology is applied, may be implemented by a head-mount display (HMD) …The XR device 100c may analyzes three-dimensional point cloud data or image data acquired from various sensors or the external devices, … and render to output the XR object to be output. The XR device may perform the above-described operations by using the learning model composed of at least one artificial neural network.” Lee teaches an XR device implemented as a head-mount display that processes image data using an artificial neural-network learning model and outputs an XR object, corresponding to the claimed processors integrated in a headset device including a display and configured to display an output image based on an output of the ML model under BRI. The rationale to combine Dao and Lee is the same as above. Regarding Claim 15, Dao discloses the device according to claim 1, however Dao does not expressly disclose, but Lee teaches: wherein the one or more processors are integrated in a vehicle, the vehicle further including one or more cameras configured to capture image data corresponding to the input data. [Lee: ¶0102-105] “The self-driving vehicle 100b, to which the AI technology is applied, may be implemented as a mobile robot, a vehicle, an unmanned flying vehicle, or the like… The self-driving vehicle 100b may include a self-driving control module… the self-driving control module may refer to a software module or a chip implementing the software module by hardware… The self-driving vehicle 100b may acquire state information about the self-driving vehicle 100b by using sensor information acquired from various kinds of sensors … may use the sensor information acquired from at least one sensor among the lidar, the radar, and the camera” Lee teaches an AI-enabled vehicle having processing components that operate using sensor information acquired from a camera, corresponding to the claimed processors integrated in a vehicle further including one or more cameras configured to capture image data corresponding to the input data under BRI. The rationale to combine Dao and Lee is the same as above. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Milakov et al. "Online normalizer calculation for softmax” Online softmax computation using running normalization information to reduce the number of passes required for computing softmax Stevens et al. “Softermax: Hardware/Software Co-Design of an Efficient Softmax for Transformers” Efficient softmax processing for Transformer architectures using separate unnormalized softmax and normalization operations to reduce computational and hardware requirements Zhouran et al. “Efficient Attention: Attention with Linear Complexities” Attention processing using reordered matrix multiplications and distributed normalization operations to reduce computational memory complexity Smolyanskiy et al. WO 2019/182974 A2 Stereo Depth Estimation Using Deep Neural Networks Machine-learning-based stereo image processing for determining depth from stereo image data, including efficient neural-network techniques for determining correspondence and disparity. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GRANT F FITCH whose telephone number is (571)270-0621. The examiner can normally be reached Monday-Thursday 7-3. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /G.F.F./Examiner, Art Unit 2124 /ALAN CHEN/Primary Examiner, Art Unit 2125
Read full office action

Prosecution Timeline

Feb 21, 2024
Application Filed
Jul 14, 2025
Response after Non-Final Action
Sep 01, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month