Prosecution Insights
Last updated: October 02, 2026
Application No. 18/323,557

System and Method for Automatically Determining Stride Values in Online Streaming Speech Processing Systems

Final Rejection §103
Filed
May 25, 2023
Examiner
VO, STEVEN
Art Unit
2148
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
15 currently pending
Career history
10
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
DETAILED ACTION This action is in response to the application filed on 05/25/2023. Claims 1-20 are pending and have been examined Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1, 2, 4-8, 15-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Doutre et al. (US 20220343894 A1) (hereafter referred to as Doutre) in view of Riad et al. (WO 2023060120 A1) (hereafter referred to as Riad) and Zhang et al. (Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition) (hereafter referred to as Zhang). Regarding claim 1, Doutre teaches performing transfer learning from the first machine learning model to a second machine learning model, wherein the second streaming machine learning model is configured as an online streaming machine learning model (Doutre, page 2, paragraph 18, “The transcripts generated by the non-streaming teacher model may then be used to distill knowledge into the streaming ASR model. In this respect, the non-streaming ASR model functions as a teacher model while the streaming ASR model that is being taught by the distillation process is a student model”) training the second machine learning model with the spectral pooling layer (Doutre, page 4, paragraph 31, “Here, the teacher model 210 distills its knowledge to the student model 152 by training the student model 152 with a plurality of student training samples 232 that include, at least in part, labels or transcriptions 212 generated by the teacher model 210.”) Doutre does not teach, but Riad does teach dynamically determining, by a first machine learning model, a stride value to be used for downsampling to be performed during the processing of the speech signal (Riad, page 9, paragraph 0057, “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer” and “S [strides] may still need to be provided as hyperparameters for each downsampling layer” (Raid, paragraph 00102). Examiner notes that Riad also discloses that DiffStride is the downsampling layer being described [0086]). the stride value being determined by the first machine learning model based on initial configuration parameters comprising an initial stride value, a smoothing parameter for smoothing a selection of a cropping mask to be used as a filter for masking the speech signal being processed, and regularization weights for different encoder layers of the first machine learning model that control a reduction in a number of frames that are processed during processing of the speech signal, which collectively control a balance between efficiency and accuracy when processing the speech signal (Riad, paragraph 0086, “DiffStride may be integrated into a ResNet architecture, which may allow maintaining consistent high performance on CIFAR10, CIFAR100 and ImageNet even when training starts from poor random stride configurations”, “applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain” (Riad, paragraph 0057), and “Moreover, formulating strides as learnable variables may allow introduction of a regularization term that controls the computational complexity of the architecture. This regularization term may allow for a tradeoff between accuracy and efficiency” (Riad, paragraph 0086). Examiner notes that the poor random stride configurations maps to the initial stride value, the input in a spatial domain to a Fourier domain maps the smoothing parameter, and the regularization term maps to the regularization weights). modifying the second machine learning model by inserting a spectral pooling layer between layers of into the second machine learning model, the spectral pooling layer configured to reduce dimensionality of intermediate representations of the speech signal within the second machine learning model during the processing of the speech signal based on using the stride value dynamically determined by the first machine learning mode (Riad, page 6, paragraph 00134, “DiffStride, a downsampling layer with learnable strides, is introduced. As described above, audio and image classification that DiffStride can be used as a drop-in replacement to strided convolutions, removing the need for cross-validating strides”, “The present disclosure includes a learnable stride downsampling layer to leam the size of a cropping mask in a Fourier domain, which may perform resizing in a differentiable way. This learnable stride may be used as a replacement for standard downsampling layers” (Riad, paragraph 0045) and “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer” (Riad, paragraph 0057)). training the second machine learning model with the spectral pooling layer (Riad, page 16, paragraph 0104, “To address the difficulty of searching stride parameters, provided herein is DiffStride, a downsampling layer that may allow spectral pooling to learn its strides through backpropagation.) the stride value controls reduction of a sequence of frames of the speech signal processed by one or more layers of the second machine learning model (Riad, page 6, paragraph 0045, “The present disclosure includes a learnable stride downsampling layer to leam the size of a cropping mask in a Fourier domain, which may perform resizing in a differentiable way. This learnable stride may be used as a replacement for standard downsampling layers”). reduced dimensionality decreases computational load of the second machine learning model during processing of the speech signal (Riad, paragraph 00134, “As the methods described herein can discover multiple equally-accurate stride configurations, a regularization term to favor the most computationally advantageous is also introduced”). Doutre and Riad are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the stride determination portion of the DiffStride layer (from Riad) to the teacher model (from Doutre) and insert the spectral pooling layer portion of the DiffStride layer to the student model (form Doutre). Riad teaches “spectral pooling … alleviates the loss of information of spatial pooling, while enabling fractional downsizing factors. Spectral pooling also preserves low frequencies without aliasing, a known weakness of spatial/temporal convnets” (Riad, page 15, paragraph 0100) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Doutre and Riad do not teach, but Zhang does teach processing the speech signal as a plurality of sequential chunks using the trained second machine learning model (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 2, Doutre, Riad, and Zhang teach the method of claim 1, Doutre further teaches the first machine learning model is a non-streaming machine learning model (Doutre, page 2, paragraph 0018, “To address the transcription performance for streaming ASR models, implementations described herein are directed toward leveraging a non-streaming ASR model as a teacher to generate transcripts for a streaming ASR student model”). Regarding claim 4, Doutre, Riad, and Zhang teaches the method of claim 1, Doutre further teaches the second machine learning model is an automated speech recognition (ASR) online streaming machine learning model (Doutre, page 2, paragraph 0018, “To address the transcription performance for streaming ASR models, implementations described herein are directed toward leveraging a non-streaming ASR model as a teacher to generate transcripts for a streaming ASR student model”). Regarding claim 5, Doutre, Riad, and Zhang teaches the method of claim 1, Riad further teaches determining the stride value includes generating and applying a cropping mask in a frequency domain to filer components of the speech signal (Riad, page 9, paragraph 0057, “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain”. Examiner notes that the downsampling layer that is referenced is the DiffStride layer). Doutre and Riad are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the DiffStride layer (from Riad). One of the ordinary skill in the art would have known to apply the known technique of using a downsampling layer (DiffStride), to create a cropping mask. Therefore, applying Riad’s technique would yield the predictable result of filtering out irrelevant background data (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 6, Doutre, Riad, and Zhang teach the method of claim 1, Doutre and Riad teach the second machine learning model with the spectral pool. Doutre and Riad do not teach, but Zhang does teach training the second machine learning model with the spectral pooling layer includes processing a period of past context for a speech signal (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre and Riad to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 7, Doutre, Riad, and Zhang the method of claim 1, Zhang further teaches training the second machine learning model with the spectral pooling layer includes determining a chunk size for processing a speech signal (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre and Riad to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 8, Doutre, Riad, and Zhang teach the method of claim 1, Zhang further teaches processing a speech signal using the trained second machine learning model by processing each chunk with a period of past context associated with one or more preceding chunks (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 15, Doutre teaches A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations (Doutre, page 8, paragraph 0052, “These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and/or object-oriented programming language, and/or in assembly/machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer readable medium, apparatus and/or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and/or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal”). performing transfer learning from the non-streaming machine learning model to an online streaming machine learning model (Doutre, page 2, paragraph 18, “The transcripts generated by the non-streaming teacher model may then be used to distill knowledge into the streaming ASR model. In this respect, the non-streaming ASR model functions as a teacher model while the streaming ASR model that is being taught by the distillation process is a student model”) training the online streaming machine learning model with the spectral pooling layer (Doutre, page 4, paragraph 31, “Here, the teacher model 210 distills its knowledge to the student model 152 by training the student model 152 with a plurality of student training samples 232 that include, at least in part, labels or transcriptions 212 generated by the teacher model 210.”) Doutre does not teach, but Riad does teach dynamically determining, by a non-streaming machine learning model, a stride value to be used for downsampling to be performed during the processing of the speech signal (Riad, page 9, paragraph 0057, “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer” and “S [strides] may still need to be provided as hyperparameters for each downsampling layer” (Raid, paragraph 00102). Examiner notes that Riad also discloses that DiffStride is the downsampling layer being described [0086]). the stride value being determined by the first machine learning model based on initial configuration parameters comprising an initial stride value, a smoothing parameter for smoothing a selection of a cropping mask to be used as a filter for masking the speech signal being processed, and regularization weights for different encoder layers of the first machine learning model that control a reduction in a number of frames that are processed during processing of the speech signal, which collectively control a balance between efficiency and accuracy when processing the speech signal (Riad, paragraph 0086, “DiffStride may be integrated into a ResNet architecture, which may allow maintaining consistent high performance on CIFAR10, CIFAR100 and ImageNet even when training starts from poor random stride configurations”, “applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain” (Riad, paragraph 0057), and “Moreover, formulating strides as learnable variables may allow introduction of a regularization term that controls the computational complexity of the architecture. This regularization term may allow for a tradeoff between accuracy and efficiency” (Riad, paragraph 0086). Examiner notes that the poor random stride configurations maps to the initial stride value, the input in a spatial domain to a Fourier domain maps the smoothing parameter, and the regularization term maps to the regularization weights). modifying the online streaming machine learning model by inserting a spectral pooling layer between layers of into the online streaming machine learning model, the spectral pooling layer configured to reduce dimensionality of intermediate representations of the speech signal within the online streaming machine learning model during the processing of the speech signal based on using the stride value dynamically determined by the non-streaming machine learning mode (Riad, page 6, paragraph 00134, “DiffStride, a downsampling layer with learnable strides, is introduced. As described above, audio and image classification that DiffStride can be used as a drop-in replacement to strided convolutions, removing the need for cross-validating strides”, “The present disclosure includes a learnable stride downsampling layer to leam the size of a cropping mask in a Fourier domain, which may perform resizing in a differentiable way. This learnable stride may be used as a replacement for standard downsampling layers” (Riad, paragraph 0045) and “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer” (Riad, paragraph 0057)). training the online streaming machine learning model with the spectral pooling layer (Riad, page 16, paragraph 0104, “To address the difficulty of searching stride parameters, provided herein is DiffStride, a downsampling layer that may allow spectral pooling to learn its strides through backpropagation.) the stride value controls reduction of a sequence of frames of the speech signal processed by one or more layers of the online streaming machine learning model (Riad, page 6, paragraph 0045, “The present disclosure includes a learnable stride downsampling layer to leam the size of a cropping mask in a Fourier domain, which may perform resizing in a differentiable way. This learnable stride may be used as a replacement for standard downsampling layers”). reduced dimensionality decreases computational load of the online streaming machine learning model during processing of the speech signal (Riad, paragraph 00134, “As the methods described herein can discover multiple equally-accurate stride configurations, a regularization term to favor the most computationally advantageous is also introduced”). Doutre and Riad are considered analogous to the claimed invention because they both deal with speech recognitions. . It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the stride determination portion of the DiffStride layer (from Riad) to the teacher model (from Doutre) and insert the spectral pooling layer portion of the DiffStride layer to the student model (form Doutre). Riad teaches “spectral pooling … alleviates the loss of information of spatial pooling, while enabling fractional downsizing factors. Spectral pooling also preserves low frequencies without aliasing, a known weakness of spatial/temporal convnets” (Riad, page 15, paragraph 0100) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Doutre and Riad do not teach, but Zhang does teach processing the speech signal as a plurality of sequential chunks using the trained online streaming machine learning model (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 16, Doutre, Riad, and Zhang teach the product of claim 15, Riad further teaches determining the stride value includes generating and applying the cropping mask in a frequency domain to filter components of the speech signal (Riad, page 9, paragraph 0057, “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain”. Examiner notes that the downsampling layer that is referenced is the DiffStride layer). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the DiffStride layer (from Riad). One of the ordinary skill in the art would have known to apply the known technique of using a downsampling layer (DiffStride), to create a cropping mask. Therefore, applying Riad’s technique would yield the predictable result of filtering out irrelevant background data (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 17, Doutre, Riad, and Zhang teach the product of claim 15, Doutre and Riad teach the second machine learning model with the spectral pool. Doutre and Riad do not teach, but Zhang does teach training the online streaming machine learning model with the spectral pooling layer includes processing a period of past context for a speech signal (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre and Riad to apply the chunk attention from Zhang into the student model. Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 18 Doutre, Riad, and Zhang teach the product of claim 15, Zhang further teaches training the online streaming machine learning model with the spectral pooling layer includes determining a chunk size for processing a speech signal (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre and Riad to apply the chunk attention from Zhang into the student model. Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 19, Doutre, Riad, and Zhang teaches the method of claim 15, Doutre further teaches the online streaming machine learning model is an automated speech recognition (ASR) online streaming machine learning model (Doutre, page 2, paragraph 0018, “To address the transcription performance for streaming ASR models, implementations described herein are directed toward leveraging a non-streaming ASR model as a teacher to generate transcripts for a streaming ASR student model”). Regarding claim 20, Doutre, Riad, and Zhang teaches the method of claim 15, Doutre further teaches processing a speech signal using the trained online machine learning model by processing each chunk with a period of past context associated with one or more preceding chunks (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Claim(s) 3, 9, 10, 12-14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Doutre, Riad, and Zhang in view of Tripathi et al. (US 20210343273 A1) (hereafter referred as Tripathi). Regarding claim 3, Doutre, Riad, and Zhang teach the method of claim 1, Doutre and Riad does not teach, but Tripathi does teach the first machine learning model is a first online streaming machine learning model (Tripathi, page 1, paragraph 0002, “ASR system is deployed on a mobile phone that experiences direct user interactivity, an application on the mobile phone using the ASR system may require the speech recognition to be streaming such that words appear on the screen as soon as they are spoken.” Examiner notes FIG. 2A illustrates audio data being inputted into a machine learning model) Doutre, Riad, Zhang and Tripathi are considered analogous to the claimed invention because they all deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to have the teacher model (from Doutre) be an online streaming machine learning model (from Tripathi). Doing so is advantageous because “when using an ASR system today there is a demand that the ASR system decode utterances in a streaming fashion that corresponds to real-time or even faster than real-time” (Tripathi, page 1, paragraph 0002) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 9, Doutre teaches a memory (Doutre, page 6, paragraph 0046, “The computing device 500 includes a processor 510 (e.g., data processing hardware 112, 144), memory 520 (e.g., memory hardware 114, 146), a storage device 530, a high-speed interface/controller 540 connecting to the memory 520 and high-speed expansion ports 550, and a low speed interface/controller 560 connecting to a low speed bus 570 and a storage device 530.”) a processor (Doutre, page 6, paragraph 0046, “The computing device 500 includes a processor 510 (e.g., data processing hardware 112, 144), memory 520 (e.g., memory hardware 114, 146), a storage device 530, a high-speed interface/controller 540 connecting to the memory 520 and high-speed expansion ports 550, and a low speed interface/controller 560 connecting to a low speed bus 570 and a storage device 530.”) performing transfer learning from the first online streaming machine learning model to a second online streaming machine learning model (Doutre, page 2, paragraph 18, “The transcripts generated by the non-streaming teacher model may then be used to distill knowledge into the streaming ASR model. In this respect, the non-streaming ASR model functions as a teacher model while the streaming ASR model that is being taught by the distillation process is a student model” Examiner notes Doutre only teaches transfer learning from a machine learning model to an online streaming model) training the online streaming machine learning model with the spectral pooling layer (Doutre, page 4, paragraph 31, “Here, the teacher model 210 distills its knowledge to the student model 152 by training the student model 152 with a plurality of student training samples 232 that include, at least in part, labels or transcriptions 212 generated by the teacher model 210.”) Doutre does not teach, but Riad does teach dynamically determining, by a non-streaming machine learning model, a stride value to be used for downsampling to be performed during the processing of the speech signal (Riad, page 9, paragraph 0057, “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer” and “S [strides] may still need to be provided as hyperparameters for each downsampling layer” (Raid, paragraph 00102). Examiner notes that Riad also discloses that DiffStride is the downsampling layer being described [0086]). the stride value being determined by the first machine learning model based on initial configuration parameters comprising an initial stride value, a smoothing parameter for smoothing a selection of a cropping mask to be used as a filter for masking the speech signal being processed, and regularization weights for different encoder layers of the first machine learning model that control a reduction in a number of frames that are processed during processing of the speech signal, which collectively control a balance between efficiency and accuracy when processing the speech signal (Riad, paragraph 0086, “DiffStride may be integrated into a ResNet architecture, which may allow maintaining consistent high performance on CIFAR10, CIFAR100 and ImageNet even when training starts from poor random stride configurations”, “applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain” (Riad, paragraph 0057), and “Moreover, formulating strides as learnable variables may allow introduction of a regularization term that controls the computational complexity of the architecture. This regularization term may allow for a tradeoff between accuracy and efficiency” (Riad, paragraph 0086). Examiner notes that the poor random stride configurations maps to the initial stride value, the input in a spatial domain to a Fourier domain maps the smoothing parameter, and the regularization term maps to the regularization weights). modifying the online streaming machine learning model by inserting a spectral pooling layer between layers of into the online streaming machine learning model, the spectral pooling layer configured to reduce dimensionality of intermediate representations of the speech signal within the online streaming machine learning model during the processing of the speech signal based on using the stride value dynamically determined by the non-streaming machine learning mode (Riad, page 6, paragraph 00134, “DiffStride, a downsampling layer with learnable strides, is introduced. As described above, audio and image classification that DiffStride can be used as a drop-in replacement to strided convolutions, removing the need for cross-validating strides”, “The present disclosure includes a learnable stride downsampling layer to leam the size of a cropping mask in a Fourier domain, which may perform resizing in a differentiable way. This learnable stride may be used as a replacement for standard downsampling layers” (Riad, paragraph 0045) and “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer” (Riad, paragraph 0057)). training the online streaming machine learning model with the spectral pooling layer (Riad, page 16, paragraph 0104, “To address the difficulty of searching stride parameters, provided herein is DiffStride, a downsampling layer that may allow spectral pooling to learn its strides through backpropagation.) the stride value controls reduction of a sequence of frames of the speech signal processed by one or more layers of the online streaming machine learning model (Riad, page 6, paragraph 0045, “The present disclosure includes a learnable stride downsampling layer to leam the size of a cropping mask in a Fourier domain, which may perform resizing in a differentiable way. This learnable stride may be used as a replacement for standard downsampling layers”). reduced dimensionality decreases computational load of the online streaming machine learning model during processing of the speech signal (Riad, paragraph 00134, “As the methods described herein can discover multiple equally-accurate stride configurations, a regularization term to favor the most computationally advantageous is also introduced”). Doutre and Riad are considered analogous to the claimed invention because they both deal with speech recognitions. . It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the stride determination portion of the DiffStride layer (from Riad) to the teacher model (from Doutre) and insert the spectral pooling layer portion of the DiffStride layer to the student model (form Doutre). Riad teaches “spectral pooling … alleviates the loss of information of spatial pooling, while enabling fractional downsizing factors. Spectral pooling also preserves low frequencies without aliasing, a known weakness of spatial/temporal convnets” (Riad, page 15, paragraph 0100) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Doutre and Riad do not teach, but Zhang does teach processing the speech signal as a plurality of sequential chunks using the trained online streaming machine learning model (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Doutre, Riad, and Zhang do not teach, but Tripathi does teach performing transfer learning from the first online streaming machine learning model to a second online streaming machine learning model (Tripathi, page 1, paragraph 0002, “ASR system is deployed on a mobile phone that experiences direct user interactivity, an application on the mobile phone using the ASR system may require the speech recognition to be streaming such that words appear on the screen as soon as they are spoken.” Examiner notes FIG. 2A illustrates audio data being inputted into a machine learning model) Doutre, Riad, Zhang and Tripathi are considered analogous to the claimed invention because they all deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to have the teacher model (from Doutre) be an online streaming machine learning model (from Tripathi). Doing so is advantageous because “when using an ASR system today there is a demand that the ASR system decode utterances in a streaming fashion that corresponds to real-time or even faster than real-time” (Tripathi, page 1, paragraph 0002) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 10, Doutre, Riad, Zhang and Tripathi teach the system of claim 9, Riad further teaches determining the stride value includes generating and applying the cropping mask in a frequency domain to filter components of the speech signal (Riad, page 9, paragraph 0057, “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain”. Examiner notes that the downsampling layer that is referenced is the DiffStride layer). Doutre, Riad, Zhang, and Tripathi are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the DiffStride layer (from Riad). One of the ordinary skill in the art would have known to apply the known technique of using a downsampling layer (DiffStride), to create a cropping mask. Therefore, applying Riad’s technique would yield the predictable result of filtering out irrelevant background data (See MPEP 2141 (III)(D) Applying a known technique to a known device ready for improvement to yield predicable results. Regarding claim 12, Doutre, Riad, and Zhang teach the system of claim 9 Doutre and Riad teach the second machine learning model with the spectral pool. Doutre and Riad do not teach, but Zhang does teach training the online streaming machine learning model with the spectral pooling layer includes processing a period of past context for a speech signal (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre and Riad to apply the chunk attention from Zhang into the student model. Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 13, Doutre, Riad and Tripathi teach the system of claim 9, Zhang further teaches training the second online streaming machine learning model with the spectral pooling layer includes determining a chunk size for processing a speech signal (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Regarding claim 14, , Doutre, Riad, Zhang and Tripathi teach the system of claim 9, Zhang further teaches processing a speech signal using the trained second machine learning model by processing each chunk with a period of past context associated with one or more preceding chunks (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Doutre, Riad, and Zhang are considered analogous to the claimed invention because they both deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre to apply the chunk attention from Zhang into the student model). Zhang teaches “the whole latency of the encoder depends on the chunk size, which is easy to control and implement. We can train the model using a fixed chunk size, we call it static chunk training, and decoding with the same chunk” (Zhang, Section 3.2.2) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Doutre, Riad, Zhang and Tripathi in view of Chen et al. (US 20190318757 A1) (hereafter referred as Chen). Regarding claim 11, Doutre, Riad, Zhang and Tripathi teach the system of claim 9, Doutre, Riad, Zhang, and Tripathi do not teach, but Chen does teach determining the stride value includes processing a period of future context from a speech signal (Chen, page 4, paragraph 0045, “BLSTM layers of the speech separation model can accommodate variable-length inputs. The output of one BLSTM layer may be fed back into the same layer, thus allowing the BLSTM layers to “remember” the past and future context when processing a given stream of audio segments”. Examiner notes BLSTM means Bidirectional Long Short Term Memory). Doutre, Riad, Zhang, Tripathi, and Chen are considered analogous to the claimed invention because they all deal with speech recognitions. It would have been obvious to one having ordinary skill in the art prior to the effective filing date to have modified Doutre ,Riad and Tripathi to add a BLSTM layer (from Chen) into the teacher model (from Doutre). Chen teaches that it “allows the network to use the surrounding context of a given segment, e.g., segments before and after the current input segment, to contribute to the determination of the mask for the current input segment” (Chen, page 4, paragraph 0045) (See MPEP 2141 (III)(G) Some teaching, suggestion, or motivation in the prior art that would have led one of ordinary skill to modify the prior art reference or to combine prior art reference teachings to arrive at the claimed invention). Response to Arguments One page 7-8, Applicant argues: Applicant respectfully traverses the 101 rejection, asserting that the pending claims are directed to patent-eligible subject matter. In particular, as now further clarified by the amendments of this paper, the claims are directed to a specific technological improvement in the operation of a neural network for streaming speech processing by modifying the model architecture to perform controlled dimensionality reduction of signal representations. The amended claims integrate the claimed processing into a practical application that transforms speech signal data and improves computational efficiency. This closely aligns with the USPTO's subject matter eligibility guidance, particularly Example 47, which recognizes that claims directed to specific techniques for training and configuring neural networks to improve their performance are patent-eligible. Like that example, the amended claims show a specific approach for determining and applying stride values within a neural network architecture. As discussed during the interview, the amended limitations directed to determination of the stride value supports eligibility under §101 by clarifying that the stride is derived through operation of a machine learning model on speech data using specific training mechanisms. The claims now recite that the stride value is dynamically determined based on initial configuration parameters including a smoothing parameter associated with a cropping mask and regularization weights that control reduction in the number of frames processed across encoder layers. As described in paragraphs [0024]-[0031], the cropping mask is applied in the frequency domain as part of a differentiable training process, and the regularization term governs a tradeoff between computational efficiency and recognition accuracy. Accordingly, the determination and application of the stride value is tied to concrete signal processing operations and model training behavior and is integrated into the functioning of the machine learning system. For at least these reasons, the claims are directed to patent-eligible subject matter and satisfy the requirements of 35 U.S.C. §101. Withdrawal of the §101 rejection is respectfully requested. Regarding the Applicant’s argument that the claim is an improvement to a technology or technical field, Examiner agrees. “Determining … the stride value” recites an abstract idea, however, “inserting a spectral pooling … configured to reduce dimensionality of intermediate representations of the speech signal within the second machine learning model during the processing of the speech signal based on the stride value dynamically determined by the first machine learning model” shows improvement over the abstract idea. Therefore, the 35 U.S.C §101 rejection has been withdrawn. On page 9-11, Applicant argues: Applicant respectfully submits that the pending claims, particularly as now amended, are distinguished from and are patentable over the cited combination of references. As discussed during the interview, the amended claims clarify the specific configuration and interaction of elements recited in the claims, including how the stride value is dynamically determined and how that stride value is applied within a modified machine learning model architecture. These clarifications distinguish the claimed invention from the cited references. Doutre, the primary reference, fails to disclose or suggest any embodiment in which a stride value is dynamically determined by using a machine learning model. Instead, Doutre relies on predefined model configurations for establishing stride values. In contrast to Doutre, the present claims require a data-driven determination of stride that is learned through operation of a first machine learning model and subsequently applied to a second model. Doutre also fails to suggest modifying a neural network by inserting a spectral pooling layer between layers or using the layer to reduce dimensionality of intermediate representations based on the learned stride value. The additional references likewise do not address these limitations or suggest this specific combination of learned stride determination and architectural modification of a neural network. While Riad is cited broadly for stride-related and spectral pooling concepts, Riad also fails to teach or suggest any embodiment for determining stride values using a first machine learning model in the claimed manner or applying those learned stride values to modify the architecture of a second model through insertion of spectral pooling layers that operate based on the learned stride value. Tripathi and Chen similarly fail to disclose or suggest the claimed relationship between learned stride determination, architectural modification using spectral pooling, and streaming, chunk-based processing of speech signals. It is also noted that the amended claims require that the learned stride value controls reduction of frames within the model and is applied in a streaming, chunk-based processing context to reduce computational load of subsequent layers. The cited references fail to teach or suggest this specific relationship between a learned stride parameter, dimensionality reduction with spectral pooling, and improved computational efficiency in a streaming system, in the manner claimed. For at least these reasons, the pending claims are distinguished from the cited references and rejections of record. It will be appreciated, however, that while the foregoing remarks are primarily focused on some of the differences between the independent claims and the references of record, this does not mean that these are necessarily the only differences. Applicant, thus, does not acquiesce to any asserted rejections that have not been specifically traversed at this time, particularly with regard to the dependent claims, and reserves the right to challenge any of the purported teachings or assertions made in the last action at any appropriate time in the future. Regarding the Applicant’s argument that the cited combination of references fail to disclose or suggest the amended claims, the Examiner respectfully disagrees. In claim 1, Riad further teaches the stride value “to be used for downsampling to be performed during the processing of the speech signal, the stride value being determined by the first machine learning model based on initial configuration parameters comprising an initial stride value, a smoothing parameter for smoothing a selection of a cropping mask to be used as a filter for masking the speech signal being processed, and regularization weights for different encoder layers of the first machine learning model that control a reduction in a number of frames that are processed during processing of the speech signal, which collectively control a balance between efficiency and accuracy when processing the speech signal”: (Riad, paragraph 0086, “DiffStride may be integrated into a ResNet architecture, which may allow maintaining consistent high performance on CIFAR10, CIFAR100 and ImageNet even when training starts from poor random stride configurations”, “applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain” (Riad, paragraph 0057), and “Moreover, formulating strides as learnable variables may allow introduction of a regularization term that controls the computational complexity of the architecture. This regularization term may allow for a tradeoff between accuracy and efficiency” (Riad, paragraph 0086). Riad’s DiffStride is configured to train with poor random stride configurations, construct a mask, and regularization term to allow a tradeoff between accuracy and efficiency. These components map direct the initial stride value, a smoothing parameter, and regularization weights, respectively. In claim 1, Raid and Doutre teaches “the spectral pooling layer configured to reduce dimensionality of intermediate representations of the speech signal within the second machine learning model during the processing of the speech signal based on the stride value dynamically determined by the first machine learning model”: (Riad, page 6, paragraph 00134, “DiffStride, a downsampling layer with learnable strides, is introduced. As described above, audio and image classification that DiffStride can be used as a drop-in replacement to strided convolutions, removing the need for cross-validating strides”, “The present disclosure includes a learnable stride downsampling layer to leam the size of a cropping mask in a Fourier domain, which may perform resizing in a differentiable way. This learnable stride may be used as a replacement for standard downsampling layers” (Riad, paragraph 0045) and “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer” (Riad, paragraph 0057)). (Doutre, page 4, paragraph 31, “Here, the teacher model 210 distills its knowledge to the student model 152 by training the student model 152 with a plurality of student training samples 232 that include, at least in part, labels or transcriptions 212 generated by the teacher model 210.”) Raid teaches the downsampling layer resizing the size of the cropping mask, which maps to the spectral pooling layer reducing dimensionality of the speech signal representation. Doutre teaches the second machine learning model In claim 1, Riad teaches “reduced dimensionality decreases computational load of the second machine learning model during processing of the speech signal”: (Riad, paragraph 00134, “As the methods described herein can discover multiple equally-accurate stride configurations, a regularization term to favor the most computationally advantageous is also introduced”). Computationally advantageous maps to a decrease computational load. In claim 1, Doutre and Zhang et al. (Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition) further teaches “processing the speech signal as a plurality of sequential chunks using the trained second machine learning model”. (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Zhang teaches the chunk with a period of past context with preceding chunks and apply that to the second machine learning model from Doutre. In claim 5, Riad further teaches “determining the stride value includes generating and applying the cropping mask in a frequency domain to filter components of the speech signal.” (Riad, page 9, paragraph 0057, “method 200 may include applying a downsampling layer of the machine learning model to the plurality of batches of the training data to determine a stride comprising a learnable parameter for the downsampling layer, where applying the downsampling layer of the machine learning model to a batch of the training data comprises: projecting an input in a spatial domain to a Fourier domain, constructing a mask in the Fourier domain based on a current value of the stride and dimensions of the input, applying the mask as a low-pass filter to the projected input to produce a tensor in the Fourier domain, cropping the tensor based on the mask, and transforming the cropped tensor to the spatial domain”). The downsampling layer comprises a mask in the Fourier domain that maps with the cropping mask in a frequency domain. In claim 8, Zhang further teaches “processing a speech signal using the trained second machine learning model by processing each chunk with a period of past context associated with one or more preceding chunks”: (Zhang, Section 3.2.2, “We adopt a chunk attention in this work … we split the input to several chunks by a fixed chunk size C … every chunk depends on itself and the all the previous chunks”). Zhang teaches the chunk with a period of past context with preceding chunks and apply that to the second machine learning model from Doutre. Claims 9 and 15 follow the same reasoning as claim 1, claims 10 and 16 follow the same reasoning as claim 5, and claims 14 and 20 follow the same reasoning as claim 8, since they are the system claims and non-transitory computer readable medium claims, respectively. Examiner respectfully directs the Applicant to the above 103 rejection. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Sukyas et al. (System and Method for Training Domain-Specific Speech Recognition Language Models) discloses a starting ASR (Automated Speech Recognition) neural network model configured to receive audio in the language and generate a starting transcript of the audio. Sypniewski et al. (End-to-End Neural Networks for Speech Recognition and Classification) discloses end-to-end neural networks for speech recognition and classification and additional machine learning techniques that may be used in conjunction or separately. Penn et al. (System and Method for Applying a Convolutional Neural Network to Speech Recognition) discloses applying a convolutional neural network (CNN) to speech recognition. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEVEN VO whose telephone number is (571)272-9622. The examiner can normally be reached Monday - Friday from 7-3 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.V./Examiner, Art Unit 2148 /MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

May 25, 2023
Application Filed
Mar 09, 2026
Non-Final Rejection mailed — §103
Apr 23, 2026
Applicant Interview (Telephonic)
Apr 27, 2026
Examiner Interview Summary
May 14, 2026
Response Filed
Aug 12, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month