DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to applicant RCE application filed on June 16, 2026 and wherein claims 1, 8, 14 amended.
In virtue of this communication, claims 1-20 are currently pending in this Office Action.
The Office appreciates the explanation of the amendment and analyses of the prior arts, and however, although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993) and MPEP 2145.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(B) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention.
Claim 1 amended as “… mapping one-dimensional discrete inputs to multidimensional intermediate states via one or more matrices, wherein each intermediate state is based on a previous intermediate state and a corresponding discrete input, and the intermediate states to outputs, …” and wherein “the intermediate states” appears to have an insufficient antecedent basis for the limitation and causes confusing because it is unclear whether this “the intermediate states” is referred to “multidimensional inte4rmediate states”, or “previous intermediate state” for each of “one-dimensional discreate inputs”, or a new term that has not been recited previously in claim 1 and “the intermediate states to outputs” appears to be missing a verb, which further causes confusing because it is unclear whether “the intermediate sates” is mapped to or is corresponded to the “outputs”, and thus, renders claim indefinite. Claim 1 further recites “encoding, via …, the audio sample based on … the intermediate states to outputs; … decoding, via the processing circuitry, the encoded audio sample” which is further confusing because it is unclear whether “encoded audio sample” is, or is from, the “outputs” from “the intermediate states” or from “multidimensional intermediate states via one or more matrices”, and thus, further renders claim indefinite. Claims 2-13 are rejected due to the dependencies to claim 1.
Claim 14 rejected for the at least similar reasons described in claim 1 above since claim 14 recited the similar deficient features as recited in claim 1. Claims 15-20 are rejected due to the dependencies to claim 14.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 5, 7-8, 12, 14, 18, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Audhkhasi (US 20220310074 A1) and in view of references Gu et al. (“Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers”, 35th Conference on Neural Information Processing Systems, Advances in Neural Information Processing Systems 34, NeurIPS 2021, p.1-14) and Steven et al. (“Sharing Low Rank Conformer Weights for Tiny Always-on Ambient Speech Recognition Models”, hereinafter Steven, IEEE Xplore, https://ieeexplore.ieee.org/document/10095006, May 5, 2023, p.1-5).
Claim 1: Audhkhasi teaches a method for generating a text representation of a speech sample (title and abstract, ln 1-16, method steps in fig. 5 and executed on a computing device in fig. 6, a computing device including mobile phones, tablets, laptops, smart watches, etc., para 28), comprising:
receiving, via processing circuitry (including computing device above, para 28), an audio sample (acoustic frames 110 captured by a capture device 16a and converted to a digital format at step 502 in fig. 5, para 29);
encoding, via the processing circuitry, the audio sample (at step 504) based on left context of the audio sample (covering limited context sequence window, including left-context or left-only context for the encoder, para 30) with a structured state-space sequence model (NNR with application of mixture model MiMo attention by a set of mixture components of softmaxes over a context window, para 26 and processing input sequence to output sequence, para 3 and with layers, para 6-7) and a conformer (conformer-based acoustic encoders, para 26, including a plurality of conformer layers, para 38), the structured state-space sequence model being initialized with recurrent weights (a set of weights w0, w1, …, wm m=0, …, M-1, applied in the softmaxm, as mixture weights, at each time step k, i.e., recurrent for attention probability distribution function PDF during training and maintained during inference, para 42) and trained with a set of training data (training data, e.g., utterance-transcription pairs, para 22);
decoding, via the processing circuitry, the encoded audio sample (via decoder 230 in fig. 2, para 34 and applied to the output from the encoder 300 and label encoder 220); and
generating, via the processing circuitry, a transcript of the audio sample based on the decoding (transcription 120 as result of a recognition with respect to the input utterance 106, para 29, corresponding to calculated and selected probability distribution function PDF values, and outputted via the softmax 240, para 40).
However, Audhkhasi does not explicitly teach wherein the structured state-space sequence model mapping one-dimensional discrete inputs to multidimensional intermediate states via one or more matrices, wherein each intermediate state is based on a previous intermediate state and a corresponding discrete input, and the intermediate states to outputs, and does not explicitly teach a diagonal matrix of the disclosed recurrent weights with which the structured state-space sequence model is initialized.
Gu teaches an analogous field of endeavor by disclosing a method for processing speech audio with a structured state-space sequence model (title and abstract, 1-19, a Linear State Space Layer LSSL layer in fig. 1) and wherein encoding audio sample is disclosed with the structured state-space sequence model (encoding the speech signal by using CNNs, session 1 Introduction, p.1, and CNNs as special case of LSSLs, session Summary of Contributions, p.3) and wherein the structured state-space sequence model is mapping one-dimensional discrete inputs (ut Є ℛ as its input of mapping, para 1, p.2 and from 1-dimension function or sequence, para 3, p.2) to multidimensional intermediate states via one or more matrices (based on equation 4-5, and derived LSSL, session 3.1, p.4-6, e.g., dimension H that is more than 1, session 3.3 Deep LSSPs, p.6), wherein each intermediate state (xt in equation 4-5, p.4) is based on a previous intermediate state (based on xt-1 in the equation 4) and a corresponding discrete input (ut as the input and discussed above and in a stacked multiple LSSLs layers, session 3.3 Deep LSSLs, p.6), and the intermediate states to outputs (yt is output from xt and ut in equation 5, discussed above) and wherein the structure state-space sequence model is initialized with a matrix of recurrent weights (through initializing matrix A and ∆t, session 3.3 Deep LSSLs, p.6) and trained with a set of training data (training and tuning the LSSL layers in much higher learning rates, para 2, p.10 and training data are inherency for training and tuning processing, structured A and time scale ∆t, session 5.4 LSSL Ablations: Learning the Memory Dynamics and Timescale, p.9) for benefits of improving the performance (by processing very long sequence in the sequence models, abstract, and with efficient inference and unbounded context and parallelizable training and irregular sampling in fig. 1).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the structured state-space sequence model included in the encoder and wherein the structured state-space sequence model mapping the one-dimensional discrete inputs to the multidimensional intermediate states via the one or more matrices, wherein each intermediate state is based on the previous intermediate state and the corresponding discrete input, and the intermediate states to outputs, as taught by Gu, to the structured state-space sequence model included in the encoder in the method for generating the text representation of the speech sample, as taught by Audhkhasi, for the benefits discussed above.
However, the combination of Audhkhasi and Gu does not explicitly teach wherein the matrix of recurrent weights is a diagonal matrix of recurrent weights.
Steven teaches an analogous field of endeavor by disclosing a method for generating an encoder (title and abstract, ln 1-22 and fig. 1) and wherein an encoder is disclosed to include conformer (conformer N blocks in the encoder, as enhanced transformer, session 3.1. Conformer Model, p.2 and session 3.2. Repeat Full Layers, p.2 and session 3.1. Conformer Model, p.2) and wherein the encoder with the conformer is initialized with a diagonal matrix of recurrent weights (weight matrix M in the conformer layers, session 3.4. Low-Rank Factorization, p.2, and weight matrix M is as a diagonal matrix, session 3.4. Low-Rank Factorization, p.3, and before a low-rank decomposition is applied to the M and decomposed into three distinct sub-matrices: U Є ℛmxk, V Є ℛnxk, and ∑ Є ℛkxk into at initial stage and the such diagonal matrix can be calculated through equation 5 by using a singular value decomposition SVD) and trained with a set of training data (fine tuning to reconstructed structure as M~UVT, session 3.4. Low-Rank Factorization, p.3 and including 960 hours of training data, and spoken-word input data structured as 80 log Mel-filterbank energy features with a window size of 25ms and a 10ms stride, session 4.1 Dataset) for benefits of higher efficiency and wider applications (by shrinking size of the model and the number of parameters of each of models, and applied variety of small appliances such as smart phones, wearables with low memory and real-time operations, with essential equivalent quality, abstract).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the diagonal matrix of recurrent weights with which the encoder with the conformer being initialized and trained with the set of training data, as taught by Steven, to the matrix of the recurrent weights with which the structured state-space sequence model being initiated and trained with the set of training data in the method, as taught by the combination of Audhkhasi and Gu, for the benefits discussed above.
Claim 8 has been analyzed and rejected according to claim 1 above and the combination of Audhkhasi, Gu, and Steven further teaches a device comprising: processing circuitry configured to implement the method steps of claim 1 (Audhkhai, computing device including mobile phone, smart watches, smart speakers, etc., para 28 and processing circuitry for those devices, para 59, and Gu, based on the equation 4-5 the processing of the speech signal, session 3.2 Expressivity of LSSLs, and session 3.3 Deep LSSLs, p.5-6, and session 4, p.6-7, and Steven, on-board low-power devices such as smart phones, wearables, with low memory available, abstract and processing circuitry is inherency for those devices above).
Claim 14 has been analyzed and rejected according to claims 1, 8 above and the combination of Audhkhasi, Gu, and Steven further teaches a non-transitory computer-readable storage medium for storing computer-readable instructions (Audhkhasi, non-transitory memory storing programs and instructions, para 51, and memory 620 and processor 610 in fig. 6 and Gu, and Steven, memory in TPU) that, when executed by a computer (Audhkhasi, the processor 610 executed instructions, para 53 and neural processors, abstract), cause the computer to perform the method of claim 1 (Audhkhasi, the discussion in claim 1 above).
Claim 5: the combination of Audhkhasi, Gu, and Steven further teaches, according to claim 1 above, wherein the diagonal matrix of recurrent weights is a real-valued matrix (Audhkhasi, the weights wo, …, wM-1 in equation 1 are inherently real values and Gu, the matrix A, B, C, D are real valued matrix, Session 3.3 Deep LSSLs, p.6, and Steven, the divided into three distinct sub-matrices U Є ℛmxk, V Є ℛnxk, and ∑ Є ℛkxk for calculating weight matrix M, which is inherently real-valued matrix).
Claim 7: the combination of Audhkhasi, Gu, and Steven further teaches, according to claim 1 above, wherein the diagonal matrix of recurrent weights is a 2×2 matrix (Audhkhasi, M softmaxes is set to two mixture components, including a first softmax and a second softmax, i.e., M=0, 1, or 2x2 diagonal matrix and discussed in claim 1 above, and Gu, the matrix parameters, A, B, C, D, session 3.3 Deep LSSLs, and session 4 Combining LSSLs with Continuous-time Memorization, p.6, and Steven, M~UVT while one vector combined with another, i.e., two dimension defined by two separate and distinct sub matrixes, session 3.4. Low-Rank Factorization, p.3).
Claim 12 has been analyzed and rejected according to claims 8, 5 above.
Claim 18 has been analyzed and rejected according to claims 14, 5 above.
Claim 20 has been analyzed and rejected according to claims 14, 7 above.
Claims 2, 9, 15 are rejected under 35 U.S.C. 103 as being unpatentable over Audhkhasi (above) and in view of references Gu (above), Steven (above), and Gulati et al. (“Conformer: Convolution-augmented Transformer for Speech Recognition, INTERSPEECH, October 25-29, 2020, Shanghai, China, pp.5036-5040).
Claim 2: the combination of Audhkhasi, Gu, and Steven further teaches, according to claim 1 above, the structured state-space sequence model (based on the disclosure of Audhkhasi, Gu, and Steven and discussed in claim 1 above), except wherein the structured state-space sequence model is preceded by a convolutional network of the conformer.
Gulati teaches an analogous field of endeavor by disclosing a method for speech recognition (title and abstract, ln 1-* and fig. 1) and wherein receiving an audio sample is disclosed (audio input with a convolution subsampling, session 2. Conformer Encoder, p.5037) and encoding the audio sample with a structured state-space sequence model (including Dropout and linear, etc. in fig. 1 or convolution module in fig. 2, as part of conformer blocks in fig. 1) and a conformer (conformer encoder in fig. 1) and wherein the structured state-space sequence model is preceded by a convolutional network of the conformer (including convolution subsampling and ahead of linear and dropout or ahead of convolution module in fig. 1) for benefits of achieving a higher efficiency with a smaller word error rate WER (by combining CNN with respect to local features effectively and transformers with respect to global feature to model both local and global dependencies of the audio sequence in a parameter-efficient manner, abstract).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied wherein the structured state-space sequence model is preceded by the convolutional network of the conformer, as taught by Gulati, to the structured state-space sequence model in the method, as taught by the combination of Audhkhasi, Gu, and Steven, for the benefits discussed above.
Claim 9 has been analyzed and rejected according to claims 8, 2 above.
Claim 15 has been analyzed and rejected according to claims 14, 2 above.
Claims 3, 10, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Audhkhasi (above) and in view of references Gu (above), Steven (above), and Jiang et al. (“Nextformer: A Convnext Augmented Conformer for End-to-End Speech Recognition”, hereinafter Jiang, source: arxiv.org/abs/2206.14747, p.1-5, 2022).
Claim 3: the combination of Audhkhasi, Gu, and Steven, further teaches, according to claim 1 above, the structured state-space sequence model (based on the disclosure of Audhkhasi. Gu, and Steven, discussed in claim 1 above), except wherein the structured state-space sequence model replaces a convolutional network of the conformer.
Jiang teaches an analogous field of endeavor by disclosing a method for speech recognition (title and abstract, ln 1-22 and fig. 1) and wherein a structured state-space sequence model is disclosed (CNTF module in fig. 1(b) in the Nextformer that is based on CTC/AED system in fig. 1(a)) and a convolutional network of the conformer is disclosed (convolution included in conformer blocks as compared with, session 4.2. Training setups, p.3) and wherein the structured state-space sequence model replaces a convolutional network of the conformer (using Nextfomer encoder to replace the conformer encoder, session 3.NEXTFORMER, P.2 and the conformer blocks including causal convolution, session 4.2. Training setups, p.3) for benefits of achieving more accurately and efficiently in E2E speech recognition (by inserting a additional downsampling in middle of conformer layers in fig. 1(a), abstract and reducing the overall computational cost by Nextformer, session 3.3. Additional downsampling module, p.3).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied wherein the structured state-space sequence model replaces the convolutional network of the conformer, as taught by Jiang, to the structured state-space sequence model and the conformer in the method, as taught by the combination of Audhkhasi, Gu, and Steven for the benefits discussed above.
Claim 10 has been analyzed and rejected according to claims 8, 3 above.
Claim 16 has been analyzed and rejected according to claims 14, 3 above.
Claims 4, 6, 11, 13, 17, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Audhkhasi (above) and in view of references Gu (above), Steven (above), and Goel et al. (“It’s Raw! Audio Generation with State-Space Models”, hereinafter Goel, Proceedings of the 39th International Conference on Machine Learning, Baltimore, Maryland, USA, PMLR 162, p.1-18, 2022).
Claim 4: the combination of Audhkhasi, Gu, and Steven, further teaches, according to claim 1 above, wherein a convolutional kernel of the conformer is disclosed (Audhkhasi, a plurality of conformer layers included in the audio encoder 300 in fig. 2, para 38 and Steven, convolution of the conformer block as convolutional kernel of the conformer in fig. 1 and represented as parameters
PNG
media_image1.png
29
52
media_image1.png
Greyscale
in equation 4, and including Pre-Conv, Conv, and Post-Conv in fig. 1 and parameterization is in equation 4, representing conformer block within the 14M parameter conformer model in fig. 1) and the structure state-space sequence model is also disclosed (based on the disclosure from Audhkhasi and Steven and discussed in claim 1 above), except wherein the convolutional kernel of the conformer is based on parameterization of the structured state-space sequence model.
Goel teaches an analogous field of endeavor by disclosing a method for speech recognition (title and abstract, ln 1-28 and a model SaShIMi in fig. 1) and wherein a structured state-space sequence model is disclosed (a state-space model digitized with parameterization represented by equations 4-5 with parameters
PNG
media_image2.png
20
17
media_image2.png
Greyscale
,
PNG
media_image3.png
20
16
media_image3.png
Greyscale
,) and the structured state-space sequence model is initiated with a diagonal matrix of recurrent weights (S4 is a particular instantiation of state-space model that parameterizes A as a diagonal plus low-rank DPLR matrix, session 3.3. S4, p.5), and wherein a convolutional kernel is disclosed (a convolution kernel
PNG
media_image4.png
19
19
media_image4.png
Greyscale
is computed by the S4 to use a special algorithm, session 3.3. S4, p.5) and based on parameterization of the structured state-space sequence model (using the special HiPPO matrices and specifying parameters A and B with a formula h’(t) = Ah(t) + Bx(t), session 3.3 S4, and solved by unrolling equation 4, session 3.2. State Space Models, p.4) for benefits of high efficiency in the model initialization (by special algorithm for faster computation, session 3.3. S4, p.5, and stabilizing the model S4, abstract).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the convolution kernel and wherein the convolutional kernel is based on parameterization of the structured state-space sequence model, as taught by Goel, to the convolutional kernel of the conformer in the method, as taught by the combination of Audhkhasi, Gu, and Steven, for the benefits discussed above.
Claim 6: the combination of Audhkhasi, Gu, Steven, and Goel further teaches, wherein the diagonal matrix of recurrent weights includes complex numbers (Audhkhasi, weights w0, …, wM-1 and the discussed in claim 1 above, and Steven, weight matrix M in the conformer layers which is a diagonal matrix, session 3.4. Low-Rank Factorization, p.2-3, and the discussion in claim 1 above, and Goel, a state-space model S4 and digitized with parameterization represented by equations 4-5 with parameters
PNG
media_image2.png
20
17
media_image2.png
Greyscale
,
PNG
media_image3.png
20
16
media_image3.png
Greyscale
, and wherein
PNG
media_image2.png
20
17
media_image2.png
Greyscale
is as Hurwitz matrix being diagonal by A = Λ – pp*, session 4.1. Stabilizing S4 for Recurrence, p.5 and with a stable Hurwitz matrix if and only if all
PNG
media_image2.png
20
17
media_image2.png
Greyscale
and real part of
PNG
media_image2.png
20
17
media_image2.png
Greyscale
mapped to the complex left half plane, session Definition 4.2, p.5 and wherein
PNG
media_image2.png
20
17
media_image2.png
Greyscale
comprises complex diagonal matrix as first matrix plus a second matrix, session A.1. S4 Stability, p.12).
Claim 11 has been analyzed and rejected according to claims 8, 4 above.
Claim 17 has been analyzed and rejected according to claims 14, 4 above.
Claim 13 has been analyzed and rejected according to claims 14, 6 above.
Claim 19 has been analyzed and rejected according to claims 14, 6 above.
Response to Arguments
Applicant's arguments filed on June 16, 2026 have been fully considered and but are moot in view of the new ground(s) of rejection necessitated by the applicant amendment. The Office has thoroughly reviewed Applicants' arguments but firmly believes that the cited references to reasonably and properly meet the claimed limitations.
In the response to this office action, the Office respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Office in prosecuting this application.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LESHUI ZHANG whose telephone number is (571)270-5589. The examiner can normally be reached Monday-Friday 6:30amp-4:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LESHUI ZHANG/
Primary Examiner,
Art Unit 2695