Prosecution Insights
Last updated: July 26, 2026
Application No. 18/539,764

MULTI-TIME-SCALE NEURAL AUDIO CODEC STREAMS

Final Rejection §103
Filed
Dec 14, 2023
Priority
Oct 18, 2023 — provisional 63/591,181
Examiner
OPSASNICK, MICHAEL N
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Cisco Technology Inc.
OA Round
2 (Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
750 granted / 916 resolved
+19.9% vs TC avg
Moderate +10% lift
Without
With
+10.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
35 currently pending
Career history
960
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
50.0%
+10.0% vs TC avg
§102
32.5%
-7.5% vs TC avg
§112
1.2%
-38.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 916 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5,7,8,10-17,19-23, 25-29 are rejected under 35 U.S.C. 103 as being unpatentable over Zeghidor et al (“Soundstream: An End-to-End Neural Audio Codec”, IEEE/ACMtransaction on Audio, speech, and Language PRicessing, vol. 30, 2022, pp 495-507), hereinafter “Soundstream:” in view of Narayanan et al (20220122622) As per claim 1, “Soundstream:” teaches a method comprising: obtaining speech audio to be encoded (pp 495, second column, last paragraph “a novel audio codec that can compress speech, music, and general audio..”; encoding the speech audio with a plurality of audio encoders the plurality of encoded streams including one or more encoded streams at a faster time scale to represent faster-varying aspects of the speech audio; and one or more encoded streams of the plurality of encoded streams at a slower time scale to represent slower-varying aspects of the speech audio (as, the encoder in Soundstream handles multi-rate signals, pg 497, first column, second paragraph – due to the residual vector quantizer and dropout scheme; for this aspect, see Section III-C – starts on bottom of pg 498, of importance – pp 499 – see the use of the exponential average on the quantizer – as well as the variable stride lengths, so as to emphasize the critical features of the signal; the adaptability of the residual vector quantizer emphasizes both slow and fast relevant artifacts; examiner notes that applicants specification discloses the concepts of varying stride lengths to ‘handle differently’, slowly varying artifacts vs faster varying artifacts – see appl. Spec, para 0045 – 0048, as well as para 0057); and and decoding the plurality of encoded streams to generate output audio (as decoded output – see Fig. 2). As per claim 1, “Soundstream:” teaches the encoding approach, that matches applicants embodiment of non-parallel processing (see applicants specification, para 0052, having a single encoder shared among streams); however, the claim specifies that the encoders are in parallel; Narayanan et al (20220122622) teaches cascaded encoders in parallel, processing the data input (see fig. 3, subblock 300 containing encoder 200b and 200c; see further in the specification, para 0036/0037 showing that model 200b can perform both streaming/non-streaming information, at a faster basis, where the application requires as little latency as possible; and model 200c operates on slower-varying artifacts, such as speech recognition for voicemail, wherein the latency demand is less – ie, higher latency is tolerable – para 0038). Narayanan et al (20220122622) teaches the concept of parallel processing of audio analysis encoding based on speed – ie, slower changing artifacts and faster changing artifacts. Therefore, it would have been obvious to one of ordinary skill in the art of audio processing to modify the processing as shown in “Soundstream:” with a parallel processing of audio information split along the lines of speed (slow vs fast) changing audio artifacts, as taught by Narayanan et al (20220122622), because it would advantageously allow for specific tailoring of encoding/decoding by the needs of the application – ie, higher latency applications can be allowed to perform more processing to generate more accurate results; rather than being forced to process faster (because of the demands of low latency applications) and forced lower accuracy across all applications (see Narayanan et al (20220122622), para 0021, last 13 lines, starting with “On the other hand…”). As per claim 2, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 1, wherein the encoded information comprises codeword indices generated using a neural network audio codec system comprising an audio encoder and an audio decoder trained end-to-end (see “Soundstream:”, using neural based end-to-end audio codecs (encoder/decoder) – see pp 497, first column “Neural audio codecs:…”. As per claim 3, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 2, wherein the audio encoder is trained with multiple versions of audio sharing the same phonetic content but different global attributes including speaker, emotion and/or prosody (see “Soundstream:”, pp 497, column 1, second paragraph .. ‘Soundstream does not make assumption on the signal it encodes – therefore, it can process audio, as well as speech combining speaker, phonetic and pitch embedding). As per claim 4, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 3, wherein the multiple versions of audio are a result of randomly changing a speed or time shift given time interval or segment of the speech audio to produce augmented segments and enforcing the encoded information to be the same for the speech audio and the augmented segments (as, changing the bitrate evaluation based on the measured parameters – see “Soundstream:”, pp497, column 1, second paragraph, starting with “Unlike[42]…”, see “the ability of a single model to operate at different bitrates at no additional cost, thanks to…training scheme”). As per claim 5, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 1, wherein respective ones of the plurality of encoded streams carry different properties of the speech audio (see “Soundstream:”, the Speechstream system can handle “diverse audio content types” – pp497, col. 1, second paragraph), in view of the opening comment, that the codec disclosed can handle audio/speech/and any combination thereof.) As per claim 7, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 6, wherein the plurality of audio encoders are configured with different encoder parameters including: one or more of: number of audio samples used to create a token in an encoded stream, frequency of tokens in the encoded stream, embedding vector dimensionality, and loss function used to train the audio encoder. (see “Soundstream:”, vector dimensionality – pg 498, second column, under “Residual Vector Quantizer”; examiner notes that the claim scope is “one or more”, and one of the elements has been met by the “Soundstream:” reference). As per claim 8, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 6, wherein the plurality of audio encoders are trained by enforcing respective ones of the plurality of audio encoders to learn only certain attributes of speech audio using one more loss functions (see “Soundstream:”, using variable loss functions, which are dependent upon the feature space – page 500, col. 1, see “feature loss” and “spectral reconstruction loss”, and further the definition of the “feature loss”). As per claim 10, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 1, further comprising adjusting a time scale of one or more of the plurality of encoded streams based on the speech audio to be encoded (see “Soundstream:”, adjustable for many bitrates – pp 499, first column, middle – “Enabling bitrate scalability). As per claim 11, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 10, wherein adjusting comprises reducing the time scale of one or more of the plurality of encoded streams based on ease of prediction of future encoded information from previously generated encoded information (see “Soundstream:”, using causal prediction – pp 496, first paragraph – as, using previously generated information to predict the current; see also p498, first paragraph). As per claim 12, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 1, further comprising: transmitting redundant information about a previous audio packet together with a current audio packet by transmitting only a higher sending rate encoded stream for the previous audio packet and transmitting the plurality of encoded streams for the current audio packet, wherein decoding comprises decoding a longer time scale encoded stream from the current audio packet with a shorter time scale encoded stream for the previous audio packet in order to reconstruct the previous audio packet (as using generative adversarial models, applying discriminators over multiple scales (see “Soundstream:”, sending rates) as well as differing periods in the audio samples – pp 496, second columns, first paragraph -- “Soundstream:” disclosed that the design of their decoder is similar to the aforementioned decoders in the paragraph – with GAN based discriminators looking at previous audio information to glean information about the current period/packets of the incoming encoded signal). As per claim 13, the combination of “Soundstream:” in view of Narayanan et al (20220122622) teaches the method of claim 1, wherein encoding comprises: encoding first speech audio for a first speaker to produce at least a first encoded stream at a first time scale and a second encoded stream at a second time scale; and encoding second speech audio for a second speaker to produce at least a first encoded stream at a first time scale and a second encoded stream at a second time scale, wherein decoding comprises decoding the first encoded stream for the first speech audio and the second encoded stream for the second speech audio to generate output audio that associates properties of the first speaker to speech content of the second speech audio of the second speaker (see “Soundstream:”, figure 3, wherein an input stream is entered into the encoder, with a decoded output (this matches the claim scope for input stream and output stream); see pp497, column 1, para’s #1 and #2, wherein Soundstream operates on variable bitrates – “i.e., the ability of a single model to operate at different bitrates” – this matches the claim scope, toward, ‘a first time scale’ and a ‘second time scale’, and the decoder operates on these same timescales – see Fig. 3 and see pp 498, first column, “decoder architecture”, wherein the strides are used to mark the timescale; ; and lastly, toward multiple speakers or audio or both, “Soundstream:” teaches the ability to work on ‘diverse audio content types’ – pp497, second paragraph; and inferring from the beginning of the second paragraph, “Soundstream:” successfully operates not only on speech with multiple speakers, but audio as well). Claims 14-17, 19, 20 are system claims performing the steps found across method claims, ranging across claims 1-5,7,8,10-13 above and as such, claims 14-17,19, 20 are similar in scope and content to the commonly found claim steps in claims 1-5,7,8,10-13 above; therefore, claims 14-17, 19, 20 are rejected under similar rationale as presented against claims 1-5,7,8,10-13 above. Claims 21-23, 25-29 are apparatus claims that perform steps found throughout, in method claims 1-5,7,8,10-13 above and as such, claims 21-23, 25-29 are similar in scope and content to the commonly found steps in method claims 1-5,7,8,10-13 above; therefore, claims 21-23, 25-29 are rejected under similar rationale as presented against claims 1-5,7,8,10-13 above. Furthermore, “Soundstream:” teaches a process performing the steps – para 505, first column, last 3 lines. Response to Arguments Applicant’s arguments with respect to the claim(s) have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Examiner notes the introduction of the Narayanan et al (20220122622) reference to introduce the concept of parallel processing of encoding the audio stream; applicants arguments against the “Soundstream:” reference, are found on the last 2/3rds of pp 10 of the response. Further, to applicants arguments toward the features of fast/slower time scale processing of the audio artifacts, examiner further notes the recitation to the Narayanan et al (20220122622) reference teaching, the parallel encoding sections, operating on low-latency and high-latency artifacts – ie, more processing on audio that does not require fast processing (increased accuracy) and less processing that requires fast turnaround times. Further, to applicants embodiment of altering stride length within the model to change the capability of handling faster/slower changing signals, see Park et al (20230123826) teaching 2D convolutional networks varying the stride length (para 0240, 02226, 122, and para366 – 368). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see the following references that teach certain features of applicants disclosure: Park et al (20230123828) teaching 2D convolutional networks varying the stride length (para 0240, 02226, 122, and para366 – 368). Zeghidour et al (20230186927) teaches feature vector codebooks in encoding audio/speech using adversarial loss and neural networks (fig. 3) Jang et al (20230267940) teaches neural network based end-to-end speech/audio compression (para 0007) Lecomte et al (20240005935) teaches variable bit rate encoding using concealment features (para 0026, 0035, 0039) Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Michael N Opsasnick/Primary Examiner, Art Unit 2658 04/22/2026
Read full office action

Prosecution Timeline

Dec 14, 2023
Application Filed
Sep 24, 2025
Non-Final Rejection mailed — §103
Nov 12, 2025
Examiner Interview Summary
Nov 12, 2025
Applicant Interview (Telephonic)
Dec 17, 2025
Response Filed
Apr 24, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12676164
System and Method for Modulation Domain-Based Audio Signal Encoding
3y 0m to grant Granted Jul 07, 2026
Patent 12658172
COMPUTING SYSTEM FOR UNSUPERVISED EMOTIONAL TEXT TO SPEECH TRAINING
4y 3m to grant Granted Jun 16, 2026
Patent 12651607
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND PROGRAM
2y 2m to grant Granted Jun 09, 2026
Patent 12619603
GENERATING A DISTILLED GENERATIVE RESPONSE ENGINE TRAINED ON DISTILLATION DATA GENERATED WITH A LANGUAGE MODEL PROGRAM
1y 11m to grant Granted May 05, 2026
Patent 12609117
APPARATUS AND METHOD FOR SPEECH RECOGNITION
2y 4m to grant Granted Apr 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
92%
With Interview (+10.1%)
3y 2m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 916 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month