Prosecution Insights
Last updated: August 15, 2026
Application No. 18/661,330

AUDIO SIGNAL PROCESSING METHOD AND APPARATUS

Final Rejection §101§103§112
Filed
May 10, 2024
Priority
May 12, 2023 — RE 10-2023-0061838 +1 more
Examiner
ZHANG, LESHUI
Art Unit
2695
Tech Center
2600 — Communications
Assignee
Gaudio Lab Inc.
OA Round
2 (Final)
78%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
740 granted / 950 resolved
+15.9% vs TC avg
Strong +35% interview lift
Without
With
+35.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
28 currently pending
Career history
984
Total Applications
across all art units

Statute-Specific Performance

§101
5.7%
-34.3% vs TC avg
§103
44.5%
+4.5% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
29.3%
-10.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 950 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Office Action is in response to the claim amendment filed on March 6, 2026 and wherein claims 1, 5, 11, 15 amended and claims 6, 16 cancelled. In virtue of this communication, claims 1-5, 7-15, 17-20 are currently pending in this Office Action. With respect to the rejection to claims 1-20 under 35 USC §112(b), as set forth in the previous Office Action, the claim amendment, including the cancelation of claims 6, 16 and argument, see paragraph 2 of page 13 and paragraph 1 of page 14 in Remarks filed on March 6, 2026, have been fully considered and the argument is persuasive. Therefore, the rejection to claims 1-20 under 35 USC § 112(b), as set forth in the previous Office Action, has been withdrawn. In the response to this office action, the Examiner respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Examiner in prosecuting this application. Claim Objections Claims 11-15, 17-20 are objected to because of the following informalities: Claim 11 recited “is cocatenated into …”, which should be -- is concatenated Claims 12-15, 17-20 are objected due to the dependencies to claim 11. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (B) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 11-15, 17-20 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention. Claim 11 recited “the audio signal processing device is configured to, …, generate the audio signal by fixing …” which is confusing because it is unclear whether “the audio signal” herein is referred back to “an audio signal” of the limitation “input the input modal to which the acquired label is concatenated into a generative model to generate an audio signal” at line 8 of claim 11, or “an audio signal” in limitation of “generates an audio signal after the training” at line 12 of claim 11 and thus, renders claim indefinite. Claims 12-15, 17-20 are rejected due to the dependencies to claim 11. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-5, 7-15, 17-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites an apparatus comprising a processor as generic components to implement all claimed operations that comprise generating an audio signal based on an input modal concatenated with acquired label and the label indicating feature or quality reference information, and the generative model is trained by using a predetermined quality and a quality that is different from the predetermined quality (e.g., quality threshold or minimum requirement of the quality), and wherein the claimed limitation of “acquiring”, “concatenating”, “inputting”, and “generating”, etc., as drafted, would be interpreted, under their BRI, to cover performance of mathematical concept and mathematical manipulation of math symbols or variables such as “label” indicating “feature”, “input modal” concatenated with “label” as a “generative model to generate an audio signal” (as a math formula or algorithm), and claimed “quality reference information” and “predetermined quality” are broad with no recitation of object(s) or no recitation of what they are and/or referred to, and thus, the claimed “generating the audio signal” is merely performed from manipulating the math formula or algorithm by inputting various math symbols as variables under their BRI. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, claim recited additional component “processor”. However, the claimed “processor” is recited in a high-level of generality to perform the abstract idea, which does not count as precluding the operations from practically being performed mathematically discussed above, i.e., a generic hardware to perform the abstract idea such that if it amounts no more than mere instructions to apply the exception using the generic computer component. Claim further recited “the generative model is trained using not only the predetermined quality, but also quality reference information indicating a quality other than the predetermined quality” which would be interpreted as an intention to indicate feature of the “trained generative model” that is trained by using two different quality indications and accordingly, these additional elements do not count as integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The additional components such as the generic “processor” above is recited in a high-level and the intention of trained “generative model” above do not count as sufficient amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application and intention of the trained “model”, there is no additional component in claim that can be considered to be significant more than abstract idea. Therefore, the claim is not patent eligible, see 2019 Revised Patent Subject Matter Eligibility Guidance, “2019 PEG”. Claim 11 recited a training method by essentially reciting the operations of claim 1 and thus, rejected for the at least similar reasons described in claim 1 above. Claim 2 depends on claim 1 and further recited additional component “label has a text form”, which does not count as an integration of the abstract idea into a practical application because representing math variables by letter and text is common, which does not impose any meaningful limits on practicing the abstract idea and such additional component “label has a text form” are insufficient to amount to significantly more than the judicial exception. Accordingly, the claim does not rectify the 101 issue of parent claim 1 and is not patent eligible. Claim 3 depends on claim 1 and further added additional components by reciting “multiple input modals” which merely modified a size of the claimed “input modal” and does not count as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea and such additional component is insufficient to amount to significantly more than the judicial exception. Accordingly, the claim does not rectify the 101 issue of parent claim 1 and is not patent eligible. Claim 4 depends on claim 1 and further recited additional component by reciting “label is text indicating …” and thus, as discussed in claim 2 above, representing mathematic variable by using letter or text is common and accordingly, it does not rectify the 101 issue of parent claim 1 and is not patent eligible. Claim 5 depends on claim 1 and further recited the limitation of the “audio signal” to be generated as “background sound” by indicating that “the feature” “comprises at least one among recording environment information …” and such additional component would not be considered as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea, and is insufficient to amount to significantly more than the judicial exception because the additional component is merely as naming the “audio signal” to be “background sound”. Accordingly, the claim does not rectify the 101 issue of parent claim 1 and is not patent eligible. Claim 7 depends on claim 1 and further recited “input modal comprises at least one text indicating a category, text”, “text” inferred from (an image or a video) which is merely a definition of the “input modal”, which does not count as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea, and is insufficient to amount to significantly more than the judicial exception, and thus, claim 7 does not rectify the 101 issue of parent claim 1 and is not patent eligible. Claim 8 depends on claim 1 and essentially further recites that “audio signal” is generated by “synthesizing an audio signal from” “a audio frequency characteristic” that is generated by “a feature vector” and the “feature vector” is generated from “acquired token”, etc., wherein the recited “acquire tokens from the input modal and the label”, “generate”, and “synthesize” are merely math manipulation under their BRI because they are recited in a high level, and these additional components do not account as a practical application because it does not impose any meaningful limits on practicing the abstract idea, and are insufficient to amount to significantly more than the judicial exception, and thus, claim 8 does not rectify the 101 issue of parent claim 1 and is not patent eligible. Claim 9 depends on claim 8 and further recites that “a first audio frequency characteristic” and “a second audio frequency characteristic having a higher resolution … than the first audio frequency characteristic” which merely defined “audio signal” to be generated by using different “resolution” in a high level, and this additional component does not account as a practical application because it does not impose any meaningful limits on practicing the abstract idea, and are insufficient to amount to significantly more than the judicial exception, and thus, claim 9 does not rectify the 101 issue of parent claim 8 and is not patent eligible. Claim 10 depends on claim 8 and further recites that internal feature of the “generative model” by reciting “skip connections and FiLM … when “an audio frequency characteristic” is generated “from the feature vector”, etc., in a high level, and this additional component does not account as a practical application because it does not impose any meaningful limits on practicing the abstract idea, and are insufficient to amount to significantly more than the judicial exception, and thus, claim 9 does not rectify the 101 issue of parent claim 8 and is not patent eligible. Claim 12 depends on claim 11 and rejected for the at least similar reason as described in claim 2 above because claim 12 recited similar deficient feature as recited in claim 2. Claim 13 depends on claim 12 and rejected for the at least similar reason as described in claim 3 above because claim 13 recited similar deficient features as recited in claim 3. Claim 14 depends on claim 11 and rejected for the at least similar reason as described in claim 4 above because claim 14 recited similar deficient feature as recited in claim 4. Claim 15 depends on claim 11 and rejected for the at least similar reason as described in claim 5 above because claim 15 recited similar deficient feature as recited in claim 5. Claim 17 depends on claim 11 and rejected for the at least similar reason as described in claim 7 above because claim 17 recited similar deficient feature as recited in claim 7. Claim 18 depends on claim 11 and rejected for the at least similar reason as described in claim 8 above because claim 18 recited similar deficient feature as recited in claim 8. Claim 19 depends on claim 18 and rejected for the at least similar reason as described in claim 9 above because claim 19 recited similar deficient feature as recited in claim 9. Claim 20 depends on claim 18 and rejected for the at least similar reason as described in claim 10 above because claim 20 recited similar deficient feature as recited in claim 10. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 1-4, 7, 11-14, 17 are rejected under 35 U.S.C. 103 as being unpatentable over Reber et al. US 20180336881 A1, hereinafter Reber) and in view of reference Koishida (above). Claim 1: Reber teaches an audio signal processing device (title and abstract, ln 1-18, a text-to-speech computing platform in fig. 1 and functionality, TTS system 200 in fig. 2) for generating an audio signal for an input modal (generating audible speech signal 204 for TTS conversion represented by at least one of parameters phoneme sequence 222, base frequency fo 224, and duration 226, as the input modal), the audio signal processing device comprising a processor (the processor 102 in fig. 1), wherein the processor is configured to: acquire a label (at least two of the phoneme sequence 222, the base frequency Fo 224 , and the duration 226, output from a front-end subsystem 220 of the TTS system 200 in fig. 2) indicating a feature of a group to which the input modal belongs (anyone of the phoneme sequence, the base frequency fo, and the duration belong to the other two, constructed as a whole into audible speech 204 from the text input 202 in fig. 4A, the input text, low frequency feature lower than sampling rate of audible stream, etc., and the group of the sequence, the base frequency fo, and the duration, with a fixed 16 kHz sampling rate, and being belong to the inputted text 202, para 30), wherein the feature of the group indicated by the label comprises quality reference information (QI, as quality reference information, derived from the phoneme sequence, the fo, and the duration through Err_acoustic_S(i) and Err_audible_F(j) in fig. 4B-4C, para 62-64, as an indication of quality of the audible speech from signal generation unit 300 in figs. 4B-4C, para 48, e.g., as total audible error signal energy represented as a quality of indicator QI 422, para 62); combining inputting the acquired label to the input modal (all three parameters of phoneme sequence, the fo, and the duration are placed into a vector, para 29 and through the upsampling unit 330, e.g., combined with a vector representative of (ph, f0, D, i), para 35); and input the input modal to which the acquired label is combined into a generative model (the vector inputted into back-end subsystem 230 as a generative model in fig. 4A-4C) to generate an audio signal (audible speech 204 is generated as output from the back-end subsystem 230 in figs. 4A-4C, para 35); and generate the audio signal by fixing the quality reference information to a predetermined quality (a non-zero quality threshold, para 64), wherein the generative model is trained (via back-end training system 400 in figs. 4B-4C) using not only the predetermined quality (the quality threshold to be compared as criteria of end of the training, para 64), but also quality reference information indicating a quality other than the predetermined quality (QI is updated iteratively by training the neural network of the signal generation unit 300 through the element 400 over time until the QI 442 is minimized and ideally converged within the quality threshold, para 64). However, Reber does not explicitly teach combining inputting the acquired label to the input modal is performed by concatenation. Koishida teaches an analogous field of endeavor by disclosing an audio signal processing device (title and abstract, ln 1-17, an speech enhancement system in fig. 1) for generating an audio signal (producing an enhanced waveform 128 in fig. 1, para 15) for an input modal (one of the mixed waveform 108 to a STFT 112/mixed magnitude 110 or video data frames from lip frames 118 in fig. 1), the audio signal processing device comprising a processor (one or more processors in the computing system 400 in fig. 4, para 45), wherein the processor is configured to: acquire a label (other one of Fvi representing features extracted from ResNet each layer, and PNG media_image1.png 23 130 media_image1.png Greyscale para 17 and video data lip frames 118 of video stream 102 in fig. 1) indicating a feature of a group (visual features of layers i with a dimension as the group, e.g., three dimensions covering the features of temporal, channel, frequency, para 22) to which the input modal belongs (through SE fusion block 106 and the audio and visual features are integrated, para 22); concatenate (channel-wise concatenation offered by squeeze-excitation SE networks, para 13 and contained in skip connections between the encoder and the decoder, para 36) the acquired label (via squeeze-excitation SE fusion blocks 107 of layers or networks, para 13) to the input modal (through STFT 112 to obtain mixed magnitude 110, then Log-Mel transformation to the encoder, etc., or through Log-Mel transformation to the encoder by taking the mixed magnitude 110 as the input modal in fig. 1, para 15, 36) for generating enhanced input the input modal (Mixed waveform processed by STFT 112 to iSTFT 132 via generated mixed spectrogram phase information 130 from the STFT 112 in fig. 1, or the mixed magnitude 110 as input modal inputted to an element-wise operation to generate enhanced magnitude 126 in fig. 1, para 15, 36) and the acquired label into a generative model (including Mask Prediction 124, etc., in fig. 1) to generate an audio signal (enhanced waveform 128 is produced through iSTFT 132 by taking the enhanced magnitude 126 and the mixed spectrogram phase information 130 from STFT 112, para 15) for benefits of obtaining a high quality of audio (by application of squeeze-excitation SE networks, abstract, and improving perceptual quality of generated sound, para 10) with improved performance implementation (para 14, 40). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the concatenation and wherein the combining of inputting the acquired label to the input modal is performed by concatenation, as taught by Koishida, to combining inputting the acquired label to the input modal as taught by Reber, for the benefits discussed above. Claim 11 has been analyzed and rejected according to claim 1 above and the combination of Reber and Koishida further teaches a training method (Reber, through a training system 400 in fig. 4B-4C, and Koishida, method steps in figs. 3A-3C) for an audio signal processing device (Reber, for generating audible speech 204 through trained neural network in figs. 4B-4C, para 47, 64, and Koishida, the system in fig. 1) for generating an audio signal for an input modal (Reber, the trained neural network, etc., are applied in back-end subsystem 230 for generating audible speech 204 in figs. 4B-4C, and Koishida, producing enhanced waveform 128 by taking mixed waveform as input in fig. 1), wherein the audio signal processing device is trained to: acquire a label indicating a feature of a group to which the input modal belongs (discussion in claim 1 above, related to feature “acquire”, Reber, and Koishida); concatenate the acquired label to the input modal (discussion in claim 1 above, related to feature “concatenate”, Koishida); and input the input modal and the acquired label into a generative model to generate an audio signal (discussion in claim 1 above, related to the feature “input the input modal”, Reber and Koishida). Claim 2: the combination of Reber and Koishida further teaches, according to claim 1 above, the label has a text form (Reber, input text 202 in fig. 1). Claim 3: the combination of Reber and Koishida further teaches, according to claim 1 above, wherein the processor is configured to, in case that the audio signal processing device receives multiple input modals (Koishida, the mixed audio waveform 108 and video frames 118 in fig. 1), independently concatenate a label corresponding to each of the multiple input modals to each of the multiple input modals (Koishida, through separated ResBlock in 120 for the video frames and Upsample/Conv2D 122 and ResBlock in encoder 116 and processed by STFT 112, Log-Mel 114, etc., for audio frames and by Spatio-Temporal 3D Conv for video, in fig. 1). Claim 4: Kim further teaches, according to claim 1 above, wherein the label is text indicating the feature of the group indicated by the label (Reber, the discussed in claim 1 above, and Koishida, mixed waveform 108 processed by STFT 112 and video frames 118 processed by spatio-temporal 3D Conv 120, audio frame and video frame are different data categories inherently). Claim 7: Kim further teaches, according to claim 1 above, wherein the input modal comprises at least one among text indicating a category, text, text inferred from an image, or text inferred from a video (Koishida, an image by providing lip frames 118 in video stream 102 in fig. 1, Markush applied, MPEP 2117). Claim 12 has been analyzed and rejected according to claims 11, 4 above. Claim 13 has been analyzed and rejected according to claims 12, 3 above. Claim 14 has been analyzed and rejected according to claims 11, 4 above. Claim 17 has been analyzed and rejected according to claims 11, 7 above. Claims 5, 15 are rejected under 35 U.S.C. 103 as being unpatentable over Reber (above) and in view of references Koishida (above) and Kim (US 20210304777 A1, hereinafter Kim7). Claim 5: the combination of Reber and Koishida teaches, according to claim 1 above, wherein the feature of the group indicated by the label comprises quality reference information (QI, as quality reference information, derived from the phoneme sequence, the fo, and the duration through Err_acoustic_S(i) and Err_audible_F(j) in fig. 4B-4C, para 62-64, as an indication of quality of the audible speech from signal generation unit 300 in figs. 4B-4C, para 48, e.g., as total audible error signal energy represented as a quality of indicator QI 422, para 62 and Koishida, sequence of feature data related to language as language input 1510, col 37, ln 12-2, features 386 of image from detectors, and contained in a descriptor, col 15, ln 40-47, or user’s posture or pose, col 15, ln 62-67, col 16, ln 1-8 or acoustic waveform 1320), except explicitly teaching wherein the feature of the group indicated by the label comprises at least one among recording environment information indicating a recording environment, or background sound information indicating whether the audio signal to be generated is background sound. Kim7 teaches an analogous field of endeavor by disclosing an audio signal processing device (title and abstract, ln 1-15 and a system figs. 2A-2E) for generating an audio signal (audio signals outputted from renderer 230 in figs. 2A-2C) for an input modal (through microphone array 205 with picked acoustic signals from different audio sources 211 in figs. 2A-2C) and wherein acquiring a label (constraints output from block 236 in fig. 2A) indicating a feature of a group to which the input modal belongs is disclosed (e.g., a spatial constraint and/or a target distribution function, for obtaining a clean signal audio estimate as desired, para 51), wherein the feature of the group indicated by the label comprises at least one among recording environment information indicating a recording environment (including reverberation associated with the target audio source, i.e., recording environment by microphone array in fig. 2A, para 79), or background sound information indicating whether the audio signal to be generated is background sound (the target can be engine noise, i.e., background noise to be generated with a given distribution function, para 51) for benefits of improving sound quality (by providing for a clean estimate of the original audio signal via removing interference sound effects, para 34; by more accurately and optimally generating local characteristics through convergence of conditionally training, para 56-57). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied wherein the feature of the group indicated by the label comprises at least one among the recording environment information indicating the recording environment, or the background sound information indicating whether the audio signal to be generated is the background sound, as taught by Kim7, to the feature of the group indicated by the label in the audio signal processing device, as taught by the combination of Reber and Koishida, for the benefits discussed above. Claim 15 has been analyzed and rejected according to claims 11, 5 above. Claim 8, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Reber (above) and in view of references Koishida (above) and Kim et al. (US 11501794 B1, hereinafter Kim). Claim 8: the combination of Reber and Koishida further teaches, according to claim 1 above, generate an audio frequency characteristic from the feature vector (Reber, including the fo from the input vectors via analysis and conversion of text input 202, para 29 and Koishida, magnitude and phase 130 from spectrogram generated by STFT 112, para 15 or generating fused feature vectors by concatenating the audio feature representations and the visual feature representaitons outputs from ResBlock of encoder 116 via skip link or output from Upsample/Conv2D 122 based on the mixed waveform input in fig. 1 and output from ResBlock 18 in visual network 120 in fig. 1, para 37-38); and synthesize an audio signal from the generated audio frequency characteristic (Reber, via neural network 310a and pre-existing knowledge base 320 in figs. 4B-4C para 51, and Koishida, through iSTFT 132, applied to enhanced magnitude 126 through the encoder-decoder in fig. 1). However, the combination of Reber and Koishida does not explicitly teach acquire tokens from the input modal and the label and generate a feature vector from the acquired tokens. Kim teaches analogous field of endeavor by disclosing an audio signal processing device (title and abstract, ln 1-16, an autonomously motile device 110 for sentiment detection in fig. 1 and details in fig. 9) for generating an audio signal (generating audio data, col 3, ln 8-13, via one or more loudspeakers 320 in fig. 3A), the audio signal processing device comprising a processor (controllers/processors 2304 in fig. 23A), and wherein the processor is to perform acquiring tokens from an input modal and a label (tokens as lexical data representing audio and text from audio input and text input in fig. 13, col 8, ln 2-11, ln 42-52 and output from LSTM1-n, as tokens, to be concatenated for concatenated vector data 1340 in fig. 13); generating a feature vector from the acquired tokens (generating concatenated vector data 1340 via LSTM1..n corresponding to at least acoustic 1320 and language 1330 in fig. 13); generating an audio frequency characteristic from the feature vector (parameters created for TTS, including frequency, volume, noise, etc., col 10, ln 16-23 and based on generated sentiment data 1360 in fig. 13); and synthesize an audio signal from the generated audio frequency characteristic (the parametric synthesis to perform TTS by varying parameters such as frequency, volume, and noise to create audio data including artificial speech waveform, col 10, ln 16-23, based on confidential user recognition data 695 as output, col 21, ln 64-67, col 24, ln 1-15 and sentiment data, col 43, ln 11-24) for benefits of improving user’s experiences (improving an interaction with the user by evaluating user’s current emotion, and via currently detected user’s emotion or sentiment detection, col 3, ln 1-7 and col 4, ln 22-33). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied tokens that acquired from the input modal and the label and generated a feature vector from the acquired tokens, as taught by Kim, to generating the audio frequency characteristic from the feature vector in the audio signal processing device for generating an audio signal, as taught by a combination of Reber and Koishida, for the benefits discussed above. Claim 18 has been analyzed and rejected according to claims 11, 8 above. Claims 9, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Reber (above) and in view of references Koishida (above), Kim (above), and Disch et al. (US 20150287417 A1, hereinafter Disch). Claim 9: the combination of Reber, Koishida, and Kim teaches, according to claim 8 above, the generative modal (Reber, and Koishida and Kim, the discussion in claims 1, 8 above) is configured to generate audio frequency characteristic (Reber, and Koishida, via STFT 112 and/or Log-Mel to generate audio spectrogram, abstract, and Kim, generating concatenated vector data 1340 via LSTM1..n corresponding to at least acoustic 1320 and language 1330 in fig. 13), except explicitly teaching generating a first audio frequency characteristic and generating, as the audio frequency characteristic, a second audio signal frequency characteristic having a higher resolution in at least one of a time axis or a frequency axis than the first audio frequency characteristic. Disch teaches an analogous field of endeavor by disclosing audio signal processing device (title and abstract, ln 1-13 and an audio decoder in fig. 2A) and wherein a generative model is disclosed (including joint channel decoding 204, IGF and tonal mask 206, including in the audio decoder in fig. 2A) and wherein the generative model is configured to generate a first audio frequency characteristic (second spectral portions represented by a second encoded representation 109 in fig. 2A, para 78 and the second spectral portion having a second spectral resolution, para 80) and generating, as the audio frequency characteristic, a second audio signal frequency characteristic (a first set of first spectral portions having first spectral resolution, para 78) having a higher resolution in at least one of a time axis or a frequency axis than the first audio frequency characteristic (the second spectral resolution being lower than the first spectral resolution, para 80) for benefits of reducing computation power (by reducing computation complexity, para 306 and by reducing memory consumption, para 204). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied the first and the second audio frequency characteristics and wherein the second audio signal characteristic having the higher resolution in at least one of the frequency axis than the first audio frequency characteristic, as taught by Disch, to the audio frequency characteristic generated by the generative model in the audio signal processing device, as taught by the combination of Koishida and Kim, for the benefits discussed above. Claim 19 has been analyzed and rejected according to claims 18, 9 above. Claims 10, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Reber (above) and in view of references Koishida (above), Kim (above), and Steinmetz et al. (US 20230352058 A1, hereinafter Steinmetz). Claim 10: the combination of Reber, Koishida, and Kim further teaches, according to claim 8 above, wherein skip connections is used when the generative model generates an audio frequency characteristic from the feature vector (Beber, the trained neural network is applied in figs. 4B-4C, and Koishida, skip connection between the encoder and the decoder in fig. 1) and synthesizes an audio signal from the generated audio frequency characteristic Reber, via neural network 310a and pre-existing knowledge base 320 in figs. 4B-4C para 51, and Koishida, through the enhanced magnitude 126 through the transformed spectrogram, as the frequency characteristic, generated from Log-Mel 114 to the encoder 116 in fig. 1, abstract, and based on the mixed magnitude 110 outputted from STFT 112 in fig. 1, and synthesize through iSTFT from enhanced magnitude 126 via the decoder 122, etc., in fig. 1 and Kim, representation of audio quantity such as log filter-bank energies LFBEs, mel-frequency cepstral coefficients MFCCs, etc., col 58, ln 18-28 and parametric synthesizing to generate artificial speech waveform, col 10, 14-23, discussed in claim 8 above), and the neural network is used for the generative model (Reber, the trained neural network is further trained for compensating imperfections in figs. 4B-4C, para 82, and Koishida, skip-connected encoder-decoder in fig. 1, and Kim, neural networks, col 23, 43-67, col 24, ln 1-6), except explicitly teaching FiLM is used for the generative model to generate an audio frequency characteristic from feature vectors. Steinmetz teaches an analogous field of endeavor by disclosing an audio signal processing device (title and abstract, ln 1-11 and a system for performing multitrack mixing in fig. 5D) and wherein a generative model is disclosed (including transformation networks 5300, etc. in fig. 5D and comprising fully connected subnetworks FC block, TCNs, and output layer, para 84) and wherein skip connections and FiLM are used (skip connection from input of TCNk k=1, 2, …, 10, to an adder and feature-wise linear modulation FiLM in fig. 8A, para 99) when the generative model generates an audio frequency characteristic (outputted from FiLM generator by taking parameters and outputs characteristics to TCNk, including energy allocation cross the frequency spectrum, para 72, and output from each of adders in fig. 8A) from the feature vector (from the parameters used for processing the waveform as input in fig. 8A) and synthesizes an audio signal from the generated audio frequency characteristic (through the adder of each of stages including TCNk in fig. 8A) for benefits of improving performance of an audio generative device (through pre-training of controller network to adapted for scaling of different signal processing, para 67, and improving efficiency of the device by sharing weights in the transformation networks, para 82 and adapted for number of channels to be added instantly, para 110). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied wherein the skip connections and FiLM are used when the generative model generates the audio frequency characteristic from the feature vector and synthesizes the audio signal from the generated audio frequency characteristic, as taught by Steinmetz, to the application of the skip connections for the generative model to generate the audio frequency characteristic from the feature vector and synthesizes the audio signal in the audio signal processing device, as taught by the combination of Reber, Koishida, and Kim, for the benefits discussed above. Claim 20 has been analyzed and rejected according to claims 18, 10 above. Response to Arguments Applicant's arguments filed on March 6, 2026 have been fully considered and but are moot in view of the new ground(s) of rejection necessitated by the applicant amendment. The Office has thoroughly reviewed Applicants' arguments but firmly believes that the cited references to reasonably and properly meet the claimed limitations. In the response to this office action, the Office respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Office in prosecuting this application. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LESHUI ZHANG whose telephone number is (571)270-5589. The examiner can normally be reached Monday-Friday 6:30amp-4:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LESHUI ZHANG/ Primary Examiner, Art Unit 2695
Read full office action

Prosecution Timeline

May 10, 2024
Application Filed
Dec 10, 2025
Non-Final Rejection mailed — §101, §103, §112
Mar 06, 2026
Response Filed
May 14, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707227
METHOD FOR PROCESSING AUDIO INPUT DATA AND A DEVICE THEREOF
2y 6m to grant Granted Aug 11, 2026
Patent 12706102
SEPARATING SPATIAL AUDIO OBJECTS
2y 10m to grant Granted Aug 11, 2026
Patent 12694886
AREA REPRODUCTION SYSTEM AND AREA REPRODUCTION METHOD
2y 6m to grant Granted Jul 28, 2026
Patent 12682897
MOTOR VEHICLE AND METHOD FOR SUMMARIZING A CONVERSATION IN A MOTOR VEHICLE
2y 9m to grant Granted Jul 14, 2026
Patent 12676162
APPARATUS AND METHOD FOR MULTICHANNEL INTERFERENCE CANCELLATION
6y 8m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+35.2%)
2y 9m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 950 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month