DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 7/30/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-6, 9-17, and 20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Seetharaman et al. US 20260105906 A1 (hereinafter Seetharaman).
Regarding independent claim 1, 12, and 20, Seetharaman teaches a method performed by at least one processor, the method comprising / an apparatus comprising / a non-transitory computer readable medium having instructions stored therein, which when executed by a processor cause the processor to execute a method comprising:
at least one memory configured to store program code (FIG. 7, 715); and
at least one processor configured to read the program code and operate as instructed by the program code (FIG. 7, 705), the program code including
receiving an input audio waveform (FIG. 1, [0031] “the input prompt includes an input audio clip”);
inputting the input audio waveform into a variational autoencoder (VAE) to generate a latent representation of the input audio waveform (FIG. 2, 210; FIG. 7, 725, [0041] “the audio encoder includes a variational autoencoder (VAE) trained to generate latent input representation representing the sound based on the input prompt”);
generating a modified latent representation of the input audio waveform by inputting the latent representation of the input audio waveform and a text prompt into a latent diffusion model (FIG. 9, 945, 930, 925; [0107] “diffusion transformer 900 receives latent input 905 and noise input 910 to generate predicted latent 945”, examiner interprets 930 as the diffusion model and 945 as the modified latent representation; [0108] “the guidance 925 includes a text embedding of a text prompt”, examiner interprets 925 as the text prompt.); and
generating a modified audio waveform by inputting the modified latent representation of the input audio waveform into a VAE decoder (FIG. 8, 850, 845; [0098] “VAE includes… a decoder (e.g., the audio decoder 845) that reconstructs the original data from samples drawn from this latent space”).
Regarding claims 2 and 13, Seetharaman teaches all of the limitations of claim 1 and 12, upon which claims 2 and 13 depend.
Additionally, Seetharaman teaches wherein the text prompt includes a text description regarding the modified audio waveform ([0108] “a text prompt describing a sound may be provided to a text encoder of the system to generate the text embedding to guide the audio generation process within the diffusion transformer 900”).
Regarding claims 3 and 14, Seetharaman teaches all of the limitations of claim 2 and 13, upon which claims 3 and 14 depend.
Additionally, Seetharaman teaches inputting the text prompt into a text encoder to generate a feature embedding ([0108] “the guidance 925 includes a text embedding of a text prompt”),
wherein the latent diffusion model generates the modified latent representation of the input audio waveform based on the feature embedding (FIG. 9, 930, 945, 925).
Regarding claims 4 and 15, Seetharaman teaches all of the limitations of claim 1 and 12, upon which claims 4 and 15 depend.
Additionally, Seetharaman teaches wherein the generating the modified latent representation of input audio waveform further comprises: performing a diffusion process on the latent representation of the input audio to generate a noisy latent representation of the input audio waveform (FIG. 9, 905, 915, 900;).
Regarding claims 5 and 16, Seetharaman teaches all of the limitations of claim 4 and 15, upon which claims 5 and 16 depend.
Additionally, Seetharaman teaches wherein the generating the modified latent representation of input audio waveform further comprises: inputting the noisy latent representation of the input audio waveform into a first diffusion transformer (FIG. 9, 915, 930;); and
inputting an output of the first diffusion transformer into a second diffusion transformer to generate the modified latent representation of the input audio waveform ([0107] “the next intermediate layer is provided to a second transformer block including a self-attention layer and a cross-attention layer to generate the predicted latent 945”).
Regarding claims 6 and 17, Seetharaman teaches all of the limitations of claim 5 and 16, upon which claims 6 and 17 depend.
Additionally, Seetharaman teaches wherein the first diffusion transformer is a first neural network, and the second diffusion transformer is a second neural network ([0107] “the self-attention layer 935 receives the noised latent 915 and generates an intermediate feature and the intermediate feature is passed to the next neural network layer (e.g., a cross-attention layer 940) to generate the next intermediate layer”).
Regarding claims 9, Seetharaman teaches all of the limitations of claim 1, upon which claims 9 depends.
Additionally, Seetharaman teaches wherein the latent diffusion model is trained based on one or more tokens masked with diffusion noise, wherein the latent diffusion model is trained to reconstruct the one or tokens masked with noise using one or more corresponding tokens that are not masked with noise ([0123] “the audio generation model is augmented with mask tokens, where the mask tokens indicates whether a token represents audio prompt or represents an extension to be generated”; [0135] “The layers, for instance, are configurable to include an input layer, an output layer, and one or more hidden layers. Calculations are performed by the nodes within the layers through the hidden states through a system of weighted connections that are “learned” during training, e.g., through the use of the selected loss function and backpropagation to optimize the performance of the machine-learning model to perform an associated task”).
Regarding claims 10, Seetharaman teaches all of the limitations of claim 1, upon which claims 10 depends.
Additionally, Seetharaman teaches wherein the latent diffusion model is trained based on annotated data set comprising synthetic captions to facilitate text-to-audio alignment learning ([0022] “The text embedding is used to condition the generated audio on the meaning of text, producing music that aligns with the given prompt”).
Regarding claims 11, Seetharaman teaches all of the limitations of claim 1, upon which claims 11 depends.
Additionally, Seetharaman teaches wherein a classifier-free guidance (CFG) is used during diffusion sampling for alignment of generated audio ([0101] “a variant of classifier-free guidance can be used to improve the system performance at test-time”).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 7 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Seetharaman in view of Entezari et al. US 12271978 B1 (hereinafter Entezari).
Regarding claim 7 and 18, Seetharaman teaches all of the limitations of claim 5 and 16, upon which claims 7 and 18 depend.
Seetharaman fails to teach wherein the generating the modified latent representation of the input audio waveform further comprises: applying an adaptive layer normalization with low-rank adjustment to each diffusion transformer.
However, Entezari teaches wherein the generating the modified latent representation of the input audio waveform further comprises: applying an adaptive layer normalization with low-rank adjustment to each diffusion transformer ([Column 8, line 67- Column 9, line 2] “An adaptive layer normalization (adaLN) can be used to condition the diffusion network on text representations, enabling parameter-efficient adaptation”; [Column 18, line 18-21] “Any number of transformer blocks may be chained together within the diffusion transformer model. The transformer blocks may include a linear layer trained using learnable low-rank (LoRA) matrices”).
Seetharamanin view of Entezari are considered to be analogous to the claimed invention because both are the same field of content generation through prompts. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the techniques of generating a synthetic audio clip including the sound based on the latent sound representation of Seetharaman with the technique of layer normalization taught by Entezari in order to synthesize content using a generative AI model (see Entezari [Abstract]).
Claims 8 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Seetharaman in view of Serra et al. US 20240395278 A1.
Regarding claim 8 and 19, Seetharaman teaches all of the limitations of claim 5 and 16, upon which claims 8 and 19 depend.
Seetharaman fails to teach wherein the generating the modified latent representation of the input audio waveform further comprises: applying one or more long-skip connections in which one or more low level features are provided from a first transformer block from the plurality of transformer blocks to a second transformer block from the plurality of transformer blocks, wherein the first transformer block and the second transformer block are non-consecutive block.
However, Serra teaches wherein the generating the modified latent representation of the input audio waveform further comprises: applying one or more long-skip connections in which one or more low level features are provided from a first transformer block from the plurality of transformer blocks to a second transformer block from the plurality of transformer blocks ([0074] “The architecture of the conditioning network 120 may be based on an encoder-decoder structure, for example using ResNets, optionally with skip connections in the encoder”), wherein the first transformer block and the second transformer block are non-consecutive block ([0059] “The generator network 110 is a neural network for generating the enhanced audio signal (clean audio signal) x, 30. It comprises a plurality of layers (e.g., convolutional layers, transformer layers, recurrent neural network layers)”).
Seetharamanin view of Serra are considered to be analogous to the claimed invention because both are the same field of speech enhancement. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the techniques of generating a synthetic audio clip including the sound based on the latent sound representation of Seetharaman with the technique of applying long-skip connections taught by Serra in order to remove a variety of artifacts from noisy audio signals containing speech, in addition to denoising the audio signals (see Serra [0002]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chemerys et al. (US 20240395028 A1) teaches a system for improving machine learning models. In some cases, the system improves such models by identifying an autoencoder for a latent diffusion machine learning model, the latent diffusion machine learning model is trained to receive text as input and output an image based on the received text. The system identifies a number of channels in a decoder of the autoencoder, the decoder being configured to receive latent features as input and output images. The system further identifies a performance characteristic of the decoder and changes the node topology of the decoder based on the performance characteristic to generate an updated decoder. The system retrains the latent diffusion machine learning model using the updated decoder by inputting latent features to the updated decoder, receiving an outputted image from the updated decoder, and updating one or more weights of the decoder based on an assessment of the outputted image.
Angkititrakul et al. (US 20250356121 A1) teaches a method for audio generation includes defining an audio input condition for an obtained input using an encoder, where the obtained input is indicative of one or more audio characteristics. The method further includes defining an audio style condition of a selected audio style profile employing an audio feature extraction neural network, and outputting a generated audio data indicative of a desired generated audio using a multi-conditioned latent diffusion model that employs the audio input condition and the audio style condition as adapters to the multi-conditioned latent diffusion model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZEESHAN SHAIKH whose telephone number is (703)756-1730. The examiner can normally be reached Monday-Friday 7:30AM-5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZEESHAN MAHMOOD SHAIKH/Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658