Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 6, 9, 18, 21-24, 26 are rejected under 35 U.S.C. 103 as being unpatentable over Seymour et al, US Patent 10,418,024 in view of Lovelace et al, US PG Pub 2025/014692.
Regarding claim 1, Seymour et al teach a method comprising: obtaining an input prompt representing a sound (Figure 1A, steps 110, 120); and generating, using the audio generation model, a synthetic audio clip including the sound based on the latent sound representation (step 110, 140, column 4 line 20 – column 6 line 17). Seymour et al suggest that noise needs to be removed to increase the quality of the audio generation in column 5 lines 1-11, however Seymour et al fails to explicitly state generating, using an audio generation model, a latent sound representation by denoising a noise input based on the input prompt. Lovelace et al disclose generating, using an audio generation model, a latent sound representation by denoising a noise input based on the input prompt (figure 8, step 830). It would have been obvious to one of ordinary skill in the art at the time of filing to modify Seymour et al per the teachings of Lovelace et al for the purpose of improving a synthetic audio generation model.
Regarding claim 2, Seymour et al teach the method of claim 1, wherein: the input prompt comprises an input audio clip and the synthetic audio clip comprises an extension of the input audio clip (Figure 1A, the input audio clip is either the first- or second-person’s input audio and the synthetic audio clip is generated in step 150; see column 6 lines 17-50 which describes the synthetic audio clip as an extension of the input audio clip).
Regarding claim 3, Lovelace et al teach the method of claim 2, further comprising: encoding the input audio clip to obtain a latent input representation, wherein the latent sound representation is generated based on the latent input representation (steps 815 and 820).
Regarding claim 4, Seymour et al teach the method of claim 1, wherein: the sound comprises a background sound (the audio generation model incorporates background noise as described in column 5 lines 1-14).
Regarding claim 6, Seymour et al teach the method of claim 1, wherein: the input prompt comprises a text description of the sound (step 110 or step 120).
Regarding claim 9, Seymour et al teach the method of claim 1, wherein: the audio generation model is trained is using a training set including an input audio clip and a text description of the input audio clip (step 110 or step 120).
Regarding claim 18, the claim represents the apparatus version of claim 1, therefore it is rejected using the same rationale as that for claim 1.
Regarding claim 21, the claim represents the computer readable medium version of claim 1, therefore it is rejected using the same rationale as that for claim 1.
Regarding claim 22, the claim represents the computer readable medium version of claim 2, therefore it is rejected using the same rationale as that for claim 2.
Regarding claim 23, the claim represents the computer readable medium version of claim 3, therefore it is rejected using the same rationale as that for claim 3.
Regarding claim 24, the claim represents the computer readable medium version of claim 4, therefore it is rejected using the same rationale as that for claim 4.
Regarding claim 26, the claim represents the computer readable medium version of claim 6, therefore it is rejected using the same rationale as that for claim 6.
Claims 5 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Seymour et al in view of Lovelace et al as applied to claims 4 and 24 above, and further in view of King et al, US PGPub 2025/0372077.
Regarding claim 5, the combination of Seymour et al and Lovelace et al fails to teach the method of claim 4, wherein obtaining the input prompt comprises: obtaining a preliminary audio clip; and extracting the background sound from the preliminary audio clip to obtain the input prompt. King et al teach obtaining the input prompt comprises: obtaining a preliminary audio clip; and extracting the background sound from the preliminary audio clip to obtain the input prompt in paragraph 55. It would have been obvious to one of ordinary skill in the art at the time of filing to modify the combination of Seymour et al and Lovelace et al per the teachings of King et al for the purpose of improve the audio synthesis quality.
Regarding claim 25, the claim represents the computer readable medium version of claim 5, therefore it is rejected using the same rationale as that for claim 5.
Claims 7, 8, 27, and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Seymour et al in view of Lovelace et al as applied to claims 1 and 21 above, and further in view of Rios, US PGPub 2023/0018661.
Regarding claim 7, the combination of Seymour et al and Lovelace et al fails to teach the method of claim 1, wherein: the synthetic audio clip comprises a plurality of spatial sound channels. However, Rios discloses a synthetic audio clip comprises a plurality of spatial sound channels (paragraph 11). It would have been obvious to one of ordinary skill in the art to modify the combination of Seymour et al and Lovelace et al per the teachings of Rios for the purpose of creating a more pleasing synthetic sound.
Regarding claim 8, the combination of Seymour et al, Lovelace et al, and Rios teachs the method of claim 7, wherein: the plurality of spatial sound channels comprises at least one mono channel and at least one side channel since the stereo sound taught in paragraph 11 of Rios inherently comprises a mono and a side channel.
Regarding claim 27, the claim represents the computer readable medium version of claim 7, therefore it is rejected using the same rationale as that for claim 7.
Regarding claim 28, the claim represents the computer readable medium version of claim 8, therefore it is rejected using the same rationale as that for claim 8.
Regarding claim 7, the combination of Seymour et al and Lovelace et al fails to teach the method of claim 1, wherein: the synthetic audio clip comprises a plurality of spatial sound channels. However, Rios discloses a synthetic audio clip comprises a plurality of spatial sound channels (paragraph 11). It would have been obvious to one of ordinary skill in the art to modify the combination of Seymour et al and Lovelace et al per the teachings of Rios for the purpose of creating a more pleasing synthetic sound.
Claims 19 is rejected under 35 U.S.C. 103 as being unpatentable over Seymour et al in view of Lovelace et al as applied to claim 18 above, and further in view of Wang et al, US PGPub 2026/0105670.
Regarding claim 19, the combination of Seymour et al and Lovelace et al fails to teach the apparatus of claim 18, wherein: the audio generation model includes a diffusion transformer (DiT) model. Wang et al teach wherein: the audio generation model includes a diffusion transformer (DiT) model in paragraphs 34 and 59, among others. It would have been obvious to one of ordinary skill at the time of filing to modify the combination of Seymour et al and Lovelace et al to include the teachings of Wang et al for the purpose of improving the accuracy of the audio generation.
Claims 20 is rejected under 35 U.S.C. 103 as being unpatentable over Seymour et al in view of Lovelace et al in further view of Wang et al, as applied to claim 19 above, and further in view of Zhang et al, US PGPub 2026/10155131.
Regarding claim 20, the combination of Seymour et al, Lovelace et al, Wang et al fails to teach wherein: the audio generation model includes a variational autoencoder (VAE). Zhang et al teach wherein: the audio generation model includes a variational autoencoder (VAE) in the abstract and paragraphs 3, 5, and 6. It would have been obvious to one of ordinary skill at the time of filing to modify the combination of Seymour et al, Lovelace et al and Wang et al, to include the teachings of Zhang et al for the purpose of improving the accuracy of the audio generation.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Yassa et al, US Patent 10,311,855.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Brian T Pendleton whose telephone number is (571)272-7527. The examiner can normally be reached M-F 8:30AM - 5:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Colleen Fauz can be reached at (571) 272-1667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Brian T. Pendleton
Supervisory Patent Examiner
Art Unit 2425
/Brian T Pendleton/ Supervisory Patent Examiner, Art Unit 2425