Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is sent in response to Applicant’s communication received on 03/07/2025 for the application number 19199782. The office hereby acknowledges receipt of the following placed of record in the file: Specification, Abstract, Oath/Declaration and claims.
Status of the claims
Claims 1-20 are presented for examination.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed
to an abstract idea without significantly more. The claim(s) does/do not include additional
elements that are sufficient to amount to significantly more than the judicial exception because as
explained below.
Claim 1 recites a method comprising:
receiving first audio having a first accuracy in a three-dimensional environment and a first number of channels;
And generating second audio based on the first audio, the second audio having a second accuracy in the three-dimensional environment and a second number of channels,
wherein the first accuracy is greater than the second accuracy and the first number of channels is greater than the second number of channels.
Step (a) comprises data gathering. This step merely receives audio information having
particular characteristics.
Step (b) comprises a mental process. This step can be performed by a human as a [person can evaluate the received audio information and determine a modified representation having reduced spatial accuracy and fewer channels.
Step (c) comprises a mental process. This step can be performed by a human as a person can compare the characteristics of the first and second audio representations and determine that the first has greater spatial accuracy and a greater number of channels.
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any
statutory category. See MPEP 2106.03. The claim recites at least method. Thus, the claim is a
process, which is one of the statutory categories of invention. (Step 1: YES).
Step 2A, Prong One: This part of the eligibility analysis evaluates whether the
claim recites a judicial exception. As explained in MPEP 2106.04, subsection II, a claim
“recites” a judicial exception when the judicial exception is “set forth” or “described” in
the claim. As discussed above, the broadest reasonable interpretation of steps (b)-(c)
recites a mental process.
Specifically, step (b) can be performed by a human as a [person can evaluate the received audio information and determine a modified representation having reduced spatial accuracy and fewer channels.
Step (c) can be performed by a human as a person can compare the characteristics of the first and second audio representations and determine that the first has greater spatial accuracy and a greater number of channels.
Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a
whole integrates the recited judicial exception into a practical application of the exception or
whether the claim is “directed to” the judicial exception. This evaluation is performed by (1)
identifying whether there are any additional elements recited in the claim beyond the judicial
exception, and (2) evaluating those additional elements individually and in combination to
determine whether the claim as a whole integrates the exception into a practical application. See
MPEP 2106.04(d).
The claim recites additional elements including first audio, a second audio, and a three dimensional environment. However, these elements are recited at a high level of generality and perform generic computer functions, such as executing instructions and storing data. The use of these elements to receive and generate audio having specified spatial accuracies and numbers of channels constitutes a generic computational technique that merely automates the mental processes described above. Such processing does not impose any meaningful limit on the judicial exception. Accordingly, these additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea. (Step 2A, Prong Two: NO), and the claim is directed to the judicial exception. (Step 2A: YES).
Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole
amounts to significantly more than the recited exception i.e., whether any additional element, or
combination of additional elements, adds an inventive concept to the claim. As explained with
respect to Step 2A, Prong Two, first audio, a second audio, and a three dimensional environment comprise additional elements that do not contribute to the patentability of the claim as a whole. They constitute at best mere instructions to “apply” the abstract ideas, which cannot provide an inventive concept. See MPEP 2106.05(f). At Step 2B, the evaluation of the insignificant extra solutional activity consideration takes into account whether or not the extra-solutional activity is well understood, routine, and conventional in the field. See MPEP 2106.05(g). As known in the art these elements are well routine and conventional. Even when considered in combination, these additional elements represent mere instructions to implement an abstract idea or other exception on a computer and insignificant extra-solutional activity which do not provide an inventive concept. The claim is not patent eligible.
Claim 2 recites a mental process as a human can associate the spatial accuracy of a sound field with the order of the sound field.
Claim 3 recites insignificant extra solutional activity as spherical harmonic coefficients merely specify the format of the information being received.
Claim 4 recites a mental process as a human can model relationships between lower-order and higher-order audio information and use the modeled relationship to generate modified audio.
Claim 5 recites a mental process as a human can consider user feedback and use the feedback to modify a previously learned relationship.
Claim 6 recites insignificant extra solutional activity as playing back the resulting audio merely outputs the result of the recited process.
Claim 7 recites generic implementation of the abstract idea using an encoder and decoder to process the audio information.
Claim 8 merely limits the audio information to ambisonics input audio and binaural output audio and does not alter the underlying mental process.
Claims 9 and 16 are analogous to claim 1 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claim 10 is analogous to claim 3 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Claims 11 and 17 are analogous to claim 4 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claims 12 and 18 are analogous to claim 4 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claims 13 and 20 are analogous to claim 6 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claim 14 is analogous to claim 7 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Claim 15 is analogous to claim 8 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Claim 19 is analogous to claim 2 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-4, 6, 8-11, 13, 15-17, 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Morrell et al. (US 9420393 B2) in view of Nawfal et al. (US 20240312468 A1).
Regarding claim 1, Morrell teaches receiving first audio having a first accuracy in a three-dimensional environment and a first number of channels (Col. 9, Ln. 1-4, “the audio playback system 32 may receive the spherical harmonic coefficients” where each order/sub-order combination “effectively define[s] a separate SHC channel” thereby teaching the audio having a first number of channels and wherein the spatial accuracy is inherent to how closely the audio represents the respective sound field); and generating second audio based on the first audio, the second audio having a second accuracy in the three-dimensional environment and a second number of channels (Col. 10, Ln. 10-14, “Binaural audio renderer 34 may then convert to binaural audio by summing the filtered arrays to obtain the binaural output signals” wherein binaural audio includes two channels, specifically a left and right and wherein the spatial accuracy is inherent to how closely the audio represents the respective sound field), wherein the first number of channels is greater than the second number of channels (Col. 9, Ln. 39-43, teaching that each combination of the SHCs “effectively define[s] a separate SHC channel”; Morrell further teaches that each SHC representation has dimensions [Length, (N+1)^2] which is a minimum of 4 channels if N = 1. The second audio is a binaural output and necessarily includes two channels, a left and a right). And wherein the first audio is based on a higher order than the second audio (Col. 25, Ln. 57 – Col. 26, Ln. 7 teaches that “order reduction unit 504 … processes inbound SHCs 422 to reduce an order or sub-order of the SHCs 422 to generate order reduced SHCs 502” and “the order-reduced SHC 502” is provided “to the binaural rendering unit 402”. Thus, the second audio is generated from an order reduced version of the first signal).
Morrell does not teach wherein the first accuracy is greater than the second accuracy.
However, Nawfal teaches wherein the first accuracy is greater than the second accuracy (Para 0016, “Each increasing order … adds spatial resolution when played back to a listener” and teaching that lower-order ambisonics “provides low spatial resolution” thereby teaching that a higher-order audio has greater resolution/accuracy as a lower order second audio).
It would have been obvious to one of ordinary skill in the art before the effective filing data to modify Morrell to incorporate the teachings of Nawfal in order to account for the spatial resolution associated with different orders of ambisonic audio (Para 0016).
Regarding claim 2, Morrell teaches wherein the first accuracy is at least one of a spatial-accuracy associated with a first sound field or a spatial-accuracy associated with an order of the first sound field (Col. 1, Ln. 22-26, “spherical harmonic coefficients representative of a sound field in three dimensions” wherein the spatial accuracy is inherent to how closely the audio represents the respective sound field), and the second accuracy is at least one of a spatial-accuracy associated with a second sound field or a spatial-accuracy associated with an order of the second sound field (Col. 25, Ln. 57 – Col. 26, Ln. 3, order reduction unit 504 “processes inbound SHCs 422 to reduce an order or sub-order of the SHCs 422 to generate order reduced SHCs 502” wherein the spatial accuracy is inherent to how closely the audio represents the respective sound field).
Regarding claim 3, Morrell teaches wherein the first audio includes spherical harmonics coefficients associated with the first number of channels (Col. 9, Ln. 39-43, teaching that each order/sub-order combination of the SHCs “effectively define[s] a separate SHC channel” thereby teaching spherical harmonic coefficients associated with the first number of channels).
Regarding claim 4, Morrell does not teach wherein the first audio is encoded audio, and the generating of the second audio includes decoding the first audio with a model configured to model complex non-linear relationships between low-order audio signals and high-order audio signals.
However, Nawfal teaches wherein the first audio is encoded audio (Para 0025, teaching that “a physical microphone array or a microphone array arranged virtually can be used to encode Ambisonics (FOA) audio” thereby teaching the first Ambisonics audio is encoded audio), and the generating of the second audio includes decoding the first audio with a model configured to model complex non-linear relationships between low-order audio signals and high-order audio signals (Para 0022 and 0027, “teaching that “low order HOA audio can also be upscaled to a higher order HOA format using a machina learning model” and that the ML model may comprise “an artificial neural network (ANN)” including a “convolutional neural network (CNN) or a deep convolutional neural network (DCNN)” thereby teaching a nonlinear machine learning model configured to model relationships between lower and higher order Ambisonics).
It would have been obvious to one of ordinary skill in the art before the effective filing data to modify Morrell to incorporate the teachings of Nawfal in order to improve the spatial resolution of the rendered audio (Para 0022).
Regarding claim 6, Morrell teaches further comprising playing back the second audio on a device including speakers configured to playback binaural audio (Col. 27, Ln. 52-59, teaching that binaural rendering unit 402 generates “a left channel 436A and a right channel 436B” that are “suitable for playback via a headset, such as headphones”).
Regarding claim 8, Morrell teaches wherein the first audio is ambisonics audio and the second audio is binaural audio (Col. 27, Ln. 38-43, teaching that SHCs 422 may be “referred to as higher order ambisonics (HOA)” and col. 27, Ln. 52-59, teaching that the binaural rendering unit renders the SHCs to generate left and right binaural output channels).
Claims 9 and 16 are analogous to claim 1 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claim 10 is analogous to claim 3 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Claims 11 and 17 are analogous to claim 4 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claims 13 and 20 are analogous to claim 6 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claim 15 is analogous to claim 8 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Claim 19 is analogous to claim 2 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Claims 5, 12, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Morrell (US 9420393 B2) in view of Nawfal (US 20240312468 A1) as above and further in view of Le et al. (US 20220366233 A1).
Regarding claim 5, Morrell does not teach wherein the model is a machine learning model trained to model the complex non-linear relationships between low-order audio signals and high-order audio signals.
However, Nawfal teaches wherein the model is a machine learning model trained to model the complex non-linear relationships between low-order audio signals and high-order audio signals (Para 0022 and 0027, “teaching that “low order HOA audio can also be upscaled to a higher order HOA format using a machina learning model” and that the ML model may comprise “an artificial neural network (ANN)” including a “convolutional neural network (CNN) or a deep convolutional neural network (DCNN)” thereby teaching a nonlinear machine learning model configured to model relationships between lower and higher order Ambisonics).
It would have been obvious to one of ordinary skill in the art before the effective filing data to modify Morrell to incorporate the teachings of Nawfal in order to improve the spatial resolution of the rendered audio (Para 0022).
Morrell modified by Nawfal does not teach and the model is a machine learning model trained in two training operations where a second training is based on user feedback.
However, Le teaches wherein the model is a machine learning model trained in two training operations where a second training is based on user feedback (Para 0059, teaching that “the first neural network is trained based on a first training regime and/or a second training regime” wherein “the first training regime comprises training an initial version of the first neural network” and “the second training regime comprises further training the initial version based on a second data set” where para 0062, teaching that “the second data may comprise feedback data on actual intents of the user” thereby teaching two training operations wherein the second training is based on user feedback).
It would have been obvious to one of ordinary skill in the art before the effective filing data to modify Morrell to incorporate the teachings of Le in order to adapt the trained model to changes in user behavior (Para 0062).
Claims 12 and 18 are analogous to claim 4 in that they recite substantially the same limitations. They are therefore rejected for the same reasons.
Claims 7, 14 are rejected under 35 U.S.C. 103 as being unpatentable over Morrell (US 9420393 B2) in view of Nawfal (US 20240312468 A1) as above and further in view of Zhu et al. ("Binaural Rendering of Ambisonic Signals by Neural Networks", arXiv:2211.02301 [cs.SD], Nov. 4, 2022, pp 1-6).
Regarding claim 7, Morrell does not teach wherein the generating of the second audio uses a transcoder including an encoder and a decoder, the encoder is a machine learning-based encoder, and the decoder is a machine learning-based decoder.
However, Zhu teaches wherein the generating of the second audio uses a transcoder including an encoder and a decoder (Pg. 2, Fig. 1(c), illustrating “ambisonic -> Neural network -> Binaural” and Pg. 3, 3.3, teaching that “The UNet consists of 6 encoder blocks and 6 decoder blocks as shown in Fig. 2”), the encoder is a machine learning-based encoder (Pg. 3, 3.3, teaching that the UNet is one of the disclosed “neural network architectures” and that “The encoder block EB (Cout, k, s) consists of a convolution block … followed by an average pooling” thereby teaching a machine-learning-based encoder), and the decoder is a machine learning-based decoder (Pg. 3, 3.3, teaching that “The decoder block DB(Cout, k, s) consists of a transpose convolution layer … followed by a convolution block” within the disclosed UNet neural network architecture, thereby teaching a machine learning based decoder).
It would have been obvious to one of ordinary skill in the art before the effective filing data to modify Morrell to incorporate the teachings of Zhu in order to improve the efficiency and performance of the audio rendering process (Pg. 3).
Claim 14 is analogous to claim 7 in that it recites substantially the same limitations. It is therefore rejected for the same reasons.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL ALAN FOSTER JR. whose telephone number is (571)272-8874. The examiner can normally be reached M - F 8:00am - 5:00pm, Alternate Fridays Off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL A FOSTER JR/ Examiner, Art Unit 2654 /Richa Sonifrank/ Primary Examiner, Art Unit 2654