DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
2. Claims 1-25 are pending and have been examined. Claims 1, 24, and 25 are independent.
This Application was published as U.S. 2025/0210052 A1.
Apparent priority: September 9, 2022.
Information Disclosure Statement
3. The information disclosure statements (IDS) submitted on March 9, 2025, June 3, 2025, October 7, 2025, October 22, 2025, and March 12, 2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Drawings
4. The drawings are objected to because:
a. FIGS. 4, 5, and 6 depict conventional technology described in the Background of the Invention at paragraphs [0007], [0010], and [0012], and are not designated by a legend such as "Prior Art."
b. Reference character 960 appears in FIG. 9 but is not mentioned anywhere in the description. See 37 CFR 1.84(p)(5).
c. Reference characters 800 and 900 are mentioned throughout the description, at paragraphs [0175] through [0192] and [0193] through [0201] respectively, but do not appear in the drawings. See 37 CFR 1.84(p)(5).
5. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as "amended." If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either "Replacement Sheet" or "New Sheet" pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
6. The disclosure is objected to because of the following informalities:
a. In paragraph [0001], "International Application No. No. PCT/EP2022/075151" should apparently read "International Application No. PCT/EP2022/075151."
b. In paragraph [0049] and [0191], "encoding the dowmixed signal" should apparently read "encoding the downmixed signal."
c. In paragraph [0099], "each transport channel of the two or more one transport channels" should apparently read "each transport channel of the two or more transport channels."
d. In paragraph [0104] and [0105], "may, e.g., be to configured to generate" should apparently read "may, e.g., be configured to generate."
e. In paragraph [0117], "may, e.g., be configured generate the control parameters" should apparently read "may, e.g., be configured to generate the control parameters."
f. In paragraph [0166], "the individual decision logic 722 detects voice activity does not detect voice activity in any of the audio input channels" should apparently read "the individual decision logic 722 does not detect voice activity in any of the audio input channels."
g. In paragraph [0177] and [0194], "being implemented a decision logic module 820" should apparently read "being implemented as a decision logic module 820."
h. In paragraph [0181], "along with a control parameters that control the spatialness" should apparently read "along with control parameters that control the spatialness."
i. In paragraph [0182], "no transmission of object indicates and power ratios may, e.g., take place" should apparently read "no transmission of object indices and power ratios may, e.g., take place."
j. In paragraph [0222], "in response to receiving seed 1, seed 2 and seed3" should apparently read "in response to receiving seed 1, seed 2 and seed 3."
k. The element numbered 802 is called a "direction information determiner" at paragraphs [0043] and [0106] and a "direction information extractor" at paragraph [0188]. One designation should apparently be adopted throughout.
Appropriate correction is required.
Claim Interpretation
7. The following is a quotation of 35 U.S.C. 112(f):
(f) ELEMENT IN CLAIM FOR A COMBINATION. An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
8. The following three-prong analysis is used to determine whether a claim limitation invokes 35 U.S.C. 112(f):
(A) the claim limitation uses the term "means" or "step" or a term used as a substitute for "means" that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term "means" or "step" or the generic placeholder is modified by functional language, typically, but not always linked by the transition word "for" (e.g., "means for") or another linking word or phrase, such as "configured to" or "so that"; and
(C) the term "means" or "step" or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Absence of the word "means" in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
9. This application includes one or more claim limitations that do not use the word "means," but are nonetheless being interpreted under 35 U.S.C. 112(f), because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
a. "a renderer for generating one or more audio output signals" in claim 1;
b. "a noise information determiner" and "a multi-channel generator" in claim 3;
c. "a random generator for generating random noise" in claim 4;
d. "a signal power computation unit" and "a direct power computation unit" in claim 18;
e. "a direct response computation unit" in claim 19;
f. "an input covariance matrix computation unit," "a target covariance matrix computation unit," and "a mixing matrix computation unit" in claim 20; and
g. "a transport signal generator," "a voice activity determiner," and "a bitstream generator" in claim 23.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
Corresponding structure for performing the claimed functions is described in the specification as follows. For the renderer of claim 1, renderer 220 of FIG. 2 as implemented by renderer 950 and units 951 through 957 of FIGS. 9 and 10, at paragraphs [0199] and [0203] through [0211]. For the noise information determiner and multi-channel generator of claim 3, silence insertion descriptor decoder 920 and mono to stereo converter 930 of FIG. 9, at paragraphs [0196] and [0197]. For the random generator of claim 4, Random Generator units 1, 2, and 3 of FIGS. 11 and 12, at paragraphs [0218] through [0222]. For the signal power computation unit and direct power computation unit of claim 18, units 951 and 952 of FIG. 10, at paragraphs [0204] and [0205]. For the direct response computation unit of claim 19, unit 953 of FIG. 10, at paragraph [0206]. For the input covariance matrix computation unit and the target covariance matrix computation unit of claim 20, units 954 and 955 of FIG. 10, at paragraphs [0207] and [0208]. For the mixing matrix computation unit of claim 20, unit 956 of FIG. 10, at paragraph [0209], performing the covariance synthesis described at paragraphs [0012], [0013] and [0210]. For the transport signal generator, voice activity determiner, and bitstream generator of claim 23, transport signal generator 110 of FIG. 1, decision logic module 720 of FIG. 7, and multiplexer 850 of FIG. 8, at paragraphs [0093] through [0095], [0163] through [0167], and [0192].
The "input interface" of claim 1 is not being interpreted under 35 U.S.C. 112(f), because the term "interface" conveys a recognized structural meaning to one of ordinary skill in this art.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f), applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f).
Claim Rejections - 35 USC § 102
10. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless -
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
11. Claims 1-8, 10, 12-14, and 21-25 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Eckert (U.S. 2023/0215445).
Regarding claim 1, Eckert discloses:
An audio decoder, comprising: (Eckert, paragraph [0028]: "The decoding unit 150 of FIG. 1 comprises a decoding module 160 which is configured to derive a reconstructed downmix signals 114 from the coded audio data 106. Furthermore, the decoding unit 150 comprises a metadata decoding module 161 which is configured to derive the SPAR metadata 105 from the coded metadata 107.");
an input interface for receiving a bitstream which depends on audio content comprising at least one of a plurality of audio objects and a plurality of audio channels (Eckert, paragraph [0022]: "In particular, the multi-channel input signal 101 may comprise (possibly a combination of) one or more mono signals, one or more stereo signals, one or more binaural signal, one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects, and/or one or more SR signals. The different signal components may be considered to be individual channels of the multi-channel input signal 101."; see also paragraph [0028] and FIG. 1.);
wherein a transport signal comprising two or more transport channels is encoded within the bitstream, and the audio content is encoded within the transport signal (Eckert, paragraph [0023]: "The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels."; paragraph [0091]: "For the two channel downmix example (4-2-4, for a first order soundfield), the comfort noise parameters for the mono dowmmix (W′) channel and for one prediction channel may be provided to the decoding unit 150."; paragraph [0098]: "For the case of a 4-3-4 downmix, mono codec CNG parameters may be generated and sent for the mono representation of the W downmix channel and for two prediction channels.");
or wherein information on a background noise is encoded within the bitstream instead of the transport signal, wherein the information on the background noise comprises information on a background noise of at least one of the two or more transport channels or information on a background noise of a derived signal which depends on at least one of the two or more transport channels; and (Eckert, paragraph [0056]: "wherein “AB” indicates an encoder bitstream for an active frame, wherein “SID” indicates a silence indicator frame, which comprises a series of bits for comfort noise generation, and wherein “ND” indicates no data frames"; paragraph [0058]: "the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) for a fraction of the inactive frames (i.e., for the SID frames)."; paragraph [0092]: "For the SID frames, two independent encoder module 140 instances encode spectral information regarding the mono (W′) channel and spectral information regarding the residual channel, respectively.");
a renderer for generating one or more audio output signals depending on the audio content being encoded with the bitstream (Eckert, paragraph [0029]: "the decoding unit 150 comprises a reconstruction module 170 which is configured to derive a reconstructed multi-channel signal 111 from the SPAR metadata 105 and from the reconstructed downmix signal 114. ... The reconstructed multi-channel signal 111 may be used for speaker rendering, for headphone rendering and/or for SR rendering.");
wherein, if the transport signal comprising the two or more transport channels is encoded within the bitstream, the renderer is configured to generate the one or more audio output signals depending on the two or more transport channels, and (Eckert, paragraph [0095]: "Since the downmix signals 103 are continuously running with the same downmix configuration in active and inactive frames, background noise typically sounds smooth even during transition frames."; paragraph [0030]: "A first mixer 211 may be configured to upmix the one or more channels of the reconstructed downmix signal 114 to an increased number of signals. The first mixer 211 depends on the SPAR metadata 105."; paragraph [0029]: "The reconstructed multi-channel signal 111 may be used for speaker rendering, for headphone rendering and/or for SR rendering.");
wherein, if the information on the background noise is encoded within the bitstream instead of the transport signal, the renderer is configured to generate the one or more audio output signals depending on the information on the background noise (Eckert, paragraph [0094]: "WCNG, PING comfort noise signals and the two decorrelated signals may then be upmixed to an FOA output using the SPAR metadata 105."; paragraph [0095]: "since the decoding unit 150 is using the prediction coefficients and the decorrelation coefficients computed by the SPAR encoder 120, spatial properties are replicated in the comfort noise which is generated by the SPAR decoder 150.").
Regarding claim 2, Eckert discloses:
2. An audio decoder according to claim 1, wherein, if the audio content exhibits voice activity, the transport signal comprising the two or more transport channels is encoded within the bitstream; and wherein, if the audio content does not exhibit voice activity, the information on the background noise is encoded within the bitstream instead of the transport signal (Eckert, paragraph [0057]: "the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames). Furthermore, the encoding unit 100 may be configured to determine and encode data for comfort noise generation within a “SID” frame"; paragraph [0058].).
Regarding claim 3, Eckert discloses:
3. An audio decoder according to claim 1, wherein the audio decoder comprises a noise information determiner and a multi-channel generator, wherein, if the information on the background noise is encoded within the bitstream, the noise information determiner is configured to determine the information on the background noise from the bitstream, the multi-channel generator is configured to generate the derived signal as an intermediate signal comprising two or more intermediate channels from the information on the background noise, and the renderer is configured to generate the one or more audio output signals depending on the two or more intermediate channels of the intermediate signal (Eckert, paragraph [0028]: "the decoding unit 150 comprises a metadata decoding module 161 which is configured to derive the SPAR metadata 105 from the coded metadata 107."; paragraph [0092]: "two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds.").
Regarding claim 4, Eckert discloses:
4. An audio decoder according to claim 3, wherein the multi-channel generator comprises a random generator for generating random noise, wherein the multi-channel generator is configured to generate the two or more intermediate channels depending on the random noise, being generated by the random generator (Eckert, paragraph [0063]: "The decoding unit 150 may be configured to generate random white noise as an excitation signal. The excitation signal may comprise multiple channels of white noise, wherein the white noise in the different channels is typically uncorrelated from one another.").
Regarding claim 5, Eckert discloses:
5. An audio decoder according to claim 4, wherein the multi-channel generator is configured to shape the random noise depending on the information on the background noise to acquire shaped noise, wherein the multi-channel generator is configured to generate the two or more intermediate channels from the shaped noise (Eckert, paragraph [0063]: "the decoding unit 150 may be configured to shape the random white noise within the different channels (spectrally and spatially) using the noise shaping parameters that have been provided within the bitstream."; paragraph [0092]: "The uncorrelated comfort noise signals may be spectrally shaped based on the representation of the W channel and the residual channel in the uncoded downmix, respectively.").
Regarding claim 6, Eckert discloses:
6. An audio decoder according to claim 4, wherein the multi-channel generator is configured to run the random generator at least twice with a different seed to acquire the random noise (Eckert, paragraph [0092]: "two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds."; paragraph [0050]: "The multi-channel excitation signal may be a multi-channel white noise signal where all channels are generated with different seed and are uncorrelated with each other.").
Regarding claim 7, Eckert discloses:
7. An audio decoder according to claim 4, wherein the multi-channel generator is configured to generate the two or more intermediate channels depending on the random noise and depending on control parameters being encoded within the bitstream, for example wherein the control parameters comprise, e.g., a scaling factor and/or, e.g., either a coherence or a correlation (Eckert, paragraph [0095]: "the decoding unit 150 is using the prediction coefficients and the decorrelation coefficients computed by the SPAR encoder 120"; paragraph [0094].).
Regarding claim 8, Eckert discloses:
8. An audio decoder according to claim 7, wherein at least one of the control parameters is encoded within the bitstream and comprises a plurality of parameter values for a plurality of subbands, and wherein the multi-channel generator is configured to generate each subband of a plurality of subbands of the two or more intermediate channels depending on a parameter value of the plurality of parameter values of the at least one of the control parameters being associated with said subband (Eckert, paragraph [0023]: "Downmixing may be performed in the subband domain or QMF domain (e.g., using 10 or more subbands)."; paragraph [0076]: "for an inactive frame, the covariance may be computed for a reduced number of bands compared to the case of an active frame (e.g., 6 bands instead of 12 bands).").
Regarding claim 10, Eckert discloses:
10. An audio decoder according to claim 4, wherein the multi-channel generator is configured to generate the two or more intermediate channels by generating a first random noise portion of the random noise using the random generator with a first seed, and by generating a first one of the two or more intermediate channels depending on the first random noise portion, by generating a second random noise portion of the random noise using the random generator with a second seed being different from the first seed, and by generating a second one of the two or more intermediate channels depending on the second random noise portion (Eckert, paragraph [0092]: "two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds."; paragraph [0050]: "all channels are generated with different seed and are uncorrelated with each other.").
Regarding claim 12, Eckert discloses:
12. An audio decoder according to claim 4, wherein the multi-channel generator is configured to generate the two or more intermediate channels by generating a first one of the two or more intermediate channels depending on the random noise, and by generating a second one of the two or more intermediate channels from the first one of the two or more intermediate channels (Eckert, paragraph [0094]: "The two decorrelated channels may be created by running WCNG through time domain or filterbank domain decorrelators.").
Regarding claim 13, Eckert discloses:
13. An audio decoder according to claim 12, wherein the multi-channel generator is configured to generate the second one of the two or more intermediate channels such that the second one of the two or more intermediate channels is identical to the first one of the two or more intermediate channels, or wherein the multi-channel generator is configured to generate the second one of the two or more intermediate channels by modifying the first one of the two or more intermediate channels (Eckert, paragraph [0094]: "The two decorrelated channels may be created by running WCNG through time domain or filterbank domain decorrelators or by generating uncorrelated comfort noise with different seed and by spectrally shaping the uncorrelated comfort noise channels as per WCNG.").
Regarding claim 14, Eckert discloses:
14. An audio decoder according to claim 1, wherein the renderer is configured to generate the two or more audio output signals as the one or more audio output signals (Eckert, paragraph [0029]: "The reconstructed multi-channel signal 111 may comprise the same types of channels as the multi-channel input signal 101."; paragraph [0022].).
Regarding claim 21, Eckert discloses:
21. An audio decoder according to claim 1, the renderer is configured to generate one or more of the two or more transport channels by applying Code-Excited Linear Prediction or by applying a Modified Discrete Cosine Transform or an inverse of the Modified Discrete Cosine Transform or by applying a combination of the Code-Excited Linear Prediction and of the Modified Discrete Cosine Transform (Eckert, paragraph [0025]: "the encoding unit 100 may comprise a coding module 140 which is configured to perform waveform encoding (e.g., EVS encoding) of the downmix signal 103, thereby providing coded audio data 106. Each channel of the downmix signal 103 may be encoded using a mono waveform encoder."; paragraph [0028].).
Regarding claim 22, Eckert discloses:
22. An audio decoder according to claim 1, wherein, if the audio content comprises the plurality of audio channels, but not the plurality of audio objects, a number of the two or more transport channels is smaller than a number of the plurality of audio channels, wherein, if the audio content comprises the plurality of audio objects, but not the plurality of audio channels, the number of the two or more transport channels is smaller than a number of the plurality of audio objects, wherein, if the audio content comprises both the plurality of audio objects and the plurality of audio channels, the number of the two or more transport channels is smaller than a sum of the number of the plurality audio channels and the number of the plurality of audio objects (Eckert, paragraph [0026]: "the number of channels of the downmix signal 103 may be dependent on the target bitrate."; paragraph [0091]: the 4-2-4 configuration reducing four channels to a two channel downmix; paragraph [0098]: the 4-3-4 configuration; paragraph [0022]: the multi-channel input signal 101 "may comprise ... one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects");
or wherein, if the audio content comprises the plurality of audio channels, but not the plurality of audio objects, a number of the two or more transport channels is smaller than or equal to a number of the plurality of audio channels, wherein, if the audio content comprises the plurality of audio objects, but not the plurality of audio channels, the number of the two or more transport channels is smaller than or equal to a number of the plurality of audio objects, wherein, if the audio content comprises both the plurality of audio objects and the plurality of audio channels, the number of the two or more transport channels is smaller than or equal to a sum of the number of the plurality audio channels and the number of the plurality of audio objects (Eckert, paragraph [0026]: "the number of channels of the downmix signal 103 may be dependent on the target bitrate."; paragraph [0091]: the 4-2-4 configuration reducing four channels to a two channel downmix; paragraph [0098]: the 4-3-4 configuration; paragraph [0022]: the multi-channel input signal 101 "may comprise ... one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects").
Regarding claim 23, Eckert discloses:
23. A system, comprising: an audio encoder, and an audio decoder according to claim 1 (Eckert, paragraph [0023]: "The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels."; paragraph [0028]: "The decoding unit 150 of FIG. 1 comprises a decoding module 160 which is configured to derive a reconstructed downmix signals 114 from the coded audio data 106."; see also FIG. 1.);
wherein the audio encoder comprises:
a transport signal generator for generating two or more transport channels of a transport signal from audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels (Eckert, paragraph [0023]: "The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels."; paragraph [0022]: the multi-channel input signal 101 "may comprise ... one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects");
a voice activity determiner for determining a voice activity decision for the transport signal, which indicates whether or not the audio input within the transport signal exhibits voice activity, and (Eckert, paragraph [0057]: "the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames)."; paragraph [0060]: "The encoding unit 100 may comprise a voice activity detector which is configured to switch the encoder to DTX mode");
a bitstream generator for generating a bitstream depending on the audio input (Eckert, paragraph [0025]: "The coded audio data 106 and the coded metadata 107 may be inserted into a bitstream.");
wherein, if the voice activity determiner has determined that the transport signal exhibits voice activity, the bitstream generator is adapted to encode the two or more transport channels within the bitstream (Eckert, paragraph [0058]: "the encoding unit 100 may be configured to send audio data 106 and encoded metadata 107 to the decoding unit 150 for every active frame.");
wherein, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, the bitstream generator is suitable to encode, instead of the two or more transport channels, information on a background noise, wherein the information on the background noise comprises information on a background noise of at least one of the two or more transport channels or information on a background noise of a derived signal which depends on at least one of the two or more transport channels (Eckert, paragraph [0058]: "the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) for a fraction of the inactive frames (i.e., for the SID frames)."; paragraph [0092].);
wherein the audio encoder is configured to generate a bitstream from audio input, and wherein the audio decoder is configured to generate one or more audio output signals from the bitstream (Eckert, paragraph [0025]: "The coded audio data 106 and the coded metadata 107 may be inserted into a bitstream."; paragraph [0029]: "The reconstructed multi-channel signal 111 may be used for speaker rendering, for headphone rendering and/or for SR rendering.").
Claim 24 is a method claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale.
Claim 25 is a computer program product claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally:
Regarding claim 25, Eckert discloses:
25. A non-transitory digital storage medium having a computer program stored thereon to perform the method for decoding, comprising: (Eckert, paragraph [0008]: "According to another aspect, a storage medium is described. The storage medium may comprise a software program adapted for execution on a processor and for performing the method steps outlined in the present document when carried out on the processor.");
when said computer program is run by a computer (Eckert, paragraph [0008]: "The storage medium may comprise a software program adapted for execution on a processor and for performing the method steps outlined in the present document when carried out on the processor.").
Claim Rejections - 35 USC § 103
12. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
13. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Eckert (U.S. 2023/0215445) in view of Oshikiri (U.S. 2013/0223633).
Regarding claim 9, Eckert discloses:
9. An audio decoder according to claim 7, wherein the control parameters are encoded within the bitstream, wherein the control parameters are single broadband control parameters (Eckert, paragraph [0095]: "the decoding unit 150 is using the prediction coefficients and the decorrelation coefficients computed by the SPAR encoder 120."; the prediction coefficients and the decorrelation coefficients that Eckert encodes within the bitstream and applies at the decoder read on the recited control parameters).
Eckert does not disclose the broadband character of the control parameters.
Oshikiri discloses:
the control parameters are single broadband control parameters (Oshikiri, paragraph [0051]: "Frame energy encoding section 301 determines the frame energy of the input L-channel signal and generates quantized L-channel signal frame energy information by performing scalar quantization (encoding) of the frame energy."; paragraph [0052]: the corresponding R-channel frame energy encoding section 302; paragraph [0067]: "Multiplexing section 312 multiplexes the quantized L-channel signal frame energy information, the quantized R-channel signal frame energy information ... to generate encoded stereo data."; paragraph [0071]: the decoder demultiplexes the same frame energy information.).
Eckert and Oshikiri pertain to discontinuous transmission and comfort noise generation for multi-channel audio. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the control parameters of Eckert to be single broadband parameters, as taught by Oshikiri. One of ordinary skill would have been motivated to make this modification in order to avoid the energy discontinuities that arise when a background noise spectrum is described by per-sub-band parameters, and to avoid the bitrate increase that follows from transmitting a parameter for each sub-band (Oshikiri, paragraph [0013]: "because sub-hands are multiplied by panning coefficients, there is a problem that energy steps occurring in the spectra between sub-bands reduce the quality ... Although narrowing the width of the sub-bands to suppress the occurrence of energy steps can be envisioned as a method of solving this problem, the number of panning coefficients that must be transmitted from the encoder side to the decoder side increases, resulting in an increase in the bit rate.").
Eckert itself confirms the expectation of success, describing that the number of bands is reduced for inactive frames precisely because background noise is broadband (Eckert, paragraph [0076]: "The assumption behind reducing the number of bands for inactive frames is that typically less frequency resolution is required for capturing noise parameters, due to the broadband nature of background noise.").
14.Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Eckert (U.S. 2023/0215445) in view of Dickins (U.S. 2016/0027447).
Regarding claim 11, Eckert discloses:
11. An audio decoder according to claim 4, wherein the multi-channel generator is configured to generate the two or more intermediate channels depending on the random noise and depending on control parameters being encoded within the bitstream, for example wherein the control parameters comprise, e.g., a scaling factor and/or, e.g., either a coherence or a correlation, wherein the multi-channel generator is configured to generate a first one the two or more intermediate channels depending on a first random noise portion and depending on a third noise portion and depending on the control parameters, for example the scaling factor and the coherence and/or correlation, wherein the multi-channel generator is configured to generate a second one of the two or more intermediate channels depending on a second random noise portion and depending on the third noise portion and depending on the control parameters (Eckert, paragraph [0092]: "two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds."; paragraph [0095]: the decorrelation coefficients computed by the SPAR encoder and used at the decoder.);
wherein the multi-channel generator is configured to generate the first random noise portion of the random noise using the random generator with a first seed, wherein the multi-channel generator is configured to generate the second random noise portion of the random noise using the random generator with a second seed, and wherein the multi-channel generator is configured to generate the third random noise portion of the random noise using the random generator with a third seed, wherein the second seed is different from the first seed, and wherein the third seed is different from the first seed and different from the second seed (Eckert, paragraph [0050]: "The multi-channel excitation signal may be a multi-channel white noise signal where all channels are generated with different seed and are uncorrelated with each other.").
Eckert does not disclose the shared third noise portion aspects of the claim.
Dickins discloses:
The recited "depending on" is read as not excluding a dependence on further noise portions. The linear mapping of Dickins forms each of the three output signals from all three noise sources 401, 402, 403 through the entries of the warping matrix M, which is set from the target covariance matrix; a first output signal therefore depends on the first noise source, on the third noise source and on the covariance matrix, and a second output signal depends on the second noise source, on the same third noise source and on the covariance matrix. The target covariance matrix received in the layered bitstream is read on the recited control parameters comprising a correlation.
wherein the multi-channel generator is configured to generate a first one the two or more intermediate channels depending on a first random noise portion and depending on a third noise portion and depending on the control parameters, for example the scaling factor and the coherence and/or correlation, wherein the multi-channel generator is configured to generate a second one of the two or more intermediate channels depending on a second random noise portion and depending on the third noise portion and depending on the control parameters (Dickins, paragraph [0071]: "Generator 119 comprises a plurality of noise sources, in the example, three noise sources 401, 402, 403, configured to generate independent and identically distributed (IID) noise samples, e.g., using a random number generator."; paragraph [0076]: "the spatial modification of stage 521 is a linear mapping defined by a 3×3 matrix, denoted M and called a warping matrix that in one embodiment combines mapping between a first and a second soundfield format with achieving at least one target spatial property indicated by a target statistical property, e.g., a target covariance matrix."; paragraph [0079]: "the matrix M operation of the spatial modification stage 521 is configured to create target statistics in the WXY domain, e.g., a desired covariance matrix in the WXY domain, denoted RT."; paragraph [0117]: "the spectral and spatial properties data are packaged as one of the fields of a layered coding method that codes layers of information (fields) and sends the layers to the receiving endpoint 111, e.g., as a multiplexed bitstream of the layers.").
and wherein the multi-channel generator is configured to generate the third random noise portion of the random noise using the random generator with a third seed, wherein the second seed is different from the first seed, and wherein the third seed is different from the first seed and different from the second seed (Dickins, paragraph [0071]: "three noise sources 401, 402, 403, configured to generate independent and identically distributed (IID) noise samples, e.g., using a random number generator.").
Eckert and Dickins pertain to the generation of comfort noise for two or more channels from random noise at a receiving decoder. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the two separately seeded comfort noise channels of Eckert to be generated, as taught by Dickins, by a warping matrix applied to three independent random noise sources and set from a covariance matrix received in the bitstream, so that each channel is formed from its own noise source and from a third noise source common to both. Under the seeding of Eckert, in which "all channels are generated with different seed" (Eckert, paragraph [0050]), the third noise source of the combination is generated with a third seed different from the first seed and from the second seed; the third seed is supplied by the combination and is not attributed to either reference alone. One of ordinary skill would have been motivated to make this modification in order to give the generated comfort noise a spatial property matching that of the signals captured at the sending endpoint, as described by the covariance the sending endpoint determines and sends (Dickins, paragraph [0110]: "one embodiment of the endpoint 111 is configured to modify the spatial properties of the generated presence noise to match the typically different respective spatial properties of the voice signals captured at, and sent from the other endpoints 105, 107, and 109 that send voice to endpoint 111."; paragraph [0113]: "Some embodiments of a sending endpoint further include determining spatial properties, e.g., estimates of the covariance matrix statistics, including at least estimates of the covariance cross terms across the spectra.").
15.Claims 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Eckert (U.S. 2023/0215445) in view of McGrath (U.S. 11,942,097).
Regarding claim 15, Eckert discloses:
15. An audio decoder according to claim 1, wherein the audio content comprises the plurality of audio objects, wherein, if the audio content exhibits voice activity, a plurality of audio object indices being associated with the plurality of audio objects, a plurality of power ratios being associated with the plurality of audio objects for a plurality of subbands and broadband direction information for the plurality of audio objects are encoded within the bitstream, and the renderer is configured to generate the one or more audio output signals depending on the plurality of audio object indices, depending on the plurality of power ratios and depending on the broadband direction information for the plurality of audio objects (Eckert, paragraph [0022]: "In particular, the multi-channel input signal 101 may comprise (possibly a combination of) ... one or more audio objects"; paragraph [0057]: "the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames)."; paragraph [0058]: "the encoding unit 100 may be configured to send audio data 106 and encoded metadata 107 to the decoding unit 150 for every active frame."; paragraph [0029]: "the decoding unit 150 comprises a reconstruction module 170 which is configured to derive a reconstructed multi-channel signal 111 from the SPAR metadata 105 and from the reconstructed downmix signal 114.").
Eckert does not disclose the audio object indices, the power ratios for a plurality of subbands or the broadband direction information of the claim, or the renderer generating the audio output signals depending on them.
The recited broadband direction information is read on direction information provided once per time segment for all frequency subbands, as distinguished from the power ratios, which the claim recites for a plurality of subbands. The recited audio object indices are read on the dominant object index with which each transmitted direction vector of McGrath is associated.
McGrath discloses:
a plurality of audio object indices being associated with the plurality of audio objects, a plurality of power ratios being associated with the plurality of audio objects for a plurality of subbands and broadband direction information for the plurality of audio objects are encoded within the bitstream, and the renderer is configured to generate the one or more audio output signals depending on the plurality of audio object indices, depending on the plurality of power ratios and depending on the broadband direction information for the plurality of audio objects (McGrath, col. 2, ll. 39-41: "The (dominant) audio elements may relate to (dominant) acoustic objects, (dominant) sound sources, or (dominant) acoustic components in the audio scene"; col. 16, ll. 36-37: "Direction vector p indicates the direction associated with dominant object index p"; col. 17, ll. 12-16: "energy band fraction information 22 can include a fraction value ek,p,b for each band b of a set of bands"; col. 3, ll. 27-31: "an indication of signal power associated with a given direction of arrival may relate to a fraction of signal power in the frequency subband for the given direction of arrival in relation to the total signal power in the frequency subband"; col. 15, ll. 29-30: "At step S650 the direction information and energy-fraction information are encoded to form encoded metadata."; col. 15, ll. 31-33: "at step S660 the encoded downmixed stream is combined with the encoded metadata to form a compact spatial audio scene."; col. 20, ll. 22-26: "a panning vector Pan(dir) for panning the audio element to the channels of the channel-based audio signal is determined, based on the direction of arrival dir of the audio element"; col. 20, ll. 32-35: "a covariance matrix S for the intermediate representation is determined based on the energy information"; col. 20, ll. 36-38: "the coefficients of the inverse mixing matrix M are determined based on the mixing matrix E and the covariance matrix S"; col. 18, l. 66 to col. 19, l. 1: "Object panner 91 takes input from dominant object signals 90 and creates panned object stream 92").
Eckert and McGrath pertain to the coding of a spatial audio scene as a downmix of a small number of channels together with metadata from which a decoder reconstructs the scene. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have encoded, in the active frames of Eckert, the metadata of McGrath for the audio objects of the input signal, namely the index and direction vector of each dominant object and its energy fraction for each subband, and to have rendered the audio output signals from that metadata by the demixing and panning of McGrath, because McGrath teaches that representing the scene by a limited number of dominant objects with per-subband energy fractions sets "a trade-off between the metadata data-rate and the quality of the reconstructed audio scene" (McGrath, col. 22, ll. 35-44), so that the modification gives the decoder of Eckert control over the metadata rate spent on the objects while the rendering follows the transmitted object directions. The modification applies a known metadata representation of audio objects to the object input that Eckert already accepts, with the predictable result of rendering those objects from their transmitted indices, energy fractions and directions.
Regarding claim 16, Eckert discloses:
16. An audio decoder according to claim 7, wherein the audio content comprises the plurality of audio objects, wherein, if the audio content does not exhibit voice activity, broadband direction information for the plurality of audio objects and the control parameters are encoded within the bitstream, and the renderer is configured to generate the one or more audio output signals depending on the broadband direction information (Eckert, paragraph [0022]: "In particular, the multi-channel input signal 101 may comprise (possibly a combination of) ... one or more audio objects"; paragraph [0057]: "the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames)."; paragraph [0058]: "the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) to the decoding unit 150 for the SID frames."; paragraph [0094]: "WCNG, PING comfort noise signals and the two decorrelated signals may then be upmixed to an FOA output using the SPAR metadata 105.").
Eckert does not disclose broadband direction information for the plurality of audio objects.
McGrath discloses:
broadband direction information for the plurality of audio objects and the control parameters are encoded within the bitstream, and the renderer is configured to generate the one or more audio output signals depending on the broadband direction information (McGrath, col. 2, ll. 39-41: "The (dominant) audio elements may relate to (dominant) acoustic objects, (dominant) sound sources, or (dominant) acoustic components in the audio scene"; col. 16, ll. 36-37: "Direction vector p indicates the direction associated with dominant object index p"; col. 15, ll. 29-30: "At step S650 the direction information and energy-fraction information are encoded to form encoded metadata."; col. 15, ll. 31-33: "at step S660 the encoded downmixed stream is combined with the encoded metadata to form a compact spatial audio scene."; col. 20, ll. 22-26: "a panning vector Pan(dir) for panning the audio element to the channels of the channel-based audio signal is determined, based on the direction of arrival dir of the audio element"; col. 18, l. 66 to col. 19, l. 1: "Object panner 91 takes input from dominant object signals 90 and creates panned object stream 92").
The rationale for combining Eckert and McGrath is the same as that provided for claim 15 above. In the frames of Eckert that do not exhibit voice activity, the encoded metadata is what the encoding unit sends (Eckert, paragraph [0058]), and the direction vectors of McGrath are encoded as part of that metadata (McGrath, col. 15, ll. 29-33).
16.Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Eckert (U.S. 2023/0215445) in view of McGrath (U.S. 11,942,097) and Atti (U.S. 10,854,209).
Regarding claim 17, the combination of Eckert and McGrath discloses the broadband direction information encoded within the bitstream in the active and inactive phases, as set forth for claims 15 and 16 above.
17. An audio decoder according to claim 15, wherein, when the audio content exhibits voice activity, a first quantization resolution of the broadband direction information being encoded within the bitstream is different from a second quantization resolution of the broadband direction information, when the audio content does not exhibit voice activity (paragraph [0057]: "the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames)."; paragraph [0058]: "the encoding unit 100 may be configured to send audio data 106 and encoded metadata 107 to the decoding unit 150 for every active frame."; paragraph [0058]: "the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) to the decoding unit 150 for the SID frames."; col. 16, ll. 36-37: "Direction vector p indicates the direction associated with dominant object index p"; col. 15, ll. 29-30: "At step S650 the direction information and energy-fraction information are encoded to form encoded metadata."; col. 15, ll. 31-33: "at step S660 the encoded downmixed stream is combined with the encoded metadata to form a compact spatial audio scene.").
The combination of Eckert and McGrath does not disclose that the quantization resolution of the direction information in the frames that exhibit voice activity differs from the quantization resolution in the frames that do not.
The recited quantization resolution is read on the number of bits with which the azimuth and elevation of the direction information are encoded.
Atti discloses:
a first quantization resolution of the broadband direction information being encoded within the bitstream is different from a second quantization resolution of the broadband direction information, when the audio content does not exhibit voice activity (Atti, col. 5, ll. 61-67: "the streams 131-133 have an independent streams (IS) format in which the two or more of the audio signals 136-139 are processed to estimate the spatial characteristics (e.g., azimuth, elevation, etc.) of the sound sources. The audio signals 136-139 are mapped to independent streams corresponding to sound sources and the corresponding spatial metadata 124"; col. 22, ll. 40-47: "a quantized version of the spatial metadata 124 may be used where an amount of quantization for each IS stream is based on the priority of the IS stream. For example, spatial metadata encoding for high-priority streams may use 4 bits for azimuth data and 4 bits for elevation data, and spatial metadata encoding for low-priority streams may use 3 bits or fewer for azimuth data and 3 bits or fewer for elevation data"; col. 14, ll. 20-23: "assigning higher priority to streams in which speech content is detected and lower priority to streams in which speech content is not detected"; col. 14, ll. 49-56: "during some periods of the conversation the user may be silent. In response to the stream having relatively low signal energy due to the user's silence, the stream priority module 110 may reduce the priority of the stream to relatively low priority").
Eckert, McGrath and Atti pertain to the coding of a spatial audio scene in which each sound source carries direction metadata and the bits spent on a frame depend on its content. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have encoded the direction vectors of McGrath in the frames of Eckert with the number of azimuth and elevation bits set by the priority of Atti, so that the frames in which no speech is detected carry the directions with fewer bits than the frames that exhibit voice activity, because Atti reduces the priority of a stream when its content is silent (Atti, col. 14, ll. 20-23 and ll. 49-56) and encodes the spatial metadata of a lower-priority stream with fewer azimuth and elevation bits (Atti, col. 22, ll. 40-47), and states that a frame with "inactive content" has its bitrate reduced (col. 21, ll. 54-60: "when external information indicates that one stream is high priority and is supposed to be encoded using a high bitrate, but the stream has inactive content in it in a specific frame, the pre-analysis can detect the inactive content and reduce the stream's bitrate for that frame despite being indicated as high priority"). The modification spends fewer bits on the direction information of the inactive frames of Eckert, which is the purpose for which Eckert sends only encoded metadata in those frames (Eckert, paragraph [0058]).
17. Claims 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Eckert (U.S. 2023/0215445) in view of Eichenseer (WO 2022/079044) and Oshikiri (U.S. 2013/0223633).
Regarding claim 18, Eckert discloses:
18. An audio decoder according to claim 1, wherein the renderer comprises a signal power computation unit for computing a reference power depending on the two or more transport channels for each of a plurality of time-frequency tiles (Eckert, paragraph [0029]: "the decoding unit 150 comprises a reconstruction module 170 which is configured to derive a reconstructed multi-channel signal 111 from the SPAR metadata 105 and from the reconstructed downmix signal 114.").
Eckert does not disclose the power computation aspects of the claim.
The scaling language of claim 18 is interpreted as follows. The direct power limitation opens with "if the audio content does not exhibit voice activity" and then recites the use of transmitted power ratios "if the audio content exhibits voice activity." Read in light of Fig. 10 and paragraph [0205] of the specification, which describe direct power computation unit 952, the limitation is interpreted as requiring the direct power computation unit to scale the reference power using the transmitted power ratios when the audio content exhibits voice activity, and using a scaling factor when the audio content does not exhibit voice activity, the scaling factor being either encoded within the bitstream or a constant. The alternative in which the scaling factor is encoded within the bitstream is applied below. The transmitted power ratio scaling is mapped to Eichenseer, and the scaling factor encoded within the bitstream, used when the audio content does not exhibit voice activity, is mapped to Oshikiri.
Eichenseer discloses:
a signal power computation unit for computing a reference power depending on the two or more transport channels for each of a plurality of time-frequency tiles (Eichenseer, description of Fig. 5: "the Fig. 5 embodiment comprises a signal power calculation block 721, a direct power calculation block 722, a covariance matrix calculation block 73, a target covariance matrix calculation block 724, an input covariance matrix calculation block 726, a mixing matrix calculation block 725 and a rendering block 727.");
wherein the renderer comprises a direct power computation unit, wherein, if the audio content does not exhibit voice activity, the direct power computation unit is configured for scaling the reference power to acquire a scaled reference power, using transmitted power ratios being encoded within the bitstream, if the audio content exhibits voice activity, and using a scaling factor being, wherein the scaling factor is encoded within the bitstream or wherein the scaling factor is a constant scaling factor, for example which depends on a number of transmitted objects, wherein the renderer is configured to generate the one or more audio output signals depending on the scaled reference power (Eichenseer, description of the target covariance matrix calculation: "For each time/frequency tile (within the parameter band), the audio signal power P(k,n) is determined. In the case of two transport channels, the signal power of the first channel is added to that of the second. To this signal power, each of the power ratio values is multiplied, thus yielding one direct power value for each relevant/dominant object.").
Eckert and Eichenseer pertain to the decoder-side rendering of a parametrically coded multi-channel audio scene from a transmitted downmix and its associated metadata. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the reconstruction module of Eckert to compute a reference power from the transport channels per time-frequency tile and to scale it by the transmitted power ratios, as taught by Eichenseer. One of ordinary skill would have been motivated to make this modification in order to determine the contribution of each dominant object to each time-frequency tile from a transmitted ratio rather than from a separately transmitted absolute power, which saves the transmission of one power data item (Eichenseer, description of the power ratio calculation: "power ratios are preferred in order to save the transmission of one power data item.").
The combination of Eckert and Eichenseer does not disclose scaling the reference power using a scaling factor being encoded within the bitstream when the audio content does not exhibit voice activity.
The recited scaling factor being encoded within the bitstream is read on a quantized frame energy that is encoded for a frame in which the audio content does not exhibit voice activity and by which the decoder multiplies a signal whose frame energy has been normalized to one.
Oshikiri discloses:
wherein the renderer comprises a direct power computation unit, wherein, if the audio content does not exhibit voice activity, the direct power computation unit is configured for scaling the reference power to acquire a scaled reference power, using transmitted power ratios being encoded within the bitstream, if the audio content exhibits voice activity, and using a scaling factor being, wherein the scaling factor is encoded within the bitstream or wherein the scaling factor is a constant scaling factor, for example which depends on a number of transmitted objects, wherein the renderer is configured to generate the one or more audio output signals depending on the scaled reference power (Oshikiri, paragraph [0033]: "VAD section 101 analyzes an input signal (a stereo signal formed by an L-channel signal and an R-channel signal) and judges whether the input signal of the current frame is a speech part or a non-speech part. ... In the following, a background noise part will be described as a typical non-speech part."; paragraph [0045]: "Stereo DTX decoding section 204 decodes the encoded stereo data input from switching section 202 (that is, the encoded stereo data generated in stereo signal encoding apparatus 100 when the stereo signal is a background noise part) to generate a decoded stereo signal (decoded L-channel signal and decoded R-channel signal)."; paragraph [0071]: "Demultiplexing section 401 demultiplexer the encoded stereo data input from switching section 202 (FIG. 2) into the quantized L-channel signal frame energy information, the quantized R-channel signal frame energy information ..."; paragraph [0083]: "Excitation generation section 409 generates an excitation signal represented by a random signal or a limited number of pulses and outputs the excitation signal to multiplication section 410. Normalization is done so that the frame energy of the excitation signal is 1."; paragraph [0084]: "Multiplication section 410 multiplies the excitation signal by the decoded L-channel signal frame energy and outputs the multiplication result to synthesis filter section 411.").
Eckert, Eichenseer, and Oshikiri pertain to the coding of multi-channel audio at a low bit rate by transmitting a signal and parameters from which a decoder reconstructs the channels. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the direct power computation of the combination of Eckert and Eichenseer so that, for a frame in which the audio content does not exhibit voice activity, the reference power is scaled by a quantized frame energy encoded within the silence indicator frame, in place of the transmitted power ratios, as Oshikiri encodes a quantized frame energy for a non-speech frame and multiplies the unit-energy excitation by the decoded frame energy at the decoder. One of ordinary skill would have been motivated to make this modification in order to describe the power of the background noise in a silence indicator frame with a single value rather than with a value for each sub-band, which Oshikiri expressly teaches avoids an increase in the bit rate (Oshikiri, paragraph [0013]: "Although narrowing the width of the sub-bands to suppress the occurrence of energy steps can be envisioned as a method of solving this problem, the number of panning coefficients that must be transmitted from the encoder side to the decoder side increases, resulting in an increase in the bit rate.").
Regarding claim 19, the combination of Eckert, Eichenseer and Oshikiri discloses the renderer and the scaled reference power as set forth for claim 18 above. Eckert further discloses the classification of the audio content into frames that exhibit voice activity and frames that do not (Eckert, paragraph [0057]: "the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames).").
Claim 19 is read as follows. The claim conditions the direction information from which the direct response is computed on whether the audio content exhibits voice activity: the quantized direction information of the dominant objects, being a proper subset of the plurality of audio objects, when it does, and the quantized direction information of all audio objects when it does not. The claim does not require that the direct response for all audio objects be computed only when the audio content does not exhibit voice activity. A renderer that computes a direct response from the quantized direction information of every audio object in every frame, and that uses the direct responses of the dominant subset when the audio content exhibits voice activity, is read as meeting both conditions. On that reading the direct response computation is mapped to Eichenseer, and the two conditions are supplied by the classification of frames of Eckert quoted above.
Eichenseer further discloses:
19. An audio decoder according to claim 18, wherein the renderer comprises a direct response computation unit for computing a direct response, wherein the renderer is configured to compute the direct response depending on quantized direction information of dominant objects being a proper subset of the plurality of audio objects of the audio content, if the audio content exhibits voice activity, wherein the renderer is configured to compute the direct response depending on quantized direction information of all audio objects of the audio content, if the audio content does not exhibit voice activity, wherein the quantized direction information is encoded within the bitstream, wherein the renderer is configured to generate the one or more audio output signals depending on the direct response (Eichenseer, description of the parametric side information read by the decoder: "Direction information as quantized azimuth and elevation values (for each frame)"; description of the output signal rendering and synthesis: "For all (input) objects, using the transmitted object directions, so-called direct response values are determined that describe the panning gains to be employed to the output channels. ... Each object has a vector of direct response values dri (containing as many elements as there are loudspeakers) associated with it. These vectors are computed once per frame."; description of the covariance synthesis substeps: "For each parameter band, the object indices, describing the subset of dominant objects among the input objects within the time/frequency tiles grouped into this parameter band, are used to extract the subset of vectors dri needed for the further processing.").
The rationale for combining Eckert, Eichenseer and Oshikiri is the same as that provided for claim 18 above.
Regarding claim 20, the combination of Eckert, Eichenseer and Oshikiri discloses the renderer, the scaled reference power, and the direct response as set forth for claims 18 and 19 above.
Eichenseer further discloses:
20. An audio decoder according to claim 19, wherein the renderer comprises an input covariance matrix computation unit for computing an input covariance matrix depending on the two or more transport channels, wherein the renderer comprises a target covariance matrix computation unit for computing a target covariance matrix depending on the direct response and depending on the scaled reference power, wherein the renderer comprises a mixing matrix computation unit for computing a mixing matrix for rendering depending on the input covariance matrix and depending on the target covariance matrix, wherein the renderer is configured to generate the one or more audio output signals depending on the mixing matrix (Eichenseer, description of the covariance synthesis: "For each (sub)frame and for each frequency band, an input covariance matrix Cx = xxT of size transport channels-by-transport channels is calculated from the decoded audio signal."; description of Fig. 5: "the Fig. 5 embodiment comprises a signal power calculation block 721, a direct power calculation block 722, a covariance matrix calculation block 73, a target covariance matrix calculation block 724, an input covariance matrix calculation block 726, a mixing matrix calculation block 725 and a rendering block 727.");
computing a mixing matrix for rendering depending on the input covariance matrix and depending on the target covariance matrix (Eichenseer, description of the direct response information: "This direct response information preferably comprises gain values either used for a covariance synthesis or an advanced covariance synthesis ... the covariance synthesis information which is, preferably, the mixing matrix, is applied to the one or more transport channels to obtain the number of audio channels.").
The rationale for combining Eckert, Eichenseer and Oshikiri is the same as that provided for claim 18 above.
Claim 18 is also rejected under 35 U.S.C. 103 as being unpatentable over Eckert (U.S. 2023/0215445) in view of Laitinen (U.S. 11,412,336) and Oshikiri (U.S. 2013/0223633).
Regarding claim 18, Eckert discloses:
18. An audio decoder according to claim 1, wherein the renderer comprises a signal power computation unit for computing a reference power depending on the two or more transport channels for each of a plurality of time-frequency tiles (Eckert, paragraph [0029]: "the decoding unit 150 comprises a reconstruction module 170 which is configured to derive a reconstructed multi-channel signal 111 from the SPAR metadata 105 and from the reconstructed downmix signal 114.").
Eckert does not disclose the signal power computation unit or the direct power computation unit of the claim.
The scaling language of claim 18 is interpreted as follows. The direct power limitation opens with "if the audio content does not exhibit voice activity" and then recites the use of transmitted power ratios "if the audio content exhibits voice activity." Read in light of Fig. 10 and paragraph [0205] of the specification, which describe direct power computation unit 952, the limitation is interpreted as requiring the direct power computation unit to scale the reference power using the transmitted power ratios when the audio content exhibits voice activity, and using a scaling factor when the audio content does not exhibit voice activity, the scaling factor being either encoded within the bitstream or a constant. The alternative in which the scaling factor is encoded within the bitstream is applied below. The transmitted power ratio scaling is mapped to Laitinen, and the scaling factor encoded within the bitstream, used when the audio content does not exhibit voice activity, is mapped to Oshikiri. The recited transmitted power ratios are read on the direct-to-total energy ratio parameter transmitted for each time-frequency tile in Laitinen, and the recited two or more transport channels on the downmix signals of Laitinen.
Laitinen discloses:
a signal power computation unit for computing a reference power depending on the two or more transport channels for each of a plurality of time-frequency tiles (Laitinen, col. 27, ll. 23-26: "The covariance matrices 1206 in frequency bands is simply determined in the covariance matrix estimator 1203 and measured from the downmix signals in frequency bands from the time-frequency domain transformer 1201"; col. 27, ll. 36-42: "estimate the overall energy E 1204 of the target covariance matrix based on the input covariance matrix" ... "determined from the sum of the diagonal elements of the input covariance matrix");
wherein the renderer comprises a direct power computation unit, wherein, if the audio content does not exhibit voice activity, the direct power computation unit is configured for scaling the reference power to acquire a scaled reference power, using transmitted power ratios being encoded within the bitstream, if the audio content exhibits voice activity, and using a scaling factor being, wherein the scaling factor is encoded within the bitstream or wherein the scaling factor is a constant scaling factor, for example which depends on a number of transmitted objects, wherein the renderer is configured to generate the one or more audio output signals depending on the scaled reference power (col. 28, ll. 4-5: "determine the direct part energy as rE"; col. 27, ll. 55-56: "r is the direct-to-total energy ratio parameter from the input metadata"; col. 27, ll. 47-48: "CT = CD + CA"; col. 32, ll. 55-57: "The optimal mixing matrix may then be determined based on estimated covariance matrix and target covariance matrix"; col. 20, ll. 35-38: "Similar processing can be also performed for audio object input, by treating the audio objects as audio channels at determined positions at each temporal parameter estimation interval").
Eckert and Laitinen pertain to the decoder-side rendering of a parametrically coded spatial audio scene from a transmitted downmix and its associated metadata. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have rendered the reconstructed downmix of Eckert by the synthesis of Laitinen, in which the overall energy of each time-frequency tile is measured from the downmix signals and the direct part energy is set by multiplying it by the transmitted ratio, because Laitinen applies that synthesis to audio object input by treating the objects as channels at their positions (Laitinen, col. 20, ll. 35-38), which is the object input that Eckert accepts, and because setting the direct energy from a transmitted ratio of a measured reference power reproduces the level of the directional part of each tile from a single transmitted parameter, a predictable result of the combination.
The combination of Eckert and Laitinen does not disclose scaling the reference power using a scaling factor being encoded within the bitstream when the audio content does not exhibit voice activity.
The recited scaling factor being encoded within the bitstream is read on a quantized frame energy that is encoded for a frame in which the audio content does not exhibit voice activity and by which the decoder multiplies a signal whose frame energy has been normalized to one.
Oshikiri discloses:
wherein the renderer comprises a direct power computation unit, wherein, if the audio content does not exhibit voice activity, the direct power computation unit is configured for scaling the reference power to acquire a scaled reference power, using transmitted power ratios being encoded within the bitstream, if the audio content exhibits voice activity, and using a scaling factor being, wherein the scaling factor is encoded within the bitstream or wherein the scaling factor is a constant scaling factor, for example which depends on a number of transmitted objects, wherein the renderer is configured to generate the one or more audio output signals depending on the scaled reference power (Oshikiri, paragraph [0033]: "VAD section 101 analyzes an input signal (a stereo signal formed by an L-channel signal and an R-channel signal) and judges whether the input signal of the current frame is a speech part or a non-speech part. ... In the following, a background noise part will be described as a typical non-speech part."; paragraph [0045]: "Stereo DTX decoding section 204 decodes the encoded stereo data input from switching section 202 (that is, the encoded stereo data generated in stereo signal encoding apparatus 100 when the stereo signal is a background noise part) to generate a decoded stereo signal (decoded L-channel signal and decoded R-channel signal)."; paragraph [0071]: "Demultiplexing section 401 demultiplexer the encoded stereo data input from switching section 202 (FIG. 2) into the quantized L-channel signal frame energy information, the quantized R-channel signal frame energy information ..."; paragraph [0083]: "Excitation generation section 409 generates an excitation signal represented by a random signal or a limited number of pulses and outputs the excitation signal to multiplication section 410. Normalization is done so that the frame energy of the excitation signal is 1."; paragraph [0084]: "Multiplication section 410 multiplies the excitation signal by the decoded L-channel signal frame energy and outputs the multiplication result to synthesis filter section 411.").
Eckert, Laitinen, and Oshikiri pertain to the coding of multi-channel audio at a low bit rate by transmitting a signal and parameters from which a decoder reconstructs the channels. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the direct power computation of the combination of Eckert and Laitinen so that, for a frame in which the audio content does not exhibit voice activity, the reference power is scaled by a quantized frame energy encoded within the silence indicator frame, in place of the transmitted power ratios, as Oshikiri encodes a quantized frame energy for a non-speech frame and multiplies the unit-energy excitation by the decoded frame energy at the decoder. One of ordinary skill would have been motivated to make this modification in order to describe the power of the background noise in a silence indicator frame with a single value rather than with a value for each sub-band, which Oshikiri expressly teaches avoids an increase in the bit rate (Oshikiri, paragraph [0013]: "Although narrowing the width of the sub-bands to suppress the occurrence of energy steps can be envisioned as a method of solving this problem, the number of panning coefficients that must be transmitted from the encoder side to the decoder side increases, resulting in an increase in the bit rate.").
Conclusion
18. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure:
a. Ramo (U.S. 2025/0140271) discloses transmitting a quantised average spatial direction value in a silence descriptor frame and signalling the use of prediction or non-prediction across the remaining frames of the interval with a one bit flag.
b. Eriksson (U.S. 2017/0047072) discloses generating a stereo comfort noise pair by cross-mixing two noise signals under a transmitted coherence.
c. U.S. 11,470,436 discloses spatial audio parameter signalling for multichannel reproduction.
d. U.S. 9,578,435 discloses transmitting object level differences for each of a plurality of audio objects.
e. J. Vilkamo, T. Backstrom and A. Kuntz, "Optimized Covariance Domain Framework for Time-Frequency Processing of Spatial Audio," Journal of the Audio Engineering Society, vol. 61, no. 6, pp. 403-411, 2013 June, discloses the least-squares optimized derivation of a mixing matrix from an input covariance matrix and a target covariance matrix.
19. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Yuval Levental whose telephone number is (571)270-0000. The examiner can normally be reached Monday through Friday, 8:00 a.m. to 4:30 p.m. Eastern Time.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir, can be reached on (571)272-0000. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YUVAL HAIM LEVENTAL/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659