DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-26 are pending. Claims 1, 25, and 26 are independent.
This Application was published as US 20250210051 A1.
Apparent priority: 9 September 2022.
Information Disclosure Statement
The IDS dated 3-9-2025, 6-3-2025, 6-10-2025, 10-8-2028, 10-21-2025, and 3-18-2026 has been considered and placed in the application file.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f), is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f):
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f), is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f), because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
a transport signal generator for in claim 1,6, 13, 14, 16;
a voice activity determiner for in claim 1, 10;
determining a voice activity decision for claim 25, 26; and
a renderer for claim 24.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f), applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f).
Claim Objections
Claim(s) 3 and 13 is/are objected to because of the following informalities:
Claim(s) 3 recites “transport channel of the two or more one transport channels” it should be “transport channel of the two or more
Claim(s) 13 recites “object indices and power ratios of the plurality of audio input objects and or of the” object indices and power ratios of the plurality of audio input objects and/or of the”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claim(s) 15 and 16 are rejected under 35 U.S.C. 112(b), as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim 15 recites the limitation "the inactive metadata generator". There is insufficient antecedent basis for this limitation in the claim.
Regarding claim 16, the phrase "for example" renders the claim indefinite because it is unclear whether the limitation(s) following the phrase are part of the claimed invention. See MPEP § 2173.05(d).
Claim(s) 16 recite “e.g., the scaling factor and/or, e.g., either the coherence or the correlation.” It is unclear what e.g. entails.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claim 1,13-16, 20, 22, 23, 24, 25, 26 is rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1,2, 7, 8, 9, 15, 16, 17, 21, 22, 23, 24, 25 of U.S. Patent No. 10971170. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of the issued patent are narrower in scope than that of the instant application and therefore anticipates the claims if the instant application. Please see the Claim Mapping in the table below.
Each of the claims of the instant application map to the co-pending application, where Claim (I) maps to Claim (P).
Claim 1 (I): Claim 23 (P); Claim 1 (I): Claim 1&2 (P); Claim 13 (I): Claim 15 (P); Claim 14&16 (I): Claim 16 (P); Claim 15 (I): Claim 17 (P); Claim 20 (I): Claim 7&8&9 (P); Claim 22 (I): Claim 21 (P); Claim 23 (I): Claim 22 (P); Claim 24 (I): Claim 23 (P); Claim 25 (I): Claim 24 (P); Claim 26 (I): Claim 25 (P);.
Instant Application: 19074413
Co-pending application: 19074416
Claim 1: 1. An audio encoder, comprising:
a transport signal generator for generating two or more transport channels of a transport signal from audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels,
a voice activity determiner for determining a voice activity decision for the transport signal, which indicates whether or not the audio input within the transport signal exhibits voice activity, and
a bitstream generator for generating a bitstream depending on the audio input,
wherein, if the voice activity determiner has determined that the transport signal exhibits voice activity, the bitstream generator is adapted to encode the two or more transport channels within the bitstream,
wherein, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, the bitstream generator is suitable to encode, instead of the two or more transport channels, information on a background noise,
wherein the information on the background noise comprises information on a background noise of at least one of the two or more transport channels or information on a background noise of a derived signal which depends on at least one of the two or more transport channels.
Claim 23: 23. A system, comprising:
an audio encoder, and
an audio decoder according to an audio decoder according to
wherein the audio encoder comprises:
a transport signal generator for generating two or more transport channels of a transport signal from audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels,
a voice activity determiner for determining a voice activity decision for the transport signal, which indicates whether or not the audio input within the transport signal exhibits voice activity, and
a bitstream generator for generating a bitstream depending on the audio input,
wherein, if the voice activity determiner has determined that the transport signal exhibits voice activity, the bitstream generator is adapted to encode the two or more transport channels within the bitstream,
wherein, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, the bitstream generator is suitable to encode, instead of the two or more transport channels, information on a background noise,
wherein the information on the background noise comprises information on a background noise of at least one of the two or more transport channels or information on a background noise of a derived signal which depends on at least one of the two or more transport channels,
wherein the audio encoder is configured to generate a bitstream from audio input,
and
wherein the audio decoder is configured to generate one or more audio output signals from the bitstream.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 5, 10,11,13,18, 22, 24-26 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US 20230215445, (ECKERT).
Claim 1, 25, and 26
Regarding Claim 1, ECKERT teach 1. An audio encoder, comprising:
a transport signal generator for generating two or more transport channels of a transport signal from audio input comprising at least one of a plurality of audio input objects and a plurality of audio input channels,
(“[0022] FIG. 1 illustrates an encoding unit 100 and a decoding unit 150 for encoding and decoding a multi-channel input signal 101, which may comprise an SR signal. In particular, the multi-channel input signal 101 may comprise (possibly a combination of) one or more mono signals, one or more stereo signals, one or more binaural signal, one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects, and/or one or more SR signals. The different signal components may be considered to be individual channels of the multi-channel input signal 101.
[0023] The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels. The downmix signal 103 may itself be an SR signal, notably a first order ambisonics (FOA) signal, if the input signal 101 comprises a HOA signal. Downmixing may be performed in the subband domain or QMF domain (e.g., using 10 or more subbands).”
“[0077] For the case of a two channel downmix 103 (which comprises e.g., the representation of W channel signal and the Ypred or Y′ channel signal), the data comprised within the bitstream from the encoding unit 100 to the decoding unit 150 may comprise (for a frame of the input signal 101):”)
a voice activity determiner for determining a voice activity decision for the transport signal, which indicates whether or not the audio input within the transport signal exhibits voice activity, and
(“[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.
[0062] Hence a VAD and/or SAD may be used to classify the frames of a sequence of frames into active frames or inactive frames. Encoding and/or generating comfort noise may be applied to the inactive frames. The encoding of the comfort noise (notably the encoding of noise shaping parameters) within the encoding unit 100 may be performed such that the decoding unit 150 is enabled to generate high quality comfort noise for a soundfield. The comfort noise that is generated by the decoding unit 150 preferably matches the spectral and/or spatial characteristics of the background noise within the input signal 101. This does not necessarily imply the waveform reconstruction of the input background noise. The comfort noise generated by a soundfield decoding unit 150 for a series of inactive frames is preferably such that the comfort noise sounds continuous with regards of the noise within the directly preceding active frames. Hence, the transition between active and inactive frames at the decoding unit 150 is preferably smooth and non-abrupt.”)
a bitstream generator for generating a bitstream depending on the audio input,
(“[0005] According to an aspect, a method for encoding a multi-channel input (audio) signal which comprises N different channels, with N>1, in particular N>2, is described. The method comprises determining whether a current frame of the multi-channel input signal is an active frame or an inactive frame, using a signal and/or a voice activity detector. Furthermore, the method comprises determining a downmix signal based on the multi-channel input signal and/or based on a target bitrate for encoding the multi-channel input signal, wherein the downmix signal comprises less than or equal to N channels. The method further comprises determining upmixing metadata comprising a set of (spatial) parameters for generating, based on the downmix signal, a reconstructed multi-channel signal comprising N channels. The upmixing metadata may be determined in dependence of whether the current frame is an active frame or an inactive frame. In addition, the method comprises encoding the upmixing metadata into a bitstream.”
“[0025] In addition, the encoding unit 100 may comprise a coding module 140 which is configured to perform waveform encoding (e.g., EVS encoding) of the downmix signal 103, thereby providing coded audio data 106. Each channel of the downmix signal 103 may be encoded using a mono waveform encoder, thereby enabling an efficient encoding. Furthermore, the encoding unit 100 comprises a quantization module 141 which is configured to quantize the SPAR metadata 105 and to perform entropy encoding of the (quantized) SPAR metadata 105, thereby providing coded metadata 107. The coded audio data 106 and the coded metadata 107 may be inserted into a bitstream. Coding of the downmix signal 103 and/or of the SPAR metadata 105 is typically controlled using a mode and/or bitrate control module 142.”)
wherein, if the voice activity determiner has determined that the transport signal exhibits voice activity, the bitstream generator is adapted to encode the two or more transport channels within the bitstream,
(“[0126] The method 600 may further comprise encoding the downmix signal 103 into the bitstream, if, in particular only if, the current frame is an active frame. The one or more channels of the downmix signal 103 may be encoded individually using (one or more instances of) a single channel audio encoder (such as an EVS (enhanced voice services) encoder) to provide audio data 106 to be inserted into the bitstream.”
“[0077] For the case of a two channel downmix 103 (which comprises e.g., the representation of W channel signal and the Ypred or Y′ channel signal), the data comprised within the bitstream from the encoding unit 100 to the decoding unit 150 may comprise (for a frame of the input signal 101):”)
wherein, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, the bitstream generator is suitable to encode, instead of the two or more transport channels, information on a background noise,
(“[0058] In other words, the encoding unit 100 may be configured to send audio data 106 and encoded metadata 107 to the decoding unit 150 for every active frame. On the other hand, the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) for a fraction of the inactive frames (i.e., for the SID frames). For the remaining inactive frames (i.e., for the ND frames), no data may be sent at all (not even encoded metadata 107).”
“[0062] Hence a VAD and/or SAD may be used to classify the frames of a sequence of frames into active frames or inactive frames. Encoding and/or generating comfort noise may be applied to the inactive frames. The encoding of the comfort noise (notably the encoding of noise shaping parameters) within the encoding unit 100 may be performed such that the decoding unit 150 is enabled to generate high quality comfort noise for a soundfield. The comfort noise that is generated by the decoding unit 150 preferably matches the spectral and/or spatial characteristics of the background noise within the input signal 101. This does not necessarily imply the waveform reconstruction of the input background noise. The comfort noise generated by a soundfield decoding unit 150 for a series of inactive frames is preferably such that the comfort noise sounds continuous with regards of the noise within the directly preceding active frames. Hence, the transition between active and inactive frames at the decoding unit 150 is preferably smooth and non-abrupt.
[0063] The decoding unit 150 may be configured to generate random white noise as an excitation signal. The excitation signal may comprise multiple channels of white noise, wherein the white noise in the different channels is typically uncorrelated from one another. The bitstream from the encoding unit 100 may only comprise noise shaping parameters (as encoded metadata 107), and the decoding unit 150 may be configured to shape the random white noise within the different channels (spectrally and spatially) using the noise shaping parameters that have been provided within the bitstream. By doing this, spatial comfort noise may be generated in an efficient manner.”)
wherein the information on the background noise comprises information on a background noise of at least one of the two or more transport channels or information on a background noise of a derived signal which depends on at least one of the two or more transport channels.
(“[0091] For the two channel downmix example (4-2-4, for a first order soundfield), the comfort noise parameters for the mono dowmmix (W′) channel and for one prediction channel may be provided to the decoding unit 150. The decoding unit 150 may apply a method for generating FOA spatial comfort noise from a two channel downmix 103 and from the SPAR metadata 105. The two downmix channels may be uncorrelated comfort noise signals, one having the spectrum shaped according to the original W channel representation and the other one having the spectrum shaped according to the original residual channel.
[0092] For the SID frames, two independent encoder module 140 instances encode spectral information regarding the mono (W′) channel and spectral information regarding the residual channel, respectively. Furthermore, two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds. The uncorrelated comfort noise signals may be spectrally shaped based on the representation of the W channel and the residual channel in the uncoded downmix, respectively. The reconstructed W channel may be referred to as W.sub.CNG and the reconstructed residual channel may be referred to as PING.
[0093] PING typically is a better approximation of the original uncorrelated residual channel compared to decorrelating W.sub.CNG and applying decorrelating coefficients (as done in the full parametric approach, which makes use of a single downmix channel only). As a result of this, the perceptual quality of the background noise is typically higher, when using a multi-channel downmix signal 103.”)
Claim 25 contains limitations similar to those found in Claim 1 and is rejected under similar rationale.
Claim 26
Regarding Claim 26, ECKERT further teaches 26. A non-transitory digital storage medium having a computer program stored thereon to perform the method for audio encoding, wherein the method comprises:
([0178] One or more of the components, blocks, processes or other functional components may be implemented through a computer program that controls execution of a processor-based computing device of the system. It should also be noted that the various functions disclosed herein may be described using any number of combinations of hardware, firmware, and/or as data and/or instructions embodied in various machine-readable or computer-readable media, in terms of their behavioral, register transfer, logic component, and/or other characteristics. Computer-readable media in which such formatted data and/or instructions may be embodied include, but are not limited to, physical (non-transitory), non-volatile storage media in various forms, such as optical, magnetic or semiconductor storage media.)
Claim 26 contains limitations similar to those found in Claim 1 and is rejected under similar rationale.
Claim 2
Regarding Claim 2, ECKERT teach 2. An audio encoder according to claim 2, wherein the voice activity determiner is configured to determine an individual voice activity decision for each transport channel of one or more transport channels of the transport signal, which indicates whether or not the audio input within the transport channel exhibits voice activity, and
(“[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.”
“[0113] Furthermore, the method 600 may comprise determining 602 a downmix signal 103 based on the multi-channel input signal 101 and/or based on the operating and/or target bitrate, wherein the downmix signal 103 typically comprises less than or equal to N channels. In particular, the downmix signal 103 comprises n channels, with typically n≤N, preferably n<N. The number n of channels of the downmix signal 103 may be equal to the number N of channels of the multi-channel input signal 101, in particular for relatively high bit rates. The downmix signal 103 may be generated by selecting one or more channels from the multi-channel input signal 101. The downmix signal 103 may e.g., comprise the W channel of a FOA signal. Furthermore, the downmix signal 103 may comprise one or more residual channels of the FOA signal (which may be derived using the prediction operations described herein).
[0114] The downmix signal 103, in particular the number n of channels of the downmix signal 103, is typically determined in dependence on the target data rate for the bitstream.”)
wherein the voice activity determiner is configured to determine the voice activity decision for the transport signal depending on the individual voice activity decision of each transport channel of the one or more transport channels.
(“[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.”)
Claim 3
Regarding Claim 3, ECKERT teach 3. An audio encoder according to claim 2, wherein the voice activity determiner is configured to determine an individual voice activity decision for each transport channel of the two or more transport channels of the transport signal, which indicates whether or not the audio input within said transport channel exhibits voice activity, and
(“[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.”
“[0113] Furthermore, the method 600 may comprise determining 602 a downmix signal 103 based on the multi-channel input signal 101 and/or based on the operating and/or target bitrate, wherein the downmix signal 103 typically comprises less than or equal to N channels. In particular, the downmix signal 103 comprises n channels, with typically n≤N, preferably n<N. The number n of channels of the downmix signal 103 may be equal to the number N of channels of the multi-channel input signal 101, in particular for relatively high bit rates. The downmix signal 103 may be generated by selecting one or more channels from the multi-channel input signal 101. The downmix signal 103 may e.g., comprise the W channel of a FOA signal. Furthermore, the downmix signal 103 may comprise one or more residual channels of the FOA signal (which may be derived using the prediction operations described herein).
[0114] The downmix signal 103, in particular the number n of channels of the downmix signal 103, is typically determined in dependence on the target data rate for the bitstream.”)
wherein the voice activity determiner is configured to determine the voice activity decision for the transport signal depending on the individual voice activity decision of each transport channel of the two or more one transport channels of the transport signal.
(“[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.”)
Claim 5
Regarding Claim 5, ECKERT teach 5. An audio encoder according to claim 1, wherein the audio encoder is configured to determine, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, whether to transmit the bitstream having encoded therein the information on the background noise, or whether to not generate and to not transmit the bitstream.
(“[0143] If the current frame is an active frame, each channel of the downmix signal 103 may be encoded individually using an instance of a mono audio encoder (such as EVS), wherein the mono audio encoder may be configured to encode the audio signal within a channel of the downmix signal 103 into an (encoded) excitation signal and into (encoded) spectral data.”
“[0057] Hence, the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames). Furthermore, the encoding unit 100 may be configured to determine and encode data for comfort noise generation within a “SID” frame (which corresponds e.g., to the current S frame of a series of S frames). The SID frames may be sent repeatedly, in particular periodically, for a series of S frames. By way of example, a SID frame may be sent every 8.sup.th frame (which corresponds to a time interval of 160 ms between subsequent SID frames, when using 20 ms frames). No data may be transmitted during the one or more following S frames of the series of S frames. Hence, the encoding unit 100 may be configured to perform DTX (discontinuous transmission) or to switch to a DTX mode.
[0058] In other words, the encoding unit 100 may be configured to send audio data 106 and encoded metadata 107 to the decoding unit 150 for every active frame. On the other hand, the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) for a fraction of the inactive frames (i.e., for the SID frames). For the remaining inactive frames (i.e., for the ND frames), no data may be sent at all (not even encoded metadata 107).
[0059] The encoded metadata 107 which is sent for a SID frame may be reduced and/or compressed with regards to the encoded metadata 107 which is sent for an active frame.”)
Claim 10
Regarding Claim 10, ECKERT teaches 10. An audio encoder according to claim 1, wherein the audio encoder comprises a direction information determiner for determining direction information depending on the audio input,
(“[0034] The encoding unit 100 may be configured to convert an FOA input signal 101 into a downmix signal 103 and parameters, i.e., SPAR metadata 105, used to regenerate the input signal 101 at the decoding unit 150. The number of channels of the downmix signal 103 may vary from 1 to 4 channels. The parameters may include prediction parameters Pr, cross-prediction parameters C and/or decorrelation parameters P. These parameters may be calculated from the covariance matrix of a windowed input signal 101. Furthermore, the parameters may be calculated in a specified number of subbands. In the case of comfort noise, a reduced number of subbands (also referred to as frequency bands) may be used, e.g., 6 subbands instead of 12 subbands.”
“[100] In a preferred example, uniform quantization of the parameters Pr (prediction coefficients) and/or P (decorrelator coefficients) may be performed. The quantization scheme may depend on the direction of the noise. In particular, the number of quantization points which is allocated to the different channels may be dependent on the direction of the noise.”
“[0115] The method 600 may further comprise determining 603 upmixing metadata 105, in particular SPAR metadata, comprising a set of parameters. The upmixing metadata 105 may be determined such that it allows generating a reconstructed multi-channel signal 111 comprising N channels based on the downmix signal 103 (or based on a corresponding reconstructed downmix signal 114). The set of parameters of the upmixing metadata 105 may describe and/or model one or more spatial characteristics of audio content, in particular of noise, comprised within the current frame of the multi-channel input signal 101.”
SPAR metadata = direction information)
wherein the audio encoder comprises a direction information quantizer for quantizing the direction information to acquire quantized direction information, and
(“[0025] In addition, the encoding unit 100 may comprise a coding module 140 which is configured to perform waveform encoding (e.g., EVS encoding) of the downmix signal 103, thereby providing coded audio data 106. Each channel of the downmix signal 103 may be encoded using a mono waveform encoder, thereby enabling an efficient encoding. Furthermore, the encoding unit 100 comprises a quantization module 141 which is configured to quantize the SPAR metadata 105 and to perform entropy encoding of the (quantized) SPAR metadata 105, thereby providing coded metadata 107. The coded audio data 106 and the coded metadata 107 may be inserted into a bitstream. Coding of the downmix signal 103 and/or of the SPAR metadata 105 is typically controlled using a mode and/or bitrate control module 142.”
SPAR metadata = direction information)
wherein the bitstream generator is configured to encode the quantized direction information within the bitstream.
(“[0025] In addition, the encoding unit 100 may comprise a coding module 140 which is configured to perform waveform encoding (e.g., EVS encoding) of the downmix signal 103, thereby providing coded audio data 106. Each channel of the downmix signal 103 may be encoded using a mono waveform encoder, thereby enabling an efficient encoding. Furthermore, the encoding unit 100 comprises a quantization module 141 which is configured to quantize the SPAR metadata 105 and to perform entropy encoding of the (quantized) SPAR metadata 105, thereby providing coded metadata 107. The coded audio data 106 and the coded metadata 107 may be inserted into a bitstream. Coding of the downmix signal 103 and/or of the SPAR metadata 105 is typically controlled using a mode and/or bitrate control module 142.”
SPAR metadata = direction information)
Claim 11
Regarding Claim 11, ECKERT teaches 11. An audio encoder according to claim 10, wherein the transport signal generator is configured to generate the two or more transport channels of the transport signal from the audio input using the direction information.
(“[0034] The encoding unit 100 may be configured to convert an FOA input signal 101 into a downmix signal 103 and parameters, i.e., SPAR metadata 105, used to regenerate the input signal 101 at the decoding unit 150. The number of channels of the downmix signal 103 may vary from 1 to 4 channels. The parameters may include prediction parameters Pr, cross-prediction parameters C and/or decorrelation parameters P. These parameters may be calculated from the covariance matrix of a windowed input signal 101. Furthermore, the parameters may be calculated in a specified number of subbands. In the case of comfort noise, a reduced number of subbands (also referred to as frequency bands) may be used, e.g., 6 subbands instead of 12 subbands.”
“[0035] An example representation of SPAR parameter extraction may be as follows (as described with reference to FIG. 3): [0036] 1. Predict all side signals (Y, Z, X) of the input signal 101 from the main W signal of the input signal 101
…
and R.sub.AB=cov(A, B) are elements of the input covariance matrix corresponding to signals A and B. Similarly, the Z′ and X′ residual channels have corresponding parameters, pr.sub.z and pr.sub.x. They may be calculated by replacing the let “Y” with the letter “Z” or “X” in the above formula. The prediction parameters Pr (also referred to as PR) may be the vector of the prediction coefficients [pr.sub.Y, pr.sub.Z, pr.sub.X].sup.T.
[0037] The prediction parameters may be determined within the prediction module 311 shown in FIG. 3, thereby providing the residual channels Y′, Z′ and X′ 301.
[0038] In an exemplary implementation, W may be an active channel (or in other words, with active prediction, hereinafter referred to as W′). As an example (but not as limitation), an active W′ channel that allows some kind of mixing of the X, Y, Z channels into the W channel may be defined as follows:
W′=W+f*pr.sub.y*Y+f*pr.sub.z*Z+f*pr.sub.x*X
[0039] Here, f is the mixing factor and can be static or dynamic across time and/or frequency. In an implementation, f may vary between active and inactive frames. In other words, the mixing factor may be dependent on whether the current frame is an active frame or an inactive frame. In yet other words, the mixing of the X, Y and/or Z channel into the W channel may be different for active frames and for inactive frames. Hence, a representation of the W channel, i.e., the W′ channel, may be determined by mixing the initial W channel with one or more of the other channels. By doing this, the perceptual quality may be further increased.”“[0042] For the example of a WABC remix 302 with 1-4 channels, d and u represent the following channels:…”)
Claim 13
Regarding Claim 13, ECKERT teaches 13. An audio encoder according to claim 1, wherein the audio encoder comprises an active metadata generator for generating metadata comprising at least one of quantized direction information,
(“[0100] In a preferred example, uniform quantization of the parameters Pr (prediction coefficients) and/or P (decorrelator coefficients) may be performed. The quantization scheme may depend on the direction of the noise. In particular, the number of quantization points which is allocated to the different channels may be dependent on the direction of the noise.”
“[0119] The upmixing metadata 105 may be determined in dependence of whether the current frame is an active frame or an inactive frame. In particular, the set of parameters, which is comprised within the upmixing metadata 105 may depend on whether the current frame is an active frame or an inactive frame. If the current frame is an active frame, the set of parameters of the upmixing parameters 105 may be larger and/or may comprise a higher number of different parameters than if the current frame is an inactive frame.”
“[0127] The method 600 may comprise quantizing the parameters from the set of parameters for encoding 604 the upmixing metadata 105 for the current frame into the bitstream, using a quantizer. In other words, a quantizer may be used to quantize the set of parameters, which is to be encoded into the bitstream. The quantizer, in particular the quantization step size and/or the number of quantization steps of the quantizer, may be dependent on whether the current frame is an active frame or an inactive frame. In particular, the quantization step size may be lower and/or the number of quantization steps may be higher for an active frame than for an inactive frame. Alternatively, or in addition, the quantizer, in particular the quantization step size and/or the number of quantization steps of the quantizer, may be dependent on the number of channels of the downmix signal. By doing this, the efficiency of encoding spatial background noise at high perceptual quality may be further increased.”)
object indices and power ratios of the plurality of audio input objects and or of the plurality of audio input channels of the audio input, if the voice activity determiner has determined that the transport signal exhibits voice activity.
(“[0113] Furthermore, the method 600 may comprise determining 602 a downmix signal 103 based on the multi-channel input signal 101 and/or based on the operating and/or target bitrate, wherein the downmix signal 103 typically comprises less than or equal to N channels. In particular, the downmix signal 103 comprises n channels, with typically n≤N, preferably n<N. The number n of channels of the downmix signal 103 may be equal to the number N of channels of the multi-channel input signal 101, in particular for relatively high bit rates. The downmix signal 103 may be generated by selecting one or more channels from the multi-channel input signal 101. The downmix signal 103 may e.g., comprise the W channel of a FOA signal. Furthermore, the downmix signal 103 may comprise one or more residual channels of the FOA signal (which may be derived using the prediction operations described herein).”
“[0116] As indicated above, the multi-channel input signal 101 may comprise an ambisonics signal, notably an FOA signal, with a W channel, a Y channel, a Z channel and an X channel. The set of parameters of the upmixing metadata 105 may comprise prediction coefficients for predicting the Y channel, the Z channel and the X channel based on the W channel, thereby providing residual channels, referred to as Y′ channel, Z′ channel and X′ channel, respectively. The prediction coefficients are referred to herein as Pr or PR. The downmix signal 103 may comprise a representation of W channel and one or more residual signals (in particular, the one or more residual signals having the highest energy).”
“[0119] The upmixing metadata 105 may be determined in dependence of whether the current frame is an active frame or an inactive frame. In particular, the set of parameters, which is comprised within the upmixing metadata 105 may depend on whether the current frame is an active frame or an inactive frame. If the current frame is an active frame, the set of parameters of the upmixing parameters 105 may be larger and/or may comprise a higher number of different parameters than if the current frame is an inactive frame.”)
Claim 18
Regarding Claim 18, ECKERT teaches 18.An audio encoder according to claim 1, wherein the transport signal generator is configured to generate the two or more transport channels of the transport signal from the audio input comprising by downmixing at least one of a plurality of audio input objects and a plurality of audio input channels to acquire a downmix as the transport signal, which comprises two or more downmix channels as the two or more transport channels.
(“[0022] FIG. 1 illustrates an encoding unit 100 and a decoding unit 150 for encoding and decoding a multi-channel input signal 101, which may comprise an SR signal. In particular, the multi-channel input signal 101 may comprise (possibly a combination of) one or more mono signals, one or more stereo signals, one or more binaural signal, one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects, and/or one or more SR signals. The different signal components may be considered to be individual channels of the multi-channel input signal 101.
[0023] The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels. The downmix signal 103 may itself be an SR signal, notably a first order ambisonics (FOA) signal, if the input signal 101 comprises a HOA signal. Downmixing may be performed in the subband domain or QMF domain (e.g., using 10 or more subbands).”
“[0091] For the two channel downmix example (4-2-4, for a first order soundfield), the comfort noise parameters for the mono dowmmix (W′) channel and for one prediction channel may be provided to the decoding unit 150. The decoding unit 150 may apply a method for generating FOA spatial comfort noise from a two channel downmix 103 and from the SPAR metadata 105. The two downmix channels may be uncorrelated comfort noise signals, one having the spectrum shaped according to the original W channel representation and the other one having the spectrum shaped according to the original residual channel.”)
Claim 22
Regarding Claim 22, ECKERT teaches 22. An audio encoder according to claim 1, wherein the transport signal generator is configured to encode the audio input by applying Code-Excited Linear Prediction or by applying a Modified Discrete Cosine Transform or by applying a combination of the Code-Excited Linear Prediction and of the Modified Discrete Cosine Transform.
(“[0025] In addition, the encoding unit 100 may comprise a coding module 140 which is configured to perform waveform encoding (e.g., EVS encoding) of the downmix signal 103, thereby providing coded audio data 106. Each channel of the downmix signal 103 may be encoded using a mono waveform encoder, thereby enabling an efficient encoding. Furthermore, the encoding unit 100 comprises a quantization module 141 which is configured to quantize the SPAR metadata 105 and to perform entropy encoding of the (quantized) SPAR metadata 105, thereby providing coded metadata 107. The coded audio data 106 and the coded metadata 107 may be inserted into a bitstream. Coding of the downmix signal 103 and/or of the SPAR metadata 105 is typically controlled using a mode and/or bitrate control module 142.”
EVS encoding teaches “… linear prediction” and “… cosine transform”)
Claim 24
Regarding Claim 24, ECKERT teaches 24. A system, comprising: an audio encoder according to an audio encoder according to
an audio decoder,
(“[0022] FIG. 1 illustrates an encoding unit 100 and a decoding unit 150 for encoding and decoding a multi-channel input signal 101, which may comprise an SR signal. In particular, the multi-channel input signal 101 may comprise (possibly a combination of) one or more mono signals, one or more stereo signals, one or more binaural signal, one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects, and/or one or more SR signals. The different signal components may be considered to be individual channels of the multi-channel input signal 101.”)
wherein the audio decoder comprises:
an input interface for receiving a bitstream which depends on audio content comprising at least one of a plurality of audio objects and a plurality of audio channels;
(“[0125] The method 600 may further comprise encoding 604 the upmixing metadata 105 into a bitstream (wherein the bitstream may be transmitted or provided to a corresponding decoding unit 150). The set of parameters of the upmixing metadata 105 may be entropy encoded to provide coded metadata 107 to be inserted into the bitstream. As a result of this, an efficient encoding of spatial background noise is provided.”
“[0022] FIG. 1 illustrates an encoding unit 100 and a decoding unit 150 for encoding and decoding a multi-channel input signal 101, which may comprise an SR signal. In particular, the multi-channel input signal 101 may comprise (possibly a combination of) one or more mono signals, one or more stereo signals, one or more binaural signal, one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects, and/or one or more SR signals. The different signal components may be considered to be individual channels of the multi-channel input signal 101.
[0023] The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels. The downmix signal 103 may itself be an SR signal, notably a first order ambisonics (FOA) signal, if the input signal 101 comprises a HOA signal. Downmixing may be performed in the subband domain or QMF domain (e.g., using 10 or more subbands).”)
wherein a transport signal comprising two or more transport channels is encoded within the bitstream, and the audio content is encoded within the transport signal; or
(“[0022] FIG. 1 illustrates an encoding unit 100 and a decoding unit 150 for encoding and decoding a multi-channel input signal 101, which may comprise an SR signal. In particular, the multi-channel input signal 101 may comprise (possibly a combination of) one or more mono signals, one or more stereo signals, one or more binaural signal, one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects, and/or one or more SR signals. The different signal components may be considered to be individual channels of the multi-channel input signal 101.
[0023] The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels. The downmix signal 103 may itself be an SR signal, notably a first order ambisonics (FOA) signal, if the input signal 101 comprises a HOA signal. Downmixing may be performed in the subband domain or QMF domain (e.g., using 10 or more subbands).”)
wherein information on a background noise is encoded within the bitstream instead of the transport signal,
(“[0057] Hence, the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames). Furthermore, the encoding unit 100 may be configured to determine and encode data for comfort noise generation within a “SID” frame (which corresponds e.g., to the current S frame of a series of S frames). The SID frames may be sent repeatedly, in particular periodically, for a series of S frames. By way of example, a SID frame may be sent every 8.sup.th frame (which corresponds to a time interval of 160 ms between subsequent SID frames, when using 20 ms frames). No data may be transmitted during the one or more following S frames of the series of S frames. Hence, the encoding unit 100 may be configured to perform DTX (discontinuous transmission) or to switch to a DTX mode.
[0058] In other words, the encoding unit 100 may be configured to send audio data 106 and encoded metadata 107 to the decoding unit 150 for every active frame. On the other hand, the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) for a fraction of the inactive frames (i.e., for the SID frames). For the remaining inactive frames (i.e., for the ND frames), no data may be sent at all (not even encoded metadata 107).”)
wherein the information on the background noise comprises information on a background noise of at least one of the two or more transport channels or information on a background noise of a derived signal which depends on at least one of the two or more transport channels; and
(“[0092] For the SID frames, two independent encoder module 140 instances encode spectral information regarding the mono (W′) channel and spectral information regarding the residual channel, respectively. Furthermore, two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds. The uncorrelated comfort noise signals may be spectrally shaped based on the representation of the W channel and the residual channel in the uncoded downmix, respectively. The reconstructed W channel may be referred to as W.sub.CNG and the reconstructed residual channel may be referred to as PING.
[0093] PING typically is a better approximation of the original uncorrelated residual channel compared to decorrelating W.sub.CNG and applying decorrelating coefficients (as done in the full parametric approach, which makes use of a single downmix channel only). As a result of this, the perceptual quality of the background noise is typically higher, when using a multi-channel downmix signal 103.”)
a renderer for generating one or more audio output signals depending on the audio content being encoded with the bitstream;
(“[0029] In addition, the decoding unit 150 comprises a reconstruction module 170 which is configured to derive a reconstructed multi-channel signal 111 from the SPAR metadata 105 and from the reconstructed downmix signal 114. The reconstructed multi-channel signal 111 may comprise a reconstructed SR signal. In particular, the reconstructed multi-channel signal 111 may comprise the same types of channels as the multi-channel input signal 101. The reconstructed multi-channel signal 111 may be used for speaker rendering, for headphone rendering and/or for SR rendering.”)
wherein, if the transport signal comprising the two or more transport channels is encoded within the bitstream, the renderer is configured to generate the one or more audio output signals depending on the two or more transport channels, and
(“[0028] The decoding unit 150 of FIG. 1 comprises a decoding module 160 which is configured to derive a reconstructed downmix signals 114 from the coded audio data 106. Furthermore, the decoding unit 150 comprises a metadata decoding module 161 which is configured to derive the SPAR metadata 105 from the coded metadata 107.
[0029] In addition, the decoding unit 150 comprises a reconstruction module 170 which is configured to derive a reconstructed multi-channel signal 111 from the SPAR metadata 105 and from the reconstructed downmix signal 114. The reconstructed multi-channel signal 111 may comprise a reconstructed SR signal. In particular, the reconstructed multi-channel signal 111 may comprise the same types of channels as the multi-channel input signal 101. The reconstructed multi-channel signal 111 may be used for speaker rendering, for headphone rendering and/or for SR rendering.”
“[0092] For the SID frames, two independent encoder module 140 instances encode spectral information regarding the mono (W′) channel and spectral information regarding the residual channel, respectively. Furthermore, two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds. The uncorrelated comfort noise signals may be spectrally shaped based on the representation of the W channel and the residual channel in the uncoded downmix, respectively. The reconstructed W channel may be referred to as W.sub.CNG and the reconstructed residual channel may be referred to as PING.”)
wherein, if the information on the background noise is encoded within the bitstream instead of the transport signal, the renderer is configured to generate the one or more audio output signals depending on the information on the background noise,
(“[0095] Since the downmix signals 103 are continuously running with the same downmix configuration in active and inactive frames, background noise typically sounds smooth even during transition frames. Furthermore, since the decoding unit 150 is using the prediction coefficients and the decorrelation coefficients computed by the SPAR encoder 120, spatial properties are replicated in the comfort noise which is generated by the SPAR decoder 150.”
See paragraphs 91-94 as well.)
wherein the audio encoder is configured to generate a bitstream from audio input,
(“[0025] In addition, the encoding unit 100 may comprise a coding module 140 which is configured to perform waveform encoding (e.g., EVS encoding) of the downmix signal 103, thereby providing coded audio data 106. Each channel of the downmix signal 103 may be encoded using a mono waveform encoder, thereby enabling an efficient encoding. Furthermore, the encoding unit 100 comprises a quantization module 141 which is configured to quantize the SPAR metadata 105 and to perform entropy encoding of the (quantized) SPAR metadata 105, thereby providing coded metadata 107. The coded audio data 106 and the coded metadata 107 may be inserted into a bitstream. Coding of the downmix signal 103 and/or of the SPAR metadata 105 is typically controlled using a mode and/or bitrate control module 142.”
See paragraphs 22-24 as well)
and wherein the audio decoder is configured to generate one or more audio output signals from the bitstream.
(“[0011] According to another aspect, a decoding unit for decoding a bitstream which is indicative of a reconstructed multi-channel signal comprising N channels is described. The reconstructed signal comprises a sequence of frames. The decoding unit is configured to determine a reconstructed downmix signal, wherein the reconstructed downmix signal comprises less than or equal to N channels. The decoding unit is further configured to determine, based on the bitstream, whether a current frame of the signal is an active frame or an inactive frame. In addition, the decoding unit is configured to generate the reconstructed multi-channel signal based on the reconstructed downmix signal and based on upmixing metadata comprised within the bitstream. The reconstructed multi-channel signal may be generated in dependence of whether the current frame is an active frame or an inactive frame.”)
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 4 are rejected under 35 U.S.C. 103 as obvious over US 20230215445, (ECKERT) in view of WO 2022042908 (RAVELLI).
Claim 4
Regarding Claim 4, ECKERT does not explicitly teach all of 4. An audio encoder according to claim 3, wherein the voice activity determiner is configured to determine that the transport signal exhibits voice activity, if at least one of the two or more transport channels of the transport signal exhibits voice activity, and wherein the voice activity determiner is configured to determine that the transport signal does not exhibit voice activity, if none of the two or more transport channels of the transport signal exhibits voice activity.
However, RAVELLI teaches 4. An audio encoder according to claim 3, wherein the voice activity determiner is configured to determine that the transport signal exhibits voice activity, if at least one of the two or more transport channels of the transport signal exhibits voice activity, and
(page 22 paragraph 3 “Notably, in examples a frame is classified (at stage 381) as inactive frame only if both channels 301 and 303 are classified as inactive by stages 380-1 and 380-3, respectively. Therefore, problems are avoided in the activity detection decision as discussed above. In particular, it is not necessary to signal the classification of active/inactive for each channel for each frame (thereby reducing the signalling), and a synchronization between the channels is inherently obtained. Further, where the decoder is as discussed in the present document, it is possible to make use of the coherence between the first and second channels 301 and 303 and to generate some noise signals, which are correlated/decorrelated according to the coherence obtained for the signal 304. Now, the elements of the encoder 300 (300a, 300b) which are used for encoding the inactive frame are discussed in detail. As explained, any other technique may be used for encoding the active frames 308, and is therefore not discussed here.”)
wherein the voice activity determiner is configured to determine that the transport signal does not exhibit voice activity, if none of the two or more transport channels of the transport signal exhibits voice activity.
(page 22 paragraph 3 “Notably, in examples a frame is classified (at stage 381) as inactive frame only if both channels 301 and 303 are classified as inactive by stages 380-1 and 380-3, respectively. Therefore, problems are avoided in the activity detection decision as discussed above. In particular, it is not necessary to signal the classification of active/inactive for each channel for each frame (thereby reducing the signalling), and a synchronization between the channels is inherently obtained. Further, where the decoder is as discussed in the present document, it is possible to make use of the coherence between the first and second channels 301 and 303 and to generate some noise signals, which are correlated/decorrelated according to the coherence obtained for the signal 304. Now, the elements of the encoder 300 (300a, 300b) which are used for encoding the inactive frame are discussed in detail. As explained, any other technique may be used for encoding the active frames 308, and is therefore not discussed here.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of RAVELLI to provide a “4. An audio encoder according to claim 3, wherein the voice activity determiner is configured to determine that the transport signal exhibits voice activity, if at least one of the two or more transport channels of the transport signal exhibits voice activity, and wherein the voice activity determiner is configured to determine that the transport signal does not exhibit voice activity, if none of the two or more transport channels of the transport signal exhibits voice activity.” Doing so would preserve the stereo image, as recognized by RAVELLI. (page 3 paragraph 3 ).
Backup Rejection for claim 4 only (see below in view of NORVELL):
Claims 4, are rejected under 35 U.S.C. 103 as obvious over US 20230215445, (ECKERT) in view of WO 2017202680 A1 (NORVELL ).
Claim 4
Regarding Claim 4, ECKERT does not explicitly teach all of 4. An audio encoder according to claim 3, wherein the voice activity determiner is configured to determine that the transport signal exhibits voice activity, if at least one of the two or more transport channels of the transport signal exhibits voice activity, and wherein the voice activity determiner is configured to determine that the transport signal does not exhibit voice activity, if none of the two or more transport channels of the transport signal exhibits voice activity.
However, NORVELL teaches 4. An audio encoder according to claim 3, wherein the voice activity determiner is configured to determine that the transport signal exhibits voice activity, if at least one of the two or more transport channels of the transport signal exhibits voice activity, and
(page 8 paragraph 4 ” Figure 7a illustrates another example of realization of a multi-channel voice activity detector using a monophonic voice activity detector 603. Monophonic voice detection is run on each channel individually, producing a voice activity decision per channel. The decision is then aggregated in the decision aggregator 701, for instance by using majority decision. The decision may also be biased towards a certain decision,e.g. if any voice activity detector signals active voice, the overall decision is active voice.”)
wherein the voice activity determiner is configured to determine that the transport signal does not exhibit voice activity, if none of the two or more transport channels of the transport signal exhibits voice activity.
(page 8 paragraph 4 ” Figure 7a illustrates another example of realization of a multi-channel voice activity detector using a monophonic voice activity detector 603. Monophonic voice detection is run on each channel individually, producing a voice activity decision per channel. The decision is then aggregated in the decision aggregator 701, for instance by using majority decision. The decision may also be biased towards a certain decision,e.g. if any voice activity detector signals active voice, the overall decision is active voice.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of NORVELL to provide a “4. An audio encoder according to claim 3, wherein the voice activity determiner is configured to determine that the transport signal exhibits voice activity, if at least one of the two or more transport channels of the transport signal exhibits voice activity, and wherein the voice activity determiner is configured to determine that the transport signal does not exhibit voice activity, if none of the two or more transport channels of the transport signal exhibits voice activity.” Doing so would improve the accuracy of the spatial SAD, as recognized by NORVELL. (page 9 paragraph 5 ).
Claims 6 are rejected under 35 U.S.C. 103 as obvious over US 20230215445, (ECKERT) in view of WO-2012066727-A1 (OSHIKIRI,).
Claim 6
Regarding Claim 6, ECKERT teaches wherein the audio encoder comprises an information generator for generating the information on the background noise as information on the background noise of the mono signal.
(“[0057] Hence, the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames). Furthermore, the encoding unit 100 may be configured to determine and encode data for comfort noise generation within a “SID” frame (which corresponds e.g., to the current S frame of a series of S frames). The SID frames may be sent repeatedly, in particular periodically, for a series of S frames. By way of example, a SID frame may be sent every 8.sup.th frame (which corresponds to a time interval of 160 ms between subsequent SID frames, when using 20 ms frames). No data may be transmitted during the one or more following S frames of the series of S frames. Hence, the encoding unit 100 may be configured to perform DTX (discontinuous transmission) or to switch to a DTX mode.”
“[0092] For the SID frames, two independent encoder module 140 instances encode spectral information regarding the mono (W′) channel and spectral information regarding the residual channel, respectively. Furthermore, two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds. The uncorrelated comfort noise signals may be spectrally shaped based on the representation of the W channel and the residual channel in the uncoded downmix, respectively. The reconstructed W channel may be referred to as W.sub.CNG and the reconstructed residual channel may be referred to as PING.”)
ECKERT does not explicitly teach all of 6. An audio encoder according to claim 1, wherein the audio encoder comprises a mono signal generator for generating, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, the derived signal as a mono signal from at least one of the two or more transport channels, and
However, OSHIKIRI, teaches
6. An audio encoder according to claim 1, wherein the audio encoder comprises a mono signal generator for generating, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, the derived signal as a mono signal from at least one of the two or more transport channels, and
(page 3 paragraph 3 “The VAD unit 101 analyzes an input signal (a stereo signal composed of an L channel signal and an R channel signal) and determines whether the input signal of the current frame is a voice part or a non-voice part. As non-speech parts, backgrounds typified by silent parts that are perceptually silent because the signal amplitude is very small, and environmental sounds that are perceived in daily life (duct operating sounds and car running sounds) This corresponds to the noise part. Hereinafter, the background noise part will be described as a representative of the non-voice part. This analysis uses at least the energy of the signal. Then, if the input signal of the current frame is determined to be a voice part as a result of the analysis, the VAD part 101 generates VAD data indicating that the input signal of the current frame is a voice part. If the input signal is determined to be the background noise part, VAD data indicating that the input signal of the current frame is the background noise part is generated. Then, the VAD unit 101 outputs the generated VAD data to the switching units 102 and 105 and the multiplexing unit 106.”
page 10 paragraph 5 “The monaural signal generation unit 501 generates a monaural signal by downmixing the L channel signal and the R channel signal constituting the stereo signal. Then, the monaural signal generation unit 501 outputs the generated monaural signal to the spectrum parameter analysis unit 502.”)
OSHIKIRI further teaches wherein the audio encoder comprises an information generator for generating the information on the background noise as information on the background noise of the mono signal.
(page 10 paragraph 6-7 “The spectrum parameter analysis unit 502 performs LPC analysis on the monaural signal and generates an LSP parameter indicating the spectrum characteristic of the monaural signal. For example, the LSP parameter of the monaural signal can be obtained by converting the LPC coefficient obtained by analyzing the monaural signal. Then, the spectral parameter analysis unit 502 outputs the LSP parameter of the monaural signal to the spectral parameter quantization unit 503. The spectral parameter quantization unit 503 quantizes (encodes) the LSP parameter of the monaural signal based on vector quantization, scalar quantization, or a combination of these quantization methods. The spectral parameter quantization unit 503 outputs the monaural signal spectral parameter quantization information obtained by the quantization process to the multiplexing unit 312.”
Page 13 paragraph 5 “Therefore, in this embodiment, even when only the LPC coefficient of a monaural signal is transmitted, high quality background noise can be generated, and the bit rate can be further reduced as compared with the first embodiment. Can do.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of OSHIKIRI, to provide a “6. An audio encoder according to claim 1, wherein the audio encoder comprises a mono signal generator for generating, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, the derived signal as a mono signal from at least one of the two or more transport channels, and wherein the audio encoder comprises an information generator for generating the information on the background noise as information on the background noise of the mono signal.” Doing so would reduce the bitrate without degrading the quality, as recognized by OSOSHIKIRI, (page 10 paragraph 2 ).
Claims 7, 8, 9 are rejected under 35 U.S.C. 103 as obvious over ECKERT in view of OSHIKIRI in further view of US 20220174443, (LAITINEN) in further view of US 20200126575 A1, (Vasilache)
Claim 7
Regarding Claim 7, ECKERT in view of OSHIKIRI does not explicitly teach all of 7. An audio encoder according to claim 6, wherein the mono signal generator is configured to generate the mono signal by adding the two or more transport channels or by adding two or more channels derived from the two or more transport channels, or wherein the mono signal generator is configured to generate the mono signal by choosing that transport channel of the two or more transport channels which exhibits a higher energy.
However, LAITINEN, teaches 7. An audio encoder according to claim 6, wherein the mono signal generator is configured to generate the mono signal by adding the two or more transport channels or by adding two or more channels derived from the two or more transport channels,
(“[0258] The prototype signals creator 1511 in some embodiments is configured to create a prototype signal for a mono transport audio signal using the two transport audio signals, based on the received transport audio signal type. For example the following may be used.
[0259] If T(n)=“spaced”,
M.sub.proto(b,n)=S.sub.0(b,n).
[0260] If T(n)=“downmix” or T(n)=“coincident”,
M.sub.proto(b,n)=S.sub.0(b,n)+S.sub.1(b,n).
[0261] In some embodiments the downmixer 1303 comprises a target energy determiner 1501. The target energy determiner 1501 is configured to receive the T/F-domain transport audio signals 504 and generate a target energy value as the sum of the energies of the transport audio signals”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT in view of OSHIKIRI to incorporate the teachings of LAITINEN, to provide a “7. An audio encoder according to claim 6, wherein the mono signal generator is configured to generate the mono signal by adding the two or more transport channels or by adding two or more channels derived from the two or more transport channels,” Doing so would avoid artefacts, as recognized by LAITINEN,. (paragraph 203).
ECKERT in view of OSHIKIRI in view of LAITINEN does not explicitly teach all of or wherein the mono signal generator is configured to generate the mono signal by choosing that transport channel of the two or more transport channels which exhibits a higher energy.
However, Vasilache, teaches or wherein the mono signal generator is configured to generate the mono signal by choosing that transport channel of the two or more transport channels which exhibits a higher energy.
(“[0043] As an example of decomposition that depends on one or more characteristics of the input audio signal 115, the signal for the first channel 223-1 may be derived on basis of the one of the left channel 115-1 signal and the right channel 115-2 signal that has a higher energy whereas the signal for the second channel 223-2 may be derived on basis of the other one of the left channel 115-1 and right channel 115-2 signals. The derivation may comprise, for example, predefined or adaptive scaling and/or filtering of the respective one of the left channel 115-1 and right channel 115-2 signals. In a variation of this example, the higher-energy one of the left channel 115-1 and the right channel 115-2 signals may be provided as such as the first channel 223-1 signal while the other one is provided as such as the second channel 223-2 signal.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT in view of OSHIKIRI in view of LAITINEN to incorporate the teachings of Vasilache, to provide a “or wherein the mono signal generator is configured to generate the mono signal by choosing that transport channel of the two or more transport channels which exhibits a higher energy.” Doing so would convey a larger portion of the energy carried by the input chanel, as recognized by Vasilache,. (paragraph 45).
Claim 8
Regarding Claim 8, ECKERT in view of OSHIKIRI in view of LAITINEN to in view of Vasilache, further ECKERT teaches 8. An audio encoder according to claim 6, wherein the information generator is configured to generate the information on a background noise of the mono signal as the information on the mono signal.
(“[0057] Hence, the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames). Furthermore, the encoding unit 100 may be configured to determine and encode data for comfort noise generation within a “SID” frame (which corresponds e.g., to the current S frame of a series of S frames). The SID frames may be sent repeatedly, in particular periodically, for a series of S frames. By way of example, a SID frame may be sent every 8.sup.th frame (which corresponds to a time interval of 160 ms between subsequent SID frames, when using 20 ms frames). No data may be transmitted during the one or more following S frames of the series of S frames. Hence, the encoding unit 100 may be configured to perform DTX (discontinuous transmission) or to switch to a DTX mode.”
“[0092] For the SID frames, two independent encoder module 140 instances encode spectral information regarding the mono (W′) channel and spectral information regarding the residual channel, respectively. Furthermore, two independent instances of the decoding unit 150 may generate uncorrelated comfort noise signals with different seeds. The uncorrelated comfort noise signals may be spectrally shaped based on the representation of the W channel and the residual channel in the uncoded downmix, respectively. The reconstructed W channel may be referred to as W.sub.CNG and the reconstructed residual channel may be referred to as PING.”)
Claim 9
Regarding Claim 9, ECKERT in view of OSHIKIRI in view of LAITINEN to in view of Vasilache, further ECKERT teaches 9. An audio encoder according to claim 8, wherein the information generator is configured to generate a silence insertion description of the background noise of the mono signal as the information on the background noise of the mono signal.
(“[0056] For a discontinuous transmission (DTX) system, where the actual bitrate of the codec may be substantially reduced during inactive frames by only sending noise shaping parameters and by assuming that background noise characteristics do not change as frequent as active speech or audio frames, the above sequence may be translated into the following sequence of frames by the encoding unit 100: AB-AB-SID-ND-ND-ND-ND-ND-ND-ND-SID-ND-ND-ND-ND-ND-ND-ND-SID-ND-ND-ND-ND-AB-AB-AB-AB
wherein “AB” indicates an encoder bitstream for an active frame, wherein “SID” indicates a silence indicator frame, which comprises a series of bits for comfort noise generation, and wherein “ND” indicates no data frames, i.e., nothing is transmitted to the decoding unit 150 during these frames.
[0057] Hence, the encoding unit 100 may be configured to classifying the different frames of the input signal 101 into active (A) or silent (S) frames (which are also referred to as inactive frames). Furthermore, the encoding unit 100 may be configured to determine and encode data for comfort noise generation within a “SID” frame (which corresponds e.g., to the current S frame of a series of S frames). The SID frames may be sent repeatedly, in particular periodically, for a series of S frames. By way of example, a SID frame may be sent every 8.sup.th frame (which corresponds to a time interval of 160 ms between subsequent SID frames, when using 20 ms frames). No data may be transmitted during the one or more following S frames of the series of S frames. Hence, the encoding unit 100 may be configured to perform DTX (discontinuous transmission) or to switch to a DTX mode.”
“[0090] For the one channel downmix example, the decoding unit 150 may be configured to generate a reconstructed downmix signal 114 based on the audio data 106. This reconstructed downmix signal 114 may be referred to as W.sub.CNG, which, during inactive frames, may include a parametric reconstruction of background noise present in the uncoded representation of the W channel in the downmix using white noise as an excitation signal and using spectral shaping parameters coded by a mono audio codec (e.g., EVS). The three decorrelated channels for reconstructing the Y, X and Z channel signals may be generated from W.sub.CNG using decorrelators 201 (e.g., time domain or filterbank domain decorrelators). Alternatively, three decorrelated channels for reconstructing the Y, X and Z channel signals may be generated by generating uncorrelated comfort noise with different seeds and spectrally shaping the uncorrected comfort noise according to W.sub.CNG The SPAR metadata 105 may be applied to W.sub.CNG and the decorrelated channels to generate comfort noise in a soundfield format, having the spectral and spatial characteristics of the original background noise.”)
Claims 12 are rejected under 35 U.S.C. 103 as obvious over ECKERT in view of US 20250238122 (Tokozume,).
Claim 12
Regarding Claim 12, ECKERT does not explicitly teach all of 12. An audio encoder according to claim 10, wherein the audio input comprises the plurality of audio input objects, wherein the direction information comprises information on an azimuth angle and on an elevation angle of an audio input object of the plurality of audio input object of the audio input.
However, Tokozume, teaches
12. An audio encoder according to claim 10, wherein the audio input comprises the plurality of audio input objects, wherein the direction information comprises information on an azimuth angle and on an elevation angle of an audio input object of the plurality of audio input object of the audio input.
(“[0105] The output parameter here is at least any of three-dimensional position information indicating a position of an object in a three-dimensional space and a gain of audio data of the object. As an example, the three-dimensional position information is, for example, polar coordinates indicating a position of the object in a polar coordinate system including an azimuth angle “azimuth” indicating a position of the object in the horizontal direction, an elevation angle “elevation” indicating a position of the object in the vertical direction, and the like.”
“[0115] In the example illustrated in FIG. 4, as illustrated on the left side in the drawing, pieces of audio data of three objects of Objects 1 to 3 are input, and an azimuth angle “azimuth” and an elevation angle “elevation” as three-dimensional position information are output as output parameters of each of the objects.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of Tokozume, to provide a “12. An audio encoder according to claim 10, wherein the audio input comprises the plurality of audio input objects, wherein the direction information comprises information on an azimuth angle and on an elevation angle of an audio input object of the plurality of audio input object of the audio input.” Doing so would create high quality content, as recognized by Tokozume,. (paragraph 64 ).
Claims 14, 15, 19, 20 are rejected under 35 U.S.C. 103 as obvious over US 20230215445, (ECKERT) in view of WO 2022022876 A1 (FUCHS ).
Claim 14
Regarding Claim 14, ECKERT teaches 14.An audio encoder according to claim 1, wherein the audio input comprises the plurality of audio input objects,
and wherein the audio encoder comprises an inactive metadata generator for generating, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, metadata comprising quantized direction information and control parameters, for example comprising a scaling factor and/or either a coherence or a correlation.
(“[0068] The method for encoding spatial comfort noise may comprise VAD and/or SAD for a frame of the mono downmix signal 103 (e.g., the W channel signal for a soundfield signal). The encoding of spatial comfort noise parameters may be performed, if the frame is detected to be an inactive frame.”
"[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.
[0062] Hence a VAD and/or SAD may be used to classify the frames of a sequence of frames into active frames or inactive frames. Encoding and/or generating comfort noise may be applied to the inactive frames. The encoding of the comfort noise (notably the encoding of noise shaping parameters) within the encoding unit 100 may be performed such that the decoding unit 150 is enabled to generate high quality comfort noise for a soundfield. The comfort noise that is generated by the decoding unit 150 preferably matches the spectral and/or spatial characteristics of the background noise within the input signal 101. This does not necessarily imply the waveform reconstruction of the input background noise. The comfort noise generated by a soundfield decoding unit 150 for a series of inactive frames is preferably such that the comfort noise sounds continuous with regards of the noise within the directly preceding active frames. Hence, the transition between active and inactive frames at the decoding unit 150 is preferably smooth and non-abrupt.”)
ECKERT does not explicitly teach all of the plurality of audio input objects, quantizing direction information, and the specific parameters.
However, FUCHS , teaches 14.An audio encoder according to claim 1, wherein the audio input comprises the plurality of audio input objects,
(page 5 paragraph 1 “The audio signal encoder may be configured to encode, for the first frame, the audio signal using a time domain or frequency domain encoding mode, the encoded audio signal com prising, for example, encoded time domain samples, encoded spectral domain samples, encoded LPC domain samples and side information obtained from components of the au dio signal or obtained from one or more transport channels derived from the components of the audio signal, for example, by a downmixing operation. --> The audio signal may comprise an input format being a first order Ambisonics format, a higher order Ambisonics format, a multi-channel format associated with a given loud speaker setup, such as 5.1 or 7.1 or 7.1 + 4, or one or more audio channels representing one or several different audio objects localized in a space as indicated by information included in associated metadata, or an input format being a metadata associated spatial audio representation, wherein the soundfield parameter generator is configured for determining the first soundfield parameter representation and the second soundfield representation so that the parameters represent a soundfield with respect to a defined listener position, or wherein the audio signal comprises a microphone signal as picked up by real mi crophone or a virtual microphone or a synthetically created microphone signal e.g. being in a first order Ambisonics format, or a higher order Ambisonics format.”)
and wherein the audio encoder comprises an inactive metadata generator for generating, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, metadata comprising quantized direction information and control parameters, for example comprising a scaling factor and/or either a coherence or a correlation.
(page 22 paragraph 2 “The Voice Activity Detector (or more in general an activity detector) 320 may then be applied on the input signal 302 and/or the transport channels 326 produced by the audio scene analyzer. The transport channels are less than the number of input channels; usually a mono-downmix, a stereo downmix, an A-format, or a First Order Ambisonics signal. Based on the VAD decision the current frame under process is defined as active (306, 326) or inactive (308, 328). In case of active frames (306, 326), a conventional speech or audio encoding of the transport channels is performed. The resulting code data are then combined with the active spatial parameters 316. In case of inactive frames (308, 328), a silence information description 328 of the transport channels 324 is produced episodically, usually at regular frame intervals during inactive phase, for example at every 8 active frames (306, 326, 346). The transport channel SID (328, 348) may then be amended in the multiplexer (encoded signal former) 370 with the inactive spatial parameters. In case the inactive spatial parameters 318 are null, only the transport channel SID 348 is then transmitted. The overall SID can usually be a very low bit-rate description, which is for example as low as 2.4 or 4.25 kbps. The average bit-rate is even more reduced in the inactive phase since most of the time no transmission is done and no data are sent.”
Page 14 paragraph 1 “…The active spatial parameters 316 may be provided to an active spatial metadata encoder 396 and the inactive spatial parameters 318 may be provided to an inactive spatial metadata encoder 398. The resulting are first and second soundfield parameter representations (316, 318, complexively indicated with 314) which may be encoded in the bitstream 304 (e.g., through the encoder signal former 370) and stored for being subsequently played back by a decoder. Whether the active spatial metadata encoder 396 or the inactive spatial parameters 318 is to encode a frame, this may be controlled by a control such as the control 321 in Fig. 3 (the deviator 322 is not shown in Fig. 2), e.g. thorough the classification operated by the activity de tector. (It is to be noted that the encoders 396, 398 may also perform a quantization, in some examples)…”
Page 23 paragraph 2 “The inactive spatial parameters 318 can consist of one of multiple directions in frequency bands and associated energy ratios in frequency bands corresponding to the ratio of one directional component over the total energy. In case of one direction, as in a preferred embodiment, the energy ratio can be replaced by the diffuseness, which is complementary to the ratio of energy and then follow the original DirAC set of parameters. Since the directional component(s) is(are) in general expected to be less relevant than the diffuse part in inactive frames, it can be also transmitted on fewer bits using a coarser quantization scheme such as in active frames and/or by averaging the direction over time or frequency for getting a coarser time and /or frequency reso lution. In a preferred embodiment, the direction may be sent every 20 ms instead of 5 ms for active frames but using the same frequency resolution of 5 non-uniform bands.”
Page 23 paragraph 5 “Moreover, one can consider to transmit an inter-channel coherence if input channels correspond to channels positioned the spatial domain. Inter-channel level differences are also an alternative to the directions.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of FUCHS , to provide a “the plurality of audio input objects, quantizing direction information, and the specific parameters.” .” Doing so would reduce the bitrate demand for transmitting conversational speech, as recognized by FUCHS ,. (page 3 paragraph 1 ).
Claim 15
Regarding Claim 15, ECKERT does not explicitly teach all of 15. An audio encoder according to claim 13, wherein the direction information that is generated by the inactive metadata generator differs in a quantization resolution from the metadata that is generated by the active metadata generator.
However, FUCHS , teaches 15. An audio encoder according to claim 13, wherein the direction information that is generated by the inactive metadata generator differs in a quantization resolution from the metadata that is generated by the active metadata generator.
(page 5 last paragraph “The soundfield parameter generator may be configured for determining the second sound- field parameter representation for the second frame --> using spatial parameters for one or more directions in frequency bands and asso ciated energy ratios in frequency bands corresponding to a ratio of one directional component over a total energy, or to determine a diffuseness parameter indicating a ratio of diffuse sound or direct sound, or to determine a direction information using a coarser quantization scheme com pared to a quantization in the first frame, or using an averaging of a direction over time or frequency for obtaining a coarser time or frequency resolution, or to determine a soundfield parameter representation for one or more inactive frames with the same frequency resolution as in the first soundfield parameter repre sentation for an active frame, and with a time occurrence that is lower than the time occurrence for active frames with respect to a direction information in the soundfield parameter representation for the inactive frame, or to determine the second soundfield parameter representation having a diffuseness parameter, where the diffuseness parameter is transmitted with the same time or frequency resolution as for active frames, but with a coarser quantization, or to quantize a diffuseness parameter for the second soundfield representation with a first number of bits, and wherein only a second number of bits of each quantization index is transmitted, the second number of bits being smaller than the first number of bits, or to determine, for the second soundfield parameter representation, an inter-channel coherence if the audio signal has input channels corresponding to channels posi tioned in a spatial domain or inter-channel level differences if the audio signal has input channels corresponding to channels positioned in the spatial domain, or to determine a surround coherence being defined as a ratio of diffuse energy being coherent in a soundfield represented by the audio signal.”)
see claim 14 for rationale.
Claim 19
Regarding Claim 19, ECKERT teaches 19. An audio encoder according to claim 10, wherein the transport signal generator is configured to generate the two or more transport channels of the transport signal from the audio input comprising by downmixing at least one of a plurality of audio input objects and a plurality of audio input channels to acquire a downmix as the transport signal, which comprises two or more downmix channels as the two or more transport channels,
(“[0022] FIG. 1 illustrates an encoding unit 100 and a decoding unit 150 for encoding and decoding a multi-channel input signal 101, which may comprise an SR signal. In particular, the multi-channel input signal 101 may comprise (possibly a combination of) one or more mono signals, one or more stereo signals, one or more binaural signal, one or more (conventional) multi-channel signals (such as a 5.1 or a 7.1 signal), one or more audio objects, and/or one or more SR signals. The different signal components may be considered to be individual channels of the multi-channel input signal 101.
[0023] The encoding unit 100 comprises a spatial analysis and downmix module 120 configured to downmix the multi-channel input signal 101 to a downmix signal 103 comprising one or more channels. The downmix signal 103 may itself be an SR signal, notably a first order ambisonics (FOA) signal, if the input signal 101 comprises a HOA signal. Downmixing may be performed in the subband domain or QMF domain (e.g., using 10 or more subbands).”
[0034] The encoding unit 100 may be configured to convert an FOA input signal 101 into a downmix signal 103 and parameters, i.e., SPAR metadata 105, used to regenerate the input signal 101 at the decoding unit 150. The number of channels of the downmix signal 103 may vary from 1 to 4 channels. The parameters may include prediction parameters Pr, cross-prediction parameters C and/or decorrelation parameters P. These parameters may be calculated from the covariance matrix of a windowed input signal 101. Furthermore, the parameters may be calculated in a specified number of subbands. In the case of comfort noise, a reduced number of subbands (also referred to as frequency bands) may be used, e.g., 6 subbands instead of 12 subbands.)
ECKERT does not explicitly teach all of wherein, if the audio input within the transport signal does not exhibit voice activity, the direction information quantizer is configured to determine the quantized direction information such that a quantization resolution of the quantized direction information
However, FUCHS , teaches wherein, if the audio input within the transport signal does not exhibit voice activity, the direction information quantizer is configured to determine the quantized direction information such that a quantization resolution of the quantized direction information is different from a quantization resolution used for computing the downmix.
(page 12 second paragraph “The apparatus 300 may include an activity detector 320. The activity detector 320 may analyze the input audio signal (either in its input version 302 or in its downmix version 324), to determine, depending on the audio signal (302 or 324) whether a frame is an active frame 306 or an inactive frame 308, hence performing a classification on the frame. As can be seen from Fig. 3, the active detector 320 can be assumed as controlling (e.g. through the control 321 ) a first deviator 322 and a second deviator 322a. The first deviator 322 may select between the active spatial parameter 316 (first soundfield parameter representation) and the inactive spatial parameters 318 (second soundfield parameter representation). Therefore, the activity detector 320 may decide whether the active spatial parameters 316 or the inactive spatial parameters 318 are to be outputted (e.g. signalled in the bitstream 304). The same control 321 may control the second deviator 322a, which may select between outputting the first frame 326 (306) in the transport channel 324, or the second frame 328 (308) (e.g. parametric description) in the transport channel 326. The activities of the first and second deviators 322 and 322a are coordinated with each other: when the active spatial parameters 316 are outputted, then the transport channels 326 of the first frame 306 are also outputted, and when the inactive spatial parameters 318 are outputted, then the transport channels 328 of the first frame 306 the transport channels are outputted. This is because the active spatial parameters 316 (first soundfield parameter representation) describe spatial charac teristics of the first frame 306, while the inactive spatial parameters 318 (second soundfield pa rameter representation) describes spatial characteristics of the second frame 308. --> The activity detector 320 may therefore basically decide which one among the first frame 306 (326, 346), and its related parameters (316), and the second frame 308 (328, 348), and its related parameters (318), are to be outputted. The activity detector 320 may also control the encoding of some signalling in the bitstream which signals whether the frame is an active or an inactive (other techniques maybe used).”
Page 22 second to last paragraph “In the preferred embodiment of the invention the transport channel SID 348 has a size of 2.4kbps and the overall SID including spatial parameters has a size of 4.25kbps. The computation of the inactive spatial parameters are described in Fig. 4 for DirAC having as input a multi-channel signal like FOA, which could directly derived from a higher order of Ambisonics (HOA), in Fig. 5 for MASA input format. As described earlier, the inactive spatial parameters 318 can be derived in parallel to the active spatial parameters 316, averaging and/or requantizing the already coded active spatial parameters 318. In case of multi-channel signal like FOA as input format 302, a filterbank analysis of the multi-channel signal 302 may be performed before computing the spatial parameters, direction and diffuseness, for each time and frequency tile. The metadata encoders 396, 398 could average the parameters 316, 318 over different frequency bands and/or time slots before applying a quantizer and a coding of the quantized parameters. Further inactive spatial metadata encoder can inherit from some of the quantized parameters derived in the active spatial metadata encoder to use them directly in the inactive spatial parameters or to requantize them. In case of MASA format (e.g. Fig. 5), first the input metadata may be read and provided the metadata encoders 396, 398 at a given time-frequency and bit depth resolution. The metadata encoder(s) 396, 398 will process then further by eventually converting some parameters, adapting --> their resolution (i.e. lowering the resolution for example averaging them) and requantizing them before coding them by an entropy coding scheme for example.”
Page 5 last paragraph “The soundfield parameter generator may be configured for determining the second sound- field parameter representation for the second frame --> using spatial parameters for one or more directions in frequency bands and asso ciated energy ratios in frequency bands corresponding to a ratio of one directional component over a total energy, or to determine a diffuseness parameter indicating a ratio of diffuse sound or direct sound, or to determine a direction information using a coarser quantization scheme com pared to a quantization in the first frame, or using an averaging of a direction over time or frequency for obtaining a coarser time or frequency resolution, or to determine a soundfield parameter representation for one or more inactive frames with the same frequency resolution as in the first soundfield parameter repre sentation for an active frame, and with a time occurrence that is lower than the time occurrence for active frames with respect to a direction information in the soundfield parameter representation for the inactive frame, or to determine the second soundfield parameter representation having a diffuseness parameter, where the diffuseness parameter is transmitted with the same time or frequency resolution as for active frames, but with a coarser quantization, or to quantize a diffuseness parameter for the second soundfield representation with a first number of bits, and wherein only a second number of bits of each quantization index is transmitted, the second number of bits being smaller than the first number of bits, or to determine, for the second soundfield parameter representation, an inter-channel coherence if the audio signal has input channels corresponding to channels posi tioned in a spatial domain or inter-channel level differences if the audio signal has input channels corresponding to channels positioned in the spatial domain, or to determine a surround coherence being defined as a ratio of diffuse energy being coherent in a soundfield represented by the audio signal.”
Also note everything after "within the transport signal does not exhibit voice activity, the direction information quantizer is configured to determine the quantized direction information" is intended result and therefore has not been considered.)
See claim 14 for rationale.
Claim 20
Regarding Claim 20, ECKERT in view of FUCHS further ECKERT teaches 20. An audio encoder according to claim 14, wherein the bitstream generator is configured to encode the control parameters within the bitstream, if the voice activity determiner has determined that the transport signal does not exhibit voice activity,
(“[0058] In other words, the encoding unit 100 may be configured to send audio data 106 and encoded metadata 107 to the decoding unit 150 for every active frame. On the other hand, the encoding unit 100 may be configured to send only encoded metadata 107 (and no audio data 106) for a fraction of the inactive frames (i.e., for the SID frames). For the remaining inactive frames (i.e., for the ND frames), no data may be sent at all (not even encoded metadata 107).
[0059] The encoded metadata 107 which is sent for a SID frame may be reduced and/or compressed with regards to the encoded metadata 107 which is sent for an active frame.”
“[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.”)
wherein the control parameters are suitable for steering a generation of an intermediate signal from random noise,
(“[0063] The decoding unit 150 may be configured to generate random white noise as an excitation signal. The excitation signal may comprise multiple channels of white noise, wherein the white noise in the different channels is typically uncorrelated from one another. The bitstream from the encoding unit 100 may only comprise noise shaping parameters (as encoded metadata 107), and the decoding unit 150 may be configured to shape the random white noise within the different channels (spectrally and spatially) using the noise shaping parameters that have been provided within the bitstream. By doing this, spatial comfort noise may be generated in an efficient manner.”)
wherein the control parameters either comprises a plurality of parameter values for a plurality of subbands, or wherein the control parameters are single broadband control parameters.
(“[0123] The set of parameters may comprise corresponding parameters for a number of different frequency bands. In other words, the parameters of a given type (e.g., the Pr, the C and/or the P parameters) may be determined for a plurality of different frequency bands (also referred to herein as subbands). The number of different frequency bands, for which the parameters are determined, may depend on whether the current frame is an active frame or an inactive frame. In particular, if the current frame is an active frame, the number of different frequency bands may be higher than if the current frame is an inactive frame.”)
Claims 17 are rejected under 35 U.S.C. 103 as obvious over US 20230215445, (ECKERT) in view of WO 2021019126 (VASILACHE ).
Claim 17
Regarding Claim 17, ECKERT does not explicitly teach all of 17. An audio encoder according to claim 1, wherein the audio input comprises a plurality of audio input objects and metadata being associated with the audio input objects.
However, VASILACHE, teaches 17. An audio encoder according to claim 1, wherein the audio input comprises a plurality of audio input objects and metadata being associated with the audio input objects.
(page 8 paragraph 2 “In addition to multi-channel input format audio signals an encoding system may also be required to encode audio objects representing various sound sources within a physical space. Each audio object can be accompanied, whether it is in the form of metadata or some other mechanism, by directional data in the form of azimuth and elevation values which indicate the position of an audio object within a physical space.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of VASILACHE , to provide a “17. An audio encoder according to claim 1, wherein the audio input comprises a plurality of audio input objects and metadata being associated with the audio input objects.” Doing so would define a parametric immersive format, as recognized by VASILACHE ,. (page 8 paragraph 5 ).
Claims 16 are rejected under 35 U.S.C. 103 as obvious over ECKERT in view of FUCHS in further view of WO 2022079049 A2, (EICHENSEER )
Claim 16
Regarding Claim 15, ECKERT teaches 16. An audio encoder according to claim 13, wherein the audio input comprises the plurality of audio input objects, and
wherein the audio encoder comprises an inactive metadata generator for generating, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, metadata comprising quantized direction information and control parameters, for example comprising a scaling factor and/or either a coherence or a correlation,
(“[0060] The encoding unit 100 may comprise a voice activity detector which is configured to switch the encoder to DTX mode. If the DTX flag (e.g., the CombinedVAD flag mentioned below) is set, then packets may be generated in a discontinuous mode based on an input frame, otherwise a frame may be coded as a speech and/or audio active frame.
[0061] The encoding unit 100 may be configured to determine a mono downmix signal 103 and the mono downmix signal 103 may be used to detect an inactive frame by operating a Signal Activity Detector or Voice Activity detector (SAD/VAD) on the mono downmix signal 103. For the example of a soundfield B-format input signal 101, the SAD/VAD may be operated on the representation of the W channel signal. In an alternative example, the SAD/VAD may be operated on multiple (notably all) channel signals of the input signal 101. The individual results for the individual channel signals may then be combined into a single CombinedVAD flag. If the CombinedVAD flag is set, a frame may be considered to be inactive. On the other hand, if the CombinedVAD flag is not set, the frame may be considered to be active.
[0062] Hence a VAD and/or SAD may be used to classify the frames of a sequence of frames into active frames or inactive frames. Encoding and/or generating comfort noise may be applied to the inactive frames. The encoding of the comfort noise (notably the encoding of noise shaping parameters) within the encoding unit 100 may be performed such that the decoding unit 150 is enabled to generate high quality comfort noise for a soundfield. The comfort noise that is generated by the decoding unit 150 preferably matches the spectral and/or spatial characteristics of the background noise within the input signal 101. This does not necessarily imply the waveform reconstruction of the input background noise. The comfort noise generated by a soundfield decoding unit 150 for a series of inactive frames is preferably such that the comfort noise sounds continuous with regards of the noise within the directly preceding active frames. Hence, the transition between active and inactive frames at the decoding unit 150 is preferably smooth and non-abrupt.”)
wherein the inactive metadata generator is configured to generate the control parameters such that the control parameters differ in characteristics from a characteristics of power ratios and object indices that are generated by the active metadata generator, for example wherein the control parameters comprise, e.g., the scaling factor and/or, e.g., either the coherence or the correlation.
(“[0119] The upmixing metadata 105 may be determined in dependence of whether the current frame is an active frame or an inactive frame. In particular, the set of parameters, which is comprised within the upmixing metadata 105 may depend on whether the current frame is an active frame or an inactive frame. If the current frame is an active frame, the set of parameters of the upmixing parameters 105 may be larger and/or may comprise a higher number of different parameters than if the current frame is an inactive frame.
[0120] In particular, the cross-prediction parameter may not be part of the upmixing metadata 105 for the current frame, if the current frame is an inactive frame. On the other hand, the cross-prediction parameter may be part of the upmixing metadata 105 for the current frame, if the current frame is an active frame.
[0121] Alternatively, or in addition, if more than one residual channel is included into the downmix signal 103, the set of parameters of the upmixing metadata 105 for the current frame may comprise a decorrelation parameter for each possible combination of a non-included residual channel either with itself or with another one of the non-included residual channels, if the current frame is an active frame. On the other hand, the set of parameters of the upmixing metadata 105 for the current frame may comprise a decorrelation parameter only for the combination of a non-included residual channel with itself, if the current frame is an inactive frame.
[0122] Hence, the type of parameters which are included into the upmixing metadata 105 may be different for an active frame and for an inactive frame. In particular, one or more parameters which are less relevant for reconstructing the spatial characteristics of background noise may be omitted for an inactive frame. As a result of this, the data rate for encoding background noise may be reduced without impacting the perceptional quality.”)
ECKERT does not explicitly teach all of the plurality of input objects, quantized input direction with the specific control parameters, furthermore the power ratio and object indices.
However, FUCHS, teaches 16. An audio encoder according to claim 13, wherein the audio input comprises the plurality of audio input objects, and
(page 5 paragraph 1 “The audio signal encoder may be configured to encode, for the first frame, the audio signal using a time domain or frequency domain encoding mode, the encoded audio signal com prising, for example, encoded time domain samples, encoded spectral domain samples, encoded LPC domain samples and side information obtained from components of the au dio signal or obtained from one or more transport channels derived from the components of the audio signal, for example, by a downmixing operation. --> The audio signal may comprise an input format being a first order Ambisonics format, a higher order Ambisonics format, a multi-channel format associated with a given loud speaker setup, such as 5.1 or 7.1 or 7.1 + 4, or one or more audio channels representing one or several different audio objects localized in a space as indicated by information included in associated metadata, or an input format being a metadata associated spatial audio representation,...”)
wherein the audio encoder comprises an inactive metadata generator for generating, if the voice activity determiner has determined that the transport signal does not exhibit voice activity, metadata comprising quantized direction information and control parameters, for example comprising a scaling factor and/or either a coherence or a correlation,
(page 12 paragraoh 2 “The apparatus 300 may include an activity detector 320. The activity detector 320 may analyze the input audio signal (either in its input version 302 or in its downmix version 324), to determine, depending on the audio signal (302 or 324) whether a frame is an active frame 306 or an inactive frame 308, hence performing a classification on the frame. As can be seen from Fig. 3, the active detector 320 can be assumed as controlling (e.g. through the control 321 ) a first deviator 322 and a second deviator 322a. The first deviator 322 may select between the active spatial parameter 316 (first soundfield parameter representation) and the inactive spatial parameters 318 (second soundfield parameter representation). Therefore, the activity detector 320 may decide whether the active spatial parameters 316 or the inactive spatial parameters 318 are to be outputted (e.g. signalled in the bitstream 304)…”
Page 13 last paragraph and page 14 first paragraph “The soundfield parameter generator 310 of Fig. 4 may include a filterbank analysis block 390 (which may be the same of the filterbank analysis block 390 of Fig. 2). The filterbank analysis block 390 may provide frequency domain information 391 for each frame and for each beam (frequency tile). The frequency domain information 391 may be provided to a diffuseness analysis block 392a and/or a direction analysis block 392b, which may be those shown in Fig. 3. The diffuseness analysis block 392a and/or direction analysis block 392b may provide diffuseness information 314a and/or direction information 314b. These can be provided for each first frame 306 (346) and for each second frame 308 (348). Complexively, the information provided by the block 392a and 392b is considered soundfield parameters 314 which encompass both first sound- field parameters 316 (active spatial parameters) and second soundfield parameters 318 (inactive spatial parameters). The active spatial parameters 316 may be provided to an active spatial metadata encoder 396 and the inactive spatial parameters 318 may be provided to an inactive spatial metadata encoder 398. The resulting are first and second soundfield parameter represen tations (316, 318, complexively indicated with 314) which may be encoded in the bitstream 304 (e.g., through the encoder signal former 370) and stored for being subsequently played back by a decoder. Whether the active spatial metadata encoder 396 or the inactive spatial parameters 318 is to encode a frame, this may be controlled by a control such as the control 321 in Fig. 3 (the deviator 322 is not shown in Fig. 2), e.g. thorough the classification operated by the activity de tector. (It is to be noted that the encoders 396, 398 may also perform a quantization, in some examples). --> Fig. 5 shows another example of possible soundfield parameter generator 310, which may be alternative to that of Fig. 4, and which may also be implemented in the examples of Figs. 2 and 3. In this example, the input audio signal 302 can already be in MASA format, in which spatial parameters are already part of the input audio signal 302 (e.g., as spatial metadata), e.g. for each frequency bin of a plurality of frequency bins. Accordingly, there is no need for having a diffuse ness analysis block and/or a directional block, but they can be substituted by a MASA reader 390M. The MASA reader 390M may read specific data fields in the audio signal 302, which al ready contain information such as the active spatial parameter(s) 316 and the inactive spatial parameter(s) 318 (according to the fact whether the frame of the signal 302 is a first frame 306 or a second frame 308). Examples of parameters that may be encoded in the signal 302 (and which may be read by the MASA reader 390M) may include at least one of a direction, energy ratio, surround coherence, spread coherence, and so on.”
Page 5 last paragraph “… and wherein only a second number of bits of each quantization index is transmitted, the second number of bits being smaller than the first number of bits, or to determine, for the second soundfield parameter representation, an inter-channel coherence if the audio signal has input channels corresponding to channels positioned in a spatial domain or inter-channel level differences if the audio signal has input channels corresponding to channels positioned in the spatial domain, or to determine a surround coherence being defined as a ratio of diffuse energy being coherent in a soundfield represented by the audio signal.”)
wherein the inactive metadata generator is configured to generate the control parameters such that the control parameters differ in characteristics from a characteristics of power ratios and object indices that are generated by the active metadata generator, for example wherein the control parameters comprise, e.g., the scaling factor and/or, e.g., either the coherence or the correlation.
(Page 5 last paragraph “… and wherein only a second number of bits of each quantization index is transmitted, the second number of bits being smaller than the first number of bits, or to determine, for the second soundfield parameter representation, an inter-channel coherence if the audio signal has input channels corresponding to channels positioned in a spatial domain or inter-channel level differences if the audio signal has input channels corresponding to channels positioned in the spatial domain, or to determine a surround coherence being defined as a ratio of diffuse energy being coherent in a soundfield represented by the audio signal.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of FUCHS , to provide a “all of the plurality of input objects, quantized input direction with the specific control parameters.” Doing so would reduce the bitrate demand for transmitting conversational speech, as recognized by FUCHS ,. (page 3 paragraph 1 ).
ECKERT in view of FUCHS does not explicitly teach all of the power ratio and object indices
However, EICHENSEER , teaches wherein the inactive metadata generator is configured to generate the control parameters such that the control parameters differ in characteristics from a characteristics of power ratios and object indices that are generated by the active metadata generator, for example wherein the control parameters comprise, e.g., the scaling factor and/or, e.g., either the coherence or the correlation.
(page 9 paragraph 4-6 “The parametric side information transmitted to the decoder thus comprises:
• The power ratios calculated for a subset of relevant (dominant) objects for each time/frequency tile (or parameter band).
• Object indices that represent the subset of relevant objects for each time/frequency tile (or parameter band).”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT in view of FUCHS to incorporate the teachings of EICHENSEER , to provide a “the power ratio and object indices” Doing so would improve the audio quality, as recognized by EICHENSEER ,. (page 6 paragraph 2 ).
Claims 21 are rejected under 35 U.S.C. 103 as obvious over ECKERT in view of in view of WO 2022022876 A1 (FUCHS ) in further view of US 20160005407, (FRIEDRICH)
Claim 21
Regarding Claim 21, ECKERT in view of FUCHS does not explicitly teach 21.An audio encoder according to claim 20, wherein the audio encoder is configured generate the control parameters, by selecting, whether the control parameters either comprises the plurality of parameter values for the plurality of subbands, or whether the control parameters are the single broadband control parameters, depending on an available bitrate.
However, FRIEDRICH, teaches 21.An audio encoder according to claim 20, wherein the audio encoder is configured generate the control parameters, by selecting, whether the control parameters either comprises the plurality of parameter values for the plurality of subbands, or whether the control parameters are the single broadband control parameters, depending on an available bitrate.
(“[0084] The system 500 may comprise a configuration unit 540 which is configured to determine one or more control settings 552, 554 for the parameter coding unit 520 and/or for downmix coding unit 510. The one or more control settings 552, 554 may be determined based on one or more external settings 551 of the system 500. By way of example, the one or more external settings 551 may comprise an overall (maximum or fixed) data-rate of the bitstream 564. The configuration unit 540 may be configured to determine one or more control settings 552 in dependence on the one or more external settings 551. The one or more control settings 552 for the parameter coding unit 520 may comprise one or more of the following: [0085] a maximum data-rate for the encoded spatial parameters 562. This control setting is referred to herein as the metadata data-rate setting). [0086] a maximum number and/or a specific number of parameter sets to be determined by the parameter coding unit 520 per frame of the audio signal 561. This control setting is referred to herein as the temporal resolution setting, as it allows influencing the temporal resolution of the spatial parameters. [0087] a number of parameter bands for which spatial parameters are to be determined by the parameter coding unit 520. This control setting is referred to herein as the frequency resolution setting, as it allows influencing the frequency resolution of the spatial parameters. [0088] a resolution of the quantizer used for quantizing the spatial parameters. This control setting is referred to herein as the quantizer setting.”
“[0092] A parameter determination unit 523 of the parameter coding unit 520 (and in particular, the parameter extractor 420) may be configured to determine one or more sets of mixing parameters α.sub.1, α.sub.2, α.sub.3, β.sub.1, β.sub.2, β.sub.3, g, k.sub.1, k.sub.2 for each of the frequency bands 572. Due to this, the frequency bands 572 may also be referred to as parameter bands. The mixing parameters α.sub.1, α.sub.2, α.sub.3, β.sub.1, β.sub.2, β.sub.3, g, k.sub.1, k.sub.2 for a frequency band 572 may be referred to as the band parameters. As such, a complete set of mixing parameters typically comprises band parameters for each frequency band 572. The band parameters may be applied in the mixing matrix 130 of FIG. 3 to determine subband versions of the decoded upmix signal.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT in view of FUCHS to incorporate the teachings of FRIEDRICH, to provide a “21.An audio encoder according to claim 20, wherein the audio encoder is configured generate the control parameters, by selecting, whether the control parameters either comprises the plurality of parameter values for the plurality of subbands, or whether the control parameters are the single broadband control parameters, depending on an available bitrate.” Doing so would moderately affect the audio quality, as recognized by FRIEDRICH. (paragraph 111)
Claims 23 are rejected under 35 U.S.C. 103 as obvious over ECKERT in in view of US 20160142847 A1, (DISCH)
Claim 23
Regarding Claim 23, ECKERT does not explicitly teach 23. An audio encoder according to claim 1, wherein, if the audio input comprises the plurality of audio input channels, but not the plurality of audio input objects, a number of the two or more transport channels is smaller than a number of the plurality of audio input channels,
wherein, if the audio input comprises the plurality of audio input objects, but not the plurality of audio input channels, the number of the two or more transport channels is smaller than a number of the plurality of audio input objects,
wherein, if the audio input comprises both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transport channels is smaller than a sum of the number of the plurality of audio input channels and the number of the plurality of the audio input objects;
or
wherein, if the audio input comprises the plurality of audio input channels, but not the plurality of audio input objects, a number of the two or more transport channels is smaller than or equal to a number of the plurality of audio input channels,
wherein, if the audio input comprises the plurality of audio input objects, but not the plurality of audio input channels, the number of the two or more transport channels is smaller than or equal to a number of the plurality of audio input objects,
wherein, if the audio input comprises both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transport channels is smaller than or equal to a sum of the number of the plurality of audio input channels and the number of the plurality of the audio input objects.
However, DISCH teaches
23. An audio encoder according to claim 1, wherein, if the audio input comprises the plurality of audio input channels, but not the plurality of audio input objects, a number of the two or more transport channels is smaller than a number of the plurality of audio input channels,
(“[0055] FIG. 8 illustrates a further embodiment of the 3D audio encoder, where in contrast to FIG. 6, the SAOC encoder can be configured to either encode, with the SAOC encoding algorithm, the channels provided at the pre-renderer/mixer 200 not being active in this mode or, alternatively, to SAOC encode the pre-rendered channels plus objects. Thus, in FIG. 8, the SAOC encoder 800 can operate on three different kinds of input data, i.e., channels without any pre-rendered objects, channels and pre-rendered objects or objects alone. Furthermore, it is advantageous to provide an additional OAM decoder 420 in FIG. 8 so that the SAOC encoder 800 uses, for its processing, the same data as on the decoder side, i.e., data obtained by a lossy compression rather than the original OAM data.”
“[0102] In FIGS. 6 and 8, a SAOC encoder 800 is depicted. The SAOC encoder 800 is used to parametrically encode a number of input objects/channels by downmixing them to a lower number of transport channels and extracting the auxiliary information that may be used which is embedded into the 3D-Audio bitstream.”)
wherein, if the audio input comprises the plurality of audio input objects, but not the plurality of audio input channels, the number of the two or more transport channels is smaller than a number of the plurality of audio input objects,
(“[0055] FIG. 8 illustrates a further embodiment of the 3D audio encoder, where in contrast to FIG. 6, the SAOC encoder can be configured to either encode, with the SAOC encoding algorithm, the channels provided at the pre-renderer/mixer 200 not being active in this mode or, alternatively, to SAOC encode the pre-rendered channels plus objects. Thus, in FIG. 8, the SAOC encoder 800 can operate on three different kinds of input data, i.e., channels without any pre-rendered objects, channels and pre-rendered objects or objects alone. Furthermore, it is advantageous to provide an additional OAM decoder 420 in FIG. 8 so that the SAOC encoder 800 uses, for its processing, the same data as on the decoder side, i.e., data obtained by a lossy compression rather than the original OAM data.”
“[0107] The apparatus comprises an object mixer 210 for generating the audio transport signal comprising the one or more audio transport channels from two or more audio object signals, such that the two or more audio object signals are mixed within the audio transport signal, and wherein the number of the one or more audio transport channels is smaller than the number of the two or more audio object signals.”)
wherein, if the audio input comprises both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transport channels is smaller than a sum of the number of the plurality of audio input channels and the number of the plurality of the audio input objects;
(“[0055] FIG. 8 illustrates a further embodiment of the 3D audio encoder, where in contrast to FIG. 6, the SAOC encoder can be configured to either encode, with the SAOC encoding algorithm, the channels provided at the pre-renderer/mixer 200 not being active in this mode or, alternatively, to SAOC encode the pre-rendered channels plus objects. Thus, in FIG. 8, the SAOC encoder 800 can operate on three different kinds of input data, i.e., channels without any pre-rendered objects, channels and pre-rendered objects or objects alone. Furthermore, it is advantageous to provide an additional OAM decoder 420 in FIG. 8 so that the SAOC encoder 800 uses, for its processing, the same data as on the decoder side, i.e., data obtained by a lossy compression rather than the original OAM data.”
“[0102] In FIGS. 6 and 8, a SAOC encoder 800 is depicted. The SAOC encoder 800 is used to parametrically encode a number of input objects/channels by downmixing them to a lower number of transport channels and extracting the auxiliary information that may be used which is embedded into the 3D-Audio bitstream.”)
or
wherein, if the audio input comprises the plurality of audio input channels, but not the plurality of audio input objects, a number of the two or more transport channels is smaller than or equal to a number of the plurality of audio input channels,
(“[0055] FIG. 8 illustrates a further embodiment of the 3D audio encoder, where in contrast to FIG. 6, the SAOC encoder can be configured to either encode, with the SAOC encoding algorithm, the channels provided at the pre-renderer/mixer 200 not being active in this mode or, alternatively, to SAOC encode the pre-rendered channels plus objects. Thus, in FIG. 8, the SAOC encoder 800 can operate on three different kinds of input data, i.e., channels without any pre-rendered objects, channels and pre-rendered objects or objects alone. Furthermore, it is advantageous to provide an additional OAM decoder 420 in FIG. 8 so that the SAOC encoder 800 uses, for its processing, the same data as on the decoder side, i.e., data obtained by a lossy compression rather than the original OAM data.”
“[0102] In FIGS. 6 and 8, a SAOC encoder 800 is depicted. The SAOC encoder 800 is used to parametrically encode a number of input objects/channels by downmixing them to a lower number of transport channels and extracting the auxiliary information that may be used which is embedded into the 3D-Audio bitstream.”)
wherein, if the audio input comprises the plurality of audio input objects, but not the plurality of audio input channels, the number of the two or more transport channels is smaller than or equal to a number of the plurality of audio input objects,
(“[0055] FIG. 8 illustrates a further embodiment of the 3D audio encoder, where in contrast to FIG. 6, the SAOC encoder can be configured to either encode, with the SAOC encoding algorithm, the channels provided at the pre-renderer/mixer 200 not being active in this mode or, alternatively, to SAOC encode the pre-rendered channels plus objects. Thus, in FIG. 8, the SAOC encoder 800 can operate on three different kinds of input data, i.e., channels without any pre-rendered objects, channels and pre-rendered objects or objects alone. Furthermore, it is advantageous to provide an additional OAM decoder 420 in FIG. 8 so that the SAOC encoder 800 uses, for its processing, the same data as on the decoder side, i.e., data obtained by a lossy compression rather than the original OAM data.”
“[0107] The apparatus comprises an object mixer 210 for generating the audio transport signal comprising the one or more audio transport channels from two or more audio object signals, such that the two or more audio object signals are mixed within the audio transport signal, and wherein the number of the one or more audio transport channels is smaller than the number of the two or more audio object signals.”)
wherein, if the audio input comprises both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transport channels is smaller than or equal to a sum of the number of the plurality of audio input channels and the number of the plurality of the audio input objects.
(“[0055] FIG. 8 illustrates a further embodiment of the 3D audio encoder, where in contrast to FIG. 6, the SAOC encoder can be configured to either encode, with the SAOC encoding algorithm, the channels provided at the pre-renderer/mixer 200 not being active in this mode or, alternatively, to SAOC encode the pre-rendered channels plus objects. Thus, in FIG. 8, the SAOC encoder 800 can operate on three different kinds of input data, i.e., channels without any pre-rendered objects, channels and pre-rendered objects or objects alone. Furthermore, it is advantageous to provide an additional OAM decoder 420 in FIG. 8 so that the SAOC encoder 800 uses, for its processing, the same data as on the decoder side, i.e., data obtained by a lossy compression rather than the original OAM data.”
“[0102] In FIGS. 6 and 8, a SAOC encoder 800 is depicted. The SAOC encoder 800 is used to parametrically encode a number of input objects/channels by downmixing them to a lower number of transport channels and extracting the auxiliary information that may be used which is embedded into the 3D-Audio bitstream.”)
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified ECKERT to incorporate the teachings of DISCH, to provide a “23. An audio encoder according to claim 1, wherein, if the audio input comprises the plurality of audio input channels, but not the plurality of audio input objects, a number of the two or more transport channels is smaller than a number of the plurality of audio input channels, wherein, if the audio input comprises the plurality of audio input objects, but not the plurality of audio input channels, the number of the two or more transport channels is smaller than a number of the plurality of audio input objects, wherein, if the audio input comprises both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transport channels is smaller than a sum of the number of the plurality of audio input channels and the number of the plurality of the audio input objects; or wherein, if the audio input comprises the plurality of audio input channels, but not the plurality of audio input objects, a number of the two or more transport channels is smaller than or equal to a number of the plurality of audio input channels, wherein, if the audio input comprises the plurality of audio input objects, but not the plurality of audio input channels, the number of the two or more transport channels is smaller than or equal to a number of the plurality of audio input objects, wherein, if the audio input comprises both the plurality of audio input objects and the plurality of audio input channels, the number of the two or more transport channels is smaller than or equal to a sum of the number of the plurality of audio input channels and the number of the plurality of the audio input objects.” Doing so would increase the efficiency for coding a large number of objects or channels, as recognized by DISCH. (paragraph 105)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALI M HASSAN whose telephone number is (571)272-5331. The examiner can normally be reached Monday - Friday 8:00am - 4:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras Shah can be reached at (571)270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALI M HASSAN/
Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
09/05/2026