Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103 is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Priority
Acknowledgment is made of applicant's claim for domestic priority based US provisional application 63/608097 filed on 12/08/2023.
Claim Rejections - 35 USC § 103
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 103 that form the basis for the rejections under this section made in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 6, 10-13, 15, and 17-20 are rejected under 35 USC 103(a) as being unpatentable over Eronen et al. (US 2023/0179947 A1) in view of Muesch (US 8972250 B2).
Regarding Claims 1, 10, and 17, Eronen discloses a system (Fig. 12 and ¶255) comprising:
one or more non-transitory computer readable storage media storing instructions (¶257, memory 2011 storing program codes); and
one or more processors coupled to the one or more non-transitory computer readable storage media and operable to execute the instructions to (¶256, processor 2007 executes program codes):
access a window of audio comprising a plurality of audio signals (¶106, receive audio signal 204; ¶146, audio signal 204 is a multi-channel input; in view of ¶100, each audio channel can have a source directivity pattern with a unique identifier);
determine, for each of the plurality of audio signals, a gain to apply to the audio signal relative to the reference audio signal, wherein the gain depends on frequency and on directionality relative to a listener (Fig. 2 and ¶108, determine directivity influenced / frequency dependent reverberation gains 202 gdir, rev(k), which describes how the directivity of the sound source affects the magnitude frequency response of the late reverberation; ¶107, directivity data is in the form of gain value gdir(i, k) for a number of directions and k the frequency; per ¶88, due to sound acoustic shadow of a human talker, direct sound is attenuated when listening from behind the talker when compared to listening in front of the talker); and
encode the audio so that bandwidth is allocated based on directionality relative to the listener by allocating an amount of bandwidth to each of the plurality of audio signals based on their respective frequency-dependent and directionality-dependent determined gains (¶99, transmit average directivity influenced reverberation gain to minimize the required bit rate for transmission; ¶170 and ¶229, directivity influenced reverberation parameter encoder 1108 obtains directivity influenced frequency dependent reverberation gains and writes the bitstream description containing the frequency dependent reverberation gain data for output to bitstream encoder 1109 to be encoded / quantized and combined into a bitstream per Fig. 11 and ¶172).
Eronen does not disclose determine, for each of the plurality of audio signals, a relative power of that respective audio signal relative to a reference audio signal.
Muesch discloses processing multichannel audio signal with speech in one channel and non-speech signals in the remaining channels / reference audio signal (Col 7, Rows 10-14) by determining, for each of the plurality of audio signals, a relative power of that respective audio signal relative to the reference audio signal (Fig. 4 and Col 8, Row 56 – Col 9, Row 25, pass band-limited signal to a power estimator to generate an estimate of signal power 403 in that frequency band and pass the signal power estimate 403 to level tracker 406 that tracks signal components in the band that are not speech to determine level estimate of non-speech components 411; i.e., generate signal power estimate of speech in one channel and level estimate of non-speech components in the remaining channels); and
determine, for each of the plurality of audio signals and based on the determined relative power of that signal, a gain to apply to the audio signal relative to the reference audio signal (Col 8, Row 56 – Col 9, Row 25, derive speech enhancement gain by passing signal power estimate to a power-to-gain transformation function or gain curve to generate band gain 405 to modify signal power in the band).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to determine a relative power of respective audio signal relative to a reference audio signal and determine a gain thereof to apply to the audio signal relative to the reference audio signal (Muesch, Col 8, Rows 65-67) in order to enhance multi-channel audio by applying a gain to the audio (Muesch, Abstract; compare Eronen, ¶146, input audio signal 204 is a multichannel input).
Further regarding Claim 10, Eronen discloses one or more non-transitory computer readable storage media storing instructions that are operable when executed to implement the method of claim 1 and system of claim 17 (¶257, memory 2011 storing program codes).
Regarding Claims 2, 11, and 18, Eronen as modified by Muesch discloses determining whether a pairwise correlation between two of the plurality of audio signals exceeds a threshold correlation prior to determining the relative power of each of those two audio signals (Muesch, Fig. 3, Col 7, Rows 4-15, compute speech intelligibility metric from the relative levels of the audio signal and a competing sound in the listening environment (e.g., aircraft cabin noise); i.e., when audio signal is a multichannel audio signal with speech in one channel and non-speech signals in the remaining channel, compute speech intelligibility metric from the relative levels of all channels (“pairwise correlation”) and the distribution of spectral energy in them; in view of Col 5, Rows 15-27 and Col 8, Rows 48-55, use spectral characteristics to track the level of non-speech audio signals to discriminate speech from other; i.e., use speech intelligibility metric computed from relative levels / pairwise correlation of all channels and spectral energy distribution to track a level (i.e., “threshold”) of non-speech audio signals in order to discriminate speech from other).
Regarding Claims 3, 12, and 19, Eronen discloses wherein each of the plurality of audio signals are associated with a particular object (¶100, each audio element (audio object or channel) can have a source directivity pattern with a unique identifier).
Regarding Claims 4 and 13, Eronen as modified by Muesch discloses wherein the plurality of audio signals comprise a plurality of audio frames (Muesch, Claims 8 and 10, receive a first portion of the audio signal, the first portion comprises a frame of the audio signal), each audio frame corresponding to a separate audio channel (Muesch, Col 7, Rows 10-12, audio signal is a multichannel audio signal with speech in one channel and non-speech signals in remaining channels; e.g., speech frames in one channel and non-speech frames in remaining channels).
Regarding Claims 6, 15, and 20, Eronen as modified by Muesch discloses wherein the gain is determined based at least in part on interpolating gain curves for one or more of (1) frequency (Muesch, Col 6, Rows 47-51 and Col 7, Rows 30-33 and Figs. 3a-c showing gain function / gain curve relating input power in a frequency band to a corresponding band gain corresponding to frequency-shaping compression amplification of speech components such that speech enhancement is applied only to high frequency portion of a signal) or (2) sound pressure level.
Claims 5 and 14 are rejected under 35 USC 103(a) as being unpatentable over Eronen et al. (US 2023/0179947 A1) and Muesch (US 8972250 B2) as applied to claims 4 and 13, in view of Muench et al. (US 2019/0362735 A1).
Regarding Claims 5 and 14, Eronen does not disclose wherein the reference audio signal comprises a center channel audio signal.
Muench discloses in a 5.1 or 7.1 audio signal the speech input channel can be a center channel (¶24).
Eronen discloses receiving audio signal and output directivity-influenced reverberated audio signals / multichannel output format in 7.1 +4 channel system format (¶106).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to use a center channel audio signal as reference audio signal in the multi-channel format in adapt loudness of N channel audio input data (Muench, ¶24).
Claims 7 and 16 are rejected under 35 USC 103(a) as being unpatentable over Eronen et al. (US 2023/0179947 A1) and Muesch (US 8972250 B2) as applied to claims 1 and 10, in view of Davidson et al. (US 12347447 B2).
Regarding Claims 7 and 16, Eronen discloses wherein encoding the audio so that bandwidth is allocated based on directionality relative to the listener comprises adjusting a bit allocation determined by encoder (¶99, determine average directivity-influenced reverberation gain data for transmission by an encoder to minimize the required bit rate for transmission; i.e., minimize or allocate bits or bandwidth based on average directivity-influenced reverberation gain data).
Eronen does not disclose the encoder uses a loudness-based psychoacoustic model.
Davidson discloses that encoders use a loudness-based psychoacoustic model to improve bit allocation for frequency bands of an audio signal (Col 1, Rows 39-40, Col 2, Rows 5-8, Col 19, Rows 23-28; typically, encoder uses psychoacoustic models / perceptual models to estimate a masking threshold for an audio spectrum where computing the masking threshold comprises applying a spreading function to transformed energy values of the frequency band, the transformation involves transforming linear energy values to the loudness domain).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to use an encoder using a loudness based psychoacoustic model to adjust the bit allocation based on directionality relative to the listener in order to improve bit allocation for frequency band of an audio signal (Davidson, Col 2, Rows 5-8).
Claim 8 is rejected under 35 USC 103(a) as being unpatentable over Eronen et al. (US 2023/0179947 A1) and Muesch (US 8972250 B2) as applied to claim 1, in view of Chen et al. (US 2005/0159946 A1).
Regarding Claim 8, Eronen does not disclose adjusting the bandwidth of one or more of the plurality of audio signals based on one or more characteristics of a reproduction space for playing the audio signals.
Chen discloses adjusting bandwidth of one or more of plurality of audio signals based on one or more characteristics of a reproduction space for playing the audio signals (¶81, encoder adaptively adjusts quantization of an audio signal based upon quality and bitrate constraints; see ¶8, Table 1, bitrates for different quality audio information / reproduction space).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to adjust the bandwidth / bitrate of one or more of the plurality of audio signals based on one or more characteristics of a reproduction space (e.g., Eronen, ¶238, headphone reproduction) for playing the audio signals to implement an encoder that regulates quality and bitrate with a control strategy (Chen, Abstract).
Claim 9 is rejected under 35 USC 103(a) as being unpatentable over Eronen et al. (US 2023/0179947 A1) and Muesch (US 8972250 B2) as applied to claim 1, in view of Fuchs (US 2016/0180602 A1).
Regarding Claim 9, Eronen does not disclose adjusting the bandwidth of one or more of the plurality of audio signals based on a restricted range of movement of the listener in an extended-reality environment.
Fuchs discloses an augmented / extended reality system downloading information to optimize bandwidth based on a restricted range of movement of a user in the extended reality environment (¶90, each system object can have an awareness radius in which is it active such that information about system objects that are beyond the radius can be downloaded but only displayed when they are within the awareness radius).
It would’ve been obvious to one ordinarily skilled in the art before the effective filing date of the invention to adjusting the bandwidth of one or more of the plurality of audio signals (Eronen, ¶100, each audio element being an audio object; compare Fuchs, ¶90, system object being an audio object) based on a restricted range of movement (i.e., awareness radius) of the listener in an extended-reality environment (e.g., Eronen, ¶231, head mounted display for AR/VR with listener having an awareness radius) in order to optimize bandwidth (Fuchs, ¶90).
Conclusion
Prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
US 10075802 B1 discloses a bitrate allocation mechanism for higher order ambisonic audio data based on audio data channel gain, direction based weighting, perceptual based analysis to perform encoding on a psychoacoustic audio encoding device (Fig. 10).
US 9219972 B2 discloses efficient audio coding having reduced bit rate for ambient signals in a 5.1 multi-channel system (Col 18, Rows 15-20) with a gain factor for respective channels with respective directions (Col 18, Rows 45-67).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to examiner Richard Z. Zhu whose telephone number is 571-270-1587 or examiner’s supervisor King Poon whose telephone number is 571-272-7440. Examiner Richard Zhu can normally be reached on M-Th, 0730:1700.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RICHARD Z ZHU/Primary Examiner, Art Unit 2654 07/11/2026