DETAILED ACTION
This communication is in response to the Application filed on 26 December 2024. Claims 1-5, 11-12, and 14-21 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged that application is a National Stage application of PCT/CN2022/103170. Priority to PCT/CN2022/103170 with a priority date of 30 June 2022 is acknowledged under 35 USC 119(e) and 37 CFR 1.78.
Information Disclosure Statement
The IDS dated 26 December 2024 has been considered and placed in the application file.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim 1 is rejected under 35 U.S.C. 102(a)(1) and (a)(2) as being anticipated by US Patent Publication 20160225377 A1 (Miyasaki et al.).
Claim 1
Regarding claim 1, Miyasaki et al. disclose an audio signal encoding method, comprising:
obtaining a scene-based audio signal (Miyasaki et al. ¶ [0043], "background sound is input as a multi-channel background object (MBO) in the form of multi-channel signals");
determining a number of audio channels of the audio signal (Miyasaki et al. ¶ [0031], " In the channel-based audio system, a group of picked-up sound sources (guitar, piano, vocal etc.) are rendered in advance according to the reproduction speaker arrangement assumed by the system. ... For example, when the speaker arrangement assumed by the system is a 5-channel speaker arrangement, a group of picked-up sound sources are assigned to the channels such that the sound sources are reproduced at appropriate sound image positions by 5-channel speakers.") and an encoding rate (Miyasaki et al. ¶ [0051], "The audio scene analysis unit is configured to extract perceptual importance information of at least the object-based audio signal, and determine a number of encoding bits allocated to each of the channel-based audio signal"); and
generating an encoded codestream by encoding the audio signal according to the number of audio channels and the encoding rate (Miyasaki et al. ¶ [0038], "when the speaker arrangement of the decoding side is a 5-channel speaker arrangement, the audio objects are assigned to channels such that the audio objects are reproduced by 5-channel speakers at positions corresponding to the respective reproduction position information." ¶ [0060], "the channel-based encoder encodes the channel-based audio signal according to the number of encoding bits, and the object-based encoder encodes the object-based audio signal according to the number of encoding bits.").
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-3, 11-12, 14-15, and 18-19 are rejected under 35 U.S.C. 103 as obvious over Miyasaki et al. as applied to claim 1 above, and further in view of US Patent Publication 20220406318 A1 (Tyagi et al.).
Claim 2
Regarding claim 2, the rejection of claim 1 is incorporated.
Miyasaki et al. further disclose performing [a down-mixed] processing on the audio signal (Miyasaki et al. ¶ [0067], "the audio decoding device determines a head related transfer function (HRTF) used for performing downmixing for speakers") according to the number of audio channels (Miyasaki et al. ¶ [0038], "when the speaker arrangement of the decoding side is a 5-channel speaker arrangement, the audio objects are assigned to channels such that the audio objects are reproduced by 5-channel speakers at positions corresponding to the respective reproduction position information.") and the encoding rate (Miyasaki et al. ¶ [0060], "the channel-based encoder encodes the channel-based audio signal according to the number of encoding bits, and the object-based encoder encodes the object-based audio signal according to the number of encoding bits.") to generate a [down-mixed] parameter (Miyasaki et al. ¶ [0073], "The audio scene analysis unit 100 determines an audio scene from an input signal composed of a channel-based audio signal and an object-based audio signal, and detects audio scene information." Audio scene information is considered analogous to a parameter) and [a down-mixed] an audio channel signal (Miyasaki et al. Figure 1 illustrates "M' channel signals" which are derived from "M channel signals." "M' channel signals" are considered analogous to an audio channel signal);
encoding the [down-mixed] audio channel signal to generate an encoding parameter (Miyasaki et al. ¶ [0073]-[0075], "The channel-based encoder 101 encodes the channel-based audio signal that is an output signal of the audio scene analysis unit 100, based on the audio scene information that is an output signal of the audio scene analysis unit 100. The audio scene encoding unit 103 encodes the audio scene information that is an output signal of the audio scene analysis unit 100." The channel-based audio signal output from the audio scene analysis unit 100 is depicted as “M’ channel signals” in Figure 1. Thus, the encoded signals produced by encoding “M’ channel signals” using a channel-based encoder are considered analogous to an encoding parameter); and
generating the encoded codestream by performing codestream multiplexing on the [down-mixed] parameter and the encoding parameter (Miyasaki et al. [0075], "The audio scene encoding unit 103 encodes the audio scene information that is an output signal of the audio scene analysis unit 100." ¶ [0076], "The multiplexing unit 104 multiplexes the channel-based encoded signal that is an output signal of the channel-based encoder 101, … and the audio scene encoded signal that is an output signal of the audio scene encoding unit 103 to generate a bit stream, and outputs the bit stream." The audio scene encoded signal is derived from the audio scene information. Thus, the audio scene encoded signal is considered analogous to the parameter. The channel-based encoded signal is considered analogous to the encoding parameter.).
Miyasaki et al. do not explicitly disclose all of performing a down-mixed processing on audio signals to generate a down-mixed parameter.
However, Tyagi et al. disclose performing a down-mixed processing on an audio signal according [to the number of audio channels and] the encoding rate to generate a down-mixed parameter and a down-mixed audio channel signal (Tyagi et al. ¶ [0115], "The system performs a lookup in the bitrate distribution control table based on the table indices and extracts ... EVS target bitrate and a bitrate ratio. The system extracts and decodes the downmix audio bits per downmix channel and spatial [metadata (MD)] bits. The system provides the extracted downmix signal bits and spatial MD bits to a downstream IVAS device." EVS target bitrate and a bitrate ratio are considered analogous to an encoding rate. Spatial metadata bits are considered analogous to a down-mixed parameter. Downmix signal bits are considered analogous to a down-mixed audio channel signal).
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify Miyasaki et al.’s audio processing to include Tyagi et al.’s down-mixing because such a modification is the result of combining prior art elements according to known methods to yield predictable results. More specifically, Miyasaki et al.’s audio processing as modified by Tyagi et al.’s down-mixing can yield a predictable result of improving audio signal compatibility since downmixing audio signals from a high number of channels to stereo or mono would make the audio easier to post-process and subsequently play on a variety of different audio system setups. Thus, a person of ordinary skill would have appreciated including in Miyasaki et al.’s audio processing the ability to do Tyagi et al.’s down-mixing since the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Claim 3
Regarding claim 3, the rejection of claim 2 is incorporated.
Miyasaki et al. further disclose determining a target control parameter for the audio signal according to the number of audio channels (Miyasaki et al. ¶ [0037]-[0038], "In the object-based audio system, a group of picked-up sound sources (guitar, piano, vocal, etc.) are directly encoded as audio objects, and the audio objects are recorded and transmitted. At this time, reproduction position information of the sound sources is also recorded and transmitted. ... For example, when the speaker arrangement of the decoding side is a 5-channel speaker arrangement, the audio objects are assigned to channels such that the audio objects are reproduced by 5-channel speakers at positions corresponding to the respective reproduction position information." Position information is considered analogous to target control parameter) [and the encoding rate];
determining a down-mixed processing algorithm according to the target control parameter (Miyasaki et al. ¶ [0067], "The audio scene information is audio object position information, and the audio decoding device determines a head related transfer function (HRTF) used for performing downmixing for speakers, from the audio object position information, reproduction-side speaker arrangement information that is provided separately, and listener position information that is provided separately or pre-supposed." A head related transfer function used for performing downmixing is considered analogous to a down-mixed processing algorithm); and
performing the down-mixed processing on the audio signal according to the downmixed processing algorithm to generate [the down-mixed parameter and] the down-mixed audio channel signal (Miyasaki et al. Figure 1 illustrates "M' channel signals" which can be inferred to be downmixed from "M channel signals" according to ¶ [0067], "the audio decoding device determines a head related transfer function (HRTF) used for performing downmixing for speakers, from the audio object position information, reproduction-side speaker arrangement information that is provided separately, and listener position information that is provided separately or pre-supposed.").
Tyagi et al. further disclose determining a target control parameter for the audio signal according to [the number of audio channels and] the encoding rate (Tyagi et al. ¶ [0088]-[0089], "In Step 501, ... signal properties are extracted from the input audio signal ... In Step 502, process 500 extracts the IVAS bitrate distribution control table indices from an IVAS bitrate distribution control table using the IVAS bitrate. In Step 503, process 500 determines the input format table indices based on the signal parameters extracted in Step 501 (i.e., BW and speech/music classification), the input audio signal format, the IVAS bitrate distribution control table indices extracted in Step 502 and an EVS mono downmix backward compatibility mode." The input format table indices determined using the extracted signal parameters and IVAS bitrate are considered analogous to target control parameters);
determining a down-mixed processing algorithm according to the target control parameter (Tyagi et al. ¶ [0089], "In Step 504, process 500 selects the spatial coding mode (i.e., FP or MR) … based on the bitrate distribution control table indices ... The spatial audio coding mode indicates either an MR coding mode, where the representation of mid or W channel (M′ or W) is accompanied with one or more residual channels in the downmixed audio signal, or an FP coding mode, where only the representation of the mid or W channel (M′ or W) is present in the downmixed audio signal." The spatial audio coding mode is considered analogous to a down-mixed processing algorithm); and
performing the down-mixed processing on the audio signal according to the downmixed processing algorithm to generate the down-mixed parameter and the down-mixed audio channel signal (Tyagi et al. ¶ [0115], "The system performs a lookup in the bitrate distribution control table based on the table indices and extracts …the spatial coding mode…. The system extracts and decodes the downmix audio bits per downmix channel and spatial [metadata (MD)] bits. The system provides the extracted downmix signal bits and spatial MD bits to a downstream IVAS device." Spatial metadata bits are considered analogous to a down-mixed parameter. Downmix signal bits are considered analogous to a down-mixed audio channel signal).
Claim 11
Regarding claim 11, Tyagi et al. disclose a processor (Tyagi et al. ¶ [0157], "These computer program codes may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus that has control circuitry"); and
a memory storing machine-readable instructions (Tyagi et al. ¶ [0156], "a machine readable medium may be any tangible medium that may contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.")….
The remaining limitations of claim 11 are similar in scope to that of claim 1 and therefore are rejected for similar reasons as described above.
Claim 12
Regarding claim 12, Tyagi et al. disclose non-transitory computer-readable storage medium having computer instructions stored thereon (Tyagi et al. ¶ [0156], "a machine readable medium may be any tangible medium that may contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.")….
The remaining limitations of claim 12 are similar in scope to that of claim 1 and therefore are rejected for similar reasons as described above.
Claim 14
Regarding claim 14, the rejection of claim 11 is incorporated. The limitations of claim 14 are similar in scope to that of claim 2 and therefore are rejected for similar reasons as described above.
Claim 15
Regarding claim 15, the rejection of claim 14 is incorporated. The limitations of claim 15 are similar in scope to that of claim 3 and therefore are rejected for similar reasons as described above.
Claim 18
Regarding claim 18, the rejection of claim 12 is incorporated. The limitations of claim 18 are similar in scope to that of claim 2 and therefore are rejected for similar reasons as described above.
Claim 19
Regarding claim 19, the rejection of claim 18 is incorporated. The limitations of claim 19 are similar in scope to that of claim 3 and therefore are rejected for similar reasons as described above.
Claims 4, 16, and 20 are rejected under 35 U.S.C. 103 as obvious over Miyasaki et al. in view of Tyagi et al. as applied to claim 3, 15, and 19 above, and further in view of US Patent Publication 20040225495 A1 (Makino).
Claim 4
Regarding claim 4, the rejection of claim 3 is incorporated.
Miyasaki et al. in view of Tyagi et al. do not explicitly disclose all of calculating an initial and target bit rate.
However, Makino discloses calculating an initial average rate of each audio channel according to the number of audio channels and the encoding rate (Makino ¶ [0034], "On the assumption that the provisional number of in-use bits determined for a channel No. n in a frame is
b
n
, the total number of bits allocable to all the channels in that frame is
B
and all the total number
B
of allocable bits is allocated as the number
B
n
of usable bits to each channel by making the inter-channel bit allocation, the number
B
n
of usable bits can be determined to be as given by the following formula (1). Thus, bits (the number
B
n
of usable bits) are allocated at a ratio in demand for bits between the channels." The number of usable bits
B
n
per channel is considered analogous to an inital average rate);
determining a target average rate according to the initial average rate (Makino ¶ [0032], "each of the encoders
10
n
includes a number-of-bits adjusting means (not shown) that adjusts the number
B
n
'
of in-use bits correspondingly to the number
B
n
of usable bits supplied from the inter-channel bit allocator 30."
B
n
'
is considered analogous to a target average rate) and a preset average rate threshold (Makino ¶ [0065]-[0068], "the formula (1) can be corrected as given by the formula (2) by allocating a part of the total number
B
of allocable bits by any other method ... One of the methods is to meet a condition given by the following formula (3) by allocating
r
B
/
N
bits fixedly to each of channels (
c
h
1
to
c
h
N
), to thereby assure a minimum number of bits" A minimum number of bits is considered analogous to a preset average rate threshold); and
determining the target control parameter for the audio signal according to the initial average rate and the target average rate (Makino ¶ [0053], "The entropy encoder 19 compares the number
B
n
'
of in-use bits and number
B
n
of usable bits, and adjust the number
B
n
'
of in-use bits, by increasing or decreasing the number of quantizing steps, so that the number
B
n
'
of in-use bits will be less than and near the number
B
n
of usable bits." The adjusted number of quantizing steps is considered analogous to a target control parameter).
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify Miyasaki et al. in view of Tyagi et al. to incorporate Makino’s bit rate calculations because such a modification is the result of combining prior art elements according to known methods to yield predictable results. More specifically, Miyasaki et al.’s encoding rates as modified by Makino’s bit rate calculations can yield a predictable result of improving efficiency since allocating only the bits necessary to each channel would alleviate the processor of unnecessary computations. Thus, a person of ordinary skill would have appreciated including in Miyasaki et al.’s encoding rates the ability to do Makino’s bit rate calculations since the claimed invention is merely a combination of old elements, and in the combination each element merely would have performed the same function as it did separately, and one of ordinary skill in the art would have recognized that the results of the combination were predictable.
Claim 16
Regarding claim 16, the rejection of claim 15 is incorporated. The limitations of claim 16 are similar in scope to that of claim 4 and therefore are rejected for similar reasons as described above.
Claim 20
Regarding claim 20, the rejection of claim 19 is incorporated. The limitations of claim 20 are similar in scope to that of claim 4 and therefore are rejected for similar reasons as described above.
Claim 5 is rejected under 35 U.S.C. 103 as obvious over Miyasaki et al. as applied to claim 1 above, and further in view of US Patent Publication 20070019803 A1 (Merks et al.).
Claim 5
Regarding claim 5, the rejection of claim 1 is incorporated.
Miyasaki et al. do not explicitly disclose all of a pre-emphasis preprocessing or high-pass filtering preprocessing.
However, Merks et al. disclose performing a pre-emphasis preprocessing (Merks et al. ¶ [0026]-[0027], "the system comprises a pre-processor comprising: an amplifier to amplify the audio signal to a sufficient Sound Pressure Level") and/or a high-pass filtering preprocessing on the audio signal (Merks et al. ¶ [0031], "In preferred embodiments the pre-processor comprises a high-pass filter.").
It would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify Miyasaki et al.’s audio downmixing to incorporate Merks et al.’s preprocessing.
The suggestion/motivation for doing so would have been that, “low frequencies cannot be reproduced by the loudspeaker. So it is better to remove, these low frequency components by applying a digital high-pass filter before the digital clipping,” as noted by the Merks et al. disclosure in paragraph [0064].
Claims 17 and 21 are rejected under 35 U.S.C. 103 as obvious over Miyasaki et al. in view of Tyagi et al. as applied to claims 11 and 12 above, and further in view of Merks et al.
Claim 17
Regarding claim 17, the rejection of claim 11 is incorporated. The limitations of claim 17 are similar in scope to that of claim 5 and therefore are rejected for similar reasons as described above.
Claim 21
Regarding claim 21, the rejection of claim 12 is incorporated. The limitations of claim 21 are similar in scope to that of claim 5 and therefore are rejected for similar reasons as described above.
Reference Cited
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US Patent Publication 20100228554 A1 to Beack et al. discloses a downmixing method involving a minimum (threshold) bit rate.
US Patent Publication 20110002393 A1 to Suzuki et al. discloses a downmixing method that maintains a minimum level of audio deterioration when determining appropriate bit rates per channel.
“Efficient Channel-aware Rate Adaptation in Dynamic Environments” to Judd et al. discloses a channel-aware bit rate control method by predicting path loss.
“Scalable Rate Control for MPEG-4 Video” to Lee et al. discloses a scalable rate control scheme that utilizes a sliding window to dynamically adjust bit allocation.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB B VOGT whose telephone number is (571)272-7028. The examiner can normally be reached Monday - Friday, 11am - 8pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PARAS D SHAH can be reached at (571)270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACOB B VOGT/ Examiner, Art Unit 2653
/Paras D Shah/ Supervisory Patent Examiner, Art Unit 2653
07/15/2026