Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Setiawan et al (US 11632549 B2).
As per claim 1, Setiawan discloses an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed with the at least one processor (processor, software and memory are required to implement the following functions), cause the apparatus at least to:
obtain spatial audio content (the encoder per para 26 encodes multichannel audio into a bitstream which is received by the decoder of fig. 9, quantized bits);
decode encoded spatial metadata associated with the spatial audio content based, at least partially, on a configuration parameter indicative of at least one of an input source format of the spatial audio content or a codec configuration used to encode, at least in part, spatial metadata (para 70: obtain mode information F/A (fixed/adaptive) defining the adaptive mode or a fixed mode. );
determine at least one prototype audio signal based, at least partially, on the spatial audio content and a configuration of at least one output type (any of the signaling associated with any of the subsequent processing as shown in fig. 9, 902-908;
determine one or more spatial audio signals based, at least partially, on the at least one prototype audio signal and the decoded spatial metadata (part of block 909); and
provide the one or more spatial audio signals (stage 910).
As per claim 2, The apparatus of claim 1, wherein obtaining the spatial audio content comprises the instructions, when executed with the at least one processor, cause the apparatus to: obtain at least one encoded transport audio signal (the bitstream in 901); decode the at least one obtained encoded transport audio signal for rendering (902); and separate the at least one decoded transport audio signal and the decoded spatial metadata(steps 903-907).
As per claim 3, the apparatus of claim 1, wherein determining the at least one prototype audio signal comprises the instructions, when executed with the at least one processor, cause the apparatus to: obtain at least one transport audio signal (902); and determine the at least one prototype audio signal based (per the claim 1 rejection), at least partially, on the at least one transport audio signal.
As per claim 4, the apparatus of claim 3, wherein the instructions, when executed with the at least one processor, cause the apparatus to: determine a type of the at least one transport audio signal (based on the mode of the claim 1 rejection), wherein the one or more spatial audio signals are determined based, at least partially, on the determined type of the at least one transport audio signal (the mode determines the subsequent processing of fig. 9).
As per claim 5, the apparatus of claim 3, wherein the at least one prototype audio signal comprises at least one of: the at least one transport audio signal (the processing between 901 and 902), or a processed version of the at least one transport audio signal.
As per claim 6, the apparatus of claim 1, wherein the codec configuration used to encode, at least in part, the spatial metadata is based, at least partially, on a source format of the spatial audio content (per the mode of the claim 1 rejection).
As per claim 7, the apparatus of claim 1, wherein a format of the spatial audio content is at least partially different from a format the at least one output type supports (in the situations where the bitstream comprises content and metadata for the first mode, and then later comprises content and metadata for the second mode).
As per claim 8, the apparatus of claim 1, wherein a number of channels the spatial audio content comprises is at least partially different from a number of channels the at least one output type supports (in the situations where the bitstream comprises content and metadata for the first mode, and then later comprises content and metadata for the second mode and or different number of channels as part of the bitstream).
As per claim 9, the apparatus of claim 1, wherein determining the one or more spatial audio signals comprises the instructions, when executed with the at least one processor, cause the apparatus to:
determine a direct stream (output of 902) based, at least partially, on a direct prototype audio signal (signaling of 902) of the at least one prototype audio signal (input to 902) and at least a first part of the decoded spatial metadata (the mode per the claim 1 rejection);
determine a diffuse stream (the bitstream can support known audio formats including 5.1 per para 2, where 5.1 comprises rear/diffuse channels) based, at least partially, on a diffuse prototype audio signal of the at least one prototype audio signal and at least a second part of the decoded spatial metadata (the mode as implemented by one of stage 903-907; and
combine the direct stream (the center channel of 5.1) and the diffuse stream to determine the one or more spatial audio signals (to create the 910 signal).
As per claim 10, the apparatus of claim 1, wherein the at least one prototype audio signal comprises at least one of: at least one diffuse prototype audio signal (per claim 9 rejection), or at least one direct prototype audio signal.
As per claim 11, the apparatus of claim 1, wherein the decoded spatial metadata comprises at least one of:
at least one direction parameter (the designation of a particular channel in the context of a multichannel signal is a parameter that indicates direction, for example the 5.1 standard cited above has a direction inherent to each defined channel),
at least one direction of arrival of audio, at least one distance to an audio source, at least one coherence parameter, at least one energy ratio, at least one direct-to-total energy ratio, or at least one diffuse-to-total energy ratio.
As per claim 12, the claim 1 rejection discloses a method comprising: obtaining spatial audio content; decoding encoded spatial metadata associated with the spatial audio content based, at least partially, on a configuration parameter indicative of at least one of an input source format of the spatial audio content or a codec configuration used to encode, at least in part, spatial metadata; determining at least one prototype audio signal based, at least partially, on the spatial audio content and a configuration of at least one output type; determining one or more spatial audio signals based, at least partially, on the at least one prototype audio signal and the decoded spatial metadata; and providing the one or more spatial audio signals (per claim 1 rejection).
As per claim 13, The method of claim 12, wherein the obtaining of the spatial audio content comprises: obtaining at least one encoded transport audio signal; decoding the at least one obtained encoded transport audio signal for rendering; and separating the at least one decoded transport audio signal and the decoded spatial metadata (per claim 2 rejection).
As per claim 14, the method of claim 12, wherein the determining of the at least one prototype audio signal comprises: obtaining at least one transport audio signal; and determining the at least one prototype audio signal based, at least partially, on the at least one transport audio signal (per claim 3 rejection).
As per claim 15, the method of claim 14, further comprising: determining a type of the at least one transport audio signal, wherein the one or more spatial audio signals are determined based, at least partially, on the determined type of the at least one transport audio signal (per claim 4 rejection).
As per claim 16, the method of claim 14, wherein the at least one prototype audio signal comprises at least one of: the at least one transport audio signal, or a processed version of the at least one transport audio signal (per claim 5 rejection).
As per claim 17, the method of claim 12, wherein the codec configuration used to encode, at least in part, the spatial metadata is based, at least partially, on a source format of the spatial audio content (per claim 6 rejection).
As per claim 18, the method of claim 12, wherein the determining of the one or more spatial audio signals comprises: determining a direct stream based, at least partially, on a direct prototype audio signal of the at least one prototype audio signal and at least a first part of the decoded spatial metadata; determining a diffuse stream based, at least partially, on a diffuse prototype audio signal of the at least one prototype audio signal and at least a second part of the decoded spatial metadata; and combining the direct stream and the diffuse stream to determine the one or more spatial audio signals (per claim 9 rejection).
As per claim 19, the method of claim 12, wherein the decoded spatial metadata comprises at least one of: at least one direction parameter, at lest one direction of arrival of audio, at least one distance to an audio source, at least one coherence parameter, at least one energy ratio, at least one direct-to-total energy ratio, or at least one diffuse-to-total energy ratio (per claim 11 rejection).
As per claim 20, a non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: causing obtaining of spatial audio content; decoding encoded spatial metadata associated with the spatial audio content based, at least partially, on a configuration parameter indicative of at least one of an input source format of the spatial audio content or a codec configuration used to encode, at least in part, spatial metadata; determining at least one prototype audio signal based, at least partially, on the spatial audio content and a configuration of at least one output type; determining one or more spatial audio signals based, at least partially, on the at least one prototype audio signal and the decoded spatial metadata; and causing providing of the one or more spatial audio signals (required by the processor of the system of the claim 12 and 1 rejections).
Response to Arguments
The submitted arguments have been considered but are moot in view of the new grounds of rejection.
As per applicant’s arguments about the source format as recited in the newly amended claims, the examiner notes that element is recited in the alternative.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER KRZYSTAN whose telephone number is 571-272-7498, and whose email address is alexander.krzystan@uspto.gov
The examiner can usually be reached on m-f 7:30-4:00 est.
If attempts to reach the examiner by telephone or email are unsuccessful, the examiner’s supervisor, Fan Tsang can be reached on (571) 272-7547.
The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications.
/ALEXANDER KRZYSTAN/Primary Examiner, Art Unit 2653
Examiner Alexander Krzystan
August 20, 2026