DETAILED ACTION
Notice of AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s amendments and arguments filed 04/27/2026, with respect to claim(s) 1-23 have been fully considered. Applicant amended claims 1, 3-5, 7-23 and added new claims 24, 25.
Applicant’s arguments filed 04/27/2026, with respect to claim(s) 1-23, under 35 U.S.C. 103 have been fully considered but they are not persuasive. Applicant argued that the cited references do not disclose or support the features in the amended claimed limitations of independent claims 1 and 10. Applicant further argued that cited reference Kim do not disclose or suggest “receiving an encoded representation of an input audio signal that includes several full-range audio channels of surround-sound format in a bitstream, where the representation is encoded using MP coding-based algorithm and the band-limited audio channels are omitted from the encoded representation”. Examiner respectfully disagrees. There is no mention of “the band-limited audio channels are omitted from the encoded representation” in the claim. Kim in para.[0085], describes how a surround format audio signal 50 with multiple channel is encoded by the encoder 14A. Examiner relied on another reference Lee for the teaching of encoding using MP based algorithm. Lee teaches in column 17, lines 29-43, encoding and decoding is done based on matching pursuit based algorithm. Applicant further argued that Kim do not describe “producing several audio driver signals by assigning each of the full-range audio channels and the at least one band-limited audio channel to a particular speaker based on received metadata”. Examiner respectfully disagrees. Kim in para. [0100], Fig. 4 illustrates that rendering unit 210 may render audio signals 26A which may include channels C1 through CL that are respectively indented for playback through loudspeakers 1 through L. Applicant further argued, that the cited references do not support or teach the amended claimed limitation of independent claim 17. Examiner respectfully disagrees. Kim in para.[0085], describes how a surround format audio signal 50 with multiple channel is encoded by the encoder 14A. Kim in para.[0089], describes the generation of HOA representation from the input audio. Examiner relied on another reference Lee for the teaching of encoding using MP based algorithm. Lee teaches in column 17, lines 29-43, encoding and decoding is done based on matching pursuit based algorithm.
For these reasons, examiner believes that the previously cited prior art of record still teaches the claimed amended limitations. Please see the rejections below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 10-13, 17, 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. ( US 20170105085 A1), hereinafter referenced as Kim, in view of Lee et al. (US 7801733 B2), hereinafter referenced as Lee.
Regarding Claim 1, Kim teaches a decoder-side method, the method comprising:
receiving a bitstream that includes: a representation of an input audio signal comprising a plurality of [full- range] audio channels of a surround-sound format, [wherein the representation of the input audio signal is encoded according to a Matching Pursuit (MP) coding-based algorithm] and metadata associated with the input audio signal ( Kim: Para.[0085],[0093],[0094], Figs. 3, 4, audio decoding device 22A received encoded audio data, bitstream 56A from the encoding device 14A. The encoding device 14A receives audio signal 50 representing multi-channel audio signal of surround sound format . Para.[0043], input audio signals could be object-based audio, which involves discrete pulse-code-modulation (PCM) data for single audio objects with associated metadata containing their location coordinates and other information);
receiving at least one band-limited audio channel associated with the surround-sound format ( Kim: Para.[0085], input audio signal 50 may be a six-channel audio signal for a source loudspeaker configuration of 5.1 (i.e., a front-left channel, a center channel, a front-right channel, a surround back left channel, a surround back right channel, and a low-frequency effects (LFE) channel ( band limited audio channel));
producing a plurality of audio driver signals by rendering the input audio signal based on the metadata( Kim: Para.[0096],[0098], Fig. 4, Audio decoding unit 204 may be configured to decode coded audio signal 62 into audio signal 70. Audio decoding unit 204 may decode channels C1-CN of audio signal 62 into channels C1-CN of decoded audio signal 70. HOA generation unit 208A may be configured to generate an HOA sound field based on multi-channel audio data and spatial positioning vectors ( metadata) and provide HOA coefficients 212A to rendering unit 210),
wherein producing the plurality of audio driver signals comprises assigning each of the plurality of full-range audio channels and the at least one band-limited audio channel to a particular speaker of a plurality of speakers based on the metadata( Kim: Para.[0098]-[0100], Fig. 4, HOA generation unit 208A may be configured to generate an HOA sound field based on multi-channel audio data and spatial positioning vectors ( metadata) and provide HOA coefficients 212A to rendering unit 210. Rendering unit 210 may render audio signals 26A for playback at a plurality of local loudspeakers, where the plurality of local loudspeakers includes L loudspeakers, audio signals 26A may include channels C1 through CL that are respectively intended for playback through loudspeakers 1 through L . Para.[0085], input audio signal 50 may be a six-channel audio signal for a source loudspeaker configuration of 5.1 (i.e., a front-left channel, a center channel, a front-right channel, a surround back left channel, a surround back right channel, and a low-frequency effects (LFE) channel ( band limited audio channel) );
and driving the plurality of speakers using the plurality of audio driver signals (Kim: Para.[0100], Fig. 4, rendering unit 210 may render audio signals 26A which may include channels C1 through CL that are respectively indented for playback through loudspeakers 1 through L).
Kim while teaching the method of claim 1, fails to explicitly teach the claimed, receiving a bitstream that includes: a representation of an input audio signal comprising a plurality of full- range audio channels [of a surround-sound format], wherein the representation of the input audio signal is encoded according to a Matching Pursuit (MP) coding-based algorithm[ and metadata associated with the input audio signal]; producing a decoded representation of the input audio signal by decoding the encoded representation using he MP coding-based algorithm.
However, Lee does teach the claimed, receiving a bitstream that includes: a representation of an input audio signal comprising a plurality of full- range audio channels [of a surround-sound format], wherein the representation of the input audio signal is encoded according to a Matching Pursuit (MP) coding-based algorithm[ and metadata associated with the input audio signal] ( Lee: Column 14, lines 10-22, Fig. 2, decoding apparatus 220 receives high band and low band speech signals ( full range) from the encoding apparatus 203. Column 7, lines 18-28,63-65, column 8, lines 23-32, Figs 3, 4, the encoding unit 308 includes the sine wave dictionary amplitude and phase searcher 402 which searches for the amplitude and phase of the sine wave dictionary using a matching pursuit (MP) algorithm to encode the signal);
producing a decoded representation of the input audio signal by decoding the encoded representation using he MP coding-based algorithm (Lee: Column 17, lines 29-43, the high-band speech signal is encoded and decoded based on a structure in which a harmonic structure and a stochastic structure is combined. The harmonic structure searches for an amplitude and a phase of a sine wave dictionary using a matching pursuit (MP) algorithm) ;
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Lee’s teaching of a high-band speech encoding and decoding apparatus in wideband speech encoding and decoding method, into the method coding of higher-order ambisonics audio data, taught by Kim, because, this would reproduce high quality sound even at a low bitrate in wideband speech encoding and decoding.(Lee, Column 2, lines 29-59).
Claim 10 is a decoder-side device claim comprising: at least one processor; and memory having stored instructions which when executed by the at least one processor ( Kim: Para.[0226]-[0228], various aspects of the techniques in each of the sets of examples may provide for a non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause the one or more processors to perform the method for which the audio decoding device has been configured to perform), causes the decoder-side device to perform the steps in method claim 1 above and as such, claim 10 is similar in scope and content to claim 1 and therefore, claim 10 is rejected under similar rationale as presented against claim 1 above.
Regarding Claim 2, Kim in view of Lee teach the method of claim 1. Kim further teaches, wherein the decoded representation comprises a higher- order ambisonics (HOA) representation of the input audio signal, wherein the method further comprising applying a conversion matrix to the HOA representation to reconstruct the input audio signal ( Kim: Para.[0053]-[0055], the audio decoder may generate a set of higher order ambisonics (HOA) coefficients based on the multi-channel audio signal and the spatial positioning vectors. To reconstruct, the decoder may use rendering matrix D of number of HOA coefficients and number of channels. Para.[0098], Fig. 4, HOA generation unit 208A may generate set of HOA coefficients 212A).
Regarding Claim 3, Kim in view of Lee teach the method of claim 2. Kim further teaches, wherein the representation includes one or more salient audio components associated with the input audio signal and one or more spatial descriptors that describe features of the one or more salient audio components, wherein producing the decoded representation comprises generating the HOA representation of the input audio signal using the one or more salient audio components based on the one or more spatial descriptors ( Kim: Para.[0161]-[0163], Fig.16, demultiplexing unit 202C in the audio decoding device 22C may obtain spatial vector representation data 71B from bitstream 56C. Spatial vector representation data 71B includes data representing spatial vectors for each audio object. When the data representing the spatial vectors is quantized, vector decoding unit 209 may inverse quantize the spatial vectors to determine the spatial vectors 72 ( spatial descriptor) of the audio objects. HOA generation unit 208B may generate an HOA soundfield, such HOA coefficients 212B, based on spatial vectors 72 and audio signal 70).
Claim 12 is a decoder-side device claim performing the steps in method claim 3 above and as such, claim 12 is similar in scope and content to claim 3 and therefore, claim 12 is rejected under similar rationale as presented against claim 3 above.
Regarding Claim 4, Kim in view of Lee teach the method of claim [[2]] 1. Kim further teaches, wherein the input audio signal further comprises a set of one or more audio objects and the metadata comprises positional information relating to the set of one or more audio objects ( Kim: Para.[0043], input audio signals could be object-based audio, which involves discrete pulse-code-modulation (PCM) data for audio objects with associated metadata containing their location coordinates ( positional information) and other information),
wherein the plurality of audio driver signals are produced by spatially rendering the set of one or more audio objects according to the positional information ( Kim: Para.[0098], [0099], Fig. 4, HOA generation unit 208A generate plurality of audio signals, based on spatial positioning vectors ).
Regarding Claim 5, Kim in view of Lee teach the method of claim 4. Kim further teaches, wherein the decoded representation of the input audio signal comprises a first decoded representation of the plurality of [full-range] audio channels, wherein the method further comprises producing a second decoded representation ( Kim: Para.[0191] - [0194], Fig. 18, audio decoding unit 204 of audio decoding device 22D, obtain, from the bitstream, six-channels of audio data in the 5.1 surround sound format and generate multi-channel audio signal 70 ( first decoded presentation, surround format). Inverse quantization unit 550 of audio decoding device 22D , inverse quantize, quantized vector data 554 to generate spatial positioning vectors 72, which are based on source loudspeaker setup information. HOA generation unit 208A of audio decoding device 22D may generate HOA coefficients 212A ( second decoded presentation, surround format) based on multi-channel audio signal 70 and spatial positioning vectors 72. ).
Lee further teaches, wherein the decoded representation of the input audio signal comprises a first decoded representation of the plurality of full-range audio channels ( Lee: Column 16, lines 4-19, Fig. 2, The band combining unit 223 outputs a decoded speech signal by combining the decoded high-band speech signal output by the high-band speech decoding apparatus 221 and the decoded low-band speech signal output by the low-band speech decoding apparatus 222).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Lee’s teaching of a high-band speech encoding and decoding apparatus in wideband speech encoding and decoding method, into the method coding of higher-order ambisonics audio data, taught by Kim, because, this would reproduce high quality sound even at a low bitrate in wideband speech encoding and decoding.(Lee, Column 2, lines 29-59).
Regarding Claim 6, Kim in view of Lee teach the method of claim 4. Kim further teaches, further comprising receiving an output speaker layout for the plurality of speakers, wherein the set of one or more audio objects are spatially rendered according to the output speaker layout ( Kim: Para.[0121], Fig.6, source loudspeaker setup information 48 specifies a CICP speaker layout index. Rendering format unit 110 may determine, based on this CICP speaker layout index, locations of loudspeakers in the source loudspeaker setup. Accordingly, representation unit 115 may include, in spatial vector representation data 71A, an indication of the CICP speaker layout index).
Regarding Claim 11, Kim in view of Lee teach the decoder-side device of claim 10. Kim further teaches, wherein the decoded representation comprises a higher-order ambisonics (HOA) representation of the input audio signal, wherein the memory has further instructions to apply a conversion matrix to the HOA representation to reconstruct the input audio signal, ( Kim: Para.[0053]-[0055], the audio decoder may generate a set of higher order ambisonics (HOA) coefficients based on the multi-channel audio signal and the spatial positioning vectors. To reconstruct, the decoder may use rendering matrix D of number of HOA coefficients and number of channels. Para.[0098], Fig. 4, HOA generation unit 208A may generate set of HOA coefficients 212A. Para.[0043], input audio signals could be object-based audio, which involves discrete pulse-code-modulation (PCM) data for single audio objects with associated metadata containing their location coordinates and other information).
Regarding Claim 13, Kim in view of Lee teach the decoder-side device of claim [[11]] 10. Kim further teaches, wherein the input audio signal further comprises a set of one or more audio objects and the metadata comprises positional information relating to the set of one or more audio objects ( Kim: Para.[0043], input audio signals could be object-based audio, which involves discrete pulse-code-modulation (PCM) data for audio objects with associated metadata containing their location coordinates ( positional information) and other information),
wherein the plurality of audio driver signals are produced by spatially rendering the set of one or more audio objects according to the positional information and a layout of the plurality of speakers ( Kim: Para.[0098], [0099], Fig. 4, HOA generation unit 208A generate plurality of audio signals, based on spatial positioning vectors. Para.[0121], Fig.6, source loudspeaker setup information 48 specifies a CICP speaker layout index. Rendering format unit 110 may determine, based on this CICP speaker layout index, locations of loudspeakers in the source loudspeaker setup. Accordingly, representation unit 115 may include, in spatial vector representation data 71A, an indication of the CICP speaker layout index).
Regarding Claim 17, Kim teaches an encoder-side method, the method comprising:
receiving an input audio signal of a piece of audio content and metadata relating to the input audio signal ( Kim: Para.[0084],[0085], Fig. 3, audio encoding device 14A receives audio signal 50. Para.[0043], input audio signals could be object-based audio, which involves discrete pulse-code-modulation (PCM) data for single audio objects with associated metadata containing their location coordinates and other information),
wherein the input audio signal comprises a plurality of surround-sound audio channels that includes a sound source, the plurality of surround-sound audio channels comprises a first set of one or more full-range audio channels and a second set of one or more band-limited audio channels ( Kim: Para.[0085], Fig. 3, input audio signal 50 may include N channels of audio data denoted as channel C1 through channel CN, with surround sound format. As one example, audio signal 50 may be a six-channel audio signal for a source loudspeaker configuration of 5.1 (i.e., a front-left channel, a center channel, a front-right channel, a surround back left channel, a surround back right channel( first set), and a low-frequency effects (LFE) channel ( second set, band limited audio channel));
converting the first set into a higher-order ambisonics (HOA) representation of the sound source ( Kim: Para.[0089], Fig. 3, bitstream generation unit 52A may determine and encode an indication of how many HOA coefficients are to be used, when converting audio signal 50 into an HOA sound field.) ;
encoding, using [a Matching Pursuit (MP) coding-based algorithm], the HOA representation ( Kim: Para.[0092], audio encoding device 14A may receive a multi-channel audio signal for a source loudspeaker configuration, obtain, based on the source loudspeaker configuration, a plurality of spatial positioning vectors in the Higher-Order Ambisonics (HOA) domain that, in combination with the multi-channel audio signal, represent a set of higher-order ambisonic (HOA) coefficients that represent the multi-channel audio signal; and encode, in a coded audio bitstream (e.g., bitstream 56A)),
and transmitting a bitstream that comprises the encoded HOA representation( Kim: Para.[0090]-[0092], Figs. 3, 4, bitstream 56A ( represent a set of higher-order ambisonic (HOA) coefficients that represent the multi-channel audio signal; and encode, in a coded audio bitstream) generated by the encoding device 14A is transmitted to decoding device 22A. Para.[0096],[0100], Fig. 4, Audio decoding unit 204 may be configured to decode coded audio signal 62 into audio signal 70. Audio decoding unit 204 may decode channels C1-CN of audio signal 62 into channels C1-CN of decoded audio signal 70. Rendering unit 210 may render audio signals 26A for playback at a plurality of local loudspeakers).
Kim while teaching the method of claim 17, fails to explicitly teach the claimed, encoding, using a Matching Pursuit (MP) coding-based algorithm,[ the HOA representation ]
However, Lee does teach the claimed, encoding, using a Matching Pursuit (MP) coding-based algorithm, [ the HOA representation ](Lee: Column 17, lines 29-43, the high-band speech signal is encoded and decoded based on a structure in which a harmonic structure and a stochastic structure is combined. The harmonic structure searches for an amplitude and a phase of a sine wave dictionary using a matching pursuit (MP) algorithm. Hence, the wideband speech encoding and decoding system according to the present invention can reproduce high-quality sound at a low bitrate and with low complexity).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Lee’s teaching of a high-band speech encoding and decoding apparatus in wideband speech encoding and decoding method, into the method coding of higher-order ambisonics audio data, taught by Kim, because, this would reproduce high quality sound even at a low bitrate in wideband speech encoding and decoding.(Lee, Column 2, lines 29-59).
Regarding Claim 18, Kim in view of Lee teach the method of claim 17. Kim further teaches, wherein the encoded HOA representation comprises one or more salient audio components associated with the HOA representation of the sound source and one or more spatial descriptors based on the sound source, wherein the one or more spatial descriptors describe features of the one or more salient audio components ( Kim: Para.[0161], Fig.16, audio decoding device 22C obtains bitstream 56C which may include an encoded object-based audio signal of an audio object and data representative of a spatial vector ( spatial descriptor) of the audio object and in the HOA domain).
Regarding Claim 20, Kim in view of Lee teach the method of claim [[18]] 17. Kim further teaches, wherein the metadata comprises surround-sound speaker layout information for the plurality of surround-sound audio channels ( Kim: Para.[0168], Fig. 17, audio encoding device 14D is encoding channel based audio, may obtain source loudspeaker setup information, source location of an audio object. Para.[0191], audio encoding device 14 may receive six-channels of audio data in the 5.1 surround sound format. Para.[0043], input audio signals could be object-based audio, with associated metadata containing their location coordinates ( positional information) and other information),
wherein the first set is converted into HOA representation according to the surround-sound speaker layout information ( Kim: Para.[0177], Fig. 17, audio encoding device 14D may obtain, based on the source loudspeaker configuration, a plurality of spatial positioning vectors in the Higher-Order Ambisonics (HOA) domain that, in combination with the multi-channel audio signal, represent a set of higher-order ambisonics (HOA) coefficients that represent the multi-channel audio signal).
Claims 7-9, 14-16 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. ( US 20170105085 A1), hereinafter referenced as Kim, in view of Lee et al. (US 7801733 B2), hereinafter referenced as Lee, further in view of Peters et al. ( US 20150127354 A1), hereinafter referenced as Peters.
Regarding Claim 7, Kim in view of Lee teach the method of claim 2. Kim in view of Lee fail to explicitly teach the claimed, wherein the bitstream is received from an encoder-side device, wherein the conversion matrix is an inverse matrix of a matrix used by the encoder-side device to produce the
However, Peters does teach the claimed, wherein the bitstream is received from an encoder-side device, wherein the conversion matrix is an inverse matrix of a matrix used by the encoder-side device to produce the ( Peters: Para.[0785]-[0786], Fig. 60, the encoded bitstream generated by audio encoding device 20 after step 854, is generated by inverse nearfield filtering).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Peters’s teaching of compressing higher order ambisonics (HOA) audio data, into the method, taught by Kim in view of Lee, because, this would improve the coding efficiency.(Peters, Para.[0087]).
Claim 14 is a decoder-side device claim performing the steps in method claim 7 above and as such, claim 14 is similar in scope and content to claim 7 and therefore, claim 14 is rejected under similar rationale as presented against claim 7 above.
Regarding Claim 8, Kim in view of Lee teach the method of claim 1. Kim in view of Lee fail to explicitly teach the claimed, wherein the decoded representation of the input audio signal comprises a first decoded representation and the representation comprises a first encoded representation, wherein the method further comprises: producing a second decoded representation of a mixed audio signal by decoding a second encoded representation from the bitstream using the MP coding-based algorithm; and,audio signal into: a plurality of surround-sound channels of [[a]] surround-sound format, one or more audio objects that include one or more audio signals, and HOA data that includes a plurality of HOA signals.
However, Peters does teach the claimed, wherein the decoded representation of the input audio signal comprises a first decoded representation and the representation comprises a first encoded representation, wherein the method further comprises: producing a second decoded representation of a mixed audio signal by decoding a second encoded representation from the bitstream using the MP coding-based algorithm ( Peters: Para.[0387], Fig. 18 illustrates audio decoding device 354, where the decoded representation 356, from the renderer 355 shows mixed signals such as for 5.1 speakers, 22.2 speakers…Binaural headphones),
and,audio signal into: a plurality of surround-sound channels of [[a]] surround-sound format, one or more audio objects that include one or more audio signals, and HOA data that includes a plurality of HOA signals ( Peters: Para.[0387],[0388], Figs. 18, 19 illustrates audio decoding device 354 and the output representation of the audio signals 356, where 5.1 speakers, 22.2 speakers ( surround sound format), Binaural headphones ( audio object), HOA representation),
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Peters’s teaching of compressing higher order ambisonics (HOA) audio data, into the method, taught by Kim in view of Lee, because, this would improve the coding efficiency.(Peters, Para.[0087]).
Claim 15 is a decoder-side device claim performing the steps in method claim 8 above and as such, claim 15 is similar in scope and content to claim 8 and therefore, claim 15 is rejected under similar rationale as presented against claim 8 above.
Regarding Claim 9, Kim in view of Lee further in view of Peters teach the method of claim 8. Peters further teaches, wherein producing the plurality of audio driver signals further comprises: rendering the plurality of surround-sound channels, the one or more audio signals ( Peters: Para.[0387], Fig. 18 illustrates audio decoding device 354, where the decoded representation 356, from the renderer 355 shows audio drive signals such as for 5.1 speakers, 22.2 speakers ( surround sound), Binaural headphones),
and mixing the renderings into the plurality of audio driver signals ( Peters: Para.[0387], Fig. 18 illustrates audio decoding device 354, where 366 represents the mixing of plurality of audio drive signals),
Kim further teaches, and the plurality of HOA signals according to the metadata and an output speaker layout of the plurality of speakers( Kim: Para.[0098], [0099], Fig. 4, HOA generation unit 208A generate plurality of audio signals, based on spatial positioning vectors. Para.[0121], Fig.6, source loudspeaker setup information 48 specifies a CICP speaker layout index. Rendering format unit 110 may determine, based on this CICP speaker layout index, locations of loudspeakers in the source loudspeaker setup. Accordingly, representation unit 115 may include, in spatial vector representation data 71A, an indication of the CICP speaker layout index ).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Peters’s teaching of compressing higher order ambisonics (HOA) audio data, into the method, taught by Kim in view of Lee, because, this would improve the coding efficiency.(Peters, Para.[0087]).
Claim 16 is a decoder-side device claim performing the steps in method claim 9 above and as such, claim 16 is similar in scope and content to claim 9 and therefore, claim 16 is rejected under similar rationale as presented against claim 9 above.
Regarding Claim 23, Kim in view of Lee teach the method of claim 17. Kim further teaches, wherein the input audio signal further comprises: HOA representation of the piece of audio content and a set of one or more audio objects of the piece of audio content ( Kim: Para.[0085], Fig. 3, audio signal 50 may be a six-channel audio signal for a source loudspeaker configuration of 5.1 (i.e., a front-left channel, a center channel, a front-right channel, a surround back left channel, a surround back right channel, and a low-frequency effects (LFE) channel). Para.[0142], Fig. 13, audio encoding device 14C determine, based on the data indicating the virtual source location for the audio object and data indicating a plurality of loudspeaker locations, a spatial vector of the audio object in a HOA domain),
Lee further teaches,the bitstream for transmission to the audio playback device (Lee: Column 17, lines 29-43, the high-band speech signal is encoded and decoded based on a structure in which a harmonic structure and a stochastic structure is combined. The harmonic structure searches for an amplitude and a phase of a sine wave dictionary using a matching pursuit (MP) algorithm. Hence, the wideband speech encoding and decoding system according to the present invention can reproduce high-quality sound at a low bitrate and with low complexity).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Lee’s teaching of a high-band speech encoding and decoding apparatus in wideband speech encoding and decoding method, into the method coding of higher-order ambisonics audio data, taught by Kim, because, this would reproduce high quality sound even at a low bitrate in wideband speech encoding and decoding.(Lee, Column 2, lines 29-59).
Kim in view of Lee while teaching the method of claim 23, fail to explicitly teach the claimed, wherein the method further comprises: producing a mixed audio signal that includes the plurality of surround-sound audio channels, the HOA representation, and the set of one or more audio objects; and
However, Peters does teach the claimed, wherein the method further comprises producing a mixed audio signal that includes the plurality of surround-sound audio channels, the HOA representation, and the set of one or more audio objects ( Peters: Para.[0387],[0388], Figs. 18, 19 illustrates audio decoding device 354 and the output representation of the audio signals 356, where 5.1 speakers, 22.2 speakers ( surround sound format), Binaural headphones ( audio object), HOA representation).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Peters’s teaching of compressing higher order ambisonics (HOA) audio data, into the method, taught by Kim in view of Lee, because, this would improve the coding efficiency.(Peters, Para.[0087]).
Claims 21, 22 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. ( US 20170105085 A1), hereinafter referenced as Kim, in view of Lee et al. (US 7801733 B2), hereinafter referenced as Lee, further in view of Routray et al. ( Upscaling HOA Signals using Order Recursive Matching Pursuit in Spherical Harmonics Domain, IEEE, 2022), hereinafter referenced as Routray.
Regarding Claim 21, Kim in view of Lee teach the method of claim 17. Kim further teaches, wherein ( Kim: Para.[0139],[0141], Fig. 13, audio encoding device 14C is configured to encode object-based audio data. Bitstream generation unit 52C obtains an audio signal 50B for the audio object),
wherein the method further comprises producing a second HOA representation of the set of one or more audio objects ( Kim: Para.[0142], Fig. 13, audio encoding device 14C determine, based on the data indicating the virtual source location for the audio object and data indicating a plurality of loudspeaker locations, a spatial vector of the audio object in a HOA domain),
Kim in view of Lee while teaching the claim of 21, fail to explicitly teach the claimed, and second HOA representation into the bitstream for transmission to the audio playback device.
However, Routray does teach the claimed, wherein encoding comprises encoding, using the MP coding-based algorithm, the HOA representation into a bitstream for transmission to the audio playback device ( Routray: Section III.C, source is encoded using MP ( matching pursuit) and proposed ORMP ( order recursive matching pursuit) algorithm. Encoded signals are decoded using a regular spaced loudspeaker array. The decoded HOA signals are rendered using 10, 25, and 64 numbers of regularly spaced loudspeakers).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Routray’s teaching of upscaling HOA signals using order recursive matching pursuit in spherical harmonics domain, into the method, taught by Kim in view of Lee, because, by using order recursive matching pursuit, the spatial resolution during the reproduction of the signal, would be improved. (Routray, Section IV).
Regarding Claim 22, Kim in view of Lee, further in view of Routray teach the method of claim 21. Kim further teaches, further comprising: determining a number of the set of one or more audio objects ( Kim: Para.[0100], Fig. 4, audio signals 26A may include audio signals for channels C1 through CL ( set of audio objects));
and determining a conversion matrix based on the number ( Kim: Para.[0101], Fig. 4, Eq. 29 illustrates determining of rendering matrix D ( conversion matrix) based on the number of audio signal 26A),
wherein the second HOA representation is produced by applying the conversion matrix to the set of one or more audio objects ( Kim: Para.[0051], an audio encoder may determine and encode one or more spatial positioning vectors (SPVs) that enable conversion of the encoded audio data into HOA coefficients. Para.[0101], in eq. 29, H represents HOA coefficients and D is the rendering matrix ).
Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. ( US 20170105085 A1), hereinafter referenced as Kim, in view of Lee et al. (US 7801733 B2), hereinafter referenced as Lee, further in view of Routray et al. ( Upscaling HOA Signals using Order Recursive Matching Pursuit in Spherical Harmonics Domain, IEEE, 2022), hereinafter referenced as Routray, further in view of Sun et al. ( Immersive audio, capture, transport, and rendering: a review, Cambridge University, 24th August, 2021).
Regarding Claim 19, Kim in view of Lee teach the method of claim [[18]] 17. Kim in view of Lee fail to explicitly teach the claimed, further comprising encoding the second set into the bitstream separately from the encoded HOA representation.
However, Sun does teach the claimed, further comprising encoding the second set into the bitstream separately from the encoded HOA representation ( Sun: Section III, D, Fig.26 illustrates encoding different sets into bitstreams separately ).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Sun’s teaching of immersive audio background, end to end workflow, covering audio capture, compression, and rendering, into the method, taught by Kim in view of Lee, further in view of Routray, because, by better understanding of the immersive audio system the overall coding efficiency and customer’s experience could be improved. (Sun, Section III, B,IV, E) ).
Claims 24, 25 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. ( US 20170105085 A1), hereinafter referenced as Kim, in view of Lee et al. (US 7801733 B2), hereinafter referenced as Lee, further in view of Multrus et al. (US 20190156843 A1), hereinafter referenced as Multrus.
Regarding Claim 24, Kim in view of Lee teach the method of claim 1. Kim in view of Lee fail to explicitly teach the claimed, wherein the at least one band-limited audio channel is received as part of the bitstream and is encoded according to a different coding-based algorithm, wherein the method further comprises reconstructing the at least one band-limited audio channel by decoding the encoded at least one band-limited audio channel using the different coding- based algorithm.
However, Multrus does teach the claimed, wherein the at least one band-limited audio channel is received as part of the bitstream and is encoded according to a different coding-based algorithm, wherein the method further comprises reconstructing the at least one band-limited audio channel by decoding the encoded at least one band-limited audio channel using the different coding- based algorithm ( Multrus: Para.[0076]-[0078], Fig. 8 illustrates an audio encoder receiving audio signal 103 having a lower frequency band and an upper frequency band ( band limited audio) and encoded audio signal 814 is received after several intermediate steps such as shaper (804), quantizer and coder ( 806). Para.[0123],the resulting decoded signal after attenuation is perceptually significantly more pleasant than before, where huge parts of the spectrum were zeroed out completely. Para. [0210], two different ALFE algorithms are selected consistently in encoder and decoder).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Multrus’s teaching of An audio encoder for encoding an audio signal having a lower frequency band and an upper frequency band, taught by Kim in view of Lee, because, by attenuating the upper frequency band of the audio signal, the overall perceptual quality of the result of the encoding operation can be improved. (Multrus, Para.[0038]).
Regarding Claim 25, Kim in view of Lee teach the method of claim 1. Kim in view of Lee fail to explicitly teach the claimed, wherein the at least one band-limited audio channel is received separately from the bitstream.
However, Multrus does teach the claimed, wherein the at least one band-limited audio channel is received separately from the bitstream ( Multrus: Para.[0076], Fig. 8, audio signal 103 having a lower frequency band and an upper frequency band ( band limited audio)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Multrus’s teaching of An audio encoder for encoding an audio signal having a lower frequency band and an upper frequency band, taught by Kim in view of Lee, because, by attenuating the upper frequency band of the audio signal, the overall perceptual quality of the result of the encoding operation can be improved. (Multrus, Para.[0038]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NADIRA SULTANA whose telephone number is (571)272-4048. The examiner can normally be reached M-F,7:30 am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached on (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NADIRA SULTANA/Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
07/21/2026