DETAILED ACTION
This communication is in response to the Application filed on 03/04/2024 (provisional). Claims 1-20 have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on March 4, 2025 was filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The disclosure is objected to because of the following informalities:
¶ [12] should read: "The techniques of this disclosure provide multi-channel (3 or more) ambience extraction ..."
¶ [27] should read: “... compared to existing techniques, such as Dolby Surround™ ...”
Appropriate correction is required.
The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. Applicant’s cooperation is requested in correcting any errors of which applicant may become aware in the specification.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 9, 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Merimaa et al. (US 8107631 B2) in view of Goodwin et al. (US 20080232617 A1).
Regarding claim 1, Merimaa teaches:
determining a time frequency representation of each of the m channels (see Column 9, lines 16-20, where the method begins with the receipt of a stereo input signal in operation 102. Next, in operation 104, the input signals are converted to a frequency-domain or subband representation using any known method, for example a short-time Fourier transform. Also see Column 13, lines 3-8, where time-to-frequency transform module 504 is configured to convert multichannel input signal 502 into time-frequency representations for any number of channels of the multichannel input signal. Accordingly, the left and right channels are converted into time-frequency representations and outputted from module 504);
performing
simultaneous multichannel surround ambience extraction (see Column 1, lines 21-29, where the stereo signal may be decomposed into a primary component and an ambience component. One common application of these methods is listening enhancement systems where ambient signal components are modified and/or spatially redistributed over multichannel loudspeakers, while primary signal components are unmodified or processed differently. In these systems, the ambience components are typically directed to surround speakers Also see Column 5, lines 55-66, where any input signals at a single frequency band and within a time period of interest {
X
→
L
,
X
→
R
} are assumed to be composed of a single primary component and ambience:
X
→
L
=
P
→
L
+
A
→
L
;
X
→
R
=
P
→
R
+
A
→
R
,
where
P
→
L
and
P
→
R
are the primary components and
A
→
L
and
A
→
R
are the ambient components) and
primary component extraction on the m-channel audio signal, using generalized equal-levels ambience extraction (see Column 14, lines 5-9, where a mixer can be used to subtract the ambience components from multichannel input signal 502 (which includes the primary and ambience components for the right and left channels) in order to extract the primary components from multichannel input signal 502. Also see Column 4, lines 30-37, where in a second embodiment, equal levels of ambience in the respective channels (e.g., left and right channels) of the input signal are assumed. In general, channels of a two-channel input signal are referred to as "left" and "right" channels. These methods provide a further improvement in extracting ambience from input content wherein the dominant (non-ambient) sources are panned to any particular channel. Also see Column 5, lines 55-66, where any input signals at a single frequency band and within a time period of interest {
X
→
L
,
X
→
R
} are assumed to be composed of a single primary component and ambience:
X
→
L
=
P
→
L
+
A
→
L
;
X
→
R
=
P
→
R
+
A
→
R
,
where
P
→
L
and
P
→
R
are the primary components and
A
→
L
and
A
→
R
are the ambient components. Also see Column 8, lines 8-12, where logical assumption for ambience extraction is therefore:
A
→
L
=
A
→
R
=
I
A
,
where the notation
I
A
2
is introduced to denote the ambience level. Also see Column 8, lines 23-25, where the total ambience energy is less than or equal to the total signal energy. This limits the number of solutions to one, yielding
I
A
2
=
1
2
(
r
L
L
+
r
R
R
-
r
L
L
-
r
R
R
2
+
4
r
L
R
2
)
);
Merimaa fails to teach obtaining an encoded signal with 3 or more m-channels, and upmixing these input signals into an n-channel greater than m, assigning an extracted primary component from an m-channel to an n-channel, and obtaining an upmixed signal based on the primary and ambience components, and applying an inverse time frequency operation to the output.
However, Goodwin does teach:
accessing an audio signal encoded in an m-channel format, wherein m is 3 or more (see [0022], where the input signals 101 comprise an ensemble of audio signals, for example a five-channel signal as shown or a two-channel stereo signal. The received input signals 101 are intended for reproduction over a pre-defined loudspeaker layout such as the standard five-channel layout 103);
diffusing the extracted surround ambience to an n-channel audio signal, where n is greater than m (see [0029], where in general, an M-channel to N-channel passive format conversion process can be expressed as an N by M matrix C that generates a set of N output signals from M input signals. Also see [0022], where the actual layout 105 depicts a seven-channel reproduction system with arbitrary loudspeaker positions not configured according to any established standard. Though seven speakers are shown, this is not intended to be limiting. That is, the diagram should be taken as a general representation of the output layout without limitation, including but not limited to limitations as to number or layout of speakers);
assigning each extracted primary component of the m-channel audio signal to a primary component of a channel in the n-channel audio signal (see [0050], where block 903 provides primary components 905 and ambience components 907 as outputs. These are supplied respectively to primary format conversion block 909 and ambience format conversion block 911, which operate in accordance with embodiments of the current invention); and
obtaining an upmixed n-channel audio signal by
performing channel-wise addition of each primary component and ambience component of the n-channel audio signal (see [0050], where blocks 909 and 911 provide format-converted primary channels 913 and format-converted ambience channels 915 to mixer block 917, which combines the primary and ambient channels, in one embodiment as a direct sum and in other embodiments using alternate weights, to determine output signals 919) and
applying an inverse time frequency operation to the channel-wise addition (see [0026], where FIG. 3 depicts a preferred embodiment wherein the format conversion is carried out in the STFT domain. Time-domain input signals 301 are converted to a frequency-domain representation by the short-time Fourier transform block 303. The STFT-domain input signals 305 are then provided to block 307, which implements format conversion based on spatial analysis and synthesis as depicted in block 200 of FIG. 2 and provides STFT-domain output signals 309 to block 311, which generates time-domain output signals 313 via an inverse short-time Fourier transform and overlap-add process).
Merimaa, and Goodwin are considered to be analogous to the claimed invention because they are in the same field of speech analysis and audio data processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Merimaa to incorporate the teachings of Goodwin in order to apply Merimaa’s equal-levels ambience extraction to a multi-channel surround input and generate a higher-channel immersive output to improve immersiveness (see [0007] where provided is a frequency-domain method for format conversion of a multichannel audio signal, intended for playback over a pre-defined loudspeaker layout, in order to achieve accurate spatial reproduction over a different layout potentially comprising a different number of loudspeakers).
Regarding claim 9, Merimaa teaches one or more non-transitory computer readable storage media storing instructions that are operable when executed (see Column 12, lines 56-59, where modules 504, 506, 508, 510, 512 can be implemented as program subroutines that are programmed into a memory and executed by a processor of a computer system).
Regarding the remainder of claim 9, which recites a media, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 1. As detailed in the rejection of claim 1, the disclosed method teaches each step of the apparatus recited in claim 9. Accordingly, claim 9 is rejected for the same reasons set forth in the rejection of claim 1.
Regarding claim 14, Merimaa teaches a system comprising: one or more non-transitory computer readable storage media storing instructions; and one or more processors coupled to the one or more non-transitory computer readable storage media and operable to execute the instructions (see Column 12, lines 56-59, where modules 504, 506, 508, 510, 512 can be implemented as program subroutines that are programmed into a memory and executed by a processor of a computer system).
Regarding the remainder of claim 14, which recites a system, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 1. As detailed in the rejection of claim 1, the disclosed method teaches each step of the system recited in claim 14. Accordingly, claim 14 is rejected for the same reasons set forth in the rejection of claim 1.
Claim(s) 2, 10, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Merimaa et al. (US 8107631 B2) in view of Goodwin et al. (US 20080232617 A1), and further in view of Mitchell (“DTS Neo:X 11.1 Surround Sound Format Debuts”, 2015).
Regarding claim 2, which depends on claim 1, Merimaa in view of Goodwin teaches all the limitations of claim 1 but fails to teach wherein the m-channel format comprises a 5-channel surround format and the n-channel format comprises an 11-channel immersive format.
However, Mitchell does teach wherein the m-channel format comprises a 5-channel surround format and the n-channel format comprises an 11-channel immersive format (see Page 1, lines 1-5¸ “if five speakers are not enough to satisfy your surround sound needs, DTS has developed a new surround sound format to separate sound into 11 speakers, plus a subwoofer. The new 11.1 surround technology is called DTS Neo:X, claiming to be the next-generation of immersive 3D entertainment for home theater enthusiasts and industry audio professionals”. Also see FIG. 2 “DTS Neo:X Inputs and Outputs,” where a source audio signal can be a 5.1 LPCM which can be upmixed to a variety of increased audio channels such as 11.1 LPCM using DTS Neo:X processing).
Merimaa, Goodwin, and Mitchell are considered to be analogous to the claimed invention because they are in the same field of speech analysis and audio data processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Merimaa and Goodwin to incorporate the teachings of Mitchell in order to provide a semi-spherical 3D sound field with height and wide speakers for a more immersive and lifelike audio experience (see Page, lines 5-8, “for cinema, music or gaming entertainment, DTS Neo:X provides a semi-spherical sound field using an 11.1 speaker configuration adding height/wide speakers to create a natural, immersive, spacious and lifelike 3D surround soundscape”).
Regarding claim 10, which depends on claim 9 and recites a media, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 2. As detailed in the rejection of claim 2, the disclosed method teaches each step of the media recited in claim 10. Accordingly, claim 10 is rejected for the same reasons set forth in the rejection of claim 2.
Regarding claim 15, which depends on claim 14 and recites a system, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 2. As detailed in the rejection of claim 2, the disclosed method teaches each step of the system recited in claim 15. Accordingly, claim 15 is rejected for the same reasons set forth in the rejection of claim 2.
Claim(s) 3-4, 11-12, 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Merimaa et al. (US 8107631 B2) in view of Goodwin et al. (US 20080232617 A1), further in view of Mitchell (“DTS Neo:X 11.1 Surround Sound Format Debuts”, 2015), and further in view of Tracey (US 20220139403 A1).
Regarding claim 3, which depends on claim 2, Merimaa in view of Goodwin and further in view of Mitchell teaches all the limitations of claim 2 but fails to teach generating immersive rear and side channels in the n-channel audio signal using immersive rear channels extraction.
However, Tracey does teach generating immersive rear and side channels in the n-channel audio signal using immersive rear channels extraction (See [0031], where left and right channel signals are further divided into left and right front, left and right surround, left and right front height, and left and right back height signals. These divisions are based on the inter-aural correlation coefficient and the degree to which inputs are panned left or right).
Merimaa, Goodwin, Mitchell, and Tracey are considered to be analogous to the claimed invention because they are in the same field of speech analysis and audio data processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Merimaa, Goodwin, and Mitchell to incorporate the teachings of Tracey in order to improve the immersiveness and user experience by incorporating a variety of front, rear, left, right, and overhead channels to the upmixing (see [0018], where legacy audio sources often include only two channels—left and right. Such sources do not have the information that allows height channels to be developed by current sound technologies. Accordingly, the listener cannot enjoy the full immersive surround sound experience from legacy audio sources. Also see [0019], where he present disclosure comprises an up-mixer that is configured to develop two (or more) height channels from audio sources that do not include height-related encoding, e.g., stereo sources with left and right audio signals. Accordingly, the present up-mixing allows a listener to enjoy a more immersive audio experience than is otherwise available in a stereo input. The up-mixing involves determining correlations and normalized channel energies between input audio signals. At least two height channels (e.g., left and right height audio signals) are developed from the correlations and normalized energies).
Regarding claim 4, which depends on claim 3, Merimaa in view of Goodwin, further in view of Mitchell, and further in view of Tracey teaches all the limitations of claim 3. Furthermore, Tracey teaches:
frontal ambience for an overhead front left channel and for an overhead front right channel from a left channel and a right channel in the m-channel format (see [0023], where the height components are used to develop left height and right height channels from input stereo or traditional surround sound content. In some examples the height components are used to develop left front height, right front height, left rear height, and right rear height channels from input stereo or traditional surround sound content. Also see [0030], where the input left and right audio signals are up-mixed by the audio system processor to create a 5.1.4 channel output. The five horizontal channels include left and right front, center, and left and right surround channels. The four height channels include left and right front height and left and right back height channels. Also see [0031], where left and right channel signals are further divided into left and right front, left and right surround, left and right front height, and left and right back height signals. These divisions are based on the inter-aural correlation coefficient and the degree to which inputs are panned left or right).
rear ambience for an overhead rear left channel and an overhead rear right channel from a left surround channel and a right surround channel in the m-channel format (see [0023], where the height components are used to develop left height and right height channels from input stereo or traditional surround sound content. In some examples the height components are used to develop left front height, right front height, left rear height, and right rear height channels from input stereo or traditional surround sound content. Also see [0030], where the input left and right audio signals are up-mixed by the audio system processor to create a 5.1.4 channel output. The five horizontal channels include left and right front, center, and left and right surround channels. The four height channels include left and right front height and left and right back height channels. Also see [0031], where left and right channel signals are further divided into left and right front, left and right surround, left and right front height, and left and right back height signals. These divisions are based on the inter-aural correlation coefficient and the degree to which inputs are panned left or right).
Regarding claim 11, which depends on claim 10 and recites a media, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 3. As detailed in the rejection of claim 3, the disclosed method teaches each step of the media recited in claim 11. Accordingly, claim 11 is rejected for the same reasons set forth in the rejection of claim 3.
Regarding claim 12, which depends on claim 11 and recites a media, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 4. As detailed in the rejection of claim 4, the disclosed method teaches each step of the media recited in claim 12. Accordingly, claim 12 is rejected for the same reasons set forth in the rejection of claim 4.
Regarding claim 16, which depends on claim 15 and recites a system, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 3. As detailed in the rejection of claim 3, the disclosed method teaches each step of the system recited in claim 16. Accordingly, claim 16 is rejected for the same reasons set forth in the rejection of claim 3.
Regarding claim 17, which depends on claim 16 and recites a system, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 4. As detailed in the rejection of claim 4, the disclosed method teaches each step of the system recited in claim 17. Accordingly, claim 17 is rejected for the same reasons set forth in the rejection of claim 4.
Claim(s) 5-6, 13, 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Merimaa et al. (US 8107631 B2) in view of Goodwin et al. (US 20080232617 A1), further in view of Mitchell (“DTS Neo:X 11.1 Surround Sound Format Debuts”, 2015), further in view of Tracey (US 20220139403 A1), and further in view of Kyriakakis et al. (US 20220400351 A1).
Regarding claim 5, which depends on claim 3, Merimaa in view of Goodwin, further in view of Mitchell, and further in view of Tracey teaches all the limitations of claim 3 but fails to teach where the method occurs in response to a request to play the audio signal.
However, Kyriakakis does teach where the method occurs in response to a request to play the audio signal (see [0048], where a track may need to be upmixed into a higher number of channels immediately with as little lag as possible. Systems and methods described herein can upmix audio tracks to higher channel formats in near real time. Also see [0052], where audio upmixing processes described herein can operate in real time. For example, processes described herein can upmix a stereo audio stream to a 5.1 channel stream which is played back using speakers designed and/or placed to render 5.1 channel audio without noticeable latency to the user. As can be readily appreciated, a stereo to 5.1 upmix is merely an example, and any arbitrary number of channels can be upmixed using processes described herein. However, in order to provide a concrete example to enhance understanding, an upmix from stereo to 5.1 channel surround sound is used as an example below).
Merimaa, Goodwin, Mitchell, Tracey, and Kyriakakis are considered to be analogous to the claimed invention because they are in the same field of speech analysis and audio data processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Merimaa, Goodwin, Mitchell, and Tracey to incorporate the teachings of Kyriakakis in order to provide the upmixed audio immediately and prevent user-perceptible delay that could degrade the listening experience (see [0048], where home surround sound systems are often provided music as a source input that is not in 1:1 channel format with the speaker layout, but the listener expects for the music they've selected to be immediately played back from all the loudspeakers in their system. As such, a track may need to be upmixed into a higher number of channels immediately with as little lag as possible. Systems and methods described herein can upmix audio tracks to higher channel formats in near real time).
Regarding claim 6, which depends on claim 5, Merimaa in view of Goodwin, further in view of Mitchell, further in view of Tracey, and further in view of Kyriakakis teaches all the limitations of claim 5. Furthermore, Kyriakakis teaches wherein the method occurs substantially in real time with the request (see [0048], where a track may need to be upmixed into a higher number of channels immediately with as little lag as possible. Systems and methods described herein can upmix audio tracks to higher channel formats in near real time. Also see [0052], where audio upmixing processes described herein can operate in real time. For example, processes described herein can upmix a stereo audio stream to a 5.1 channel stream which is played back using speakers designed and/or placed to render 5.1 channel audio without noticeable latency to the user. As can be readily appreciated, a stereo to 5.1 upmix is merely an example, and any arbitrary number of channels can be upmixed using processes described herein. However, in order to provide a concrete example to enhance understanding, an upmix from stereo to 5.1 channel surround sound is used as an example below).
Regarding claim 13, which depends on claim 11 and recites a media, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 5. As detailed in the rejection of claim 5, the disclosed method teaches each step of the media recited in claim 13. Accordingly, claim 13 is rejected for the same reasons set forth in the rejection of claim 5.
Regarding claim 18, which depends on claim 16 and recites a system, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 5. As detailed in the rejection of claim 5, the disclosed method teaches each step of the system recited in claim 18. Accordingly, claim 18 is rejected for the same reasons set forth in the rejection of claim 5.
Regarding claim 19, which depends on claim 18 and recites a system, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 6. As detailed in the rejection of claim 6, the disclosed method teaches each step of the system recited in claim 19. Accordingly, claim 19 is rejected for the same reasons set forth in the rejection of claim 6.
Claim(s) 7, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Merimaa et al. (US 8107631 B2) in view of Goodwin et al. (US 20080232617 A1), and further in view of Kim et al. (US 20200221242 A1).
Regarding claim 7, which depends on claim 1, Merimaa in view of Goodwin teaches all the limitations of claim 1 but fail to teach wherein the method is performed by a smart TV.
However, Kim does teach wherein the method is performed by a smart TV (see [0058], where the image display apparatus 100 in FIG. 1 may be a TV, a monitor, a tablet PC, a mobile terminal, a display for a vehicle, etc. Also see [0059], where the image display apparatus 100 may upmix an input audio signal of stereo channel into an audio signal of multichannel using a deep neural network. Also see [0060], where the image display apparatus 100 according to an embodiment of the present disclosure includes a converter 1010 for frequency converting an input stereo audio signal, a primary component analyzer 1030 that performs primary component analysis based on the signal from the converter 1010, a feature extractor 1040 for extracting a feature of the primary component signal based on the signal from the primary component analyzer 1030, an envelope adjustor 1060 for performing envelope adjustment based on prediction performed on the basis of a deep neural network model, and an inverse converter 1070 for inversely converting a signal from the envelope adjustor 1060 to output an upmix audio signal of multichannel).
For purposes of this rejection, the term “smart TV” is a known term in the consumer electronics field and is interpreted under broadest reasonable interpretation as a television with an integrated processor and internet connectivity that enables advanced processing capabilities, including the ability to run applications and perform video/audio processing tasks. While the claim language uses the term “smart TV,” Kim does not use the same terminology when referring to a television with processing capabilities and instead uses “TV.” However, this mere difference in terminology does not create a substantive distinction and serve the same function of a television capable of performing video and audio processing.
Merimaa, Goodwin, and Kim are considered to be analogous to the claimed invention because they are in the same field of speech analysis and audio data processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Merimaa and Goodwin to incorporate the teachings of Kim in order to enable the TV to perform multi-channel upmixing, because content is often delivered in a downmixed form which can hinder the user’s viewing experience (see [0007], where recently, for realistic audio for ultra-high definition such as UHDTV, it has been further deepened into a multi-channel capable of expressing three-dimensional space such as a 5.1.2 channel or a 22.2 channel. [0008] However, low-quality stereo sound source or multi-channel is downmixed and delivered to consumers, due to problems such as high cost of contents production, transmission equipment for transmitting contents to consumers, constraints of wired and wireless environments, and price competitiveness of audio play apparatus. [0009] In order to effectively play such a downmix two-channel stereo sound source in a multichannel audio play apparatus, a multichannel upmix method is required).
Regarding claim 20, which depends on claim 14, Merimaa in view of Goodwin teaches all the limitations of claim 1 but fails to teach a smart TV that contains the media and the one or more processors.
However, Kim does teach a smart TV that contains the media and the one or more processors (see [0062], where referring to FIG. 2, the image display apparatus 100 according to an embodiment of the present disclosure includes a broadcast receiving unit 105, a storage unit 140, a user input interface 150, a sensor unit (not shown), a signal processing device 170, a display 180, an audio output unit 185, and an illumination sensor 197. Also see [0077], where the storage unit 140 may store a program for each signal processing and control in the signal processing device 170, and may store signal-processed image, audio, or data signal. Also see [0104], where the signal processing device 170 according to an embodiment of the present disclosure may include a demultiplexer 310, an image processing unit 320, a processor 330, and an audio processing unit 370. In addition, it may further include a data processing unit (not shown)).
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Merimaa et al. (US 8107631 B2) in view of Goodwin et al. (US 20080232617 A1), and further in view of Oh et al. (US 10158964 B2).
Regarding claim 8, which depends on claim 1, Merimaa in view of Goodwin teaches all of the limitations of claim 1 but fails to teach multimedia content identification.
However, Oh does teach wherein the audio signal is part of a multimedia content comprising audio and one or more images (see Abstract, where provided are an audio signal processing method and apparatus for adjusting a location of an audio object in correspondence to a location of a visual object. Also see Column 5, lines 33-35, the video signal from which the visual object V obj is extracted and the audio signal A_sig may be signals included in the same multimedia content); and the method further comprises:
identifying one or more objects in at least some of the one or more images (see Column 7, lines 28-40, where the object extracting unit 140 may receive a video signal V_sig or an audio signal A_sig and obtain at least one visual object V_obj from the video signal V_sig and obtain at least one audio object A_sig from the audio signal A_sig. Here, the video signal V_sig may directly include a visual object that exists separately, and the object extracting unit 140 may obtain a visual object by separating or distinguishing the visual object from the video signal. Alternatively, the object extracting unit 140 may extract a visual object V_obj from a video signal through various image signal processing techniques. The extraction of the visual object V_obj may be performed based on the visual feature of each part of the image of the video signal V_sig. Also see Column 10, lines 48-51, where the candidate visual object estimator 240 may receive the video signal V _sig and extract at least one candidate visual object CVO from the video signal V _sig);
determining an image position of each of the identified one or more objects (see Column 7, lines 55-58, where the object extracting unit 140 may calculate the location of a visual object based on the visual feature of a video signal and calculate the location of an audio object based on the acoustic feature of an audio signal. Also see Column 5, lines 48-55, where the visual object V _obj and the audio object A_obj may have a location value (e.g., information) for a predetermined reference location. That is, the location of the visual object V obj may be calculated during a process of extracting the visual object V _obj from a video signal or when the image pattern related to the visual object V obj is generated, the location may be directly assigned);
determining, for at least some of the one or more objects, a portion of the audio signal that corresponds to that object (see Column 5, lines 63-67 and Column 6, line 1, where the matching unit 110 may receive at least one of the visual object (V_obj) information and at least one audio object (A_obj) information, and may select the related visual object and audio object. Alternatively, the matching unit 110 may select the audio object A_obj corresponding to the visual object V_obj. Also see Column 11, lines 13-21, where the matching unit 210 may compare the candidate visual object CVO and the candidate audio object CAO received from the candidate visual object estimator 240 and the candidate audio object estimator 250 and select visual objects and audio objects matching (or related to) each other based on the comparison result. Also see Column 14, lines 42-47, where an audio signal processing apparatus according to an embodiment of the present invention may compare the locational change of the visual object with the locational change of the audio object and match the visual object and the audio object that represent the same or similar locational change);
determining, for each of the least some one or more objects, and based on
the image position of the object and (2) the portion of the audio signal that corresponds to that object, a spatial rendering for the portion of the audio signal (see Column 6, lines 22-28, where the location adjusting unit 120 may adjust the location of a sound image of the audio signal A_sig based on the location A_obj_loc of the selected audio object and the location V_obj_loc of a visual object corresponding to the selected audio object. Also see Column 17, lines 44-54, where the audio signal processing apparatus may adjust the sound image of the audio signal based on the angular difference between the two objects. According to FIG. 6B, the audio signal processing apparatus may rotate the sound image of the audio signal around a predetermined reference location in a virtual acoustic space according to the audio signal. Accordingly, the location of the audio object A_obj may be changed to the location of A_obj', that is, the location of the visual object. Here, the audio signal processing apparatus may rotate the entire sound image of the audio signal); and
enhancing the n-channel audio signal with each determined spatial rendering (see Column 6, lines 45-51, where the output unit 130 may output an audio signal. The output unit 130 may include an audio output module for generating sound ( or audio) that is a physical phenomenon based on an audio signal that is an electrical signal. According to a preferred embodiment of the present invention, the output unit 130 may output the audio signal A_sig' whose location of the sound image is adjusted).
Merimaa, Goodwin, and Oh are considered to be analogous to the claimed invention because they are in the same field of speech analysis and audio data processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Merimaa and Goodwin to incorporate the teachings of Oh in order to improve the accuracy and immersiveness of the upmixed audio by aligning spatial position of audio objects with the corresponding visual objects in the video (see Column 1, lines 22-36, where the sense of immersion is an important factor in next generation contents such as 360-degree contents or VR contents. The content having excellent sense of immersion may make a user feel as if he is present in the virtual world in the content, and provide a user with a near-real experience. In order to give a sense of immersion to contents during the production of the contents, various issues should be considered. First, the video and audio of the multimedia contents should basically harmonize with each other. That is, the moment when video content changes and the moment when audio content changes are required to coincide with each other temporally, and audio content related to video content should be located at the location where the video content exists).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure:
Avendano et al. (U.S. Patent 7412380): Ambience Extraction And Modification For Enhancement And Upmix Of Audio Signals.
Seldess (U.S. PG Pub. 2019/0297447): Multi-channel Subband Spatial Processing For Loudspeakers.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHN HONG FANG-WU whose telephone number is (571)270-0607. The examiner can normally be reached Monday - Friday, 9AM to 5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras Shah can be reached at (571)-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JOHN HONG FANG-WU/Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
09/20/2026