DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the response to this office action, the Examiner respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Examiner in prosecuting this application.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) ELEMENT IN CLAIM FOR A COMBINATION.—An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claim in this application is given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f):
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f). The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) because the claim limitations use a generic placeholder such as “unit” that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: a first output unit” in claim 8.
Because this claim limitations are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
However, a review of the specification shows none of structure described in the specification for the 35 U.S.C. 112(f) that achieves the claimed functions “outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering”. For example, metadata comprising “speaker zone constraint data by “The metadata may, for example, include data indicating the position and/or trajectory of an audio object in a three-dimensional space, speaker zone constraint data, etc.”, and there is no disclosure to support “outputting” both “the decorrelated audio object audio signals and the speaker zone constraint data for rendering” and this triggers 35 U.S.C. 112(b) issue as set forth below.
If applicant wishes to provide further explanation or dispute the examiner’s interpretation of the corresponding structure, applicant must identify the corresponding structure with reference to the specification by page and line number, and to the drawing, if any, by reference characters in response to this Office action.
If applicant does not intend to have the claim limitation(s) treated under 35 U.S.C. 112(f) applicant may amend the claim(s) so that it/they will clearly not invoke 35 U.S.C. 112(f) or present a sufficient showing that the claim recites/recite sufficient structure, material, or acts for performing the claimed function to preclude application of 35 U.S.C. 112(f).
For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S.C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9, 2011).
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement.
Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b).
Claim 1-8 rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-3, 5-9 of U.S. Patent No. 12,212,953 B2 and in view of Kuech et al. ( US 20120076308 A1, hereinafter Kuech). The conflicting claims 1-3, 5-9 of U.S. Patent No. 12,212,953 B2 do not explicitly teach the limitation “receiving, from an interface system, speaker zone constraint data” and “outputting … the speaker zone constraint data for rendering” as recited in instant claims 1, 7-8. However, Kuech teaches this limitation (side information 320 received by a parameter processor 480 and outputting corresponding matrix information representing the side information to the pre-mixing matrix M1 and post-mixing matrix M2 in fig. 4 and wherein speaker zone constraint data is included in the side information, para 76, or inputting playback configuration 570 and object positions 580 to scene rendering engine 540 and outputting rendering matrix representing the to the loudspeaker configuration 570 and object positions 580 to the scene rendering engine 540 in fig. 6B) for benefits of improving the audio playback performance (by assuring high audio quality with efficient representation of the multichannel audio signals, para 10 and efficient combination of sound effect and noise cancellation, para 41). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied receiving, from an interface system, speaker zone constraint data and outputting … the speaker zone constraint data for rendering, as taught by Kuech, to the audio processing method, as taught by conflicting claims 1-3, 5-9 of U.S. Patent No. 12,212,953 B2, for benefits discussed above. The following is the comparison between claims 1-8 of the instant application and the conflicting claims 1-3, 5-9 of the U.S. Patent No. 12,212,953 B2:
Claims 1-8 in the current application
Conflicting claims 1-3, 5-9 of U.S. Patent No. 12,212,953 B2
1. A method for audio processing, the method comprising: receiving audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; receiving, from an interface system, speaker zone constraint data; determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
4. The method of claim 1, wherein performing decorrelation filtering includes at least one of a delay and a filter.
5. The method of claim 1, wherein performing decorrelation filtering includes at least one of an all-pass filter and a pseudo-random filter.
6. The method of claim 1, wherein performing decorrelation filtering includes a reverberation process.
7. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
8. An apparatus for audio processing, the apparatus comprising: a receiver configured to receive audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; and a decorrelator configured to perform, based on a determination that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; a first output unit for outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
1. A method for audio processing, the method comprising: receiving audio data comprising at least one low frequency (LFE) channel, at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object, wherein the audio object metadata comprises a size of the at least one audio object and a flag indicating whether the at least one audio object is spatially diffuse; performing, based on a determination that the at least one audio object is spatially diffuse indicating that the at least one audio object has a perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals; and outputting the LFE channel.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
5. The method of claim 1, wherein performing decorrelation filtering includes at least one of a delay and a filter.
6. The method of claim 1, wherein performing decorrelation filtering includes at least one of an all-pass filter and a pseudo-random filter.
7. The method of claim 1, wherein performing decorrelation filtering includes a reverberation process.
8. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
9. An apparatus for audio processing, the apparatus comprising: a receiver configured to receive audio data comprising at least one low frequency (LFE) channel, at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object, wherein the audio object metadata comprises a size of the at least one audio object and a flag indicating whether the at least one audio object is spatially diffuse; and a decorrelator configured to perform, based on a determination that the at least one audio object is spatially diffuse indicating that the at least one audio object has a perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; a first output unit for outputting the decorrelated audio object audio signals; and a second output unit for outputting the LFE channel.
Claim 1-8 rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-4, 7-9, 12-13 of U.S. Patent No. 11,736,890 B2 and in view of Kuech et al. ( US 20120076308 A1, hereinafter Kuech). The conflicting claims 1-4, 7-9, 12-13 of U.S. Patent No. 11,736,890 B2 do not explicitly teach the limitation “receiving, from an interface system, speaker zone constraint data” and “outputting … the speaker zone constraint data for rendering” as recited in claims 1, 7-8 of the instant application. However, Kuech teaches this limitation (side information 320 received by a parameter processor 480 and outputting corresponding matrix information representing the side information to the pre-mixing matrix M1 and post-mixing matrix M2 in fig. 4 and wherein speaker zone constraint data is included in the side information, para 76, or inputting playback configuration 570 and object positions 580 to scene rendering engine 540 and outputting rendering matrix representing the to the loudspeaker configuration 570 and object positions 580 to the scene rendering engine 540 in fig. 6B) for benefits of improving the audio playback performance (by assuring high audio quality with efficient representation of the multichannel audio signals, para 10 and efficient combination of sound effect and noise cancellation, para 41). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied receiving, from an interface system, speaker zone constraint data and outputting … the speaker zone constraint data for rendering, as taught by Kuech, to the audio processing method, as taught by conflicting claims 1-4, 7-9, 12-13 of U.S. Patent No. 11,736,890 B2, for benefits discussed above. The following is the comparison between claims of the instant application and the conflicting claims of the U.S. Patent No. 11,736,890 B2:
Claim(s) in the current application
Conflicting U.S. Patent No. 11,736,890 B2
1. A method for audio processing, the method comprising: receiving audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; receiving, from an interface system, speaker zone constraint data; determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
4. The method of claim 1, wherein performing decorrelation filtering includes at least one of a delay and a filter.
5. The method of claim 1, wherein performing decorrelation filtering includes at least one of an all-pass filter and a pseudo-random filter.
6. The method of claim 1, wherein performing decorrelation filtering includes a reverberation process.
7. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
8. An apparatus for audio processing, the apparatus comprising: a receiver configured to receive audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; and a decorrelator configured to perform, based on a determination that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; a first output unit for outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
1. A method, comprising: receiving audio data comprising at least one audio object, wherein the audio data includes at least one audio signal and audio object metadata, wherein the at least one audio signal is associated with the at least one audio object and the audio object metadata is associated with the at least one audio object, wherein the audio object metadata comprises a size of the at least one audio object and a flag indicating whether the at least one audio object is spatially diffuse; performing, based on a determination that the at least one audio object is spatially diffuse indicating that the at least one audio object has a perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; and outputting the decorrelated audio object audio signals.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
4. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
7. The method of claim 1, wherein performing decorrelation includes at least one of a delay and a filter.
8. The method of claim 1, wherein performing decorrelation includes at least one of an all-pass filter and a pseudo-random filter.
9. The method of claim 1, wherein performing decorrelation includes a reverberation process.
12. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
13. An apparatus, comprising: a receiver configured to receive audio data comprising at least one audio object, wherein the audio data includes at least one audio signal and audio object metadata, wherein the at least one audio signal is associated with the at least one audio object and the audio object metadata is associated with the at least one audio object, wherein the audio object metadata comprises a size of the at least one audio object and a flag indicating whether the at least one audio object is spatially diffuse; a decorrelator configured to perform, based on a determination that the at least one audio object is spatially diffuse indicating that the at least one audio object has a perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers, and output the decorrelated audio object audio signals.
Claim 1-8 rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claim 1-4, 6-8, 12-13 of U.S. Patent No. 11,064,310 B2 and in view of Kuech et al. ( US 20120076308 A1, hereinafter Kuech). The conflicting claims 1-3, 6-8, 12-13 of U.S. Patent No. 11,064,310 B2 do not explicitly teach the limitation “receiving, from an interface system, speaker zone constraint data” and “outputting … the speaker zone constraint data for rendering” as recited in claims 1, 7-8 of the instant application. However, Kuech teaches this limitation (side information 320 received by a parameter processor 480 and outputting corresponding matrix information representing the side information to the pre-mixing matrix M1 and post-mixing matrix M2 in fig. 4 and wherein speaker zone constraint data is included in the side information, para 76, or inputting playback configuration 570 and object positions 580 to scene rendering engine 540 and outputting rendering matrix representing the to the loudspeaker configuration 570 and object positions 580 to the scene rendering engine 540 in fig. 6B) for benefits of improving the audio playback performance (by assuring high audio quality with efficient representation of the multichannel audio signals, para 10 and efficient combination of sound effect and noise cancellation, para 41). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied receiving, from an interface system, speaker zone constraint data and outputting … the speaker zone constraint data for rendering, as taught by Kuech, to the audio processing method, as taught by conflicting claims 1-3, 6-8, 12-13 of U.S. Patent No. 11,064,310 B2, for benefits discussed above. The following is the comparison between claims of the instant application and the conflicting claims of the U.S. Patent No. 11,064,310 B2:
Claim(s) in the current application
Conflicting U.S. Patent No. 11,064,310 B2
1. A method for audio processing, the method comprising: receiving audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; receiving, from an interface system, speaker zone constraint data; determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
4. The method of claim 1, wherein performing decorrelation filtering includes at least one of a delay and a filter.
5. The method of claim 1, wherein performing decorrelation filtering includes at least one of an all-pass filter and a pseudo-random filter.
6. The method of claim 1, wherein performing decorrelation filtering includes a reverberation process.
7. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
8. An apparatus for audio processing, the apparatus comprising: a receiver configured to receive audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; and a decorrelator configured to perform, based on a determination that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; a first output unit for outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
1. A method, comprising: receiving audio data comprising at least one audio object, wherein the audio data include at least one audio signal and audio object metadata, wherein the at least one audio signal is associated with the at least one audio object and the audio object metadata is associated with the at least one audio object, wherein the metadata includes a flag relating to size of the at least one audio object, and wherein the metadata further includes additional metadata that has at least one of audio object position data, audio object gain data, audio object size data, audio object trajectory data, and speaker zone constraint data; determining based on the flag that the size of the at least one audio object is greater than a threshold size value; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals; and mixing, based on the additional metadata, the decorrelated audio object audio signals with the at least one audio signal to determine a mixed audio signal for rendering.2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
6. The method of claim 1, wherein performing decorrelation includes at least one of a delay and a filter.
7. The method of claim 1, wherein performing decorrelation includes at least one of an all-pass filter and a pseudo-random filter.
8. The method of claim 1, wherein performing decorrelation includes a reverberation process.
12. A non-transitory medium having software stored thereon, the software including instructions implemented by at least one apparatus to perform the method of claim 1.
13. An apparatus, comprising: a receiver configured to receive audio data comprising at least one audio object, wherein the audio data include at least one audio signal and audio object metadata, wherein the at least one audio signal is associated with the at least one audio object and the audio object metadata is associated with the at least one audio object, wherein the metadata includes a flag relating to size of the at least one audio object, and wherein the metadata further includes additional metadata that has at least one of audio object position data, audio object gain data, audio object size data, audio object trajectory data, and speaker zone constraint data; a processor configured to determining based on the flag that the size of the at least one audio object is greater than a threshold size value; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals; and mixing, based on the additional metadata, the decorrelated audio object audio signals with the at least a one audio signal to determine a mixed audio signal for rendering.
Claim 1-8 rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claim 1-4, 6-8, 10, 18 of U.S. Patent No. 10,595,152 B2 and in view of references Kuech et al. ( US 20120076308 A1, hereinafter Kuech). The conflicting claims 1-3, 6-8, 12-13 of U.S. Patent No. 10,595,152 B2 do not explicitly teach the limitation “receiving, from an interface system, speaker zone constraint data” and “outputting … the speaker zone constraint data for rendering” as recited in claims 1, 7-8 of the instant application. However, Kuech teaches this limitation (side information 320 received by a parameter processor 480 and outputting corresponding matrix information representing the side information to the pre-mixing matrix M1 and post-mixing matrix M2 in fig. 4 and wherein speaker zone constraint data is included in the side information, para 76, or inputting playback configuration 570 and object positions 580 to scene rendering engine 540 and outputting rendering matrix representing the to the loudspeaker configuration 570 and object positions 580 to the scene rendering engine 540 in fig. 6B) for benefits of improving the audio playback performance (by assuring high audio quality with efficient representation of the multichannel audio signals, para 10 and efficient combination of sound effect and noise cancellation, para 41). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied receiving, from an interface system, speaker zone constraint data and outputting … the speaker zone constraint data for rendering, as taught by Kuech, to the audio processing method, as taught by conflicting claims 1-4, 6-8, 10, 18 of U.S. Patent No. 10,595,152 B2, for benefits discussed above. The following is the comparison between claims of the instant application and the conflicting claims of U.S. Patent No. 10,595,152 B2:
Claim(s) in the current application
Conflicting U.S. Patent No. 10,595,152 B2
1. A method for audio processing, the method comprising: receiving audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; receiving, from an interface system, speaker zone constraint data; determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
4. The method of claim 1, wherein performing decorrelation filtering includes at least one of a delay and a filter.
5. The method of claim 1, wherein performing decorrelation filtering includes at least one of an all-pass filter and a pseudo-random filter.
6. The method of claim 1, wherein performing decorrelation filtering includes a reverberation process.
7. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
8. An apparatus for audio processing, the apparatus comprising: a receiver configured to receive audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; and a decorrelator configured to perform, based on a determination that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; a first output unit for outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
1. A method, comprising: receiving audio data comprising at least one audio object and metadata associated with the at least one audio object, the metadata including data relating to size of the at least one audio object; determining that the size of the at least one audio object is greater than a threshold size value based on a flag of the metadata; performing decorrelation on the at least one audio object to determine decorrelated audio object audio signals; and mixing the decorrelated audio object audio signals with at least an audio signal for the at least one audio object to determine a mixed audio signal for rendering.
4. The method of claim 1, wherein an actual playback speaker configuration is used to render the mixed audio signal to speakers of a playback environment.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.6. The method of claim 1, whereinperforming decorrelation includes at least one of a delay and a filter.
7. The method of claim 1, wherein performing decorrelation includes at least one of an all-pass filter and a pseudo-random filter.
8. The method of claim 1, wherein performing decorrelation includes a reverberation process.18. A non-transitory medium having software stored thereon, the software including instructions for controlling at least one apparatus to: receive audio data comprising at least one audio object and metadata associated with the at least one audio object, the metadata including data relating to size of the at least one audio object; determine that the size of the at least one audio object is greater than a threshold size value based on a flag of the metadata; perform decorrelation on the at least one audio object to determine decorrelated audio object audio signals; and mix the decorrelated audio object audio signals with at least an audio signal for the at least one audio object to determine a mixed audio signal for rendering.
10. An apparatus, comprising: an interface system; and a logic system configured to: receive, via the interface system, audio data comprising at least one audio object and metadata associated with the at least one audio object, the metadata including data relating to size of the at least one audio object; determine that the size of the at least one audio object is greater than a threshold size value based on a flag of the metadata; perform decorrelation on the at least one audio object to determine decorrelated audio object audio signals; and mix the decorrelated audio object audio signals with at least an audio signal for the at least one audio object to determine a mixed audio signal for rendering.
Claims 1-8 rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1, 3, 5, 18, 35 of U.S. Patent No. US 10,003,907 B2 and in view of reference Kuech et al. ( US 20120076308 A1, hereinafter Kuech). The conflicting claims 1, 3, 5, 18, 35 of U.S. Patent No. 10,003,907 B2 do not explicitly teach the limitation “receiving, from an interface system, speaker zone constraint data” and “outputting … the speaker zone constraint data for rendering” as recited in claims 1, 7-8 of the instant application. However, Kuech teaches this limitation (side information 320 received by a parameter processor 480 and outputting corresponding matrix information representing the side information to the pre-mixing matrix M1 and post-mixing matrix M2 in fig. 4 and wherein speaker zone constraint data is included in the side information, para 76, or inputting playback configuration 570 and object positions 580 to scene rendering engine 540 and outputting rendering matrix representing the to the loudspeaker configuration 570 and object positions 580 to the scene rendering engine 540 in fig. 6B) for benefits of improving the audio playback performance (by assuring high audio quality with efficient representation of the multichannel audio signals, para 10 and efficient combination of sound effect and noise cancellation, para 41). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied receiving, from an interface system, speaker zone constraint data and outputting … the speaker zone constraint data for rendering, as taught by Kuech, to the audio processing method, as taught by conflicting claims 1, 3, 5, 18, 35 of U.S. Patent No. 10,003,907 B2, for benefits discussed above. An comparison of claims 1-8 of the instant application with the conflicting claims 1, 3, 5, 18, 35 of U.S. Patent No. US 10,003,907 B2, is listed below:
Claim(s) in the current application
Conflicting claim(s) in US Patent No. US 10,003,907 B2
1. A method for audio processing, the method comprising: receiving audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; receiving, from an interface system, speaker zone constraint data; determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
4. The method of claim 1, wherein performing decorrelation filtering includes at least one of a delay and a filter.
5. The method of claim 1, wherein performing decorrelation filtering includes at least one of an all-pass filter and a pseudo-random filter.
6. The method of claim 1, wherein performing decorrelation filtering includes a reverberation process.
7. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
8. An apparatus for audio processing, the apparatus comprising: a receiver configured to receive audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; and a decorrelator configured to perform, based on a determination that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; a first output unit for outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
1. A method, comprising: receiving, in an input interface to an encoder component of an audio rendering system, audio data comprising audio objects, the audio objects comprising audio object signals and associated metadata, the associated metadata including at least audio object size data; determining, by a large object detection component based on the audio object size data, a large audio object having an audio object size that is greater than a threshold size, wherein the large audio object is spatially diffuse and requires a plurality of speakers to reproduce the large audio object; and performing, in a decorrelator component coupled to the input interface, a decorrelation process on audio signals of the large audio object to produce decorrelated large audio object audio signals that are dependent on a defined location of the large audio object and other information, wherein the decorrelated large audio object audio signals are mutually independent of one another, and the decorrelation process comprises adjusting a level of each of the audio signals by adjusting a respective audio gain for each of the audio signals to generate the decorrelated large audio object audio signals corresponding to a speaker feed to each speaker of the plurality of speakers, and further wherein the plurality of speakers covers a large spatial area, wherein the decorrelated large audio object audio signals are mixed with at least one audio signal for other audio objects that are spatially separated by a second threshold amount of distance from the large audio object.
3. The method of claim 1, wherein the large audio object has a plurality of object locations, wherein at least some of the plurality of object locations are one of: stationary locations or locations that vary over time.
3. The method of claim 1, wherein the large audio object has a plurality of object locations, wherein at least some of the plurality of object locations are one of: stationary locations or locations that vary over time.
5. The method of claim 1, wherein the decorrelation process comprises one of: a delay process, an all-pass filter process, a pseudo-random filter process, and a reverberation process.
5. The method of claim 1, wherein the decorrelation process comprises one of: a delay process, an all-pass filter process, a pseudo-random filter process, and a reverberation process.
5. The method of claim 1, wherein the decorrelation process comprises one of: a delay process, an all-pass filter process, a pseudo-random filter process, and a reverberation process.
35. A non-transitory medium having stored thereon programming instructions, which when executed by a processing component in an audio rendering system cause the audio rendering system to: receiving, in an input interface of the audio rendering system to an encoder component of the audio rendering system, audio data comprising audio objects, the audio objects comprising audio object signals and associated metadata, the associated metadata including at least audio object size data; determining, by a large object detection component of the audio rendering system based on the audio object size data, a large audio object having an audio object size that is greater than a threshold size, wherein the large audio object is spatially diffuse and requires a plurality of speakers to reproduce the large audio object; and performing, in a decorrelator component of the audio rendering system coupled to the input interface, a decorrelation process on audio signals of the large audio object to produce decorrelated large audio object audio signals that are dependent on a defined location of the large audio object and other information, wherein the decorrelated large audio object audio signals are mutually independent of one another, and the decorrelation process comprises adjusting a level of each of the audio signals by adjusting a respective audio gain for each of the audio signals to generate the decorrelated large audio object audio signals corresponding to a speaker feed to each speaker of the plurality of speakers, and further wherein the plurality of speakers covers a large spatial area, wherein the decorrelated large audio object audio signals are mixed with at least one audio signal for other audio objects that are spatially separated by a second threshold amount of distance from the large audio object.
18. An apparatus including an audio rendering system, the apparatus, comprising: an input interface of the audio rendering system receiving audio data comprising audio objects, the audio objects comprising audio object signals and associated metadata, the associated metadata including at least audio object size data; a processing component determining, based on the audio object size data, a large audio object having an audio object size that is greater than a threshold size, wherein the large audio object is spatially diffuse and requires a plurality of speakers to reproduce the large audio object; and a decorrelator component coupled to the input interface, performing a decorrelation process on audio signals of the large audio object to produce decorrelated large audio object audio signals that are dependent on a defined location of the large audio object and other information, wherein the decorrelated large audio object audio signals are mutually independent of one another, and the decorrelation process comprises adjusting a level of each of the audio signals by adjusting a respective audio gain for each of the audio signals to generate the decorrelated large audio object audio signals corresponding to a speaker feed to each speaker of the plurality of speakers, and further wherein the plurality of speakers covers a large spatial area, wherein the decorrelated large audio object audio signals are mixed with at least one audio signal for other audio objects that are spatially separated by a second threshold amount of distance from the large audio object.
Claims 1-8 rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1, 3, 5, 18, 19 of U.S. Patent No. US 9,654,895 B2 for the similar reasons as discussed and in view of Kuech et al. ( US 20120076308 A1, hereinafter Kuech). The conflicting claims 1, 3, 5, 18, 19 of U.S. Patent No. 9,654,895 B2 do not explicitly teach the limitation “receiving, from an interface system, speaker zone constraint data” and “outputting … the speaker zone constraint data for rendering” as recited in claims 1, 7-8 of the instant application. However, Kuech teaches this limitation (side information 320 received by a parameter processor 480 and outputting corresponding matrix information representing the side information to the pre-mixing matrix M1 and post-mixing matrix M2 in fig. 4 and wherein speaker zone constraint data is included in the side information, para 76, or inputting playback configuration 570 and object positions 580 to scene rendering engine 540 and outputting rendering matrix representing the to the loudspeaker configuration 570 and object positions 580 to the scene rendering engine 540 in fig. 6B) for benefits of improving the audio playback performance (by assuring high audio quality with efficient representation of the multichannel audio signals, para 10 and efficient combination of sound effect and noise cancellation, para 41). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied receiving, from an interface system, speaker zone constraint data and outputting … the speaker zone constraint data for rendering, as taught by Kuech, to the audio processing method, as taught by conflicting claims 1, 3, 5, 18-19 of U.S. Patent No. 9,654,895 B2, for benefits discussed above. An comparison of claims 1-8 of the instant application with the conflicting claims 1, 3, 5, 18-19 of U.S. Patent No. US 9,654,895 B2, is listed below:
Claim(s) in the current application
Conflicting claim(s) of U.S. Patent No. US 9,654,895 B2
1. A method for audio processing, the method comprising: receiving audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; receiving, from an interface system, speaker zone constraint data; determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment; performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
2. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location is stationary.
3. The method of claim 1, wherein the at least one audio object is associated with at least one object location, wherein at least one of the at least one object location varies over time.
4. The method of claim 1, wherein performing decorrelation filtering includes at least one of a delay and a filter.
5. The method of claim 1, wherein performing decorrelation filtering includes at least one of an all-pass filter and a pseudo-random filter.
6. The method of claim 1, wherein performing decorrelation filtering includes a reverberation process.
7. A computer program product comprising a physical, non-transitory computer-readable medium storing instructions for performing the method of claim 1.
8. An apparatus for audio processing, the apparatus comprising: a receiver configured to receive audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object; and a decorrelator configured to perform, based on a determination that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment, decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals, wherein each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; a first output unit for outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering.
1. A method, comprising: receiving, in an input interface to an encoder component of an audio rendering system, audio data comprising audio objects, the audio objects comprising audio object signals and associated metadata, the associated metadata including at least audio object size data; determining, by a large object detection component based on the audio object size data, a large audio object having an audio object size that is greater than a threshold size, wherein the large audio object is spatially diffuse and requires a plurality of speakers to reproduce the large audio object; and performing, in a decorrelator component coupled to the input interface, a decorrelation process on audio signals of the large audio object to produce decorrelated large audio object audio signals that are dependent on a defined location of the large audio object and other information, wherein the decorrelated large audio object signals are mutually independent of one another, and the decorrelation process comprises adjusting a level of each of the audio signals by adjusting a respective audio gain for each of the audio signals to generate the decorrelated large audio object audio signals corresponding to a speaker feed to each speaker of the plurality of speakers, and further wherein the plurality of speakers covers a large spatial area.
3. The method of claim 1, wherein the large audio object has a plurality of object locations, wherein at least some of the plurality of object locations are one of: stationary locations or locations that vary over time.
3. The method of claim 1, wherein the large audio object has a plurality of object locations, wherein at least some of the plurality of object locations are one of: stationary locations or locations that vary over time.
5. The method of claim 1, wherein the decorrelation process comprises one of: a delay process, an all-pass filter process, a pseudo-random filter process, and a reverberation process.
5. The method of claim 1, wherein the decorrelation process comprises one of: a delay process, an all-pass filter process, a pseudo-random filter process, and a reverberation process.
5. The method of claim 1, wherein the decorrelation process comprises one of: a delay process, an all-pass filter process, a pseudo-random filter process, and a reverberation process.
19. A non-transitory medium having stored thereon programming instructions, which when executed by a processing component in an audio rendering system cause the audio rendering system to: receive, in an input interface to an encoder component of the audio rendering system, audio data comprising audio objects, the audio objects comprising audio object signals and associated metadata, the associated metadata including at least audio object size data; determine, by a large object detection component based on the audio object size data, a large audio object having an audio object size that is greater than a threshold size, wherein the large audio object is spatially diffuse and requires a plurality of speakers to reproduce the large audio object; and perform, in a decorrelator component coupled to the input interface, a decorrelation process on audio signals of the large audio object to produce decorrelated large audio object audio signals that are dependent on a defined location of the large audio object and other information, wherein the decorrelated large audio object signals are mutually independent of one another, and the decorrelation process comprises adjusting a level of each of the audio signals by adjusting a respective audio gain for each of the audio signals to generate the decorrelated large audio object audio signals corresponding to a speaker feed to each speaker of the plurality of speakers, and further wherein the plurality of speakers covers a large spatial area.
18. An apparatus including an audio rendering system, the apparatus comprising: an input interface of the audio rendering system receiving audio data comprising audio objects, the audio objects comprising audio object signals and associated metadata, the associated metadata including at least audio object size data; a processing component determining, based on the audio object size data, a large audio object having an audio object size that is greater than a threshold size, wherein the large audio object is spatially diffuse and requires a plurality of speakers to reproduce the large audio object; and a decorrelator component coupled to the input interface, performing a decorrelation process on audio signals of the large audio object to produce decorrelated large audio object audio signals that are dependent on a defined location of the large audio object and other information, wherein the decorrelated large audio object signals are mutually independent of one another, and the decorrelation process comprises adjusting a level of each of the audio signals by adjusting a respective audio gain for each of the audio signals to generate the decorrelated large audio object audio signals corresponding to a speaker feed to each speaker of the plurality of speakers, and further wherein the plurality of speakers covers a large spatial area.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(B) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claim 8 is rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention.
Claim 8 recites “a first output unit for outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering” and wherein (1) “the speaker zone constraint data” has an insufficient antecedent basis for the limitation and causes confusing because it is unclear what it is referred to and it is unclear what it is and (2) as discussed in title Claim Interpretation above, the application failed to disclose sufficient structure for claimed limitation “a first output unit” for the function of “outputting the decorrelated audio object audio signals and the speaker zone constraint data for rendering” and this caused indefiniteness because it is unclear what the “unit” as claimed is so that the function as claimed can be practiced and thus, claim 8 renders indefiniteness.
Reasons for Allowance
Prior art Xiang et al. (US 20140233917 A1) and Dressler et al. (US 20120232910 A1, hereinafter Dressler) are considered to be closed prior arts because
Xiang teaches a method for audio processing (title and abstract, ln 1-2, fig. 1), the method comprising:
receiving audio data comprising at least at least one audio object (receiving, by spatial audio rendering unit 60A, 60B, …, 60N, objects 34A’, 34B’, …, 34N’ in figs. 1-3), and audio object metadata (metadata 56A, 56B, …, 56N, respectively with the received objects 34A’, 34B’, …, 34N’ in fig. 3), wherein the audio object metadata is associated with the at least one audio object (association between 34A’, 34B’, …, 34N’ and 56A, 56B, …, 56N, respectively, and discussed above);
receiving, from an interface system, speaker zone constraint data (speaker configuration data for playback is assumed for rendering, i.e., received for rendering, para 45, 22);
determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size (object size and location are included in the metadata, para 112); performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals (by decorrelation, diffuse sound in the foreground and background, determined, para 63).
Dressler teaches a method for audio processing (title and abstract, ln 1-6, fig. 1-2), the method comprising:
receiving audio data comprising at least at least one audio object, and audio object metadata, wherein the audio object metadata is associated with the at least one audio object (audio objects received by extension selector 210 in fig. 2, and the audio objects received include attribute metadata and audio signal data, para 6);
receiving, from an interface system, speaker zone constraint data (the extension selector uses selection rules including surround configurations, as extension object, para 63 and thus, received for implementing the selection rules);
determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size (the mechanisms enabled to receive audio streams with variable number or size of objects in fig. 7, para 28);
performing decorrelation filtering on the at least one audio object to determine decorrelated audio object audio signals (the extension selector 210 determines diffuse by using decorrelation to determine how diffuse the objects are, para 70).
However, both Xiang and Dressler do not teach claimed limitations including “… determining that the metadata that indicates the at least one audio object is spatially diffuse having perceived size larger than a threshold in a playback environment…” and “… each of the decorrelated audio object audio signals corresponds to at least a reproduction loudspeaker of a plurality of reproduction loudspeakers; outputting the decorrelated audio object audio signals and the speaker zone constraint data …” as recited in claim 1. The Office has not found prior arts that teaches or suggests the modification of disclosures by Xiang and Dressler, and other references listed in attached PTO-892 in the fields to meet the claimed limitations above and combined other claimed features as a whole as recited in claim 1. Therefore, claim 1 is in condition for allowance.
For at least the similar reasons above, the other independent claim 7 and the dependent claims 2-6 are in condition for allowance. Applicants are advised to file Terminal Disclaimer(s) in order to overcome the non-statutory obvious-type double patenting rejection of claims 1-8 as set forth above.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LESHUI ZHANG whose telephone number is (571)270-5589. The examiner can normally be reached on Monday-Friday 6:30amp-4:00pm EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached on 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LESHUI ZHANG/
Primary Examiner, Art Unit 2695