DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/22/2026 has been entered.
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 1, 20, 21, 23, 27-29, 31 and 35-37 is/are rejected under 35 U.S.C. 103 as being unpatentable over Delikaris Manias et al. (US012363492 B1; hereafter Delikaris Manias) in view of Honma et al. (US 20210168548 A1; hereafter Honma).
Regarding claim 1, Delikaris Manias discloses a computer-implemented method (516 shown in Figs. 5 and 7, col. 13, lines 13-15, e.g.) comprising:
obtaining audio layers (108 and 124 represents different layers) that represent different types of audio assets (108 is one type of asset and 124 is another type of asset) that are to be output at the same time (by 116), the audio layers including a first audio layer representing one or more objects (e.g., 108) within a scene (as captured by microphones 106 in Fig. 1) and a second audio layer that represents background music, a soundtrack, dialogue, or sound effects (e.g., 124) that are associated with the scene;
selecting a respective Head-Related Transfer Function (HRTF) for each of the
first and second audio layers of the input audio signal (select HRTF from 112 for 108, select HRTF from 114 for 124; as implied by col. 1, lines 36-38; col. 7, lines 2-6 and 25-30), the selected HRTFs being different from each other (far-field vs near-field);
applying the respective HRTF to each of the first and second audio layers (“applying …” in col. 1, lines 36-38);
generating an output audio signal (“spatial audio” in col. 1, line 39 for 118) including the first and second audio layers to which the selected HRTFs have been applied, and which are to be output at the same time; and
providing the output audio signal for output (by 118).
Delikaris Manias fails to show an input audio signal comprises audio layers. Delikaris Manias teaches that the separated layers could be stored or transmitted to a separate device for further processing (col. 11, lines 13-17). Delikaris Manias teaches an embodiment in Fig. 2 wherein the sound scene is captured in one environment while the rendering is at a different environment. Delikaris Manias also teaches metadata defining the corresponding layer (col. 11, lines 19-20). In the same field of endeavor, Honma teaches a data stream as an input signal encoded with objects and the corresponding metadata (Fig. 4 or 10, [0100], [0137], [0228]); wherein the metadata provides information about the corresponding objects and the metadata would be utilized for rendering the objects, such as whether the object is within near-field. Thus, it would have been obvious to one of ordinary skill in the art to modify Delikaris Manias in view of Honma by encoding the multiple audio signals in separate layers and the corresponding metadata into a data stream in order to transmit the encoding audio data stream to a separate device for rendering object sing near-field HRTF and sound field using far-field HRTF, such as the embodiment as shown in Fig. 2 of Delikaris Manias.
Regarding claim 20, Delikaris Manias teaches VR application (col. 11, lines 65, e.g.).
Regarding claim 21, Delikaris Manias teaches selecting different HRTFs according to the determined type to apply (506, 508), but fails to show extracting and determining. As discussed above with respect to claim 1, the combination of Delikaris Manias and Honma teaches an encoded input signal with embedded objects in separate layers. Honma further teaches how to decode the encoded input signal (by 21 in Fig. 2, [0065], [0067]), how to determine, based on the extracted metadata, a type of HRTF to apply (S14 in Fig. 9). Such action would direct the decoded objects to different signal processing paths based on the corresponding metadata. Such teaching, one skilled in the art would have expected, would work reasonably well for decoding the encoded input signal and directing the separated layer to the corresponding path for near-field HRTF processing or far-field HRTF processing based on the corresponding metadata. Thus, it would have been obvious to one of ordinary skill in the art to modify Delikaris Manias in view of Honma by incorporating a decoding process for separating the encoded data stream into the layers for different processing paths in order to allow an user who receives an encoded input signal to render the embedded sound in different layers with the corresponding near-field HRTF or far-field HRTF.
Regarding claim 23, Delikaris Manias teaches convolution (col. 7, lines 33-37).
Most of limitations in claims 27-29 and 31 correspond to those in claims 1, 20, 21 and 23 discussed above. The VR application taught in Delikaris Manias (col. 4, line 63) is implemented by a system comprising one or more computer (col. 5, line 4) and one more storage devices storing instructions (col. 13, lines 54-58).
Most of limitations in claims 35-37 correspond to those in claims 1, 20 and 21 discussed above. Delikaris Manias teaches a non-transitory computer-readable medium (col. 13, lines 54-58; col. 15, line 23).
Claim(s) 22, 24-26, 30, 32-34 and 38 is/are rejected under 35 U.S.C. 103 as being unpatentable over Delikaris Manias and Honma as applied to claims 1, 21, 27, 29, 35 and 37 above, and further in view of Johnson et al. (US 20210084414 A1; hereafter Johnson).
Regarding claims 22, 30 and 38, the combination of Delikaris Manias and Honma teaches some claimed features. Delikaris Manias teaches extracting metadata associated with each audio layer of the two or more different audio layers comprises extracting data that represents (definition: to serve as a sign or symbol of; REPRESENT Definition & Meaning - Merriam-Webster) (ii) a type of the HRTF to apply (depending whether the layer is near-field or far-field), (iii) spatial priority information (near-field or far-field?), (iv) a virtual sound source (206 for user B), (v) a position of the virtual sound source (self-evidence as shown in Fig. 2), and (vi) an audio category of each audio layer (indicating near-field or far-field). Delikaris Manias fails to show (i) whether to apply an HRTF. Johnson teaches that the position of an object could be defined in more than 2 layers. For example, at location of 110, it is being treated as a non-spatialized audio layer that HRTF is not applicable ([0012]). Although Delikaris Manias does not show an object located in such region as shown in Johnson, Delikaris Manias does not place any limit on an object location. Similar to metadata defining layer for near-field object or far-field sound field, non-spatialized audio could be defined by corresponding metadata and being encoded in the audio stream to be processed/decoded by a separate device as discussed before without generating any unexpected result. Thus, it would have been obvious to one of ordinary skill in the art to modify the combination of Delikaris Manias and Honma in view of Johnson by defining a sound scene with more layers, such as having a layer defining non spatialized audio layer, in order to enhancing sound scene by simulating a sound scene with objects located at three different layers, including one object located inside user’s head.
Regarding claims 24 and 32, the combination of Delikaris Manias, Honma and Johnson, as discussed above, teaches identifying non-spatialized audio layers (by decoder decoding the corresponding metadata).
Regarding claims 25 and 33, Delikaris Manias teaches combining near-field layer and far-field layer (536 in Fig. 5), but fails to show combining non-spatialized audio layers. However, it is within the level of one skilled in the art to present multiple sound objects and/or sound field to the user at the same time by combining all signals to drive the necessary speakers. Examiner takes Official Notice that this feature is notoriously well known in the art. See Fig. 3 of Honma, wherein multiple lines are input to HRTF processing units and multiples lines are output to a mixing device. Thus, it would have been obvious to one of ordinary skill in the art to modify the combination of Delikaris Manias, Honma and Johnson by combining processed signal at each path, such as one path for near-field HRTF processing, one-path for far-field HRTF processing and another path for in head region processing, in order to be able to simulate a sound scene with an object in near-field, an object in far-field and another object located with the head, as object at different location requires separate processing for realistic sound effect.
Regarding claims 26 and 34, Delikaris Manias teaches 3-D audio signal (col. 9, lines 39-43).
Response to Arguments
Applicant’s arguments with respect to claim(s) 1, 27 and 35 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PING LEE whose telephone number is (571)272-7522. The examiner can normally be reached Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PING LEE/Primary Examiner, Art Unit 2695