Prosecution Insights
Last updated: August 17, 2026
Application No. 18/212,409

METHOD AND SYSTEM OF BINAURAL AUDIO EMULATION

Non-Final OA §102§103
Filed
Jun 21, 2023
Examiner
TRAN, CON P
Art Unit
Tech Center
Assignee
Intel Corporation
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
377 granted / 547 resolved
+8.9% vs TC avg
Strong +24% interview lift
Without
With
+23.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
23 currently pending
Career history
562
Total Applications
across all art units

Statute-Specific Performance

§101
5.8%
-34.2% vs TC avg
§103
55.4%
+15.4% vs TC avg
§102
13.2%
-26.8% vs TC avg
§112
18.7%
-21.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 547 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the response to this office action, the Examiner respectfully requests that support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line numbers in the specification and/or drawing figure(s). This will assist the Examiner in prosecuting this application. Claim Rejections - 35 USC § 102 2. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. 3. Claims 1, 6-9, and 14 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Delikaris Manias et al. U.S. Patent 11490218 (hereinafter, “Delikaris Manias”). Regarding claim 1, Delikaris Manias teaches a computer-implemented method of audio processing (the machine learning model 302 may have been trained to output multichannel audio output signals 303 for rendering in binaural sound (e.g., headphones), surround sound (e.g., in a multi-speaker configuration), or generally in any rendering configuration, Fig. 3B, col. 6, lines 50-55, see Delikaris Manias), comprising: receiving, by processor circuitry (host processors 202a, Fig. 2), multiple audio signals from multiple microphones ((microphone arrays 208a, Fig. 2), In one or more implementations, one or more components of the host processors 202a-b, the memories 204a-b, the RF circuitries 206a-b, the microphones of the microphone arrays 208a-b, and/or the specialized processor 210, Fig. 2, col. 5, lines 64-67, see Delikaris Manias) and overlapping in a same time and associated with a same at least one audio source (An electronic device may include multiple microphones. The microphones may produce audio signals captured from a sound scene. The audio signals may contain sounds from one or more sound sources of the sound scene, such as speech of one or more users, background noise (e.g., appliance noise, wind, traffic), and the like. The sound sources may have spatial and/or directional properties with respect to the electronic device. For example, a first user may be speaking from a position to the right of the electronic device while a second user may be speaking from a position to the left of the electronic device. When capturing the sound scene using multiple microphones the spatial properties of the sound sources can effectively be preserved for spatial reproduction in certain multichannel audio formats. For example, the spatial properties of the sound sources can be reproduced when playing back the audio signals via, e.g., binaural sound, surround sound, ambisonics and the like, col. 2, lines 8-24, see Delikaris Manias); and generating binaural audio signals (For example, the spatial properties of the sound sources can be reproduced when playing back the audio signals via, e.g., binaural sound, surround sound, ambisonics and the like (col. 2, lines 21-24, see Delikaris Manias)) comprising inputting at least one version of the multiple audio signals into a neural network (In one or more implementations, the machine learning model 302 may be, or may include, a deep neural network (DNN), and/or any other neural network, col. 9, lines 34-36, see Delikaris Manias) (The subject system provides for spatial audio reproduction of a sound scene captured using multiple microphones of an electronic device. The subject system utilizes a full waveform to waveform machine learning model (e.g., a deep neural network) that takes time domain input audio signals of a sound scene captured by the microphones of the electronic device and outputs multichannel (time domain) audio signals that spatially reproduce the sound scene in accordance with a particular target rendering configuration (e.g., binaural, surround sound, ambisonics, and the like), col. 2, lines 25-34, see Delikaris Manias)). Delikaris Manias thus teaches all the claimed limitations. Regarding claim 6, Delikaris Manias teaches the method of claim 1. Delikaris Manias further disclose wherein the multiple microphones are positioned in an array or pattern that corresponds in shape, number of microphones, or both to that of an array or pattern of a set of microphones used to train the neural network (The machine learning model 302 may be configured to output the spatial reproduction based on the expected positions of the microphones of microphone array 208a in the electronic device 102, and/or the temporal and spectral properties of the input audio signals 301. In this regard, the machine learning model 302 may be trained based on expected positions of the microphones and/or a geometry of the microphone array 208a of the electronic device 102 and further based on the target rendering configuration, Fig. 3B, col. 7, line 61-col. 8, line 3, see Delikaris Manias). Regarding claim 7, Delikaris Manias teaches the method of claim 1. Delikaris Manias further disclose wherein the multiple microphones comprise a number of microphones more than a number of microphones in a set of microphones used to train the neural network (For explanatory purposes, the beamforming module 501 is illustrated as providing all of the input signals (e.g., the beamformed signals 502) to the machine learning model 302. However, the beamforming module 501 may be configurable to selectively provide one or more of the beamformed signals 502 as input, Fig. 5, col. 14, lines 26-32, see Delikaris Manias), and wherein the inputting comprises inputting multiple audio signals from a target number of microphones of the multiple microphones that is the same as the number of microphones in the set of microphones (where the corresponding input audio signals 301 are provided as input for the beamforming module. The beamforming module 501 may be configured to form multiple beamformed signals depending on the target rendering configuration Fig. 5, col. 14, lines 32-36, see Delikaris Manias). Regarding claim 8, Delikaris Manias teaches the method of claim 1. Delikaris Manias further disclose wherein the inputting comprises inputting both time domain and frequency domain versions of the multiple audio signals into the neural network (The machine learning model 302 is configured to receive the input audio signals 301 in the time-domain without conversion or transformation into a different domain (e.g., the frequency domain) prior to being provided as an input to the machine learning model 302. For example, the input audio signals 301 may be raw audio data captured/recorded by the microphones of the electronic device 102, Fig. 3B, col. 7, lines 31-37, see Delikaris Manias). Regarding claim 9, Delikaris Manias teaches the method of claim 1. Delikaris Manias further disclose wherein the neural network is trained by using at least two overlapping audio sources (An electronic device may include multiple microphones. The microphones may produce audio signals captured from a sound scene. The audio signals may contain sounds from one or more sound sources of the sound scene, such as speech of one or more users, background noise (e.g., appliance noise, wind, traffic), and the like, col. 2, lines 8-13, see Delikaris Manias. In addition to spatial information, the machine learning model 302 may also be trained to use the spectral and/or temporal signatures of the acoustic sound field. For example, diffused noise and speech may have different spectral contents and also different spatial contents (e.g., diffused vs directional), col. 8, lines 56-61, see Delikaris Manias). Regarding claim 14, Delikaris Manias teaches a computer-implemented system, (the machine learning model 302 may have been trained to output multichannel audio output signals 303 for rendering in binaural sound (e.g., headphones), surround sound (e.g., in a multi-speaker configuration), or generally in any rendering configuration, Fig. 3B, col. 6, lines 50-55, see Delikaris Manias), comprising: memory to hold multiple audio signals from multiple microphones (In one or more implementations, one or more components of the machine learning models 302 may be implemented as software instructions, stored in a memory of the electronic device 102 (e.g., memory 204a), which when executed by the host processor 202a, cause the host processor 202a to perform particular function(s) (Fig. 2, col. 9, lines 40-45, see Delikaris Manias); In one or more implementations, the machine learning model 302, to output the multichannel audio output signals 303, may be configured to transform the raw microphone captured audio data of the input audio signals 301 from the time-domain into a different transform domain, (Fig. 3B, col. 8, lines 4-8, see Delikaris Manias)), wherein the multiple audio signals overlap in time and are associated with a same at least one audio source (An electronic device may include multiple microphones. The microphones may produce audio signals captured from a sound scene. The audio signals may contain sounds from one or more sound sources of the sound scene, such as speech of one or more users, background noise (e.g., appliance noise, wind, traffic), and the like. The sound sources may have spatial and/or directional properties with respect to the electronic device. For example, a first user may be speaking from a position to the right of the electronic device while a second user may be speaking from a position to the left of the electronic device. When capturing the sound scene using multiple microphones the spatial properties of the sound sources can effectively be preserved for spatial reproduction in certain multichannel audio formats. For example, the spatial properties of the sound sources can be reproduced when playing back the audio signals via, e.g., binaural sound, surround sound, ambisonics and the like, col. 2, lines 8-24, see Delikaris Manias); and processor circuitry communicatively connected to the memory (In one or more implementations, one or more components of the machine learning models 302 may be implemented as software instructions, stored in a memory of the electronic device 102 (e.g., memory 204a), which when executed by the host processor 202a, cause the host processor 202a to perform particular function(s) (Fig. 2, col. 9, lines 40-45, see Delikaris Manias); the processor circuitry being arranged to operate by (): generating binaural audio signals (For example, the spatial properties of the sound sources can be reproduced when playing back the audio signals via, e.g., binaural sound, surround sound, ambisonics and the like (col. 2, lines 21-24, see Delikaris Manias)) comprising inputting at least one version of the multiple audio signals into at least one neural network (In one or more implementations, the machine learning model 302 may be, or may include, a deep neural network (DNN), and/or any other neural network, col. 9, lines 34-36, see Delikaris Manias) (The subject system provides for spatial audio reproduction of a sound scene captured using multiple microphones of an electronic device. The subject system utilizes a full waveform to waveform machine learning model (e.g., a deep neural network) that takes time domain input audio signals of a sound scene captured by the microphones of the electronic device and outputs multichannel (time domain) audio signals that spatially reproduce the sound scene in accordance with a particular target rendering configuration (e.g., binaural, surround sound, ambisonics, and the like), col. 2, lines 25-34, see Delikaris Manias). Delikaris Manias thus teaches all the claimed limitations. 4. Claims 10 -12 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Nongpiur U.S. Patent Application Publication 20180124510. Regarding claim 10, Nongpiur teaches at least one non-transitory computer readable medium comprising a plurality of instructions that in response to being executed on a computing device, causes the computing device to operate (In a further example, provided is a first non-transitory computer-readable medium, comprising second processor-executable instructions stored thereon. The processor-executable instructions are configured to cause a processor to initiate executing one or more parts of the first method, par [0006], see Nongpiur) by: receiving multiple audio signals of a microphone array and of audio emitted from a same one or more sources at a same time (In block 402, first data is captured by receiving audio with the multi-microphone device. The multi-microphone device has at least two mechanical filters external to at least two respective microphones. While receiving the audio, the multi-microphone device can be located in an anechoic chamber. The at least two mechanical filters can include at least one of the microphone device 200, at least one of the microphone device 230, at least one of the microphone device 260, the like, or a combination thereof. The first data describes effects of one or more variations in notch frequency relative to a direction between a sound source and the multi-microphone device. The variations, and thus the effects, are due to one or more audio diffraction patterns due to (e.g., about) the at least two mechanical filters, Fig. 4, par [0068], see Nongpiur); receiving target binaural audio signals of the audio (In block 404, second data is captured by receiving the audio with the simulated human head microphone device. The second data describes effects of one or more audio diffraction patterns around the simulated human head microphone device, Fig. 4, par [0072], see Nongpiur); and training a neural network comprising inputting at least one version of the multiple audio signals into the neural network (The auralized multichannel recordings 502 can include the first data described in block 402. The one or more auralized multichannel recordings 502 are input to a neural network model 504, Fig. 5, par [0079], see Nongpiur), outputting output binaural audio signals (The neural network model 504 processes the auralized multichannel recordings 502 and creates a binaural output 506 by using neural network training data 508 to device-shape the received auralized multichannel recordings 502. The neural network can weigh and combine components of the auralized multichannel recordings 502 to create the binaural output 506, Fig. 5, par [0080], see Nongpiur), and comparing the output binaural audio signals to the target binaural audio signals (The binaural output 506 is compared with the one or more binaural recordings 510 to identify differences 512 between the binaural output 506 and the one or more binaural recordings 510. The differences 512 are input to a neural network training algorithm 514. The auralized multichannel recordings 502 can be input to the neural network training algorithm 514. The one or more binaural recordings 510 can be input to the neural network training algorithm 514. The neural network training algorithm 514 creates the neural network training data 508, which can include neural network coefficients used by the neural network model 504. The neural network training algorithm 514 can adjust the neural network coefficients in the neural network training data 508 to reduce the differences 512 to a substantially minimal amount, Fig. 5, par [0082], see Nongpiur). Nongpiur thus teaches al the claimed limitations. Regarding claim 11, Nongpiur teaches the medium of claim 10, wherein the training comprises using at least one audio source at a randomly selected (e.g., by moving, see par [0084]) angle and distance relative to a location of the microphone array (the capturing the first data can include, while audio is provided by a fixed-location loudspeaker 552, moving a multi-microphone device 554. The multi-microphone device 554 can be a multi-microphone device described herein, the first microphone device 200, the second microphone device 230, the third microphone device 260, the multi-microphone device 302, the like, or a combination thereof. The moving (e.g., randomly) the multi-microphone device 554 causes specific changes in direction between the fixed-location loudspeaker 552 and the multi-microphone device 554. The moving can include a horizontal-axis rotation 556, a vertical axis rotation 558, a distance variation 560, the like, or a combination thereof (moving, variation: correspond to a randomly) (see Fig. 5, par [0084], see Nongpiur). Regarding claim 12, Nongpiur teaches the medium of claim 11, wherein the training comprises simultaneously using multiple audio sources positioned at randomly selected angles and distances (the capturing the first data can include, while audio is provided by a fixed-location loudspeaker 552, moving a multi-microphone device 554. The multi-microphone device 554 can be a multi-microphone device described herein, the first microphone device 200, the second microphone device 230, the third microphone device 260, the multi-microphone device 302, the like, or a combination thereof. The moving (e.g., randomly) the multi-microphone device 554 causes specific changes in direction between the fixed-location loudspeaker 552 and the multi-microphone device 554. The moving can include a horizontal-axis rotation 556, a vertical axis rotation 558, a distance variation 560, the like, or a combination thereof (moving, variation: correspond to a randomly) (see Fig. 5, par [0084], see Nongpiur)) different from audio source to audio source to generate overlapping audio sources (In examples, the audio provided by the fixed-location loudspeaker 552 can include human speech, a sound typically occurring in a home, a sound not typically occurring in a home, a sound occurring in a home during an emergency situation, the like, or a combination thereof, par [0086], see Nongpiur). The second direction data can be substantially simultaneously recorded to correlate the direction with the direction-related variations in the notch frequency. The orientation can include direction, distance, the like, and combinations thereof. The correlation can include correlating changes in direction, changes in distance, the like, and combinations thereof with changes in the notch frequency (par [0085], see Nongpiur). Claim Rejections - 35 USC § 103 5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 6. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 7. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. 8. Claims 2-5 are rejected under 35 U.S.C. 103 as being unpatentable over Delikaris Manias et al. U.S. Patent 11490218 (hereinafter, “Delikaris Manias”) in view of Shumard et al. U.S. Patent Application Publication 20210058702 (hereinafter, “Shumard”). Regarding claim 2, Delikaris Manias teaches the method of claim 1. However, Delikaris Manias does not explicitly disclose wherein the multiple microphones provide sensitivity with a directional pattern. Shumard teaches one-dimensional array microphone with improved directivity (see Title) in which the polar patterns (i.e., sensitivity with a directional pattern) that can be formed by the array microphone 100 may depend on the placement of the microphones 102 within the array 100, as well as the type of beamformer(s) used to process the audio signals generated by the microphones 102 (Fig. 1, par [0037], see Shumard). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the sensitivity with a directional pattern taught by Shumard with the method of Delikaris Manias such that to obtain wherein the multiple microphones provide sensitivity with a directional pattern in order to improve the steered directionality of the array microphone without relying on computationally-heavy signal processing as suggested by Shumard in paragraph [0060]. Regarding claim 3, Delikaris Manias teaches the method of claim 1. However, Delikaris Manias does not explicitly disclose wherein the multiple microphones are in a linear array. Shumard teaches one-dimensional array microphone with improved directivity (see Title) in which wherein the multiple microphones are in a linear array (As shown in FIG. 1, the microphones 102 include a first plurality of microphones 104 linearly arranged along a length of the array microphone 100 and perpendicular to a preferred or expected direction of arrival for incoming sound waves. The first plurality of microphones 104 (also referred to herein as “first microphones”) are disposed along a common axis of the array microphone 100, such as first axis 105; Fig. 1, par [0040], see Shumard). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the sensitivity with a directional pattern taught by Shumard with the method of Delikaris Manias such that to obtain wherein the multiple microphones are in a linear array in order to improve the steered directionality of the array microphone without relying on computationally-heavy signal processing as suggested by Shumard in paragraph [0060]. Regarding claim 4, Delikaris Manias teaches the method of claim 1. However, Delikaris Manias does not explicitly disclose wherein the multiple microphones are in a circular array. Shumard teaches one-dimensional array microphone with improved directivity (see Title) in which disclose wherein the multiple microphones are in a circular array (Array microphones can provide several benefits over traditional microphones. Array microphones are comprised of multiple microphone elements aligned in a specific pattern or geometry (e.g., linear, circular, etc.) to operate as a single microphone device, par [0007], see Shumard). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the sensitivity with a directional pattern taught by Shumard with the method of Delikaris Manias such that to obtain wherein the multiple microphones are in a circular array in order to improve the steered directionality of the array microphone without relying on computationally-heavy signal processing as suggested by Shumard in paragraph [0060]. Regarding claim 5, Delikaris Manias teaches the method of claim 1. Delikaris Manias further teaches wherein the binaural audio signals are arranged to be used on headphones (For example, the machine learning model 302 may have been trained to output multichannel audio output signals 303 for rendering in binaural sound (e.g., headphones), surround sound (e.g., in a multi-speaker configuration), or generally in any rendering configuration, Fig. 3B, col. 6, lines 50-55, see Delikaris Manias). However, Delikaris Manias does not explicitly disclose wherein the multiple microphones comprises a linear array of at least four microphones. Shumard teaches one-dimensional array microphone with improved directivity (see Title) in which disclose wherein the multiple microphones comprises a linear array of at least four microphones (Moreover, in each pattern 300, 400, the third group of microphone sets 318, 418 includes only six microphone pairs, while the third group of microphone sets 118 in the pattern 200 includes seven microphone pairs, Figs. 3, 4. par [0061], see Shumard). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the sensitivity with a directional pattern taught by Shumard with the method of Delikaris Manias such that to obtain wherein the multiple microphones comprises a linear array of at least four microphones in order to improve the steered directionality of the array microphone without relying on computationally-heavy signal processing as suggested by Shumard in paragraph [0060]. 9. Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Nongpiur U.S. Patent Application Publication 20180124510 in view of Lluís et al., “Points2Sound: from mono to binaural audio using 3D point cloud scenes”, EURASIP Journal on Audio, Speech, and Music Processing, 2022, 15 pages, (hereinafter, “Lluís”). Regarding claim 13, Nongpiur teaches the medium of claim 10. Nongpiur further teaches a difference between output values of the output binaural audio signals and target values of the target binaural audio signals (The binaural output 506 is compared with the one or more binaural recordings 510 to identify differences 512 between the binaural output 506 and the one or more binaural recordings 510. The differences 512 are input to a neural network training algorithm 514. The auralized multichannel recordings 502 can be input to the neural network training algorithm 514, Fig. 5B, par [0082], see Nongpiur). However, Nongpiur does not explicitly disclose wherein the comparing comprises determining both a frequency domain loss and a time domain loss, and determining for both the frequency domain loss and the time domain loss: a difference between (1) differences between left and right values of the output binaural audio signals and (2) differences between left and right values of the target binaural audio signals. Lluis teaches Points2Sound: from mono to binaural audio using 3D point cloud scenes (see Title) in which as in previous work, we measure the quality of the predicted binaural audio assessing the short-time Fourier transform (STFT) Distance (i.e., time domain). Using the STFT Distance, we assess how similar the frequency components (i.e., frequency domain) of each predicted binaural channel are to the ground truth. STFT Distance ( dSTFT ) between a binaural signal sb and its estimate (time domain and frequency domain, page 7, right column, second paragraph, see Lluis). Note that in this case, Points2Sound predicts the full binaural signal, i.e., predicts both left and right binaural channels. Several methods in the literature propose to optimize the models by predicting the difference of the two binaural channels. To this end, we consider another loss function for Points2Sound, i.e., Ldiff , which optimizes the parameters to reduce the L1 loss between the estimated binaural difference channels sˆb diff and the ground truth binaural difference channels. Point2Sound is forced to learn the differences between the left and right binaural channels and predicts a one-channel signal sˆb diff . Then, considering the mono signal represented both predicted binaural channels are recovered (page 10, right column, last paragraph, to page 11, left column, first paragraph, see Lluis). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the taught by Lluis with the medium of Nongpiur such that to obtain wherein the comparing comprises determining both a frequency domain loss and a time domain loss, and determining for both the frequency domain loss and the time domain loss: a difference between (1) differences between left and right values of the output binaural audio signals and (2) differences between left and right values of the target binaural audio signals for purpose of operating in the waveform domain may allow for a more accurate auralization of a virtual audio scene, as suggested by Lluis in Abstract. 10. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Delikaris Manias et al. U.S. Patent 11490218 (hereinafter, “Delikaris Manias”) in view of Nongpiur U.S. Patent Application Publication 20180124510. Regarding claim 15, Delikaris Manias teaches the system of claim 14. Delikaris Manias further teaches the subject system provides for spatial audio reproduction of a sound scene captured using multiple microphones of an electronic device. The subject system utilizes a full waveform to waveform machine learning model (e.g., a deep neural network) that takes time domain input audio signals of a sound scene captured by the microphones of the electronic device and outputs multichannel (time domain) audio signals that spatially reproduce the sound scene in accordance with a particular target rendering configuration (e.g., binaural, surround sound, ambisonics, and the like, col. 2, lines 25-34, see Delikaris Manias). However, Delikaris Manias does not explicitly disclose wherein the binaural audio signals are associated with an interaural time difference and level difference set by the neural network. Nongpiur teaches directional microphone device and signal processing techniques (see Title) in which humans sense a single combined audio image originating from one or more specific directions. The single combined audio image is a combination of what human ears receive—two similar sounds having (1) a slight time delay (i.e., interaural time difference); and (2) direction-dependent variations in both spectral notches and frequency response of the sound (i.e., interaural level difference) (par [0025], see Nongpiur). The second method includes applying the neural network training data to a neural network, as well as creating a binaural output by device-shaping the received auralized multi-microphone input with the neural network (Fig. 5A, par [0080], see Nongpiur). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the directional microphone device and signal processing techniques taught by Nongpiur with the system of Delikaris Manias such that to obtain wherein the binaural audio signals are associated with an interaural time difference and level difference set by the neural network in order to improve source localization performance as suggested by Nongpiur in paragraph [0066]. 11. Claim 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Delikaris Manias et al. U.S. Patent 11490218 (hereinafter, “Delikaris Manias”) in view of Lester et al. U.S. Patent Application Publication 20220369031 (hereinafter, “Lester”). Regarding claim 16, Delikaris Manias teaches the system of claim 14. Delikaris Manias further teaches wherein the neural network comprises a time domain encoder (Returning to FIG. 4A, at the synthesis network 403, the machine learning model 302 may be trained to take the encoded and modulated signals and optimally transform them into multichannel waveforms in the time-domain, col. 12, lines 59-63, Delikaris Manias); a frequency domain encoder (transform the raw audio data of the input audio signals 301 (i.e., via encoder) into a different transform domain (i.e., frequency domain); in one or more implementations, the machine learning model 302, to output the multichannel audio output signals 303, may be configured to transform the raw microphone captured audio data of the input audio signals 301 from the time-domain into a different transform domain. The machine learning model 302 may be configured to transform the raw audio data of the input audio signals 301 into a different transform domain based on the raw audio data and/or on the application for which the multichannel audio output signals 303 may be utilized (Fig. 3B, col. 8, lines 4-13, see Delikaris Manias)). However, Delikaris Manias does not explicitly disclose a time domain decoder, wherein the frequency domain encoder feeds into a bottleneck between the time domain encoder and the time domain decoder. Lester teaches deep neural network denoiser mask generation system for audio processing (see Title) in which in certain embodiments, the time/frequency transform 204 can include a Fourier transform (e.g., a fast Fourier transform, a short-time Fourier transform, etc.) that transforms the audio signal sample 106 into the frequency domain audio signal sample 210. In certain embodiments, the time/frequency transform 204 can include a discrete cosine transform that transforms the audio signal sample 106 into the frequency domain audio signal sample 210. In certain embodiments, the time/frequency transform 204 can include a cochleargram transform that transforms the audio signal sample 106 into the frequency domain audio signal sample 210 (Fig. 2, par [0109], see Lester); in an aspect, the DNN model 306″ can include a set of convolutional layers configured in a U-Net architecture. In another aspect, the DNN model 306″ can include an encoder/decoder network structure with skip connections (Fig. 12, par [0162], see Lester); in one embodiment, an encoder branch of the DNN model 306″ is formed by the set of downsampling layers 1204a-n that downsample the input 1202 in a frequency axis (i.e., frequency domain encoder) by a factor of two while keeping a time axis (i.e., time domain encoder) at a same resolution to reduce latency during real-time implementation of the DNN model 306″. In another embodiment, a decoder branch of the DNN model 306″ is formed by the set of upsampling layers 1208a-n that upsample the input 1202 back an original size of the input 1202. Each gated linear unit can include a convolutional layer gated by another parallel convolutional layer with a sigmoid layer configured as an activation function. Additionally or alternatively, batch normalization and/or parametric rectified linear unit activation can be performed after the gating (Fig. 12,par [0163], see Lester); in certain embodiments, a bottleneck portion of the DNN model 306″ between the set of downsampling layers 1204a-n and the set of upsampling layers 1108a-n can include a set of convolutional layers 1206a-n. A first convolutional layer 1206a from the set of convolutional layers 1206a-n can include first downsampling, first batch normalization and/or first parametric rectified linear unit activation (Fig. 12, par [0165], see Lester). FIG. 13 illustrates a DNN model 306′″ according to one or more embodiments of the present disclosure. The DNN model 306′″ can illustrate an exemplary embodiment of the DNN model 306. Furthermore, the DNN model 306′″ can be a fully convolutional DNN. In an aspect, the DNN model 306′″ can include a set of convolutional layers configured with three levels in a U-Net architecture. In another aspect, the DNN model 306′″ can include an encoder/decoder network structure with skip connections. Input 1302 is provided to the DNN model 306′″. The input 1302 can correspond to the set of data samples 310 and/or the magnitude spectrogram 1102, for example (Fig. 13, par [0166], see Lester). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the deep neural network denoiser mask generation system for audio processing taught by Lester with the system of Delikaris Manias such that to obtain a time domain decoder, wherein the frequency domain encoder feeds into a bottleneck between the time domain encoder and the time domain decoder in order to provide improved denoising of an audio signal, as suggested by Lester in paragraph [0037]. Regarding claim 17, Delikaris Manias teaches the system of claim 14. Delikaris Manias further teaches wherein the neural network comprises a time domain encoder (Returning to FIG. 4A, at the synthesis network 403, the machine learning model 302 may be trained to take the encoded and modulated signals and optimally transform them into multichannel waveforms in the time-domain, col. 12, lines 59-63, Delikaris Manias) providing time domain output (The use of a time domain output allows the machine learning model to be directly trained using the objective function of interest. For example, the machine learning model may have been trained with information regarding the target output configuration (e.g., headphones vs. speakers, speaker placement, etc.), and to optimize a cost function to output the multichannel time domain audio signal (col. 2, lines 35-41, see Delikaris Manias)), and the processor circuitry (In one or more implementations, one or more components of the host processors 202a-b, the memories 204a-b, the RF circuitries 206a-b, the microphones of the microphone arrays 208a-b, and/or the specialized processor 210, Fig. 2, col. 5, lines 64-67, see Delikaris Manias). However, Delikaris Manias does not explicitly disclose a frequency domain encoder providing frequency domain output, and a decoder, and wherein the processor circuitry operates by combining the frequency domain output and the time domain output to generate input of the decoder. Lester teaches deep neural network denoiser mask generation system for audio processing (see Title) in which in certain embodiments, the time/frequency transform 204 can include a Fourier transform (e.g., a fast Fourier transform, a short-time Fourier transform, etc.) that transforms the audio signal sample 106 into the frequency domain audio signal sample 210. In certain embodiments, the time/frequency transform 204 can include a discrete cosine transform that transforms the audio signal sample 106 into the frequency domain audio signal sample 210. In certain embodiments, the time/frequency transform 204 can include a cochleargram transform that transforms the audio signal sample 106 into the frequency domain audio signal sample 210 (Fig. 2, par [0109], see Lester); in an aspect, the DNN model 306″ can include a set of convolutional layers configured in a U-Net architecture. In another aspect, the DNN model 306″ can include an encoder/decoder network structure with skip connections (Fig. 12, par [0162], see Lester); in one embodiment, an encoder branch of the DNN model 306″ is formed by the set of downsampling layers 1204a-n that downsample the input 1202 in a frequency axis (i.e., frequency domain encoder) by a factor of two while keeping a time axis (i.e., time domain encoder) at a same resolution to reduce latency during real-time implementation of the DNN model 306″. In another embodiment, a decoder branch of the DNN model 306″ is formed by the set of upsampling layers 1208a-n that upsample the input 1202 back an original size of the input 1202. Each gated linear unit can include a convolutional layer gated by another parallel convolutional layer with a sigmoid layer configured as an activation function. Additionally or alternatively, batch normalization and/or parametric rectified linear unit activation can be performed after the gating (Fig. 12,par [0163], see Lester); The DSP apparatus 1702 may include or otherwise be in communication with a processor 1704, a memory 1706, AI denoiser circuitry 1708, DSP circuitry 1710, input/output circuitry 1712, and/or communications circuitry 1714. In some embodiments, the processor 1704 (which may include multiple or co-processors or any other processing circuitry associated with the processor) may be in communication with the memory 1706 (Fig. 17, par [0171}, see Lester). he post-processing pipeline 902 can include one or more audio processing elements to facilitate post-processing of the denoised audio signal sample 110. In an embodiment, the post-processing pipeline 902 can include one or more parametric equalizers to modify and/or balance audio frequencies of the denoised audio signal sample 110, one or more compressors to compress a dynamic range of the denoised audio signal sample 110, one or more delays to add delay to the denoised audio signal sample 110, one or more compression codecs (e.g., one or more audio compression codecs and/or one or more video compression codecs), a dynamics processor, a matrix mixer, one or more communication codecs, and/or one or more other audio processing components to enhance the denoised audio signal sample 110 (Fig. 9, par [0153}, see Lester) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the deep neural network denoiser mask generation system for audio processing taught by Lester with the system of Delikaris Manias such that to obtain a frequency domain encoder providing frequency domain output, and a decoder, and wherein the processor circuitry operates by combining the frequency domain output and the time domain output to generate input of the decoder in order to provide improved denoising of an audio signal, as suggested by Lester in paragraph [0037]. 12. Claims 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Delikaris Manias et al. U.S. Patent 11490218 (hereinafter, “Delikaris Manias”) in view of Défossez, “Hybrid Spectrogram and Waveform Source Separation”, Proceedings of the MDX Workshop, 2021, 13 pages. Regarding claim 18, Delikaris Manias teaches the system of claim 14. Delikaris Manias further teaches a neural network (In one or more implementations, the machine learning model 302 may be, or may include, a deep neural network (DNN), and/or any other neural network. For example, the machine learning model 302 may be and/or may include, a cascade of a convolutional layer, one or more recurrent layers, one or more dense layers, or a combination thereof, 9, lines 34-39, see Delikaris Manias), a frequency domain encoder (transform the raw audio data of the input audio signals 301 (i.e., via encoder) into a different transform domain (i.e., frequency domain); in one or more implementations, the machine learning model 302, to output the multichannel audio output signals 303, may be configured to transform the raw microphone captured audio data of the input audio signals 301 from the time-domain into a different transform domain. The machine learning model 302 may be configured to transform the raw audio data of the input audio signals 301 into a different transform domain based on the raw audio data and/or on the application for which the multichannel audio output signals 303 may be utilized (Fig. 3B, col. 8, lines 4-13, see Delikaris Manias)), a time domain encoder (Returning to FIG. 4A, at the synthesis network 403, the machine learning model 302 may be trained to take the encoded and modulated signals and optimally transform them into multichannel waveforms in the time-domain, col. 12, lines 59-63, Delikaris Manias). However, Delikaris Manias does not explicitly disclose a frequency domain encoder and a time domain encoder both having a sequence of repeating encoder blocks, wherein the encoder blocks individually comprise, in order, a first convolutional layer, a rectified linear unit layer, a second convolutional layer, and a gated linear unit layer. Défossez teaches hybrid spectrogram and waveform source separation (see Title) in which The original Demucs architecture is a U-Net encoder/decoder structure (i.e., block). The original U-Net architecture is extended to provide two parallel branches: one in the time (temporal) and one in the frequency (spectral) domain (page 2, second paragraph, see Défossez); the encoder and decoder have a symmetric structure (page 3, lines 1-2 if first paragraph, see Défossez). In other words, Défossez teaches frequency domain encoder, time domain encoder, frequency domain decoder, time domain decoder. Hybrid Demucs has a dual U-Net structure, with the temporal and spectral branches (thus correspond to blocks) having their respective skip connections. Fig. 1 shows Hybrid Demucs architecture. The input waveform is processed both through a temporal encoder, and first through the STFT followed by a spectral encoder. See TEncoder repeating on lower right side of Fig. 1 (see Fig. 1 in page 4, see Défossez). Each encoder layer is composed of a convolution with a kernel size of 8, stride of 4 and doubling the number of channels (except for the first layer, which sets it to a fix value, typically 48 or 64). It is followed by a ReLU, and a so called 1x1 convolution with Gated Linear Unit activation (see page 3, first paragraph, lines 1-5). In other words, the order as: convolution layer, rectified linear unit (ReLU) layer, convolution layer, Gated Linear Unit layer. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the hybrid spectrogram and waveform source separation taught by Défossez with the system of Delikaris Manias such that to obtain a frequency domain encoder and a time domain encoder both having a sequence of repeating encoder blocks, wherein the encoder blocks individually comprise, in order, a first convolutional layer, a rectified linear unit layer, a second convolutional layer, and a gated linear unit layer in order to provide additional improvements, such as compressed residual branches, local attention or singular value regularization, as suggested by Défossez in Abstract. Regarding claim 19, Delikaris Manias teaches the system of claim 14. Delikaris Manias further teaches wherein the neural network comprises a time domain encoder (Returning to FIG. 4A, at the synthesis network 403, the machine learning model 302 may be trained to take the encoded and modulated signals and optimally transform them into multichannel waveforms in the time-domain, col. 12, lines 59-63, Delikaris Manias). However, Delikaris Manias does not explicitly disclose having a sequence of repeating decoder blocks, wherein the decoder blocks individually comprise, in order, a convolutional layer, a gated linear unit layer, a transpose convolutional layer, and a rectified linear unit layer. Défossez teaches hybrid spectrogram and waveform source separation (see Title) in which The original Demucs architecture is a U-Net encoder/decoder structure (i.e., block). The original U-Net architecture is extended to provide two parallel branches: one in the time (temporal) and one in the frequency (spectral) domain (page 2, second paragraph, see Défossez); the encoder and decoder have a symmetric structure (page 3, lines 1-2 if first paragraph, see Défossez). In other words, Défossez teaches frequency domain encoder, time domain encoder, frequency domain decoder, time domain decoder. Hybrid Demucs has a dual U-Net structure, with the temporal and spectral branches (thus correspond to blocks) having their respective skip connections. Fig. 1 shows Hybrid Demucs architecture. The input waveform is processed both through a temporal encoder, and first through the STFT followed by a spectral encoder. The two representations are summed when their dimensions align. The decoder is built symmetrically. The output spectrogram go through the ISTFT and is summed with the waveform outputs, giving the final model output. The Z prefix is used for spectral layers, and T prefix for the temporal ones. See TDecoder repeating on upper right side of Fig. 1 (see Fig. 1 in page 4, lines 8-14, see Défossez). Symetrically, a decoder layer sums the contribution from the U-Net skip connection and the previous layer, apply a 1x1 convolution with GLU, then a transposed convolution that halves the number of channels (except for the outermost layer), with a kernel size of 8 and stride of 4, and a ReLU (except for the outermost layer). There are 6 encoder layers, and 6 decoder layers, for processing 44.1 kHz audio. In order to limit the impact of aliasing from the outermost layers, the input audio is upsampled by a factor of 2 before entering the encoder and downsampled by a factor of 2 when leaving the decoder (see page 3, first paragraph). In other words, the order as: convolution layer, GLU layer (Gated Linear Unit layer), transposed convolution, and ReLU layer (rectified linear unit layer). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the hybrid spectrogram and waveform source separation taught by Défossez with the system of Delikaris Manias such that to obtain having a sequence of repeating decoder blocks, wherein the decoder blocks individually comprise, in order, a convolutional layer, a gated linear unit layer, a transpose convolutional layer, and a rectified linear unit layer in order to provide additional improvements, such as compressed residual branches, local attention or singular value regularization, as suggested by Défossez in Abstract. 13. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Delikaris Manias et al. U.S. Patent 11490218 (hereinafter, “Delikaris Manias”) in view of Messingher et al. U.S. Patent 11678111 (hereinafter, “Messingher”). Regarding claim 20, Delikaris Manias teaches the system of claim 14. However, Delikaris Manias does not explicitly disclose wherein the multiple microphones are in a same pattern shape as a pattern shape of an array of microphones used for training the neural network. Messingher teaches deep-learning based beam forming synthesis for spatial audio (see Title) in which FIG. 4 shows training of a machine learning model, according to some aspects. The training can be performed using a sufficiently large database of simulated recordings (e.g., greater than 500 recordings). The number of recordings can vary based on complexity (e.g., number of microphone signals, output channels, and spatial resolution (Fig. 4, col. 6, lines 15-20, see Messingher). A recording device 52 having a plurality of microphones arranged that match or resemble a geometrical arrangement of a particular recording device can generate the training set of recordings. For example, if the machine learning model is going to be used to map recordings captured by a smart phone model ABC, then the recording device 52 can either be a) the smart phone model ABC, or b) a set of microphones that resembles the make and geometrical arrangement of the microphones of smart phone model ABC (Fig. 4, col. 6, lines 21-29, see Messingher). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the deep-learning based beam forming synthesis for spatial audio taught by Messingher with the system of Delikaris Manias such that to obtain wherein the multiple microphones are in a same pattern shape as a pattern shape of an array of microphones used for training the neural network in order to provide immersive and improved results by utilizing non-linear techniques, as suggested by Messingher in column 1, lines 49-50. Conclusion 14. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Inventor Publication Number Disclosure Arik et al. US Patent Application Publication 20190122651 One-dimensional (1D) convolution with gated linear unit (par 0058]); Rectified linear unit (ReLU) nonlinearities to preprocess input mel-spectrograms (par [0064]). Any inquiry concerning this communication or earlier communications from the examiner should be directed to CON P TRAN whose telephone number is (571) 272-7532. The examiner can normally be reached M-F (08:30 AM- 05:00 PM) ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, VIVIAN C. CHIN can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.P.T/Examiner, Art Unit 2695 /VIVIAN C CHIN/Supervisory Patent Examiner, Art Unit 2695
Read full office action

Prosecution Timeline

Jun 21, 2023
Application Filed
Aug 10, 2023
Response after Non-Final Action
Jul 28, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707200
ELECTRONIC DEVICE HAVING MULTIPLE SPEAKERS CONTROLLED BY A SINGLE FUNCTIONAL CHIP
1y 11m to grant Granted Aug 11, 2026
Patent 12696026
SEMICONDUCTOR DEVICE PACKAGE AND ACOUSTIC DEVICE INCLUDING THE SAME
3y 1m to grant Granted Jul 28, 2026
Patent 12671956
AUDIO PROCESSING METHOD, WIRELESS EARPHONE, AND COMPUTER-READABLE MEDIUM
2y 8m to grant Granted Jun 30, 2026
Patent 12659657
HEARING PROTECTION APPARATUS
3y 0m to grant Granted Jun 16, 2026
Patent 12619306
WEARABLE CONTROL SYSTEM AND METHOD TO CONTROL AN EAR-WORN DEVICE
2y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
93%
With Interview (+23.7%)
3y 7m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 547 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month