Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 10/02/2025 and 03/11/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 26-31 and 36-45 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US Patent No. 10,362,385 (“Di Censo et al.”).
Regarding claim 26, Di Censo et al. discloses an apparatus, comprising:
at least one processor; and at least one memory storing instructions that, when executed by the at least one processor (fig. 1), cause the apparatus at least to:
output a first set of audio data via one or more loudspeakers of the apparatus (col. 5, For each of the speakers 120(i), the focus application 150 then generates the speaker signal 122(i) based on the corresponding ambient adjustment signal and requested playback signal (not shown in FIG. 1) representing audio content (e.g., music) targeted to the speaker 120(i));
capture, via one or more microphones of the apparatus, a real-world audio scene which is external to the apparatus to provide a second set of audio data for output via the one or more loudspeakers (col. 10, generate the awareness signals 252 based on the microphone signals 132 [from ambient signals]);
identify which of the first set of audio data and at least part of the real-world audio scene has the auditory attention of a user based on a measured neural activity of the user; control output of at least some of the first, some of the second or some of the first and second set of audio data via the one or more loudspeakers based on the identification (col. 8, the sensing engine 210 receives the biometric signals 142. At step 404, the sensing engine 210 determines the focus level 220 based on the biometric signals 142. At step 406, the tradeoff engine 232 computes the ambient awareness level 240 based on the focus level 220 and, optionally, any number of the configuration inputs 234); and
generate a noise cancelling signal based on the captured real-world audio scene (col, 10, if the ambience subsystem 290 generates the ambient adjustment signals 280, then at any given time, the ambient adjustment signals 280 comprise either awareness signals 252 generated by the acoustic transparency engine 250 or noise cancellation signals 262 generated by the noise cancellation engine 260),
wherein, in response to identifying that the first set of audio data has the auditory attention of the user, controlling output comprises enabling or increasing a gain associated with the noise cancelling signal for output to the one or more loudspeakers (col. 10, if the configuration input 234 indicates that the user is currently in a library, then the user is likely to be concentrating on an important task. Consequently, the tradeoff engine 230 could select the mapping 232 that specifies a threshold disable with step and set the threshold to a relative low value. By contrast, if the configuration input 234 indicates that the user is currently at the beach, then the user is likely to be enjoying the surroundings. Consequently, the tradeoff engine 230 could select the mapping 232 that specifies a threshold enable with a step and set the threshold to a relatively low value).
Regarding claim 39, Di Censo et al. discloses a method, comprising:
outputting a first set of audio data via one or more loudspeakers (col. 5, For each of the speakers 120(i), the focus application 150 then generates the speaker signal 122(i) based on the corresponding ambient adjustment signal and requested playback signal (not shown in FIG. 1) representing audio content (e.g., music) targeted to the speaker 120(i));
capturing, via one or more microphones, a real-world audio scene to provide a second set of audio data for output via the one or more loudspeakers (col. 10, generate the awareness signals 252 based on the microphone signals 132 [from ambient signals]);
identifying which of the first set of audio data and at least part of the real-world audio scene has the auditory attention of a user based on a measured neural activity of the user;controlling output of at least some of the first, some of the second or some of the first and second set of audio data via the one or more loudspeakers based on the identification (col. 8, the sensing engine 210 receives the biometric signals 142. At step 404, the sensing engine 210 determines the focus level 220 based on the biometric signals 142. At step 406, the tradeoff engine 232 computes the ambient awareness level 240 based on the focus level 220 and, optionally, any number of the configuration inputs 234); and
generating a noise cancelling signal based on the captured real-world audio scene (col, 10, if the ambience subsystem 290 generates the ambient adjustment signals 280, then at any given time, the ambient adjustment signals 280 comprise either awareness signals 252 generated by the acoustic transparency engine 250 or noise cancellation signals 262 generated by the noise cancellation engine 260),
wherein, in response to identifying that the first set of audio data has the auditory attention of the user, controlling comprises enabling or increasing a gain associated with the noise cancelling signal for output to the one or more loudspeakers (col. 10, if the configuration input 234 indicates that the user is currently in a library, then the user is likely to be concentrating on an important task. Consequently, the tradeoff engine 230 could select the mapping 232 that specifies a threshold disable with step and set the threshold to a relative low value. By contrast, if the configuration input 234 indicates that the user is currently at the beach, then the user is likely to be enjoying the surroundings. Consequently, the tradeoff engine 230 could select the mapping 232 that specifies a threshold enable with a step and set the threshold to a relatively low value).
Regarding claim 45, Di Censo et al. discloses a non-transitory computer readable medium (col. 2, a computer-readable medium configured to implement the method) comprising program instructions stored thereon for performing the method of claim 39. The claim is rejected in view of Di Cesnso as discussed with respect to claim 39.
Regarding claims 27 and 40, Di Censo et al. discloses the apparatus of claim 26, wherein the first set of audio data represents an audio track or communications session received from a user device associated with the apparatus (col. 6, the system 100 may comprise any type of audio system that enables any number of users to receive music and other requested sounds from any number and type of listening and communications systems while controlling the ambient sounds that the user perceives. Examples of listening and communication systems include, without limitation, MP3 players, CD players, streaming audio players, smartphones, etc.).
Regarding claims 28 and 41, Di Censo t al. discloses the apparatus of claim 26, wherein:
the first set of audio data represents a plurality of audio sources; identifying comprises identifying that a first audio source of the plurality of audio sources has the auditory attention of the user; and controlling output comprises amplifying the first audio source relative to at least one other audio source (the ambient audio representing the plurality of audio sources, col. 7, the ambient awareness level associated with the occupant and the microphone signals 132. Finally, for each occupant, the focus application 150 composites the requested playback signal representing requested audio content targeted to the occupant with the ambient awareness signals targeted to the occupant to generate the speaker signal 122 associated with the occupant).
Regarding claims 29 and 42, Di Censo discloses the apparatus of claim 26, wherein, in response to identifying that at least part of the real-world audio scene has the auditory attention of the user, controlling output comprises disabling or decreasing a gain associated with the first set of audio data (col, 10, if the ambience subsystem 290 generates the ambient adjustment signals 280, then at any given time, the ambient adjustment signals 280 comprise either awareness signals 252 generated by the acoustic transparency engine 250 or noise cancellation signals 262 generated by the noise cancellation engine 260).
Regarding claims 30 and 43, Di Censo et al. discloses the apparatus of claim 26, wherein, in response to identifying that at least part of the real-world scene has the auditory attention of the user, controlling output comprises enabling or increasing a gain associated with the at least some of the second set of audio data for output to the one or more loudspeakers (col. 10, if the configuration input 234 indicates that the user is currently at the beach, then the user is likely to be enjoying the surroundings. Consequently, the tradeoff engine 230 could select the mapping 232 that specifies a threshold enable with a step and set the threshold to a relatively low value).
Regarding claims 31 and 44, Di Censo et al. discloses the apparatus of claim 30, wherein controlling output comprises, in further response to identifying that at least part of the real-world audio scene has the auditory attention of the user, disabling or decreasing a gain associated with the noise cancelling signal for output to the one or more loudspeakers (col, 10, if the ambience subsystem 290 generates the ambient adjustment signals 280, then at any given time, the ambient adjustment signals 280 comprise either awareness signals 252 generated by the acoustic transparency engine 250 or noise cancellation signals 262 generated by the noise cancellation engine 260).
Regarding claim 36, Di Censo et al. discloses the apparatus of claim 26, wherein the apparatus is further caused to measure the neural activity of the user (col. 4, the biometric sensor 140 specifies neural activity associated with the user via a biometric signal 142).
Regarding claim 37, Di Censo et al. discloses the apparatus of claim 26, wherein the apparatus is comprised by an earphones device (fig. 1, headphone).
Regarding claim 38, Di Censo et al. discloses the earphones device of claim 37, wherein:
the apparatus comprises an active noise cancelling function operable in a transparency mode for outputting the at least some of the second set of audio data via the one or more loudspeakers; and controlling output of the at least some of the second set of audio data comprises at least enabling, or controlling a gain associated with, at least the transparency mode (col, 10, if the ambience subsystem 290 generates the ambient adjustment signals 280, then at any given time, the ambient adjustment signals 280 comprise either awareness signals 252 generated by the acoustic transparency engine 250 or noise cancellation signals 262 generated by the noise cancellation engine 260).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 32 is rejected under 35 U.S.C. 103 as being unpatentable over US Patent No. 10,362,385 (“Di Censo et al.”) in view of US Publication No. 2020/0201435 (“Ciccarelli et al.”).
Regarding claim 32, Di Censo et al. does not specify the apparatus of claim 30, wherein: the real-world audio scene comprises a plurality of real-world audio sources; identifying comprises identifying that a first real-world audio source of the plurality of real- world audio sources has the auditory attention of the user; and the apparatus is further caused to steer a sound capture beam of the one or more microphones towards a direction of the first real-world audio source such that audio signals of the first-world audio source are output to the one or more loudspeakers with a higher gain than those at least one of other real-world audio sources.
In the same field of endeavor, Ciccarelli et al. discloses receiving neural data responsive to a listener's auditory attention; receiving an acoustic signal responsive to a plurality of acoustic sources; for each of the plurality of acoustic sources: generating, from the received acoustic signal, audio data comprising one or more features of the acoustic source. companion device 120 can display or otherwise indicate to what acoustic source and how strongly the listener is attending, and may allow the listener to correct, teach, or override the algorithm or algorithm's decision for the attended source (e.g., the attended talker). As another example, companion device 120 may identify the attended source (e.g., the attended speaker) and/or determine the spatial location of the attended source ([0033]). Beamforming techniques capitalize on the difference in spatial location of sources and use multiple microphones. Candidate signal 206a, 206b are responsive to audio streams 202a, 202b, respectively ([0036]).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to incorporate spatial and spectral algorithms for noise reduction as disclosed by Ciccarelli et al. in Di Censo to improve sound capture of real world audio sources in environments with multiple acoustic sources.
Claims 33-35 are rejected under 35 U.S.C. 103 as being unpatentable over US Patent No. 10,362,385 (“Di Censo et al.”) in view of WO 2023/078809 (“Yasin et al.”).
Regarding claim 33, Di Censo et al. does not specify the apparatus of claim 30, wherein the apparatus is further caused to divide the second set of audio data into a plurality of frequency sub-bands, wherein: identifying comprises identifying that a first frequency sub-band of the plurality of frequency sub-bands has the auditory attention of the user; and controlling output comprises amplifying output of the first frequency sub-band with a higher gain than for the other frequency sub-bands.
In the same field of endeavor, Yasin et al. discloses an audio-signal processor (100) for filtering an audio signal-of-interest from an input audio signal (111) comprising a mixture of the signal-of-interest and background noise. A frontend unit 120 receives: an unfiltered input audio signal 111 from the receiving unit 110 and the human-derived NIFS from the HLAI processing module 140. The frontend unit 120 extracts sound level estimates from an output of the one or more bandpass filters of the filterbank 121 , using the sound level estimator 123. In this way, the sound level estimator 123 estimates a sound level output from the array of overlapping filters. At a later time, the frontend unit 120 will retrieve the modified gain value and the modified compression value from its memory 122 and apply them to the unfiltered input audio signal 111at the level of the filterbank 121, and determine a filtered output audio signal 112 (PAGE 21).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to apply the sub-band adjustment disclosed by Yasin et al. in the disclosure of Di Censo et al. in order to apply different gain value to each of the plurality of audio signals.
Regarding claim 34, Di Censo et al. discloses the apparatus of claim 33, wherein controlling output comprises, in further response to identifying that the real-world audio scene has the auditory attention of the user, disabling or decreasing the gain associated with the noise cancelling signal corresponding to the first frequency sub-band (col. 10, the ambience subsystem 290 receives the ambient awareness level 240 and generates ambient adjustment signals 280. The ambience awareness subsystem 290 includes, without limitation, an acoustic transparency engine 250 and a noise cancellation engine 260. At any given time, the ambiance subsystem 290 may or may not generate the ambient adjustment signals 280).
Regarding claim 35, Di Censo et al. discloses the apparatus of claim 34, wherein the first frequency sub-band corresponds to speech audio (col. 11, voices represented by the microphone signals 132).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIRAPON TULOP whose telephone number is (571)270-7491. The examiner can normally be reached Monday to Friday, 10:00AM-6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ahmad Matar can be reached at 571-272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JIRAPON TULOP/Examiner, Art Unit 2693