Prosecution Insights
Last updated: October 02, 2026
Application No. 17/814,665

AUDIO EVENT DATA PROCESSING

Non-Final OA §102§103
Filed
Jul 25, 2022
Priority
Jul 27, 2021 — provisional 63/203,562
Examiner
KRZYSTAN, ALEXANDER J
Art Unit
2694
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
4 (Non-Final)
81%
Grant Probability
Favorable
4-5
OA Rounds
0m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
925 granted / 1138 resolved
+19.3% vs TC avg
Moderate +7% lift
Without
With
+7.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
49 currently pending
Career history
1176
Total Applications
across all art units

Statute-Specific Performance

§101
2.8%
-37.2% vs TC avg
§103
41.9%
+1.9% vs TC avg
§102
20.2%
-19.8% vs TC avg
§112
18.1%
-21.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1138 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In view of the appeal brief filed on 12-6-2024, PROSECUTION IS HEREBY REOPENED. New grounds of rejection are set forth below. To avoid abandonment of the application, appellant must exercise one of the following two options: (1) file a reply under 37 CFR 1.111 (if this Office action is non-final) or a reply under 37 CFR 1.113 (if this Office action is final); or, (2) initiate a new appeal by filing a notice of appeal under 37 CFR 41.31 followed by an appeal brief under 37 CFR 41.37. The previously paid notice of appeal fee and appeal brief fee can be applied to the new appeal. If, however, the appeal fees set forth in 37 CFR 41.20 have been increased since they were previously paid, then appellant must pay the difference between the increased fees and the amount previously paid. A Supervisory Patent Examiner (SPE) has approved of reopening prosecution by signing at the end of this action. The examiner notes the interview summary where applicant agreed to a correction to the set of claim rejections under appeal, however the examiner was required to vacate final rejection and submit a new nonfinal rejection with a corrected set of claim rejections which now include new grounds of rejections to claim 20 for the same reasons as cited for previously rejected claim 6 and new grounds of rejection to claim 12. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-3,5-9,11,13-21,24-30 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Stamenovic et al ( US 20210289296 A1). As per claim 1, Stamenovic discloses a second device 302 (fig. 3) comprising: a memory configured to store instructions/data (required by processor 306 to implement the disclosed functions ; and one or more processors coupled to the memory (306,308) configured to: receive, from a first device 304, a first audio classification corresponding to an audio event (metadata 317 comprises metadata 317 can be used to communicate information about the user's voice from processor per para. 59; where information about a user’s voice is disclosed in the analogous system of fig. 2 per para. 47: time-series audio data to learn multiple aspects of the user's environment, while onboard ML model 232 may be configured to perform more simple supervised learning, e.g., using Naïve Bayes classification to identify a user's voice) (para. 20 the metadata identifies an environment within which the source audio signal is captured) (metadata 317 fig. 3 is an indication of an audio class when used as an input in the ML model comprising a classifier as per para. 43: remote ML model 112, 312, noting fig. 3 functions analagous to fig. 1 and fig. 2) can be trained to recognize and classify different types of acoustic signatures in sounds such as environmental sounds, a ; further noting the output metadata from one ML can be used as input to another learning model as per claim 13: wherein an output of the remote machine learning model is passed as input to the onboard machine learning model.(models 312 and 332 in fig. 3) ) (to summarize, any portion of the audio data being sent as metadata 317 is a classification of the event of detecting audio at device 304, further noting that said data or results from processing said data are used in two different locations as part of two different machine learning processes 300 and 330, and further noting that results/classifications from stage 300 are sent back to 330 which is then used to process said audio data which is then used on metadata 317 ). As per claim 16, a method comprising: receiving, at one or more processors of a second device, an indication of an audio class/first audio classification, the indication/first audio classification received from a first device and corresponding to an audio event (per claim 1 rejection, metadata 317); and processing, at the one or more processors of the second device, audio data to verify that a sound represented in the audio data corresponds to the audio event (para. 59: Metadata 317 could then be used by processor 306 to facilitate generation of metadata 316, noting that metadata can comprise a classification/verify a sound corresponding per para. 43: For example, remote ML model 112 can be trained to recognize and classify different types of acoustic signatures in sounds such as environmental sounds, a user's voice, other speakers, etc. Based on a recognized sound (e.g., acoustic signature),) (noting that 312 functions analagous to 112). As per claim 26, the system and method of the claim 1 and 16 rejections require a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a second device, cause the one or more processors to receive, from a first device, an indication of an audio class/first audio classification corresponding to an audio event (per the claim 1 and 16 rejections). As per claims 2,17,27, the second device of claim 1, wherein the one or more processors are further configured to: receive, from the first device, audio data representing a sound associated with the audio event (para. 59: metadata may be bi-directional to allow both devices 302, 304 to provide data to the other, where the metadata can include, per para. 59: metadata 317 can be used to communicate information about the user's voice/audio data from processor 320 to processor 306, ); and process the audio data at one or more classifiers to verify that the sound corresponds to the audio event/determine a classification/second audio classification (para. 6, the ML model in the accessory device comprises a classifier, where classifying a sound is identifying the sound, for example identifying an emergency sound verifies that the sound corresponds to an audio event/emergency, or identifying a user’s voice corresponds to the audio event of that user speaking). As per claims 3,18,28, the second device of claim 2, wherein the one or more processors are configured to provide the audio data and the indication of the audio class/first audio classification (metadata 317 per para. 59 and per the claim 2 rejection) as inputs/first and second inputs to the one or more classifiers to determine a second audio classification associated with the audio data (para. 59, Metadata 317 could then be used by processor 306 to facilitate generation of metadata 316; where the generation of metadata 316 comprises inputting the data into classifiers; also per para. 18, an output of the remote machine learning model is passed as input to the onboard machine learning model ) and compare the first audio classification to the second audio classification to verify the sound corresponds to the event (the meta data 317 is compared to metadata 316 per the denoising described in para. 57, because the metadata is used to provide feedback/input to the process that creates metadata 316 which is used to denoise the audio data which produces the classification/metadata 317) As per claims 5,19, the second device of claim 2, wherein the one or more processors are further configured to send a control signal to the first device based on an output of the one or more classifiers ( the classifiers produce the metadata per para. 6 where control signals in signal 316 are sent to the first device per fig. 3). As per claims 6,21,20 the control signal instructs the first device to perform an audio zoom operation/instruction (para. 58: Metadata 316 could also act on non-ML processing, such as the selection of a beamformer/audio zoom appropriate for a current environment,). As per claim 7, the second device of claim 6, wherein the control signal instructs the first device to perform spatial processing based on a direction of a source of the sound (para. 58: Metadata 316 could also act on non-ML processing, such as the selection of a beamformer/spatial direction based processing appropriate for a current environment,). As per claims 8,29, the second device of claim 2, wherein the one or more processors are further configured to: receive, from the first device, direction data corresponding to a source of the sound; and provide the audio data(para. 59: metadata may be bi-directional to allow both devices 302, 304 to provide data to the other, where the metadata can include, per para. 59: metadata 317 can be used to communicate information about the user's voice/audio data from processor 320 to processor 306), the direction data (the output of the beamforming used in one of the onboard and remote learning models per para. 18, where an output of the remote machine learning model is passed as input to the onboard machine learning model ), and the indication of the audio class/first audio classification (para. 20 the metadata identifies an environment within which the source audio signal is captured) (metadata 316 fig. 3) as inputs to the one or more classifiers to determine a second audio classification associated with the audio data (all of the metadata are inputs used by the classifiers to determine a classification). As per claim 9, the second device of claim 2, wherein the audio data includes one or more beamformed signal. (para. 19: In some implementations, at least one of the onboard machine learning model and remote machine learning model performs beamforming.) in view of para. 18 (where an output of the remote machine learning model is passed as input to the onboard machine learning model).: As per claim 11, the second device of claim 1, wherein the memory and the one or more processors are integrated into a mobile phone, and wherein the first device corresponds to a headset device (para. 35,38). As per claim 13, the second device of claim 1, further comprising a modem, wherein the indication of the audio class/first audio classification is received via the modem (the comm system in fig. 3 requires modems at each device in order to transmit the metadata). As per claim 14,24, the second device of claim 1, wherein the one or more processors are configured to selectively bypass direction-of-arrival processing on received audio data corresponding to the audio event based on whether direction-of-arrival information is received from the first device (the beamforming/direction of arrival processing used in the ML processing per the claim 8 rejection can only occur if or when it is available, as such is it selectively bypassed by the second device when it is not received from the first device). As per claims 15,25, The second device of claim 1, wherein the one or more processors are configured to selectively bypass a beamforming operation based on whether received audio data corresponds to multi-channel microphone signals from the first device or corresponds to beamformed signals from the first device (the beamforming used in the ML processing per the claim 8 rejection can only occur if or when it is available, as such it is selectively bypassed by the second device when it is not received from the first device and is also based on beamformed signals per the configuration described in the claim 8 rejection). As per claim 30, the system and method of the claim 1 and 16 rejections disclose an apparatus comprising: means for receiving an indication of an audio class/first audio classification, the indication/first audio classification received from a remote device and corresponding to an audio event; and means for processing audio data to verify that a sound represented in the audio data corresponds to the audio event. The claimed second device with the one or more processors, and apparatus ‘means for’ claim 30 are each drawn to the structure of a mobile phone and the first device is a headset in communication with the mobile phone. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 4,10,22,23, is/are rejected under 35 U.S.C. 103 as being unpatentable over Stamenovic et al ( US 20210289296 A1) as applied to claim 1,16, above, and further in view of Martin et al (US 20200059776 A1). As per claims 4,22, Stamenovic discloses the second devices of claim 2,16 wherein the audio class/first audio classification corresponds to a vehicle event (detecting an emergency sound per para. 40). However, Stamenovic does not disclose wherein the one or more processors are further configured to send a notification of the vehicle event to one or more third devices based on a location of the first device and locations of the one or more third devices. Martin teaches that electronic devices can be configured to send a broadcast/notification of a emergency to other electronic devices in order to determine/indicating that they are in the vicinity para.103. It would have been obvious to one of ordinary skill in the art at the time of filing, upon detecting an emergency vehicle sound, for the mobile phone of Stamenovic to send a notification to phones in the vicinity (which is based on the locations of the devices) as cited immediately above in Martin, for the purpose of determining if there are nearby devices during the emergency. As per claims 10,23, the second device of claims 1,16 wherein the one or more processors are further configured to: receive, from the first device, direction data corresponding to a source of a sound associated with the audio event (output of the beamforming in para. 19 Stam, in view of the function in para. 18 which requires the passing of metadata, audio data and directional/beamforming data between the first and second devices); However, Stamenovic fails to disclose update, based on the audio event, a map of directional sound sources in an audio scene to generate an updated map; and send data corresponding to the updated map to one or more third devices that are geographically remote from the first device. Martin teaches to: update, based on the audio event, a map of directional sound sources in an audio scene to generate an updated map (para. 10: es: a) displaying a virtual map comprising at least one of the second set of multimedia contents or the set of relevant sensors shown in relation to the location associated with the emergency at an electronic device associated with the ESP ); and send data corresponding to the updated map to one or more third devices that are geographically remote from the first device (the data sent cited in the claim 4 rejection comprises data corresponding to the updated map since they are both responsive to the same emergency event). It would have been obvious to one skilled in the art at the time of filing to implement the multimedia/audio map of Martin during emergency events of Stamenovic for the purpose of providing improved knowledge to all parties involved with the emergency. Claim(s) 12, is/are rejected under 35 U.S.C. 103 as being unpatentable over Stamenovic et al ( US 20210289296 A1) as applied to claim 1,16, above, and further in view of Farmer et al (US 20160150020 A1). As per claim 12, Stamenovic discloses the second device of claim 1. Stamenovic does not disclose wherein the memory and the one or more processors are integrated into a vehicle. Farmer teaches that phones can be integrated into vehicles for users to leverage the functionality of their mobile phone to promote an in-vehicle experience (abstract). It would have been obvious to one skilled in the art at the time of filing that the processors and memory of the phone along with the phone of Stamenovic could be placed in/integrated with vehicles per Farmer for the purpose of users to leverage the functionality of their mobile phone to promote an in-vehicle experience. Response to Arguments The submitted arguments related to claim 12 have been considered but are moot in view of the new grounds of rejection. The submitted arguments related to other pending claims have been considered with the following response: Appellant’s argument: Stamenovic does not describe metadata 317 as an audio classification (brief page 7). Response: Stamenovic discloses that metadata 317 (and also metadata 316) of Fig. 3 is an audio classification corresponding to an audio event per the following: -para. 20, the metadata identifies an environment within which the source audio signal is captured; where the identity of an environment is a classification of said environment, which corresponds to the audio event of the source audio signal being recorded in said environment. -para.22, the metadata includes an identity of a person speaking to a user of the wearable audio device, where the identity of a person is a classification of said person that corresponds to the audio event of said person speaking. -Para. 23, the metadata determines whether the person speaking is a user of the wearable audio device, where the determination is a classification as to the speaking state of said person that corresponds to the audio event of said person speaking being detected by the first device. -Para. 59, metadata 317 can be part of a bidirectional implementation of metadata 316, where 317 is used to communicate information about the user's voice, where said metadata 316 is disclosed as comprising any ML based process to classify a captured input signal into metadata 316 per para. 56, noting additional examples in para 56: example a characteristic of the environment (e.g., indoors versus outdoors, a social gathering, on a plane, in a quiet space, at an entertainment venue, in an auditorium, etc.), features of speech (e.g., the identity of the user, another person, a family member, etc., spectral energy, cepstral coefficients, etc.), features of noise (e.g., loud versus soft, high frequency versus low frequency, spectral energy, cepstral coefficients, etc.), which are each a classification corresponding to the audio event of the device receiving the detected audio signal, and further noting additional classification in the designation of the signal as speech or noise in the above classifications. -Claim 15 of Stamenovic, recites the machine learning model comprises a classifier configured to generate metadata associated with the input signal, where the output of a classifier is a classification, which corresponds to the audio event of one of the respective devices receiving audio. Appellant’s Argument: Stamenovic does not disclose what inputs are provided to the ML model 312 (brief, bottom of page 7) Response: Appellant’s argument is unclear as to which claim or particular claim language is being referred to as claim 1 does not recite an ML model. The examiner attempts to respond as best understood. Stamenovic, in para. 18 discloses: a second supervisory process that combines the remote machine learning model and the onboard machine learning model to form an ensemble model, where an output of the remote machine learning model is passed as input to the onboard machine learning model. The examiner notes para. 6 of Stamenovic discloses that the machine learning model comprises a classifier configured to generate metadata associated with the input signal. Hence the outputs of the models comprise metadata/classification which is passed from one ML model to another. Appellant’s argument: the metadata 317 could hypothetically include any other information that could be obtained from processing the user's voice received from the microphone 350 of the wearable audio device 304, such as a frequency envelope, volume, signal-to-noise ratio, directionality; and also: metadata generated by a de-noise system could instead include information "about the user's voice" related to de-noising, such as a noise level, signal-to-noise ratio, or the like (brief, middle of page 8, top of page 9). Response: None of those hypothetical items are relied upon in the rejection. Instead actual, real examples of classifications in the metadata are disclosed in the prior art and cited by the examiner in the original rejection and above. Appellant’s Argument: Stamenovic does not disclose metadata 316 as comprising control signals. Response: fig. 3 of Stamenovic, as originally cited by the examiner discloses a communications path between 302 and 304 to carry signal 316. Metadata 316 comprises control signals because metadata 316 is used to control functions in device 304, such as per para. 58 Stamenovic: metadata 316 could also act on non-ML processing, such as the selection of a beamformer/audio zoom appropriate for a current environment Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER KRZYSTAN whose telephone number is 571-272-7498, and whose email address is alexander.krzystan@uspto.gov The examiner can usually be reached on m-f 7:30-4:00 est. If attempts to reach the examiner by telephone or email are unsuccessful, the examiner’s supervisor, Fan Tsang can be reached on (571) 272-7547. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications. /ALEXANDER KRZYSTAN/Primary Examiner, Art Unit 2653 Examiner Alexander Krzystan /FAN S TSANG/Supervisory Patent Examiner, Art Unit 2694 January 28, 2025
Read full office action

Prosecution Timeline

Show 17 earlier events
May 08, 2025
Response after Non-Final Action
Jul 02, 2025
Response after Non-Final Action
Jul 03, 2025
Response after Non-Final Action
Jul 03, 2025
Response after Non-Final Action
Mar 13, 2026
Response after Non-Final Action
May 04, 2026
Request for Continued Examination
May 06, 2026
Response after Non-Final Action
Sep 28, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12739584
Method and Apparatus for Obtaining a Higher-Order Ambisonics (HOA) Coefficient
3y 0m to grant Granted Sep 15, 2026
Patent 12707221
METHODS, APPARATUS AND SYSTEMS FOR A PRE-RENDERED SIGNAL FOR AUDIO RENDERING
3y 7m to grant Granted Aug 11, 2026
Patent 12705016
SYSTEMS AND METHODS FOR MULTI-POINT CONTEXTUAL CONNECTIVITY FOR WIRELESS AUDIO DEVICES
2y 7m to grant Granted Aug 11, 2026
Patent 12707219
ELECTRONIC SYSTEM WITH EARPHONE/HEADPHONE RECOMMENDATION BASED ON EXPECTED CONTEXT OF DAILY USAGE OF WEARABLE AUDIO OUTPUT DEVICE(S)
2y 4m to grant Granted Aug 11, 2026
Patent 12689693
SPATIALLY BASED ADAPTIVE FILTER FOR A DOUBLE TALK RECOVERY METHOD
2y 8m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
81%
Grant Probability
88%
With Interview (+7.2%)
2y 12m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 1138 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month