Prosecution Insights
Last updated: October 02, 2026
Application No. 19/077,489

AUDIO SCENE MODIFICATION

Non-Final OA §102§103
Filed
Mar 12, 2025
Priority
Mar 27, 2024 — GB 2404370.5
Examiner
ESPINAS, KYLENINO TAGALOG
Art Unit
Tech Center
Assignee
Nokia Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
7 currently pending
Career history
6
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 26-35, 39-41, 43-45 is/are rejected under 35 U.S.C. 102(a)(1) and 35 U.S.C. 102(a)(2) as being anticipated by Ramsay (US20200329322A1). Regarding claim 26, Ramsay discloses: An apparatus, comprising: at least one processor (one or more computers (e.g., servers, network hosts, client computers, integrated circuits, microcontrollers, controllers, microprocessors, field-programmable-gate arrays, personal computers, digital computers, driver circuits, or analog computers) are programmed or specially adapted to perform one or more of the following tasks [0114]); and at least one memory storing (In some cases, the machine-accessible medium comprises (a) a memory unit or (b) an auxiliary memory storage device [0115]) instructions that, when executed by the at least one processor, cause the apparatus at least to (a machine-accessible medium has instructions encoded thereon that specify steps [0115]): measure neural activity of a user during at least one of output or capture of an audio scene (a system modifies sounds from external sources, presents the modified sound to a user, takes EEG measurements of the user's neural response to the modified sound, and determines—based on the EEG measurements—which of the sounds the user is paying attention to [0032]); identify, based on the measured neural activity, (a system modifies sounds from external sources, presents the modified sounds to a user, takes EEG measurements of the user's neural response to the modified sounds, and determines—based on the EEG measurements—which of the sounds the user is paying attention to [0048]) at least one target audio source of the audio scene which has the auditory attention of the user (In some implementations, an algorithm predicts which sound source a user is paying attention to, or engaged in, or focusing on, or is likely to remember [0064]); and cause modification of the at least one of output or capture of at least part of the audio scene based on the identification (after the system determines which external sound a user is paying attention to, the system alters external sounds that the user would otherwise hear, in order to make sound of interest easier to hear, or easier to pay attention to, or more memorable [0089]); Regarding claim 27, Ramsay discloses: The apparatus of claim 26, wherein the modifying comprises, during output of the audio scene, (For example, if two people are talking at the same time and at the same volume in a room, then: [0058]) emphasizing output of audio associated with the at least one target audio source relative to audio associated with one or more other audio sources of the audio scene ((a) the system may amplify the first voice, actively suppress the second voice and reduce background noise; and (b) the source separation algorithm may separate the external sounds into an amplified stream of the first voice, an attenuated stream of the second voice, and a background noise source. [0058]); Regarding claim 28, Ramsay discloses: The apparatus of claim 27, wherein the emphasizing comprises amplifying the audio associated with the at least one target audio source relative to the audio associated with the one or more other audio sources of the audio scene ((a) the system may amplify the first voice, actively suppress the second voice and reduce background noise [0058]); Regarding claim 29, Ramsay discloses: The apparatus of claim 27, wherein the emphasizing comprises attenuating the audio associated with the one or more other audio sources relative to the audio associated with the at least one target audio source of the audio scene ((b) the source separation algorithm may separate the external sounds into an amplified stream of the first voice, an attenuated stream of the second voice, and a background noise source [0058]); Regarding claim 30, Ramsay discloses: The apparatus of claim 26, wherein the audio scene comprises a plurality of audio sources (In some implementations, this invention is a system comprising: (a) a plurality of microphones; (b) one or more signal processors; (c) one or more speakers [0159]) and wherein audio associated with the plurality of audio sources is output such that the audio sources will be perceived at different respective positions with respect to the user (In FIGS. 2 and 9, the microphones may comprise directional microphones, positioned and oriented in such a way that the direction in which each microphone is most sensitive to sound is different [0037]); Regarding claim 31, Ramsay discloses: The apparatus of claim 30, wherein the modifying comprises, during output of the audio scene, attenuating audio associated with a background audio source ((b) the source separation algorithm may separate the external sounds into an amplified stream of the first voice, an attenuated stream of the second voice, and a background noise source [0058]) in a direction which corresponds to the position of the at least one target audio source (For instance, the sound separation algorithm may treat various angles and pickup patterns as virtual sources, selectively amplifying a particular pattern of directional sound capture over others [0053]); Regarding claim 32, Ramsay discloses: The apparatus of claim 31, wherein the audio associated with the background audio source is captured by the apparatus or a capture device (the method includes at least the following steps: One or more microphones record sound in a user's environment (Step 101). A computer separates the environmental sound into different audio channels (Step 102) [0034]) associated with the apparatus during output of the audio scene (In illustrative implementations of this invention, a system modifies sounds from external sources, presents the modified sounds to a user [0034]); Regarding claim 33, Ramsay discloses: The apparatus of claim 26, wherein the audio scene comprises a plurality of audio sources and is received as part of a communications session (In illustrative implementations of this invention, one or more devices (e.g., 208, 230, 231, 232, 930, 991, 992, 993) are configured for wireless or wired communication with other devices in a network. [0120]) in which the plurality of audio sources represents respective participants of the communications session (For instance, this invention may identify which human speaker a user is attending to in a multi-speaker environment [0101]); Regarding claim 34, Ramsay discloses: The apparatus of claim 26, wherein the audio scene is a real-world audio scene in which audio associated with one or more audio sources of the audio scene (modified sounds are amplified and mixed-in over external, real-world sounds with low latency (e.g., latency of less than 5 milliseconds, or less than 20 milliseconds) [0056]) is captured by the apparatus or a capture device associated with the apparatus (may take as inputs different audio channels from different microphones or from beamforming [0047]); Regarding claim 35, Ramsay discloses: The apparatus of claim 34, wherein the apparatus is further caused to: determine, during capture, a direction of the at least one target audio source with respect to the user (For instance, the sound separation algorithm may treat various angles and pickup patterns as virtual sources, selectively amplifying a particular pattern of directional sound capture over others [0053]); wherein the modifying comprises providing or steering a sound capture beam towards the direction of the at least one target audio source (performs acoustic beamforming and real-time compression/equalization of incoming sounds [0112]) such as to amplify audio coming from the direction of the at least one target audio source relative to audio coming from the direction of one or more other audio sources of the real-world audio scene (The system hardware may modify sound sources by direction of arrival to identify the beam of interest, and selectively amplify it [0112]); Regarding claim 39, Ramsay discloses: The apparatus of claim 26, wherein the apparatus is further caused to: determine respective confidence values associated with a plurality of audio sources of the audio scene, (In some implementations, an algorithm predicts which sound source a user is paying attention to, or engaged in, or focusing on, or is likely to remember [0064] For instance, the attention predication algorithm may calculate the average P300 amplitude elicited by the last 10 modifications as a prediction of attention paid, which effectively averages the attention prediction across the time taken to make 10 modifications [0067]) wherein the confidence value associated with a particular audio source indicates a likelihood that the particular audio source has the auditory attention of the user, wherein the at least one target audio source is identified, based at least in part, on the respective confidence values (and (b) may output a prediction of attention or memorability for one or more of the separated sound signals [0066] These instantaneous predictions may be fed into a state-space model, for example an HMM (hidden Markov model), which may predict the user's likely attentional state given noisy observations over time [0070]); Regarding claim 40, Ramsay discloses: The apparatus of claim 39, wherein the apparatus is further caused to: identify an ambiguity between two or more of the audio sources (the sound modifications may be performed in bursts to bolster prediction only when attentional state is uncertain (e.g., upon entering a new environment), when a transition of attention is expected, or when the environment is especially challenging [0074]) having the highest respective confidence values based on said respective confidence values being within a predetermined range of one another (in some implementations, the attention prediction algorithm predicts one signal (EEG or audio) from the other, compares them, and uses this to update an estimate of attention predictions over time [0070]); and resolve the ambiguity based on further measured neural activity to identify which of the two or more identified audio sources is the target audio source (how the system—even after it has been operating for a while and has already made predictions of attention—may continue to monitor the user's EEG readings and continue to make new predictions regarding which sound the user is paying attention to [0076]); Regarding claim 41, Ramsay discloses: The apparatus of claim 40, wherein the apparatus is further caused to: responsive to identifying the ambiguity, (how the system—even after it has been operating for a while and has already made predictions of attention—may continue to monitor the user's EEG readings and continue to make new predictions regarding which sound the user is paying attention to [0076]) output a reference sound in the direction of at least one of the two or more identified audio sources (The system may iterate through different beams of incoming sound and elicit changes in one sound at a time correlating the changes with measured ERPs to find the attended sound [0105] The system hardware may modify sound sources by direction of arrival to identify the beam of interest, and selectively amplify it [0112] The examiner interprets “reference sound” broadly as an additional sound (not caused by the audio source itself) presented to the user deliberately modified to be associated with an audio source specifically including its direction for purposes of eliciting a neural response as an identification method. Thus, Ramsay’s modification of a sound corresponding to a particular direction/source reads on the claimed “reference sound” as the modification has the same purpose to resolve ambiguity while also not being a sound created by the audio source it is attempting to identify); wherein resolving the ambiguity comprises identifying, based on measured neural activity when the reference sound is played to the user, (incoming sounds are modified in slight, unexpected ways, to elicit an event-related potential (ERP). This potential may be measured with an EEG and may occur in response to a stimulus (e.g., sound), which changes when the user is paying attention to it. This ERP may be a P300. [0102]) which of the two or more identified audio sources is the target audio source (This invention overlays natural sounds with slight, unexpected modifications to evoke the ERP P300, and then, based on the ERP P300, identifies the sound source that a user is attending to [0103] system may iterate through different beams of incoming sound and elicit changes in one sound at a time [0105]); Regarding claim 43, Ramsay discloses: The apparatus of claim 26, wherein the apparatus is comprised by an earphones device comprising one or more sensors for sensing biosignals of the user for measuring the user's neural activity, or the apparatus is comprised by a user device in communication with an earphones device (the EEG electrodes may be housed in or attached to any one or more of the following: (a) eyeglasses frame; (b) headband; (c) hat or other object worn on the head; (d) necklace; (e) headphones (e.g., circum-aural or supra-aural headphones); (f) earphones that face but are not inserted into the ear canal; (g) in-ear headphones (or in-ear monitors or canalphones); or (h) earbuds or earpieces that are inserted into the ear canal [0042]); Claim 44 contains similar limitations to claim 26 and therefore is rejected for the same reasons. Claim 45 contains similar limitations to claim 26 and therefore is rejected for the same reasons. Additionally, Ramsay discloses: A non-transitory computer readable medium comprising program instructions stored thereon (In illustrative implementations, one or more computers execute programs according to instructions encoded in one or more tangible, non-transitory computer-readable media [0116]); Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 36-38, 42 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ramsay (US20200329322A1), in view of Rosenwein et al. (US20220312128A1). Regarding claim 36, Ramsay discloses: The apparatus of claim 35, (a system modifies sounds from external sources, presents the modified sound to a user, takes EEG measurements of the user's neural response to the modified sound, and determines—based on the EEG measurements—which of the sounds the user is paying attention to [0032]) wherein the apparatus is further caused to: Ramsay fails to disclose a camera of the apparatus. While Ramsay discloses an apparatus that is an EEG based auditory attention system that identifies which external sound/source the user is attending to and can modify the human computer interaction based on that sound, they fail to disclose a camera of the apparatus. However, Rosenwein teaches a camera of the apparatus (As discussed above, apparatus 110 may include an image sensor 220 for capturing image data [0054 of Rosenwein]). Ramsay and Rosenwein both considered to be analogous to the claimed invention because they are in the same field of hearing aid systems. Therefore, it would have been obvious to one of ordinary skill in the art to modify the apparatus of Ramsay to include the camera and image processing functionality taught by Rosenwein to provide visual information corresponding to the user’s environment and thereby improve identification, localization, and interaction with a source of interest. Rosenwein teaches that such image processing would improve the apparatus (Therefore, there is a need for improved hearing aid systems that more accurately condition audio signals associated with different sources in the environment of a user. Solutions may include the use of an image capture device for automatically capturing and processing images from the environment of the user [0004 of Rosenwein]). It would have predictably improved the environmental knowledge in which the apparatus had access to, thereby improving the apparatus’ ability to distinguish which sound the user is paying attention to. The combination further discloses: capture, via a camera of the apparatus, an image of the real-world audio scene (apparatus 110 may include an image sensor system 220 for capturing real-time image data of the field-of-view of user 100 [0053 of Rosenwein]); display the captured image (For example, user 100 may view on display 260 data (e.g., images, video clips, extracted information, feedback information, etc.) that originate from or are triggered by apparatus 110 [0058 of Rosenwein]); determine, based on the direction of the at least one target audio source with respect to the user, (For instance, the sound separation algorithm may treat various angles and pickup patterns as virtual sources, selectively amplifying a particular pattern of directional sound capture over others [0053 of Ramsay] For example, in one embodiment, processor 210 may detect an interaction with another individual and sense that the individual is not fully in view, because image sensor 220 is tilted down. Responsive thereto, processor 210 may adjust the aiming direction of image sensor 220 to capture image data of the individual [0078 of Rosenwein]) a sub-portion of the captured image corresponding to the at least one target audio source (For example, apparatus 110 may store a portion of an image that includes a face of a person who appears in the image [0094 of Rosenwein]); and modify the determined sub-portion or cause the camera to focus on the determined sub-portion (include a processing unit 210 for controlling and performing the disclosed functionality of apparatus 110, such as to control the capture of image data, analyze the image data, and perform an action and/or output a feedback [0053 of Rosenwein]); Regarding claim 37, Ramsay discloses: The apparatus of claim 35, wherein the apparatus is further caused to: While Ramsay discloses an apparatus that is an EEG based auditory attention system that identifies which external sound/source the user is attending to and can modify the human computer interaction based on that sound, they fail to disclose a camera and the claimed image capture display. However, Rosenwein teaches a camera of the apparatus (As discussed above, apparatus 110 may include an image sensor 220 for capturing image data [0054 of Rosenwein]) as well as a display for the captured image data (For example, user 100 may view on display 260 data (e.g., images, video clips, extracted information, feedback information, etc.) that originate from or are triggered by apparatus 110 [0058 of Rosenwein]). Ramsay and Rosenwein both considered to be analogous to the claimed invention because they are in the same field of hearing aid systems. Therefore, it would have been obvious to one of ordinary skill in the art to modify the apparatus of Ramsay to include the camera and image processing functionality taught by Rosenwein to provide visual information corresponding to the user’s environment and thereby improve identification, localization, and interaction with a source of interest. Rosenwein teaches that such image processing would improve the apparatus (Therefore, there is a need for improved hearing aid systems that more accurately condition audio signals associated with different sources in the environment of a user. Solutions may include the use of an image capture device for automatically capturing and processing images from the environment of the user [0004 of Rosenwein]). Rosenwein also teaches that a visual display would improve the users interaction with their environment (In some embodiments, optional computing device 120 and/or server 250 may provide additional functionality to enhance interactions of user 100 with his or her environment [0052 of Rosenwein] Information gathered from the image capture device may be leveraged to improve the function a hearing aid device [0004 of Rosenwein]) It would have predictably improved the environmental knowledge in which the apparatus had access to, thereby improving the apparatus’ ability to distinguish which sound the user is paying attention to by allowing the user further interaction with a visual display. The combination further discloses: capture, via a camera of the apparatus, an image of the real-world audio scene (apparatus 110 may include an image sensor system 220 for capturing real-time image data of the field-of-view of user 100 [0053 of Rosenwein]); display the captured image (For example, user 100 may view on display 260 data (e.g., images, video clips, extracted information, feedback information, etc.) that originate from or are triggered by apparatus 110 [0058 of Rosenwein]); determine, based on the direction of the at least one target audio source with respect to the user, that the captured image does not include the at least one target audio source (For example, apparatus 110 may store a portion of an image that includes a face of a person who appears in the image [0094 of Rosenwein] For example, in one embodiment, processor 210 may detect an interaction with another individual and sense that the individual is not fully in view, because image sensor 220 is tilted down. Responsive thereto, processor 210 may adjust the aiming direction of image sensor 220 to capture image data of the individual [0078 of Rosenwein])); and change a lens of the camera such that the captured image will include the at least one target audio source (Some embodiments may permit alteration of the orientation of an image sensor of the capture unit, for example to better capture images of interest [0096 of Rosenwein]); Regarding claim 38, the combination of Ramsay in view of Rosenwein discloses: The apparatus of claim 26, wherein the apparatus is further caused to: While Ramsay discloses an apparatus that is an EEG based auditory attention system that identifies which external sound/source the user is attending to and can modify the human computer interaction based on that sound, they fail to disclose a gesture trigger that the user can use to indicate the direction of the audio source. However, Rosenwein teaches a gesture (a hand-related trigger may include a gesture performed by user 100 involving a portion of a hand of user 100 [0053 of Rosenwein]) and it would provide an additional user input mechanism for selectively initiating or controlling an action (and perform an action and/or output a feedback based on a hand-related trigger identified in the image data [0053 of Rosenwein]). Ramsay and Rosenwein both considered to be analogous to the claimed invention because they are in the same field of hearing aid systems. Therefore, it would have been obvious to one of ordinary skill in the art to modify the apparatus of Ramsay to include a gesture that could be used by the camera and image processing functionality taught by Rosenwein. It would have predictably improved the environmental knowledge in which the apparatus had access to by adding trigger-gesture functionality for the user to give additional information when performing a gesture (and perform an action and/or output a feedback based on a hand-related trigger identified in the image data [0053 of Rosenwein] actions may be taken based on the identified objects, gestures, or other information [0095 of Rosenwein] Information gathered from the image capture device may be leveraged to improve the function a hearing aid device [0004 of Rosenwein]). The combination further discloses: identify a predetermined trigger gesture of the user, (a hand-related trigger may include a gesture performed by user 100 involving a portion of a hand of user 100 [0053 of Rosenwein]) wherein at least the modifying is performed responsive to identifying the predetermined trigger gesture (and perform an action and/or output a feedback based on a hand-related trigger identified in the image data [0053 of Rosenwein]); and wherein the predetermined trigger gesture is identified based at least in part on the measured neural activity of the user (takes EEG measurements of the user's neural response to the modified sounds [0048 of Ramsay]) when said predetermined trigger gesture is performed (actions may be taken based on the identified objects, gestures, or other information [0095 of Rosenwein]); Regarding claim 42, Ramsay discloses: The apparatus of claim 40, wherein the apparatus is further caused to: responsive to identifying the ambiguity, sources (the sound modifications may be performed in bursts to bolster prediction only when attentional state is uncertain (e.g., upon entering a new environment), when a transition of attention is expected, or when the environment is especially challenging [0074 of Ramsay]) While Ramsay teaches methods of identifying a direction of a target audio and identifying ambiguity during said identification, it does not disclose requesting a directional gesture towards the position of the target audio source. However, Rosenwein teaches that adding a gesture would give the apparatus a cue to help decide whether or not to perform an action, which would provide additional user input mechanism for the processing apparatus (determine whether the trigger is associated with a person other than the user of the wearable apparatus to selectively determine whether to perform an action associated with the trigger [0089 of Rosenwein]). Ramsay and Rosenwein both considered to be analogous to the claimed invention because they are in the same field of hearing aid systems. Therefore, it would have been obvious to one of ordinary skill in the art to modify the apparatus of Ramsay to include the image processing and gesture identification. It would have predictably improved the environmental knowledge in which the apparatus had access to by adding trigger-gesture functionality thereby improving the apparatus’ ability to distinguish which sound the user is paying attention to by allowing the user further interaction with the apparatus (actions may be taken based on the identified objects, gestures, or other information [0095 of Rosenwein] Information gathered from the image capture device may be leveraged to improve the function a hearing aid device [0004 of Rosenwein]). The combination further discloses: request a directional gesture towards the position of the target audio source (indicating that user 100 is looking in the direction of the wrist strap 160. Wrist strap 160 may also include an accelerometer, a gyroscope, or other sensor for determining movement or orientation of a user's 100 hand for identifying a hand-related trigger [0051 of Rosenwein]); wherein the directional gesture is determined based on the measured neural activity of the user when (how the system—even after it has been operating for a while and has already made predictions of attention—may continue to monitor the user's EEG readings and continue to make new predictions regarding which sound the user is paying attention to [0076 of Ramsay]) said directional gesture is performed (actions may be taken based on the identified objects, gestures, or other information [0095 of Rosenwein] and perform an action and/or output a feedback based on a hand-related trigger identified in the image data [0053] of Rosenwein]); Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Kyle Espinas whose telephone number is (571)270-0596. The examiner can normally be reached Monday Friday, 8 a.m. 5 p.m. ET.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571) 272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Kylenino Espinas/ Patent Examiner Art Unit 2655 9/14/2026 /ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Mar 12, 2025
Application Filed
Sep 24, 2026
Non-Final Rejection mailed — §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month