Prosecution Insights
Last updated: July 29, 2026
Application No. 18/595,221

SOUND MODIFICATION USING MACHINE LEARING CLASSIFICATION

Non-Final OA §103§112
Filed
Mar 04, 2024
Examiner
MANOHARAN, SHASHIDHAR SHANKAR
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Sony Group Corporation
OA Round
2 (Non-Final)
100%
Grant Probability
Favorable
2-3
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
3 granted / 3 resolved
+38.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
23 currently pending
Career history
27
Total Applications
across all art units

Statute-Specific Performance

§103
98.3%
+58.3% vs TC avg
§102
1.8%
-38.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 3 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendments filed 02/17/2026 have been accepted and considered in this office action. Claims 1-2, 4-7, 9-10, 12-16, 18-23 have been considered. Claim 7 has been maintained. Claims 1-2, 4-6, 9-10, 12-16, 18-29, and 20 have been amended. New claims 21-23 have been added. Claims 3, 8, 11, and 17 have been cancelled. Response to Arguments Applicant’s arguments with respect to claims 1-20 have been considered but are moot in view of new grounds of rejection necessitated by the applicant’s amendments to the claims. Specification The title of the invention contains a typo in the word “Learning”. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: Sound Modification Using Machine Learning Classification Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claim 4 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 4 states “(Currently Amended) The method of claim [[3]]” with an amendment removing the claim 3 dependency but does not replace the claim 3 dependency with an alternative claim number. For the purposes of examination, this claim will be interpreted as “(Currently Amended) The method of claim [[3]]2” based on the dependencies of mirrored claims 12 and 18.” Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 7, 9, 15 are rejected under 35 U.S.C. 103 as being unpatentable over Rubenstein et al. (hereinafter Rubenstein) (AudioPaLM: A Large Language Model That Can Speak and Listen.) in view of Efros et al. (hereinafter Efros) (WO 2022071959 A1) and in further view of Benattar et al. (hereinafter Benattar) (US 20190028803 A1). Regarding claim 1, Rubenstein discloses: (Currently Amended) A computer-implemented method comprising (Rubenstein, Abstract: “We introduce AudioPaLM, a large language model for speech understanding and generation.”, “AudioPaLM fuses text-based and speech-based language models, PaLM-2 [Anil et al., 2023] and AudioLM [Borsos et al., 2022], into a unified multimodal architecture that can process and generate text and speech with applica tions including speech recognition and speech-to-speech translation.” (Specific LLM and methodology behind how it works is a computer-implemented method)): Outputting, with a large language model (Rubenstein, Abstract: “AudioPaLM, a large language model for speech understanding and generation.”, “a unified multimodal architecture that can process and generate text and speech” (supplies claimed language model)) Rubenstein does not disclose: receiving input from a user that identifies categorizes descriptions of different individual types of audio as positive audio or negative audio associated with one or more non- user sources and specifies a degree of positivity of the positive audio or a degree of negativity of the negative audio; a classification of the individual types of audio that the user identified categorizes as positive audio or negative audio; identifying capturing, with a microphone, audio in a physical environment; splitting, with an audio machine-learning model, the audio into audio sources; one or more matched audio sources that are matched with the descriptions of the individual one or more types of audio that the user categorized categorizes as positive audio or negative audio; and modifying output of an auditory device based, at least in part, on the one or more matched audio sources, wherein a positive audio source is amplified based, at least in part, on the degree of positivity and and/or a negative audio source is reduced based, at least in part, on the degree of negativity or cancelled However, Efros discloses: a classification of the individual types of audio that the user identified categorizes as positive audio or negative audio (Efros, Detailed Description: “The speech separation engine 130 implements one or more neural networks configured to process an input video of one or more speakers and generate isolated speech signals for each speaker”, Summary: “user device configured according to techniques described in this specification provides an interface for selecting different speakers”, “The system can be applied to a variety of different settings in which clean audio of a particular speaker is desired” (addresses classification/selection of individual audio types/speakers, which when combined with Rubenstein’s LLM yields the claimed LLM-based classification of positive/negative audio types (desired reads on positive, and there must be undesired (negative) audio that gets filtered out if you can get clean audio of desired (positive) audio)); identifying capturing, with a microphone, audio in a physical environment (Efros, Detailed Description: “the user device 105 records an audio soundtrack, i.e., using the microphone 120”); splitting, with an audio machine-learning model, the audio into audio sources (Efros, Detailed Description: “The speech separation engine 130 implements one or more neural networks configured to process an input video of one or more speakers and generate isolated speech signals for each speaker, from joint audio-visual features of each speaker.“, Summary: “Automatic speech separation is the problem of separating an audio soundtrack of speech of one or more speakers into isolated speech signals of each respective speaker” (addresses splitting audio into audio sources with an audio machine-learning model)); one or more matched audio sources (Efros, Summary: “generating corresponding isolated speech signals for playback in real time and corresponding to the different selected speakers.” (teaches outputting matched audio sources)) It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to modify Rubenstein in view of Efros because Rubenstein. Doing so would have combined Rubenstein’s processing of environmental audio, separating/classifying audio content and controlling playback based on identified sound sources (Rubenstein, Abstract) with Efros’ use of machine-learning models for recognizing and categorizing audio content from user-provided preferences (Efros, Abstract, P[0082-P0090]) thus, improving source-identification and personalization capabilities, and enabling more accurate and adaptive audio modification based on user preferences The combination of Rubenstein and Efros does not disclose: receiving input from a user that identifies categorizes descriptions of different individual types of audio as positive audio or negative audio associated with one or more non- user sources and specifies a degree of positivity of the positive audio or a degree of negativity of the negative audio that are matched with the descriptions of the individual one or more types of audio categorized categorizes as positive audio or negative audio modifying output of an auditory device based, at least in part, on the one or more matched audio sources, wherein a positive audio source is amplified based, at least in part, on the degree of positivity and and/or a negative audio source is reduced based, at least in part, on the degree of negativity or cancelled However, Benattar discloses: receiving input from a user that identifies categorizes descriptions of different individual types of audio as positive audio or negative audio associated with one or more non- user sources (Benattar, P[0082]: “enable listeners to actively characterize elements of the ambient sound environments in which they find themselves into desirable sound and undesirable noise” (addresses receiving input from user that categorizes different types of non-user ambient/source audio into positive/desirable and negative/undesirable audio), P[0088]: “create and add to their own library of desirable ambient sounds” (further addresses user categorization of individual types of audio), Benattar, P[0083]: “allow users to utilize a library of predetermined desirable ambient sounds and ambient profiles or “experiences”” (teaches descriptions/stored identifiable types of audio that user categorizes)) and specifies a degree of positivity of the positive audio or a degree of negativity of the negative audio (Benattar, P[0088]: “apply a variety of adjustments/mixing controls to that combined sound environment to ensure the appropriate blending of the sounds, such adjustments to include, but are not limited to, relative volume,” (addresses specifying a degree of positivity/negativity because the user-adjusted relative volume is the claimed degree of positive enhancement or negative reduction), P[0087]: “adjustments to the characteristics of the noise cancelling experience” (further addresses adjusting the degree of treatment)); that are matched with the descriptions of the individual one or more types of audio (Benattar, P[0083], P[0088]: “library of predetermined desirable ambient sounds”, “create and add to their own library of desirable ambient sounds” (teach descriptions/identifiable audio types)) that the user categorized categorizes as positive audio or negative audio (Benattar, P[0082]: “desirable sound and undesirable noise” (teaches the user’s positive/negative categorization); and modifying output of an auditory device based, at least in part, on the one or more matched audio sources, wherein a positive audio source is amplified based, at least in part, on the degree of positivity and and/or a negative audio source is reduced based, at least in part, on the degree of negativity or cancelled (Benattar, P[0090], P[0091]: “Directional microphones may be used to isolate and enhance or damp audio originating from a particular direction.”, “enhance delivery of desirable audio and damp delivery of undesirable audio” (this addresses modifying output based on matched audio sources so that positive/desirable audio is amplified and negative/undesirable audio is reduced), P[0085] “relative volume” (this addresses that the amount of amplification/reduction is based on degree of positivity/negativity)). It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to combine Rubenstein in view of Efros, and in further view of Benattar. Doing so would have provided the architectures for training using aligned audio, text represntations, tokenized, audio , and tokenized textual descriptions of Benattar (Benattar, Abstract, P[0052]) with Rubenstein’s processing of environmental audio, separating/classifying audio content and controlling playback based on identified sound sources (Rubenstein, Abstract) and Efros’ use of machine-learning models for recognizing and categorizing audio content from user-provided preferences (Efros, Abstract, P[0082-P0090]) thus enabling the combined system to perform ore robust semantic classification, matching of sound descriptions to captured audio sources, and multimodal inference using a large language mode, which means improving flexibility and classification accuracy using known techniques for their established purpose. Regarding claim 7, the combination of Rubenstein, Efros, and Benattar discloses the method of claim 1. Rubenstein further discloses: receiving training data that includes audio files and descriptions of the audio files (Rubenstein, Page 6: “All datasets used in this report are speech-text datasets which contain a subset of the following fields. • Audio: speech in the source language. • Transcript: a transcript of the speech in Audio.”)); outputting, with an audio encoder, embedded audio (Rubenstein, Page 4: “extracting embeddings from an existing speech representation model” and “e embeddings from the w2v-BERT model” (teaches outputting, with an audio encoder, embedded audio)); generating audio tokens (Rubenstein, Page 4: “convert raw waveforms into tokens”); tokenizing the descriptions of the audio files (Rubenstein, Page 4: “a SentencePiece [Kudo and Richardson, 2018b] one used to represent text.” And “text tokens” (teaches tokenizing the descriptions of the audio files)); outputting, with a text embedding layer, text tokens (Rubenstein, Page 5: “first t tokens (from zero to t) correspond to the SentencePiece text tokens” and “embeddings matrix” (teaches outputting, with a text embedding layer, text tokens)); and providing pairs of audio tokens and corresponding text tokens to the large language model for training (Rubenstein, Page 6: “All datasets used in this report are speech-text datasets which contain a subset of the following fields. • Audio: speech in the source language. • Transcript: a transcript of the speech in Audio. • Translated audio: the spoken translation of the speech in Audio. • Translated transcript: the written translation of the speech in Audio. The component tasks that we consider in this report are: • ASR(automatic speech recognition): transcribing the audio to obtain the transcript.” And Page 4: “We use a decoder-only Transformer to model sequences consisting of text and audio tokens.” (paired audio/text training examples are tokenized into corresponding audio tokens and text tokens and provided to the large language model for training)). Regarding claim 9, claim 9 recites the system corresponding to the computer-implemented method presented in claim 1 and is rejected under the same grounds as above. Efros, in combination with Rubenstein and Benattar, further discloses: A system comprising (Efros, Background, “relates to a system”): one or more processors (Efros, “a programmable processor, a computer, or multiple processors”); and logic encoded in one or more non-transitory media for execution by the one or more processors and when executed are operable to (Efros, “one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus”): Regarding claim 15, claim 15 recites the software encoded in one or more non-transitory computer-readable media corresponding to the computer-implemented method presented in claim 1 and is rejected under the same grounds as above. Efros, in combination with Rubenstein and Benattar, further discloses: (Currently Amended) Software encoded in one or more non-transitory computer-readable media for execution by one or more processors of an auditory device and when executed is operable to (Efros, Claim 11): Claims 2, 4-6, 10, 12-14, 16, 18-20, 21-23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Rubenstein et al. (hereinafter Rubenstein) (AudioPaLM: A Large Language Model That Can Speak and Listen.) in view of Efros et al. (hereinafter Efros) (WO 2022071959 A1), in further view of Benattar et al. (hereinafter Benattar) (US 20190028803 A1), and in furthest view of Cheung et al. (hereinafter Cheung) (US 8582789 B2) (See attached for Paragraph numbers). Regarding claim 2, the combination of Rubenstein, Efros, and Benattar discloses the method of claim 1. Efros further discloses: Further comprising: responsive to receiving the input from the user, generating a user interface audio (Efros, Brief Description of Drawings: “FIG. 2 illustrates an example of a user interface 200 for obtaining isolated speech signals.”, Detailed Description: “The user interface 200 is configured to receive input, e.g., tactile input, for selecting which isolated speech signals are output for speakers detected in the current scene 205.”(addresses generating a user interface responsive to user input, with selectable audio/speaker types) The combination of Rubenstein and Efros does not disclose: that includes a list of the types of audio that the user identified categorizes as positive audio or negative, wherein the user interface includes and options for specifying [[a]]of the degree of positivity positive audio or a and/or the degree of negativity negative audio; and generating a user profile based on the input from the user; while capturing the audio in the physical environment, receiving verbal instruction from the user to adjust the degree of positivity or degree of negativity of a particular type of audio; and changing the amplifying or the reducing of the output of the particular type audio of the captured audio according to the verbal instruction However, Benattar further discloses: that includes a list of the types of audio that the user identified categorizes as positive audio or negative, (Benattar, P[0082]-[0083]: “desirable sound and undesirable noise” (teaches the user’s positive/negative categorization)“ allow users to utilize a library of predetermined desirable ambient sounds and ambient profiles or “experiences” (addresses list/library of audio types the user categorizes)), wherein the user interface includes and options for specifying [[a]]of the degree of positivity positive audio or a and/or the degree of negativity negative audio (Benattar, P[0085]: “apply a variety of adjustments/mixing controls to that combined sound environment to ensure the appropriate blending of the sounds, such adjustments to include, but are not limited to, relative volume” and P[0327]: “sliders to change or customize audible parameters in an audio library.” (addresses options on UI for specifying degree of positivity/negativity)); and generating a user profile based on the input from the user (Benattar, P[0083]: “Those settings could be saved as an “experience” within their library” (addresses generating/storing a user profile based on user input/preferences)); and changing the amplifying or the reducing of the output of the particular type audio (Benattar, P[0091]: “enhance delivery of desirable audio and damp delivery of undesirable audio.” (addresses that the changed treatment is amplification/reduction of a particular type of audio) The combination of Rubenstein, Efros, and Benattar does not disclose: while capturing the audio in the physical environment, receiving verbal instruction from the user to adjust the degree of positivity or degree of negativity of a particular type of audio of the captured audio according to the verbal instruction However, Cheung discloses: while capturing the audio in the physical environment, receiving verbal instruction from the user to adjust the degree of positivity or degree of negativity of a particular type of audio (Cheung, Description, P(31): “the user has the option of manually changing the amplification of the system. The system can also have a general volume controller”, Claims 45 and 64: “the apparatus includes a volume control that is configured to respond to a voice command of the user.” and “configured to recognize at least a voice command from the user and operate according to the voice command.” (addresses receiving a verbal instruction from the user during live operation to adjust the degree of treatment of a particular type of audio)); of the captured audio according to the verbal instruction (Cheung, Claims 45: “volume control that is configured to respond to a voice command of the user.”). It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined Rubenstein in view of Efros, in further view of Benattar, and in furthest view of Cheung. Doing so would have provided the audio output adaption based on user-specific hearing profiles and adjustments of output levels according to user hearing characteristics (Cheun, Abstract, Claims) with the architectures for training using aligned audio, text representations, tokenized, audio , and tokenized textual descriptions of Benattar (Benattar, Abstract, P[0052]), Rubenstein’s processing of environmental audio, separating/classifying audio content and controlling playback based on identified sound sources (Rubenstein, Abstract), and Efros’ use of machine-learning models for recognizing and categorizing audio content from user-provided preferences (Efros, Abstract, P[0082-P0090]) thus improving the personalization of output volume and amplification levels for different users and sound types which would yield a system better tailored to user hearing needs through known hearing-enhancement techniques. Regarding claim 4, the combination of Rubenstein, Efros, Benattar, and Cheung discloses the method of claim 2. Efros further discloses: displaying on the user interface (Efros, Brief Description of Drawings: “illustrates an example of a user interface 200 for obtaining isolated speech signals.”, Detailed Description: “The user interface 200 can indicate visually which speakers are currently selected” (this addresses a visual indicator on the user interface)), The combination of Rubenstein and Efros does not disclose: updating the user profile with the adjusted degree of positivity or degree of negativity of the particular type of audio to categorize the type of sound as positive audio or negative audio based on the instruction from the user; and a visual indicator of the adjusted degree of positivity or degree of negativityHowever, Benattar further discloses: updating the user profile with the adjusted degree of positivity or degree of negativity of the particular type of audio to categorize the type of sound as positive audio or negative audio based on the instruction from the user (Benattar, P[0083]: “Those settings could be saved as an “experience” within their library”, P[0085]: “apply a variety of adjustments/mixing controls to that combined sound environment to ensure the appropriate blending of the sounds, such adjustments to include, but are not limited to, relative volume” (addresses updating user profile with the adjusted degree based on user instruction)); and a visual indicator of the adjusted degree of positivity or degree of negativity (Benattar, P[0327]: “P[0327]: “sliders to change or customize audible parameters in an audio library.” (this addresses that the indicator reflects the adjusted degree)). Regarding claim 5, the combination of Rubenstein, Efros, and Benattar discloses the method of claim 1. Rubenstein further discloses: classifying, with the large language model (Rubenstein, Abstract: “AudioPaLM, a large language model for speech understanding and generation.”, “a unified multimodal architecture that can process and generate text and speech”, Page 2: “a single decoder-only model on a mixture of tasks that involve arbitrarily interleaved speech and text” (supplies claimed language model)) Rubenstein does not disclose: detecting, with the microphone, an a volume instruction from a user to increase a volume of a particular person in the physical environment of the user; a particular type of audio associated with the particular based on the volume instruction from the user; and modifying output of the auditory device to increase the volume of the particular person and decreasing an initial volume of audio associated with at least one other person in the environment of the user; ceasing to detect the audio of the particular person; and resuming the initial volume of the audio associated with the at least one other person. However, Efros further discloses: of a particular person in the physical environment of the user (Efros, Abstract: “receiving, by a user device, a first indication of one or more first speakers visible in a current view” (addresses that the instruction is directed to a particular person in the physical environment)); a particular type of audio associated with the particular person (Efros, Description, “generate isolated speech signals for each speaker” and “selected speakers” (addresses particular type of audio associated with particular person)) and modifying output of the auditory device to increase the volume of the particular person and decreasing an initial volume of audio associated with at least one other person in the environment of the user (Efros, Summary: “Automatic speech separation is the problem of separating an audio soundtrack of speech of one or more speakers into isolated speech signals of each respective speaker, to enhance the speech of a particular speaker and/or to mask the speech of other speakers so that only particular speakers are heard.” (teaches increasing the selected person’s audio and decreasing audio of at least one other person)); ceasing to detect the audio of the particular person (Efros, see mapping below); and resuming the initial volume of the audio associated with the at least one other person (Efros, Detailed Description: “When a speaker is not in the current view, the speech from the speaker is filtered out, allowing a user to focus the user device to aim at speakers of interest with the camera at the exclusion of other speakers.” And “moving from one speaker in a room to another speaker in the room.” (teaches ceasing detection of the particular person and returning output focus to the other speakers)). The combination of Rubenstein, Efros, and Benattar does not disclose: detecting, with the microphone, an a volume instruction from a user to increase a volume based on the volume instruction from the user However, Cheung discloses: detecting, with the microphone, an a volume instruction from a user to increase a volume (Cheung, claim 45: “volume control that is configured to respond to a voice command of the user.”, claim 64: “recognize at least a voice command from the user and operate according to the voice command.”) based on the volume instruction from the user (Cheung, claim 45, (addresses the classification/selection is based on user’s volume instruction); Regarding claim 6, the combination of Rubenstein, Efros, Benattar, and Cheung discloses the method of claim 5. Rubenstein further discloses: as identified by the large language model (Rubenstein, Abstract: “AudioPaLM, a large language model for speech understanding and generation.”, “a unified multimodal architecture that can process and generate text and speech” (supplies claimed language model)) Rubenstein does not disclose: The method of claim 5, wherein the volume instruction from the user further includes an identification of a side location of the person and classifying the particular type of audio is further based on the side location of the person, and wherein modifying output of the auditory device includes modifying the auditory device associated with the side location. However, Efros further discloses: The method of claim 5, wherein the volume instruction from the user further includes an identification of a side location of the person (Efros, Detailed Description: “if the user device 105 is tracking speech on either side of a user of the device, then the user device can send the isolated speech signals to match the location of the speaker.” (teaches identification of a side location of the person)) and classifying the particular type of audio is further based on the side location of the person (Efros, Detailed Description: “send the isolated speech signals to match the location of the speaker.” (teaches classification/selection based on side location)), and wherein modifying output of the auditory device includes modifying the auditory device associated with the side location (Efros, Detailed Description: “As described above, the user device can send different signals to different audio channels to a listening device that supports multiple audio channels, e.g., supports stereo sound. The user device is configured to send an isolated speech signal for the speaker 210a to sound as though the speaker 210a is on the left side of a user listening through a corresponding listening device, and to send an isolated speech signal for the speaker 210b to sound as though the speaker 210b is on the right side of a user listening through the listening device.” And “send the isolated speech signals to match the location of the speaker.“ (teaches modifying the auditory device/channel associated with the side location)). Regarding claim 10, claim 10 recites the system corresponding to the computer-implemented method presented in claim 2 and is rejected under the same grounds as above. Regarding claim 12, claim 12 recites the system corresponding to the computer-implemented method presented in claim 4 and is rejected under the same grounds as above. Regarding claim 13, claim 13 recites the system corresponding to the computer-implemented method presented in claim 5 and is rejected under the same grounds as above. Regarding claim 14, claim 14 recites the system corresponding to the computer-implemented method presented in claim 6 and is rejected under the same grounds as above. Regarding claim 16, claim 16 recites the software encoded in one or more non-transitory computer-readable media corresponding to the computer-implemented method presented in claim 2 and is rejected under the same grounds as above. Regarding claim 18, claim 18 recites the software encoded in one or more non-transitory computer-readable media corresponding to the computer-implemented method presented in claim 4 and is rejected under the same grounds as above. Regarding claim 19, claim 19 recites the software encoded in one or more non-transitory computer-readable media corresponding to the computer-implemented method presented in claim 5 and is rejected under the same grounds as above. Regarding claim 20, claim 20 recites the software encoded in one or more non-transitory computer-readable media corresponding to the computer-implemented method presented in claim 6 and is rejected under the same grounds as above. Regarding claim 21, the combination of Rubenstein, Efros, and Benattar discloses the method of claim 1. Benattar, in combination with Rubeinstein and Efros, discloses: further comprising: determining an initial volume output level of a particular type of audio based, at least in part, on a type of the audio device (Benattar, P[0150]: “Adjustments for reproduction device characteristic may be based on pre-established profiles or user preference. The profiles may be generic to a reproduction device class or may be specific to an individual reproduction device model.” (teaches basing the output level also on a type of the audio device)) wherein the modifying of the output of the auditory device includes modifying the initial volume output level based on the degree of positivity or degree of negativity (Benattar, P[0082]: “enable listeners to actively characterize elements of the ambient sound environments in which they find themselves into desirable sound and undesirable noise” and P[0085]: “apply a variety of adjustments/mixing controls to that combined sound environment to ensure the appropriate blending of the sounds, such adjustments to include, but are not limited to, relative volume” (teaches modifying the initial volume output level based on the degree of positivity or negativity)). The combination of Rubenstein, Benattar, and Efros does not disclose: and a hearing profile of the user in which one or more characteristics of the particular type of audio matches a parameter of the hearing profile, However, Cheung discloses: and a hearing profile of the user (Cheung, Summary of Invention, P(23): “In a third approach, the user's hearing is profiled so that frequency amplification is tailored to the user.” And “The system can then adjust the amplification of the audio signals across the frequencies based on the user's hearing profile” (teaches determining an initial output level based on a hearing profile of the user)) in which one or more characteristics of the particular type of audio matches a parameter of the hearing profile (Cheung, P(22), Summary of Invention: “The decrease in hearing may not be uniform across all audio frequencies.” And P(23): “frequency amplification is tailored to the user.” (teaches matching characteristics of the audio, such as frequency content, to parameters of the hearing profile)), It would have been obvious to one of ordinary skill in the art at the time of the claimed invention to have combined Rubenstein in view of Efros, in further view of Benattar, and in furthest view of Cheung. Doing so would have provided the audio output adaption based on user-specific hearing profiles and adjustments of output levels according to user hearing characteristics (Cheun, Abstract, Claims) with the architectures for training using aligned audio, text representations, tokenized, audio , and tokenized textual descriptions of Benattar (Benattar, Abstract, P[0052]), Rubenstein’s processing of environmental audio, separating/classifying audio content and controlling playback based on identified sound sources (Rubenstein, Abstract), and Efros’ use of machine-learning models for recognizing and categorizing audio content from user-provided preferences (Efros, Abstract, P[0082-P0090]) thus improving the personalization of output volume and amplification levels for different users and sound types which would yield a system better tailored to user hearing needs through known hearing-enhancement techniques. Regarding claim 22, claim 22 recites the system corresponding to the computer-implemented method presented in claim 21 and is rejected under the same grounds as above. Regarding claim 23, claim 23 recites the software encoded in one or more non-transitory computer-readable media corresponding to the computer-implemented method presented in claim 21 and is rejected under the same grounds as above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHASHIDHAR S MANOHARAN whose telephone number is (571)272-6772. The examiner can normally be reached M-F 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHASHIDHAR SHANKAR MANOHARAN/Examiner, Art Unit 2655 /ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Show 2 earlier events
Jan 30, 2026
Interview Requested
Feb 10, 2026
Examiner Interview Summary
Feb 10, 2026
Applicant Interview (Telephonic)
Feb 17, 2026
Response Filed
May 04, 2026
Final Rejection mailed — §103, §112
Jul 08, 2026
Examiner Interview Summary
Jul 08, 2026
Applicant Interview (Telephonic)
Jul 10, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682890
MASK-CONFORMER AUGMENTING CONFORMER WITH MASK-PREDICT DECODER UNIFYING SPEECH RECOGNITION AND RESCORING
2y 4m to grant Granted Jul 14, 2026
Patent 12682173
MODULAR FRAMEWORK FOR EVALUATING LANGUAGE MODELS
2y 4m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 1m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 3 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month