Prosecution Insights
Last updated: October 04, 2026
Application No. 18/939,940

METHOD FOR OPERATING A HEARING AID SYSTEM AND HEARING AID SYSTEM

Non-Final OA §101§103
Filed
Nov 07, 2024
Priority
Nov 07, 2023 — DE 10 2023 211 026.1
Examiner
JACKSON, JAKIEDA R
Art Unit
Tech Center
Assignee
Sivantos Pte. Ltd.
OA Round
2 (Non-Final)
74%
Grant Probability
Favorable
2-3
OA Rounds
1y 1m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
683 granted / 921 resolved
+14.2% vs TC avg
Strong +16% interview lift
Without
With
+15.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
40 currently pending
Career history
955
Total Applications
across all art units

Statute-Specific Performance

§101
27.1%
-12.9% vs TC avg
§103
42.3%
+2.3% vs TC avg
§102
20.7%
-19.3% vs TC avg
§112
2.8%
-37.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 921 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment In response to the Office Action mailed May 20, 2026, applicant submitted an amendment filed on July 15, 2026, in which the applicant amended and requested reconsideration. Response to Arguments Applicants arguments towards the art rejection is persuasive, however, moot in view of new grounds of rejection. Regarding the 101, Applicants amended the claims to include a transducer and a display. Merely adding the electro-acoustic transducer and a display does not make the claim eligible. The hardware does not meaningfully integrate the abstract idea into a practical application. The claims recite information collection, analysis and generation falling within mental processes and certain methods of organizing human activity. The additional elements such as the receiving unit, evaluation unit, output unit and transducer are not integrated into a practical application. In particular, the claim only recites additional elements which are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. Reciting concrete, tangible components does not establish an inventive concept and a generic machine performing conventional functions does not automatically integrate an exception into a practical application. The transducer is not improved in a new manner. It’s simply the mechanism by which information is communicated to the user, which is insignificant extra solution activity, rather than a meaningful technological application of the abstract idea. Therefore, Applicants arguments have been considered, but are not persuasive. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claims are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. The claims are directed to the abstract idea of processing speech, as explained in detail below. The limitations, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “various units” nothing in the claim element precludes the steps from practically being performed by mental processing/organizing human activity. For example, the language, receiving speech information and converting the speech information into a speech signal (can be done by a user listening to someone speak and speaking what was said); converting the speech signal into a text signal (can be done by a user transcribing data); generating a context-dependent prompt for a natural language processing unit (can be done by a user creating a prompt accordingly); triggering the prompt unit as required and specifying a stored context (can be done by a user gathering prompt data and determining the context); evaluating the prompt by using the natural language processing unit and for generating an output signal; outputting the output signal to a user (can be done by a user evaluating the prompt and outputting the data); and upon triggering the triggering unit: a) converting at least one section of the speech signal into a text signal by using the speech recognition unit (can be done by a user transcribing data), b) generating the prompt for the natural language processing unit by using the stored context and the text signal (can be done by a user generating a prompt), c) using the natural language processing unit to generate the output signal based on the prompt, and d) using the output unit to output the output signal (can be done by a user outputting speech data). The present claim language under its broadest reasonable interpretation, covers performance of mental processing and recites generic computer components, which all falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. This judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements which are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible. The dependent claims recite using a gesture, generating a section, storing profile data, adjusting output data based on physiological data, storing data, updating data, paraphrasing data, and outputting speech, which is all part of the mental processing/organizing human activity and non-statutory. The claims are directed to organizing, processing and generating content (audio) based on rules and a user input. The claims fall under the mental processing (segmenting, configuring and outputting data) along with methods for organizing information and content generation workflows. The claims recite receiving, organizing and outputting data, which is very close to generic data processing on a computer, which is non-statutory. According to Step 1, it includes determining whether the claims fall within a statutory category. The claims include a method, therefore the claims fall within a statutory category. Step 2A Prong one, includes evaluating whether the claims recite a judicial exception. The claims recite a judicial exception, therefore an evaluation is done to determine if the claims fit into one of the categories. As explained, the claims fit into the mental processing concept. Step 2A Prong two is used to evaluate whether the claims recite additional elements that integrate the exception into a practical application. As explained the judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements which are recited at a high-level of generality (i.e., as a generic processor performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. Therefore, the claims are non-statutory. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cella et al. (PGPUB 2019/0348041), hereinafter referenced as Cella and in further of Irving et al. (PGPUB 2024/0104336), hereinafter referenced Irving. Regarding claims 1 and 10, Cella discloses a method and system, hereinafter referenced as a method for operating a hearing aid system, the method comprising: using a receiving unit for receiving speech information and converting the speech information into a speech signal (the receiver can be decoded by the signal processing module to generate an electrical signal that can be converted into audible sound; p. 0180, 0186-0189, 0195); using a speech recognition unit for converting the speech signal into a text signal (generate transcripts of conversations that are captured by the audio signals; p. 0195, 0430, 0448); using a prompt unit for generating a context-dependent prompt for a natural language processing unit (In embodiments, NLP can include natural language understanding (NLU) of the audio signal (e.g., extracting meaning of the speech portion of the audio signal). In embodiments, NLP can include natural language generation (NLG) in response to the audio signal or another prompt (e.g., generating human-understandable speech given the context of a situation). In embodiments, the in-ear device 310 can obtain user feedback on the accuracy of the NLU using NLG. In embodiments, NLP can include identifying different speakers in the speech portion. Thus, in some embodiments, the in-ear device 310 may be configured to generate transcripts of conversations that are captured by the audio signals. The in-ear device 310 may be configured to process the audio signal in other suitable manners as well; p. 0195); using an evaluation unit for evaluating the prompt by using the natural language processing unit and for generating an output signal (The one or more microphones 470 may be any suitable microphones that are configured and dimensioned to fit within a housing of the in-ear device 310. In embodiments, the microphone(s) 470 may include one or more directional microphones and/or independent microphones 470 that are positioned to receive audio signals from different directions. In these embodiments, the audio signal output by the one or more microphones 470 may be comprised of composite audio signals captured from different orientations. In embodiments, each microphone may capture a sound wave and convert the sound wave into a respective digital signal. Each one of these signals may be a composite audio signal, where the composite audio signals collectively make up the audio signal. As each microphone may capture the same sound wave but at a slightly different orientation, the output of each microphone may slightly vary in amplitude and/or frequency (e.g., due to Doppler Effect). The in-ear device 310 (e.g., the signal processing module 402) may utilize these differences to extrapolate a direction of travel of the sound wave on which the audio signal is derived. Speaker 480 may be any suitable speaker device configured and dimensioned to fit in the housing of the in-ear device 310. The speaker 480 may be used to provide audible communications and/or sounds to the user; p. 0229-0230); using an output unit for outputting the output signal to a user (content item to be played for the user on a user device; p. 0080, 0084, 0195); and upon triggering the triggering unit: a) converting at least one section of the speech signal into a text signal by using the speech recognition unit (a speech portion of the audio signal, feature vectors that include features extracted from the audio signal, words that are recognized in a speech portion of the audio signal, transcripts of the speech portion of the audio signal, annotation objects that include the results of NLP performed on the audio signal; p. 0199); and driving an electro-acoustic output transducer of the output unit with the output signal, thereby converting the output signal into an acoustic sound signal output to the user (p. 0431-0433), but does not specifically teach using a triggering unit for triggering the prompt unit as required and specifying a stored context, generating the prompt for the natural language processing unit by using the stored context and the text signal and using the natural language processing unit to generate the output signal based on the prompt. Irving discloses a method comprising: using a triggering unit for triggering the prompt unit as required and specifying a stored context (determining an initial context input, wherein the context input is the prompt; p. 0125-0135 and the follow-up data; p. 0138-0141); generating the prompt for the natural language processing unit by using the stored context and the text signal (receives a natural language request from the user; p. p. 0125-0135); and using the natural language processing unit to generate the output signal based on the prompt (updated context input to include natural language request and processing updated context input using the first trained language generation neural network, wherein the neural network generates one or more samples of a natural language response, the response can answer the natural language question and the reply is provided to the user; p. 0125-0135), in order to generate responses that account for the context of the user interaction. Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to provide more relevant and contextually appropriate natural language responses to the user. Regarding claim 2, Cella discloses a method which further comprises triggering the triggering unit by using at least two different trigger types each being linked to a different, stored, context, and generating the prompt for the natural language processing unit depending on the trigger type (The illustrated in-ear device 20 includes one or more of the following: a sensor 5A in the form of one or more physiological sensors and/or one or more environmental sensors (which can include an acoustical sensor, a motion sensor, a temperature sensor, a galvanic skin response sensor, a heat flux sensor, a chemical sensor, a pressure sensor, or other sensor) (and in some instances can also be referred to as an external energy sensor, or a sensor set herein, where “set” as used herein should be understood to encompass a set that has a single member or a plurality of members) housed on or within a housing of the device 20 (and optionally or additionally houses external thereto), at least one signal processor 4, at least one transmitter/receiver 6A, 6B, or 6C, at least one power source 2, at least one body attachment component 21 which can be an inflation element or balloon, a foam tip, a polymer-based housing, or the like, and which can include one or more of the other elements of the in-ear device 20, and at least the housing. The housing can include the main body housing 23 and a stent or extension 22 as well as a flange 24 and can further include the inflation element or balloon 21. The housing can further include an end cap 25 which can further carry or incorporate a capacitive or resistive sensor 26 or optical sensor as shown in FIGS. 2B and 2C. The main housing portion 23 can also include a venting port 23A to enable additional venting between the flange 24 and the balloon 21 when the device 20 is inserted within EAC. The sensor 26 can be used to detect gestures in ad hoc or predetermined patterns or in yet another embodiment the sensor 26 can alternatively be a fingerprint type of sensor. The inflation element or balloon 21, or other enclosure (such as a foam tip, polymer-based tip, or other flexible enclosure) can include, incorporate, carry or embed one or more sensors. In embodiments, a sensor may include a surface acoustic wave or SAW sensor 21A that can be used for measuring blood pressure. In one embodiment, a balloon having conductive traces on the surface of the balloon to serve as the surface acoustic wave sensor can be used for measuring blood pressure. In embodiments, a sensor may include an optical sensor for measuring one or more characteristics of the blood of a user, such as blood glucose levels; p. 0167, 0195). Regarding claim 3, Cella discloses a method which further comprises using a gesture of a hearing aid system user as at least one of the trigger types (The sensor 26 can be used to detect gestures in ad hoc or predetermined patterns or in yet another embodiment the sensor 26 can alternatively be a fingerprint type of sensor; p. 0167). Regarding claim 4, Cella discloses a method which further comprises generating a section of the speech signal or a combination of at least two sections of the speech signal as the output signal (In embodiments, the in-ear device may parse the speech portion of the audio signal to identify a sequence of phonemes. The in-ear device may determine potential utterances (e.g., words) based on the phonemes. In some implementations, the in-ear device generates various n-grams (unigrams, bi-grams, tri-grams, etc.) of sequential phonemes; p. 0302). Regarding claim 5, Cella discloses a method which further comprises storing a hearing profile for a user of the hearing aid system, and generating the prompt in dependence on the hearing profile (The database 2455 may also store information and metadata obtained from the system 2400, store metadata and other information associated with the first and second users 2401, 2420, store user profiles associated with the first and second users 2401, 2420, store device profiles associated with any device in the system 2400, store communications traversing the system 2400, store user preferences, store information associated with any device or signal in the system 2400, store information relating to patterns of usage relating to the first, second, third, fourth, and fifth user devices 2402, 2406, 2410, 2421, 2425, store audio content associated with the first, second, third, fourth, and fifth user devices 2402, 2406, 2410, 2421, 2425 and/or earphone devices 2415, 2430, store audio content and/or information associated with the audio content that is captured by the ambient sound microphones, store audio content and/or information associated with audio content that is captured by ear canal microphones, store any information obtained from any of the networks in the system 2400, store audio content and/or information associated with audio content that is outputted by ear canal receivers of the system 2400, store any information and/or signals transmitted and/or received by transceivers of the system 2400, store any device and/or capability specifications relating to the earphone devices 2415, 2430, store historical data associated with the first and second users 2401, 2415, store information relating to the size (e.g. depth, height, width, curvatures, etc.) and/or shape of the first and/or second user's 2401, 2420 ear canals and/or ears, store information identifying and or describing any eartip utilized with the earphone devices 2401, 2415, store device characteristics for any of the devices in the system 2400, store information relating to any devices associated with the first and second users 2401, 2420, store any information associated with the earphone devices 2415, 2430, store log on sequences and/or authentication information for accessing any of the devices of the system 2400, store information associated with the communications networks 2416, 2431, store any information generated and/or processed by the system 2400, store any of the information disclosed for any of the operations and functions disclosed for the system 2400 herewith, store any information traversing the system 2400, or any combination thereof. Furthermore, the database 2455 may be configured to process queries sent to it by any device in the system 2400; p. 0452). Regarding claim 6, Cella discloses a method which further comprises using a physiological sensor for acquiring information about a bodily state of a user of the hearing system and outputting sensor data, and at least one of: adjusting the output signal by using the sensor data from the physiological sensor, or comparing the sensor data from the physiological sensor with a stored threshold value, and triggering the triggering unit upon reaching or exceeding the threshold value (the sensor processing module 412 is configured to draw inferences regarding a user state of the user. A user state may include a mood of the user, a health condition of the user, an activity of the user, and the like. The sensor processing module 412 may utilize a rules-based approach and/or machine-learning to infer a user state. In embodiments, the sensor processing module 412 may implement a rules-based engine that receives senor data regarding a user and outputs an inferred user state. The rules-based engine may implement one or more rules that relate specific sensor values to inferred user states. An example rule may be if the user's temperature increases by 0.1 degrees and the user's heartrate increases by 5%, then the user is likely in an agitated state. Another example rule may be if the user is moving at a speed greater than five miles per hour for more than five minutes and the user's heartrate is greater than 90 beats per minute, then the user is likely exercising. In embodiments, the sensor processing module 412 may employ one or more machine-learned models to infer a user's state. For example, the sensor processing module 412 may employ a neural network that is trained using various biometric sensor readings of people in different states. In operation, the sensor processing module 412 may vectorize the received sensor data and may input the vector into the neural network. The neural network may output one or more candidate user states and a confidence score for each candidate user state. The sensor processing module 412 may select the candidate state having the greatest score as the inferred user state; p. 0245, 0435). Regarding claim 7, Cella discloses a method which further comprises storing and continuously updating the speech signal over a predetermined period of time during operation until the triggering unit is triggered, and using the stored speech signal by the speech recognition unit (According to some embodiments of the present disclosure, a method is disclosed. The method includes receiving, by a processing device of an in-ear device, an audio signal from one or more microphones of the in-ear device, identifying, by the processing device, a speech portion of the audio signal that contains speech of a user of the in-ear device, determining, by the processing device, a plurality of tokens based on the speech portion of the audio signal that contains the speech of the user, a text corpus, and a speech recognition model, and generating, by the processing device, an annotation object based on the plurality of tokens and a natural language processor, the annotation object being indicative of a possible meaning of the speech of the user. The method also includes generating, by the processing device, an in-ear data object based on the annotation object. The method also includes determining, by the processing device, a storage plan based on one or more features of the annotation object and a decision model that is configured to output storage location recommendations based on a set of input features. Each storage location recommendation corresponds to a different storage location of a plurality of possible storage locations, and wherein the plurality of possible storage locations include a storage device of the in-ear device, a user device associated with a user of the in-ear device, and one or more external systems. The method further includes obtaining, by the processing device, user feedback regarding one or more of the plurality of possible storage locations from a user of the in-ear device, updating, by the processing device, based on the user feedback, and storing, by the processing device, the in-ear data object according to the storage plan; p. 0087, 0096, 0240, 0282). Regarding claim 8, Cella discloses a method which further comprises paraphrasing at least one section of the text signal for the output signal in a course of natural language processing (As the in-ear device 310 is configured to collect vast and varied amounts of data (e.g., audio data and/or sensor data) as well as to derive vast amounts of data (e.g., feature vectors, transcripts, NLP results, annotations, summaries, inferences, conclusions, and/or diagnoses), the in-ear device 310 may be utilized as a very powerful data collection tool for a number of different applications; p. 0203). Regarding claims 9 and 11, Cella discloses a method which further comprises using a display screen of the output unit to visually display a text representing the output signal to the user (display transcribed data; p. 0195, 0199, 0203, 0243, 0456). Regarding claim 12, it is interpreted and rejected for reasons as set forth in the combination of claims 1, 9 and 11. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. This information has been detailed in the PTO 892 attached (Notice of References Cited). Tzirkel-Hancock et al. (PGPUB 2017/028746) discloses a vehicle aware speech recognition system and method using prompts. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAKIEDA R JACKSON whose telephone number is (571)272-7619. The examiner can normally be reached Mon - Fri 6:30a-2:30p. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571.272.5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAKIEDA R JACKSON/Primary Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Nov 07, 2024
Application Filed
May 20, 2026
Non-Final Rejection mailed — §101, §103
Jul 15, 2026
Response Filed
Sep 14, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749475
STREAMING SPEECH SYNTHESIS METHOD AND SYSTEM FOR SUPPORTING REAL-TIME CONVERSATION MODEL
2y 9m to grant Granted Sep 29, 2026
Patent 12749486
Arbitration-Based Voice Recognition
2y 2m to grant Granted Sep 29, 2026
Patent 12743584
GRAPH ACQUISITION METHOD AND OBJECT GROUP EXTRACTION MODEL TRAINING METHOD
2y 3m to grant Granted Sep 22, 2026
Patent 12725604
VIDEO SCENE DESCRIBER
2y 7m to grant Granted Sep 01, 2026
Patent 12725617
SELF-LEARNING END-TO-END AUTOMATIC SPEECH RECOGNITION
2y 11m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
74%
Grant Probability
90%
With Interview (+15.6%)
3y 0m (~1y 1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 921 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month