Prosecution Insights
Last updated: August 16, 2026
Application No. 18/728,335

AUDIO PROCESSING APPARATUS, AUDIO PROCESSING METHOD, AND AUDIO PROCESSING SYSTEM

Final Rejection §101§102§103
Filed
Jul 11, 2024
Priority
Jan 21, 2022 — JP 2022-008227 +1 more
Examiner
SHAH, PARAS D
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Kyocera Corporation
OA Round
2 (Final)
73%
Grant Probability
Favorable
3-4
OA Rounds
1y 7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
480 granted / 654 resolved
+11.4% vs TC avg
Strong +31% interview lift
Without
With
+31.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
20 currently pending
Career history
677
Total Applications
across all art units

Statute-Specific Performance

§101
18.3%
-21.7% vs TC avg
§103
47.4%
+7.4% vs TC avg
§102
13.4%
-26.6% vs TC avg
§112
10.7%
-29.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 654 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION This communication is in response to the Amendments and Arguments filed on 04/13/2026. Claims 1-12 are pending and have been examined. Any objection/rejection not mentioned in this Office Action has been withdrawn by the examiner. Change in Examiner The Examiner of record has changed from Feng-Tzer Tzeng to Paras Shah. Response to Amendments and Arguments With respect to the 35 USC 101 abstract rejections, the Applicant asserts that the features cannot be practically be performed in the human mind such as continuously sampling ambient sounds, storing audio data in a storage unit, and provided notifications that an utterance satisfies a set condition in the audio data according to a notification condition corresponding to detected speech. The examiner respectfully disagrees. The process of sampling ambient sounds in the environment by way of a human detecting sounds/voices/speech from other humans can be performed mentally and is something that we do everyday in order to ascertain what is occurring around us. Humans can then mentally store these speech sounds and provide a notification/message by either writing down a message based on what was heard or orally reciting such method. This still constitutes an abstract idea. The additional elements of “storage unit” and “microphone” server the purpose of additional elements which are pre-solution activity to gather the sound data. Furthermore, the “controller” is nothing but a general purpose processor as defined in the specification para [0045]. Therefore, these additional elements do not tie the claims to a practical application as the “controller” is merely being used as a tool on which the abstract idea operates and further do not amount to significantly more due to the generalized nature of these additional elements. Hence, the 35 USC 101 abstract has been maintained. With respect to the 35 USC 103 rejections, the Applicant’s arguments are moot in view of new grounds for rejection and further based on the IDS provided by the Applicant on 01/13/2026. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The independent claims 1, 11, 12 recite an apparatus, a method and a system, thus relating to a statutory category. Claims 1, 11, 12 further recite “store audio data of ambient sound around a user…detect a speech interval by sampling the audio data… upon detecting an utterance satisfying a set condition in the audio data, notify the user that the utterance satisfying the set condition has been detected in accordance with a notification condition corresponding to the utterance ” The limitations as drafted cover mental process. More specifically, with the acquired audio, a human can mentally detect and recognize speech within the ambient environment and if satisfying certain condition, and notify a user in a specific manner based on conditions/rules being met. This judicial exception is not integrated into a practical application. In particular, independent claims 1, 11, 12 recite additional elements of “an acquisition device .. audio processing apparatus ..”, “controller”, “microphone”, “storage unit” however, this is considered general purpose computing devices – see SPECIFICATION – Fig. 2 and [0045] of the specification as filed. Further, the microphone/acquisition device are nothing more than pre-solution activity along with the storage unit in which this audio is stored.. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea. Claims 1, 11, 12 do not include additional elements that are sufficient to amount to significantly more than the judicial exception.. The claim is not patent eligible. With respect to claim 2, the claim further recites “when the utterance satisfies the first condition, play notification sound to notify the user that the audio satisfying the condition has been detected; when utterance satisfies the second condition, present visual information to the user to notify the user that the audio satisfying the condition has been detected, and a priority order of notifying the user is lower when the notification condition satisfies the second condition than when the notification condition satisfies the first condition ..” where a human can mentally determine if the recognized audio satisfies the first or the second condition for the system to perform the following ready, well-known steps. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 3, the claim further recites “when the utterance satisfies the first condition, play the notification sound and then play the utterance satisfying the condition ..” where a human can mentally determine if the recognized audio satisfies the first condition for the system to perform the following ready, well-known steps. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 4, the claim further recites “when the utterance satisfies the second condition, present a list to the user to present the visual information to the user, the list being pieces of information of the detected audio satisfying the condition ..” where a human can mentally determine if the recognized audio satisfies the second condition for the system to perform the following ready, well-known steps. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 5, the claim further recites “when the utterance satisfies the second condition, play the utterance satisfying the condition in response to an input from the user ..” where a human can mentally determine if the recognized audio satisfies the second condition, and with another input (such as user’s utterance) for the system to perform the following ready, well-known steps. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 6, the claim further recites “detect, as the utterance satisfying the set condition, audio whose feature matches the audio feature set in advance ..” where a human can mentally determine when the recognized audio satisfies the condition, if the audio feature (subject to BRI) matches the audio feature set in advance. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 7, the claim further recites “detect, as the utterance satisfying the set condition, a search word ..” where a human can mentally determine when the recognized audio satisfies the condition, if the utterance including a search (key) word. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 8, the claim further recites “when the utterance satisfies the third condition, perform control with which, immediately after the search word is detected, notification sound is played and playback of the utterance is started ..” where a human can mentally determine if the recognized audio satisfies the third condition (such as detected the keyword) for the system to perform the following ready, well-known steps. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 9, the claim further recites “when the utterance satisfies the fourth condition, perform control with which, immediately after the utterance ends, notification sound is played and playback of the utterance is started ..” where a human can mentally determine if the recognized audio satisfies the fourth condition (subject to BRI) for the system to perform the following ready, well-known steps. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. With respect to claim 10, the claim further recites “perform control with which audio data of an utterance interval including the utterance is played, and the utterance interval is an interval for which the audio data continues without interruption for a set time ..” where a human can mentally determine if the audio interval ends for the system to perform the following ready, well-known steps such as audio play control. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1, 6-10, 11, and 12 are rejected under 35 U.S.C. 102(1) as being anticipated by Gordon (US 2020/01411003), cited in IDS. As to claim 11, Gordon teaches an audio processing method comprising: storing audio data of ambient sound around a user collected by a microphone (see [0082]-[0083], where audio capture devices/microphones captures ambient sounds from the environment and stores the audio sample in a buffer of the smart speaker device); detecting a speech interval by sampling the audio data (see [0087], [0092] where wake sounds can comprise identifying emergency words/phrases as well as code/token); and upon detecting an utterance satisfying a set condition being detected in the audio data, notifying the user that the utterance satisfying the set condition has been detected, in accordance with a notification condition corresponding to the utterance (see [0094], where based on the history of the captured data and the sound identification results, the smart speaker can output an audible message to the user or trigger display of information or via visual indicator and see [0096], which describes usage of event models and satisfaction of thresholds to determine whether an event has occurred via rules to determine risk/danger and see [0100], where the actions take different forms based on the type of event which could be audible, visual, remote communication actions or local device control actions and see [0101]). As to claim 1 and 12, apparatus claims 1 and 12 and method claim 11 are related as apparatus and the method of using same, with each claimed element's function corresponding to the claimed method step. Accordingly claims 1 and 12 are similarly rejected under the same rationale as applied above with respect to method claim. As to claim 1, Gordon further teaches a storage unit (see [0083], buffer), and a controller (see [0083], smart speaker device). As to claim 12, Gordon further teaches an acquisition device (see [0082], microphone), audio processing apparatus (see [0083], smart speaker device). As to claim 6, Gordon teaches wherein the controller is configured to detect, as the utterance satisfying the set condition, audio whose feature matches the audio feature set in advance (see [0087], where sounds of a person speaking a code/token or emergency word/phrase is determined and where a registry of sounds patterns are stored for classifying). As to claim 7, Gordon teaches wherein the controller is configured to detect, as the utterance satisfying the set condition, a search word (see [0087], where sounds of a person speaking a code/token or emergency word/phrase is determined). As to claim 8, Gordon teaches wherein the notification condition includes a third condition (see [0096], where plural event models are noted as well as one or more matching rules are possible thus suggesting a 3rd or 4th condition and see [0098], where quantity of captured sounds are identified for comparing to the rules); and the controller is configured to, when the utterance satisfies the third condition, perform control (see [0099], where based on the event model being matched to then perform an action) with which, immediately after the search word is detected, notification sound is played (see [0100], [0101], where depending on the particular type of event, actions may be categorized into local audible message and an output message in an audible format indicating nature of the detected event is output and see [0087], where sounds of a person speaking a code/token or emergency word/phrase is determined) and playback of the utterance is started (see [0104], where system stores the history of captured sounds for later playback accessible to a user when desired). As to claim 9, Gordon teaches wherein the notification condition includes a fourth condition (see [0096], where plural event models are noted as well as one or more matching rules are possible thus suggesting a 3rd or 4th condition and see [0098], where quantity of captured sounds are identified for comparing to the rules); and the controller is configured to, when the utterance satisfies the fourth condition, perform control (see [0099], where based on the event model being matched to then perform an action) with which, immediately after the utterance ends, notification sound is played (see [0100], [0101], where depending on the particular type of event, actions may be categorized into local audible message and an output message in an audible format indicating nature of the detected event is output and see [0087], where sounds of a person speaking a code/token or emergency word/phrase is determined) and playback of the utterance is started (see [0104], where system stores the history of captured sounds for later playback accessible to a user when desired). As to claim 10, Gordon teaches the controller is configured to perform control with which audio data of an utterance interval including the utterance is played, and the utterance interval is an interval for which the audio data continues without interruption for a set time (see [0116], where stored audio samples are stored and replayed to the user for the duration of the recording and see [0128]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 2-3 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Gordon in view of JP (JP 5272383 B2). As to claim 2, Gordon teaches wherein the notification condition includes a first condition and a second condition (see [0096], where even models are applied to the stored information in the buffer which determines if criteria of the event model are satisfied to a threshold level and where plural rules are present which equate to conditions), and the controller is configured to: when the utterance satisfies the first condition, play notification sound to notify the user that the utterance audio satisfying the condition has been detected (see [0100], [0101], where depending on the particular type of event, actions may be categorized into local audile message and an output message in an audible format indicating nature of the detected event is output); when the utterance satisfies the second condition, present visual information to the user to notify the user that the audio satisfying the condition has been detected (see [0100], [0101], where depending on the particular type of event, actions may be categorized into visual message such as textual message and see [0094], where display of information is made). However, Gordon does not specifically teach a priority order of notifying the user is lower when the notification condition satisfies the second condition than when the notification condition satisfies the first condition. JP does teach a priority order of notifying the user is lower when the notification condition satisfies the second condition than when the notification condition satisfies the first condition (see page 2, last 3 lines going into page 3, where notification is performed using voice if high/medium importance level and message display (visual) performed when low importance). Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the notification as taught by Gordon with the priority based notification as taught by JP in order to allow the user to be able to recognize the importance of the abnormality (see JP 3rd page, 1st full paragraph). As to claim 3, Gordon in view of JP teach all of the limitations as in claim 2. Furthermore, Gordon teaches wherein the controller is configured to, when the utterance satisfies the first condition, play the notification sound (see [0100], [0101], where depending on the particular type of event, actions may be categorized into local audible message and an output message in an audible format indicating nature of the detected event is output) and then play the utterance satisfying the first condition (see [0104], where system stores the history of captured sounds for later playback accessible to a user when desired). As to claim 5, Gordon in view of JP teach all of the limitations as in claim 2. Furthermore, Gordon teaches wherein the controller is configured to, when the utterance satisfies the second condition, play the utterance satisfying the second condition in response to an input from the user (see [0051], where based on certain sounds such sounds are recorded and played back to a medical professional such as for coughing and see [0087], where emergency word/phrases considered)). Claim(s) 4 is rejected under 35 U.S.C. 103 as being unpatentable over Gordon in view of JP (JP 5272383 B2), as applied in claim 2 above, and further in view of Hampiholi et al. (US 2015/0266377). As to claim 4, Gordon in view of JP teach all of the limitations as in claim 2. Furthermore, Gordon teaches when the utterance satisfies the second condition, present to the user to present the visual information to the user, information of the utterance satisfying the second condition (see [0100], [0101], where depending on the particular type of event, actions may be categorized into visual message such as textual message and see [0094], where display of information is made). However, Gordon in view of JP do not specifically teach a presentation of a list of notifications. Hampiholi does teach present a list to the user to present the visual information to the user, the list being pieces of information of the utterance ([0056]-[0058], where messages are ordered in a list based on urgency and then presented visually or audibly). Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed inventions to have modified the notification as taught by Gordon in view of JP with the list as taught by Hampiholi in order to determine when to provide messages based on operating conditions or during safe/low risk situations (see Hampiholi [0003]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PARAS D SHAH whose telephone number is (571)270-1650. The examiner can normally be reached Monday-Thursday 7:30AM-2:30PM, 5PM-7PM (EST), Friday 8AM-noon (EST). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PARAS D SHAH can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Paras D Shah/Supervisory Patent Examiner, Art Unit 2653 06/28/2026
Read full office action

Prosecution Timeline

Jul 11, 2024
Application Filed
Jan 12, 2026
Non-Final Rejection mailed — §101, §102, §103
Apr 13, 2026
Response Filed
Jul 01, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12670920
Joint Acoustic Echo Cancellation (AEC) and Personalized Noise Suppression (PNS)
3y 4m to grant Granted Jun 30, 2026
Patent 12670924
SPATIAL REGION BASED AUDIO SEPARATION
2y 7m to grant Granted Jun 30, 2026
Patent 12633295
THREE-DIMENSIONAL AUDIO SIGNAL CODING METHOD AND APPARATUS, AND ENCODER
2y 6m to grant Granted May 19, 2026
Patent 12586591
SOUND SIGNAL DECODING METHOD, SOUND SIGNAL DECODER, PROGRAM, AND RECORDING MEDIUM
3y 3m to grant Granted Mar 24, 2026
Patent 12579367
TWO-TOWER NEURAL NETWORK FOR CONTENT-AUDIENCE RELATIONSHIP PREDICTION
2y 4m to grant Granted Mar 17, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
73%
Grant Probability
99%
With Interview (+31.2%)
3y 9m (~1y 7m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 654 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month