Prosecution Insights
Last updated: October 02, 2026
Application No. 18/244,029

SYSTEM AND METHOD TO IMPROVE PRECISION AND RECALL OF PROTOTYPICAL NETWORKS FOR SOUND EVENT DETECTION

Final Rejection §103
Filed
Sep 08, 2023
Examiner
LIEBGOTT, TYLER MICHAEL
Art Unit
2694
Tech Center
2600 — Communications
Assignee
Robert Bosch GmbH
OA Round
2 (Final)
66%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
63%
With Interview

Examiner Intelligence

Grants 66% — above average
66%
Career Allowance Rate
21 granted / 32 resolved
+3.6% vs TC avg
Minimal -3% lift
Without
With
+-3.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
15 currently pending
Career history
59
Total Applications
across all art units

Statute-Specific Performance

§101
0.8%
-39.2% vs TC avg
§103
50.0%
+10.0% vs TC avg
§102
28.5%
-11.5% vs TC avg
§112
17.9%
-22.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 32 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment In response to the non-final office action dated 05/06/2026, applicant has amended claims 1, 8, 15 and 16. Claims 1-20 are currently pending in the application. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-6, and 8-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nongpiur et al (US Pub No. 2022/0027725, hereinafter Nongpiur) in view of Lee et al (US Pub No. 2023/0169795, hereinafter Lee) and Chen et al (US Pub No. 2026/0259696, hereinafter Chen). Regarding claim 1, Nongpiur teaches a method of training a neural network for sound event detection (¶ [0075], training sound model for neural network), the method comprising: receiving samples of an audio signal (¶ [0075], sound clips), wherein the audio signal includes a first portion corresponding to a support set (¶ [0075], training set) and a second portion corresponding to a query set (¶ [0075], sound the sound model is trained to detect), and wherein the support set includes labeled samples (¶ [0075], positive examples having labels); determining, based on positive samples of the audio signal, respective positive prototypes of a plurality of classes of sound events, wherein the positive samples correspond to sound events (¶ [0078], positive examples corresponding to specific target sounds); constructing a negative support set using negative samples from the support set, wherein the negative samples do not correspond to sound events (¶ [0073], negative examples added to training data to include unwanted environmental sounds); determining, based on the negative samples of the negative support set, respective negative prototypes for respective groups of the negative samples, wherein each of the negative prototypes corresponds to a combination of a plurality of the negative samples (¶ [0078], negative examples corresponding to sounds other than target sound); and generating, based on comparisons between (i) a first sample (¶ [0074], environmental recording) and (ii) the respective positive prototypes (¶ [0078], positive examples corresponding to specific target sounds) and each of the negative prototypes (¶ [0078], negative examples corresponding to sounds other than target sound), an output signal that indicates whether a first sample belongs to one of the plurality of classes of sound events (¶ [0076], pre-trained sound model determining target sound within environment). Nongpiur does not explicitly teach a prototypical network or unlabeled samples. Lee teaches a prototypical network (See Lee ¶ [0052], prototypical network). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a prototypical network as taught by Lee with the method taught by Nongpiur. Prototypical networks offer several advantages such as accurate predictions with small training sets, flexible architecture, and human-like reasoning allowing for effective decision making with limited samples which are more easily understood by the user compared to other models. Nongpiur in view of Lee does noes explicitly teach the use of unlabeled samples. Chen teaches the use of unlabeled samples (See Chen ¶ [0088], semi-supervised training technique using a mix of labeled and unlabeled examples). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate unlabeled samples as taught by Chen with the method taught by Nongpiur in view of Lee. Doing so allows for the identification of hidden patterns which provides trends and insights from unstructured datasets as stated by Chen (¶ [0088]). Regarding claim 2, Nongpiur in view of Lee and Chen teaches the method of claim 1, further comprising: obtaining first embeddings of the positive samples and second embeddings of the negative samples (See Nongpiur ¶ [0071], positive and negative samples are categorized, an embedding is a learned vector representation of a category); determining the respective positive prototypes based on the first embeddings; and determining the respective negative prototypes based on the second embeddings (See Nongpiur ¶ [0078], positive and negative samples are categorized based on target sounds and unwanted environmental sounds). Regarding claim 3, Nongpiur in view of Lee and Chen teaches the method of claim 1, wherein determining the respective negative prototypes includes determining a negative prototype for each of the respective groups of the negative samples (See Nongpiur ¶ [0078], negative examples corresponding to sounds other than target sound). Regarding claim 4, Nongpiur in view of Lee and Chen teaches the method of claim 1, further comprising: determining, based on the comparisons, at least one probability that the first sample corresponds to the one of the plurality of classes of sound events; and generating the output based on the at least one probability (See Nongpiur fig 7, step 714 probability threshold). Regarding claim 5, Nongpiur in view of Lee and Chen teaches the method of claim 4, further comprising: determining at least one probability distribution; and generating the output based on the at least one probability distribution (See Nongpiur fig 7, step 702 preliminary label with probability them steps 706 and 714 normal and high probability thresholds). Regarding claim 6, Nongpiur in view of Lee and Chen teaches the method of claim 5, further comprising: determining at least one threshold based on the at least one probability distribution; and generating the output based on a comparison between the at least one probability and the at least one threshold (See Nongpiur fig 7, step 714 probability threshold). Regarding claim 8, Nongpiur teaches a computing device configured to train a neural network for sound event detection (¶ [0075], training sound model for neural network), the computing device including a processing device configured to execute instructions stored in memory (¶ [0057],) to: receive samples of an audio signal (¶ [0075], sound clips), wherein the audio signal includes a first portion corresponding to a support set (¶ [0075], training set) and a second portion corresponding to a query set (¶ [0075], sound the sound model is trained to detect), and wherein the support set includes labeled samples (¶ [0075], positive examples having labels); determine, based on positive samples of the audio signal, respective positive prototypes of a plurality of classes of sound events, wherein the positive samples correspond to sound events (¶ [0078], positive examples corresponding to specific target sounds); construct a negative support set using negative samples from the support set, wherein the negative samples do not correspond to sound events (¶ [0073], negative examples added to training data to include unwanted environmental sounds); determine, based on the negative samples of the negative support set, respective negative prototypes for respective groups of the negative samples, wherein each of the negative prototypes corresponds to a combination of a plurality of the negative samples (¶ [0078], negative examples corresponding to sounds other than target sound); and generate, based on comparisons between (i) a first sample (¶ [0074], environmental recording) and (ii) the respective positive prototypes (¶ [0078], positive examples corresponding to specific target sounds) and each of the negative prototypes (¶ [0078], negative examples corresponding to sounds other than target sound), an output signal that indicates whether the first sample belongs to one of the plurality of classes of sound events (¶ [0076], pre-trained sound model determining target sound within environment). Nongpiur does not explicitly teach a prototypical network or unlabeled samples. Lee teaches a prototypical network (See Lee ¶ [0052], prototypical network). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a prototypical network as taught by Lee with the computing device taught by Nongpiur. Prototypical networks offer several advantages such as accurate predictions with small training sets, flexible architecture, and human-like reasoning allowing for effective decision making with limited samples which are more easily understood by the user compared to other models. Nongpiur in view of Lee does noes explicitly teach the use of unlabeled samples. Chen teaches the use of unlabeled samples (See Chen ¶ [0088], semi-supervised training technique using a mix of labeled and unlabeled examples). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate unlabeled samples as taught by Chen with the method taught by Nongpiur in view of Lee. Doing so allows for the identification of hidden patterns which provides trends and insights from unstructured datasets as stated by Chen (¶ [0088]). Regarding claim 9, Nongpiur in view of Lee and Chen teaches the computing device of claim 8, wherein the processing device is further configured to execute the instructions to: obtain first embeddings of the positive samples and second embeddings of the negative samples (See Nongpiur ¶ [0071], positive and negative samples are categorized, an embedding is a learned vector representation of a category); determine the respective positive prototypes based on the first embeddings; and determine the respective negative prototypes based on the second embeddings (See Nongpiur ¶ [0078], positive and negative samples are categorized based on target sounds and unwanted environmental sounds). Regarding claim 10, Nongpiur in view of Lee and Chen teaches the computing device of claim 8, wherein determining the respective negative prototypes includes determining a negative prototype for each of the respective groups of the negative samples (See Nongpiur ¶ [0078], negative examples corresponding to sounds other than target sound). Regarding claim 11, Nongpiur in view of Lee and Chen teaches the computing device of claim 8, wherein the processing device is further configured to execute the instructions to: determine, based on the comparisons, at least one probability that the first sample corresponds to the one of the plurality of classes of sound events; and generate the output based on the at least one probability (See Nongpiur fig 7, step 714 probability threshold). Regarding claim 12, Nongpiur in view of Lee and Chen teaches the computing device of claim 11, wherein the processing device is further configured to execute the instructions to: determine at least one probability distribution; and generate the output based on the at least one probability distribution (See Nongpiur fig 7, step 702 preliminary label with probability them steps 706 and 714 normal and high probability thresholds). Regarding claim 13, Nongpiur in view of Lee and Chen teaches the computing device of claim 12, wherein the processing device is further configured to execute the instructions to: determine at least one threshold based on the at least one probability distribution; and generate the output based on a comparison between the at least one probability and the at least one threshold (See Nongpiur fig 7, step 714 probability threshold). Claim(s) 15-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nongpiur et al (US Pub No. 2022/0027725, hereinafter Nongpiur) in view of Chen et al (US Pub No. 2026/0259696, hereinafter Chen). Regarding claim 15, , Nongpiur teaches a computer-controlled machine (Fig 1, device 140), comprising: at least one sensor configured to generate an audio signal (Fig 1, sensor devices 100, 110, 120, and 140); a control system configured to receive a first sample of the audio signal (Fig 1, machine learning system 145), and generate, based on comparisons between the (i) the first sample (¶ [0074], environmental recording) and (ii) respective positive prototypes for each of a plurality of classes of sound events (¶ [0078], positive examples corresponding to specific target sounds) and respective negative prototypes for each of a plurality of groups of negative prototypes (¶ [0078], negative examples corresponding to sounds other than target sound), an output signal that indicates whether the first sample of the audio signal belongs to one of the plurality of classes of sound events (¶ [0076], pre-trained sound model determining target sound within environment); and an actuator configured to control an operation of the computer-controlled machine in response to the output of the control system (Fig 1, decision unit 114 control driving elements), wherein the respective positive prototypes correspond to a plurality of positive samples (¶ [0078], positive examples corresponding to specific target sounds), and each of the respective negative prototypes corresponds to a combination of a plurality of negative samples (¶ [0073], negative examples added to training data to include unwanted environmental sounds) obtained from a support set of audio samples (¶ [0075], training set) including labeled samples (¶ [0075], positive examples having labels). Nongpiur does not explicitly teach unlabeled samples. Chen teaches the use of unlabeled samples (See Chen ¶ [0088], semi-supervised training technique using a mix of labeled and unlabeled examples). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate unlabeled samples as taught by Chen with the computer-controlled machine taught by Nongpiur. Doing so allows for the identification of hidden patterns which provides trends and insights from unstructured datasets as stated by Chen (¶ [0088]). Regarding claim 16, Nongpiur in view of Chen teaches the computer-controlled machine of claim 15, further comprising memory that stores the respective positive prototypes and the respective negative prototypes (See Nongpiur fig 3, storage 147). Regarding claim 17, Nongpiur in view of Chen teaches the computer-controlled machine of claim 15, wherein generating the output includes calculating a probability that the first sample belongs to a first class of the plurality of classes of sound events or a first group of the plurality of groups (See Nongpiur fig 7, step 702 preliminary label with probability). Regarding claim 18, Nongpiur in view of Chen teaches the computer-controlled machine of claim 17, wherein generating the output includes comparing the probability to at least one threshold and generating the output based on the comparison (See Nongpiur fig 7, step 714 probability threshold). Regarding claim 19, Nongpiur in view of Chen teaches the computer-controlled machine of claim 18, wherein the at least one threshold includes a plurality of thresholds corresponding to respective classes of the plurality of classes of sound events (See Nongpiur fig 7, step 702 preliminary label with probability then steps 706 and 714 normal and high probability thresholds). Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nongpiur et al (US Pub No. 2022/0027725, hereinafter Nongpiur) in view of Chen et al (US Pub No. 2026/0259696, hereinafter Chen) as applied to the claims above, and further in view of Akotkar et al (US Pub No. 2019/0049989, hereinafter Akotkar). Regarding claim 20, Nongpiur in view of Chen teaches the computer-controlled machine of claim 15. Nongpiur in view of Chen does not explicitly teach an autonomous robot. Akotkar teaches an autonomous robot (Fig 1, autonomous driving vehicle 102). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the autonomous robot taught by Akotkar with the computer-controlled machine taught by Nongpiur in view of Chen. Doing so allows for environmental monitoring and response of autonomous robots to dangerous or unforeseen incidents with little to no human input as stated by Akotkar (¶ [0003]). Allowable Subject Matter Claims 7 and 14 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Response to Arguments Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER LIEBGOTT whose telephone number is (703)756-1818. The examiner can normally be reached Mon-Fri 10-6:30 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn Edwards can be reached at (571)270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /T.M.L./Examiner, Art Unit 2694 /CAROLYN R EDWARDS/Supervisory Patent Examiner, Art Unit 2692
Read full office action

Prosecution Timeline

Sep 08, 2023
Application Filed
May 06, 2026
Non-Final Rejection mailed — §103
Aug 04, 2026
Response Filed
Sep 16, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750633
Minimizing Echo Caused by Stereo Audio Via Position-Sensitive Acoustic Echo Cancellation
3y 5m to grant Granted Sep 29, 2026
Patent 12737417
User Interfaces for Sound Engineering Application on Touch Device
3y 4m to grant Granted Sep 15, 2026
Patent 12696023
MICROPHONE ADJUSTMENT APPARATUS AND RECORDING STRUCTURE
3y 0m to grant Granted Jul 28, 2026
Patent 12688934
WASTE IDENTIFICATION METHOD, WASTE IDENTIFICATION DEVICE, AND WASTE IDENTIFICATION PROGRAM
3y 10m to grant Granted Jul 21, 2026
Patent 12684273
DIPOLE LOUDSPEAKER ASSEMBLY
3y 5m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
66%
Grant Probability
63%
With Interview (-3.0%)
2y 10m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 32 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month