DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In response to the non-final office action dated 05/06/2026, applicant has amended claims 1, 8, 15 and 16. Claims 1-20 are currently pending in the application.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6, and 8-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nongpiur et al (US Pub No. 2022/0027725, hereinafter Nongpiur) in view of Lee et al (US Pub No. 2023/0169795, hereinafter Lee) and Chen et al (US Pub No. 2026/0259696, hereinafter Chen).
Regarding claim 1, Nongpiur teaches a method of training a neural network for sound event detection (¶ [0075], training sound model for neural network), the method comprising: receiving samples of an audio signal (¶ [0075], sound clips), wherein the audio signal includes a first portion corresponding to a support set (¶ [0075], training set) and a second portion corresponding to a query set (¶ [0075], sound the sound model is trained to detect), and wherein the support set includes labeled samples (¶ [0075], positive examples having labels); determining, based on positive samples of the audio signal, respective positive prototypes of a plurality of classes of sound events, wherein the positive samples correspond to sound events (¶ [0078], positive examples corresponding to specific target sounds); constructing a negative support set using negative samples from the support set, wherein the negative samples do not correspond to sound events (¶ [0073], negative examples added to training data to include unwanted environmental sounds); determining, based on the negative samples of the negative support set, respective negative prototypes for respective groups of the negative samples, wherein each of the negative prototypes corresponds to a combination of a plurality of the negative samples (¶ [0078], negative examples corresponding to sounds other than target sound); and generating, based on comparisons between (i) a first sample (¶ [0074], environmental recording) and (ii) the respective positive prototypes (¶ [0078], positive examples corresponding to specific target sounds) and each of the negative prototypes (¶ [0078], negative examples corresponding to sounds other than target sound), an output signal that indicates whether a first sample belongs to one of the plurality of classes of sound events (¶ [0076], pre-trained sound model determining target sound within environment).
Nongpiur does not explicitly teach a prototypical network or unlabeled samples.
Lee teaches a prototypical network (See Lee ¶ [0052], prototypical network).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a prototypical network as taught by Lee with the method taught by Nongpiur. Prototypical networks offer several advantages such as accurate predictions with small training sets, flexible architecture, and human-like reasoning allowing for effective decision making with limited samples which are more easily understood by the user compared to other models.
Nongpiur in view of Lee does noes explicitly teach the use of unlabeled samples.
Chen teaches the use of unlabeled samples (See Chen ¶ [0088], semi-supervised training technique using a mix of labeled and unlabeled examples).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate unlabeled samples as taught by Chen with the method taught by Nongpiur in view of Lee. Doing so allows for the identification of hidden patterns which provides trends and insights from unstructured datasets as stated by Chen (¶ [0088]).
Regarding claim 2, Nongpiur in view of Lee and Chen teaches the method of claim 1, further comprising: obtaining first embeddings of the positive samples and second embeddings of the negative samples (See Nongpiur ¶ [0071], positive and negative samples are categorized, an embedding is a learned vector representation of a category); determining the respective positive prototypes based on the first embeddings; and determining the respective negative prototypes based on the second embeddings (See Nongpiur ¶ [0078], positive and negative samples are categorized based on target sounds and unwanted environmental sounds).
Regarding claim 3, Nongpiur in view of Lee and Chen teaches the method of claim 1, wherein determining the respective negative prototypes includes determining a negative prototype for each of the respective groups of the negative samples (See Nongpiur ¶ [0078], negative examples corresponding to sounds other than target sound).
Regarding claim 4, Nongpiur in view of Lee and Chen teaches the method of claim 1, further comprising: determining, based on the comparisons, at least one probability that the first sample corresponds to the one of the plurality of classes of sound events; and generating the output based on the at least one probability (See Nongpiur fig 7, step 714 probability threshold).
Regarding claim 5, Nongpiur in view of Lee and Chen teaches the method of claim 4, further comprising: determining at least one probability distribution; and generating the output based on the at least one probability distribution (See Nongpiur fig 7, step 702 preliminary label with probability them steps 706 and 714 normal and high probability thresholds).
Regarding claim 6, Nongpiur in view of Lee and Chen teaches the method of claim 5, further comprising: determining at least one threshold based on the at least one probability distribution; and generating the output based on a comparison between the at least one probability and the at least one threshold (See Nongpiur fig 7, step 714 probability threshold).
Regarding claim 8, Nongpiur teaches a computing device configured to train a neural network for sound event detection (¶ [0075], training sound model for neural network), the computing device including a processing device configured to execute instructions stored in memory (¶ [0057],) to: receive samples of an audio signal (¶ [0075], sound clips), wherein the audio signal includes a first portion corresponding to a support set (¶ [0075], training set) and a second portion corresponding to a query set (¶ [0075], sound the sound model is trained to detect), and wherein the support set includes labeled samples (¶ [0075], positive examples having labels); determine, based on positive samples of the audio signal, respective positive prototypes of a plurality of classes of sound events, wherein the positive samples correspond to sound events (¶ [0078], positive examples corresponding to specific target sounds); construct a negative support set using negative samples from the support set, wherein the negative samples do not correspond to sound events (¶ [0073], negative examples added to training data to include unwanted environmental sounds); determine, based on the negative samples of the negative support set, respective negative prototypes for respective groups of the negative samples, wherein each of the negative prototypes corresponds to a combination of a plurality of the negative samples (¶ [0078], negative examples corresponding to sounds other than target sound); and generate, based on comparisons between (i) a first sample (¶ [0074], environmental recording) and (ii) the respective positive prototypes (¶ [0078], positive examples corresponding to specific target sounds) and each of the negative prototypes (¶ [0078], negative examples corresponding to sounds other than target sound), an output signal that indicates whether the first sample belongs to one of the plurality of classes of sound events (¶ [0076], pre-trained sound model determining target sound within environment).
Nongpiur does not explicitly teach a prototypical network or unlabeled samples.
Lee teaches a prototypical network (See Lee ¶ [0052], prototypical network).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a prototypical network as taught by Lee with the computing device taught by Nongpiur. Prototypical networks offer several advantages such as accurate predictions with small training sets, flexible architecture, and human-like reasoning allowing for effective decision making with limited samples which are more easily understood by the user compared to other models.
Nongpiur in view of Lee does noes explicitly teach the use of unlabeled samples.
Chen teaches the use of unlabeled samples (See Chen ¶ [0088], semi-supervised training technique using a mix of labeled and unlabeled examples).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate unlabeled samples as taught by Chen with the method taught by Nongpiur in view of Lee. Doing so allows for the identification of hidden patterns which provides trends and insights from unstructured datasets as stated by Chen (¶ [0088]).
Regarding claim 9, Nongpiur in view of Lee and Chen teaches the computing device of claim 8, wherein the processing device is further configured to execute the instructions to: obtain first embeddings of the positive samples and second embeddings of the negative samples (See Nongpiur ¶ [0071], positive and negative samples are categorized, an embedding is a learned vector representation of a category); determine the respective positive prototypes based on the first embeddings; and determine the respective negative prototypes based on the second embeddings (See Nongpiur ¶ [0078], positive and negative samples are categorized based on target sounds and unwanted environmental sounds).
Regarding claim 10, Nongpiur in view of Lee and Chen teaches the computing device of claim 8, wherein determining the respective negative prototypes includes determining a negative prototype for each of the respective groups of the negative samples (See Nongpiur ¶ [0078], negative examples corresponding to sounds other than target sound).
Regarding claim 11, Nongpiur in view of Lee and Chen teaches the computing device of claim 8, wherein the processing device is further configured to execute the instructions to: determine, based on the comparisons, at least one probability that the first sample corresponds to the one of the plurality of classes of sound events; and generate the output based on the at least one probability (See Nongpiur fig 7, step 714 probability threshold).
Regarding claim 12, Nongpiur in view of Lee and Chen teaches the computing device of claim 11, wherein the processing device is further configured to execute the instructions to: determine at least one probability distribution; and generate the output based on the at least one probability distribution (See Nongpiur fig 7, step 702 preliminary label with probability them steps 706 and 714 normal and high probability thresholds).
Regarding claim 13, Nongpiur in view of Lee and Chen teaches the computing device of claim 12, wherein the processing device is further configured to execute the instructions to: determine at least one threshold based on the at least one probability distribution; and generate the output based on a comparison between the at least one probability and the at least one threshold (See Nongpiur fig 7, step 714 probability threshold).
Claim(s) 15-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nongpiur et al (US Pub No. 2022/0027725, hereinafter Nongpiur) in view of Chen et al (US Pub No. 2026/0259696, hereinafter Chen).
Regarding claim 15, , Nongpiur teaches a computer-controlled machine (Fig 1, device 140), comprising: at least one sensor configured to generate an audio signal (Fig 1, sensor devices 100, 110, 120, and 140); a control system configured to receive a first sample of the audio signal (Fig 1, machine learning system 145), and generate, based on comparisons between the (i) the first sample (¶ [0074], environmental recording) and (ii) respective positive prototypes for each of a plurality of classes of sound events (¶ [0078], positive examples corresponding to specific target sounds) and respective negative prototypes for each of a plurality of groups of negative prototypes (¶ [0078], negative examples corresponding to sounds other than target sound), an output signal that indicates whether the first sample of the audio signal belongs to one of the plurality of classes of sound events (¶ [0076], pre-trained sound model determining target sound within environment); and an actuator configured to control an operation of the computer-controlled machine in response to the output of the control system (Fig 1, decision unit 114 control driving elements), wherein the respective positive prototypes correspond to a plurality of positive samples (¶ [0078], positive examples corresponding to specific target sounds), and each of the respective negative prototypes corresponds to a combination of a plurality of negative samples (¶ [0073], negative examples added to training data to include unwanted environmental sounds) obtained from a support set of audio samples (¶ [0075], training set) including labeled samples (¶ [0075], positive examples having labels).
Nongpiur does not explicitly teach unlabeled samples.
Chen teaches the use of unlabeled samples (See Chen ¶ [0088], semi-supervised training technique using a mix of labeled and unlabeled examples).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate unlabeled samples as taught by Chen with the computer-controlled machine taught by Nongpiur. Doing so allows for the identification of hidden patterns which provides trends and insights from unstructured datasets as stated by Chen (¶ [0088]).
Regarding claim 16, Nongpiur in view of Chen teaches the computer-controlled machine of claim 15, further comprising memory that stores the respective positive prototypes and the respective negative prototypes (See Nongpiur fig 3, storage 147).
Regarding claim 17, Nongpiur in view of Chen teaches the computer-controlled machine of claim 15, wherein generating the output includes calculating a probability that the first sample belongs to a first class of the plurality of classes of sound events or a first group of the plurality of groups (See Nongpiur fig 7, step 702 preliminary label with probability).
Regarding claim 18, Nongpiur in view of Chen teaches the computer-controlled machine of claim 17, wherein generating the output includes comparing the probability to at least one threshold and generating the output based on the comparison (See Nongpiur fig 7, step 714 probability threshold).
Regarding claim 19, Nongpiur in view of Chen teaches the computer-controlled machine of claim 18, wherein the at least one threshold includes a plurality of thresholds corresponding to respective classes of the plurality of classes of sound events (See Nongpiur fig 7, step 702 preliminary label with probability then steps 706 and 714 normal and high probability thresholds).
Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nongpiur et al (US Pub No. 2022/0027725, hereinafter Nongpiur) in view of Chen et al (US Pub No. 2026/0259696, hereinafter Chen) as applied to the claims above, and further in view of Akotkar et al (US Pub No. 2019/0049989, hereinafter Akotkar).
Regarding claim 20, Nongpiur in view of Chen teaches the computer-controlled machine of claim 15.
Nongpiur in view of Chen does not explicitly teach an autonomous robot.
Akotkar teaches an autonomous robot (Fig 1, autonomous driving vehicle 102).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the autonomous robot taught by Akotkar with the computer-controlled machine taught by Nongpiur in view of Chen. Doing so allows for environmental monitoring and response of autonomous robots to dangerous or unforeseen incidents with little to no human input as stated by Akotkar (¶ [0003]).
Allowable Subject Matter
Claims 7 and 14 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER LIEBGOTT whose telephone number is (703)756-1818. The examiner can normally be reached Mon-Fri 10-6:30 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn Edwards can be reached at (571)270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/T.M.L./Examiner, Art Unit 2694
/CAROLYN R EDWARDS/Supervisory Patent Examiner, Art Unit 2692