Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claim Objections
Applicant’s cancellation of the stray enumeration of claim 14 sufficed to obviate the objection.
Claim Rejections - 35 USC § 112
Applicant’s amendments to claims 5, 10, 20 suffice to obviate the rejection under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-9, 11, 14-20 rejected under 35 U.S.C. 103 as being unpatentable over Wang: 10418957 further in view of Perl: 20230215460.
Regarding claim 1
Wang teaches:
A computer-implemented method for event detection in time-series data, wherein the method uses a processor coupled with a memory configured to store instructions implementing the method (Wang: Figs 12, 13), wherein the instructions, when executed by the processor carry out steps of the method, comprising:
processing the time-series data (Wang: Col 3:63-4:14; Fig 1B: system receives audio data, samples and subsamples same into data with a coarser and/or finer time scale generates scores based thereon) to make a hard decision on a time span of an event indicative of continuous activity of the event within the time-series data (Wang: Col 3:29-3:42; 11:35-11:60; Fig 1A, 1B, 6: region proposal network (RPN) processes the time series data to output begin and end times of a determined interval of an audio event)
and make a soft decision on a presence of the event for the entire time span (Wang: Col 3:32-3:62; 5:20-5:38: generation of a likelihood score which “indicates how likely an audio event is detected during that interval of time,” applied to each of a plurality of frames, times therein),
wherein making the soft decision comprises computing a composite confidence score by aggregating frame-level confidence scores, and wherein the frame-level confidence scores are derived based on a probability of presence of the event in each frame of the entire time span (Wang: Col 8:23-8:67; weighted frames summed together to generate or otherwise determine a composite frame and the frame scores similarly summed to thereby aggregate, generate, etc. a composite, aggregated, etc. confidence score based on the frame level scores and corresponding overall to a frame index);
applying an event level threshold to the soft decision on the presence of the event for the entire time span to produce a result of the event detection (Wang: Col 3:32-3:62; 5:20-5:38, 11:35-12:8, 12:58-12:67: threshold applied to each frame, each of likelihood scores therein to determine and output a result; such as for filtering the overall score to reify the detection of an event by comparing the incoming scores to the threshold and operates to remove candidates not satisfying the threshold); and
outputting the result of the event detection (Wang: Col 3:32-3:62; 5:20-5:38: threshold applied to each frame, each of likelihood scores therein to determine and output a result).
Wang does not explicitly teach the aggregating of the frame level confidence scores or probabilities themselves but re-scoring a score weighted feature composite, that is Wang generates a composite from frame vectors weighted by frame scores and summing the frame scores into a cumulative score operable to re-score the composite.
In a related field of endeavor Perl teaches a system and method for the performance of audio event detection by computing window confidence values over consecutive segments of audio by aggregating the per segment probability scores to determine presence of a particular class of audio event in consecutive segments (Perl: Abstract; ¶ 2, 33-35); wherein the system computes frame level scores from per frame probabilities of the presence of a class (Perl: ¶ 2, 29-35; Figs 3A, 3B) and aggregates over a time span of consecutive segments which exceed a threshold (Perl: ¶ 2, 57, 67, 74) to thereby compute a composite confidence score based on aggregating frame level scores (Perl: ¶ 35, 57, 74).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to utilize the event class thresholding and frame level thresholding to generate frame level confidence scores and a confidence score composited based on event level thresholds as taught or suggested by Perl to thereby reify the soft decision of Wang for at least the purpose of determining the presence of an event and/or event class in each and all of the consecutive frames which represent the presence of the sound event and for at least the purpose of improving the decisioning of Wang by polyphonic detection; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 2
Wang in view of Perl teaches or suggests:
The method of claim 1, wherein the soft decision is an event bounding box that comprises the time span of the event (Wang: Col 10:28-10:50, 10:58-11:15, 11:28-11:63; Fig 8, 9: such as by determining with respect to particular start and end times of an event a bounding box, such as different feature matrices with particular feature size and stride and/or to segment window of particular lengths into a determined number of segments); a type of the event (Wang: Col 13:20-13:27; Fig 6: such as by detecting a particular event of a particular type); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as particular sound event classes), and a confidence score for the presence of the event (Wang: Abstract, etc.: such as a likelihood score that individual, pluralities, etc. of frames correspond to a sound event); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B). The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 3
Wang in view of Perl teaches or suggests:
The method of claim 2, wherein the time-series data comprises an audio stream, wherein the event is a sound event (Wang: Abstract); (Perl: Abstract), and wherein the soft decision is a sound event bounding box that comprises a sound class of the sound event as the type of the event (Wang: Col 10:28-10:50, 10:58-11:15, 11:28-11:63; Fig 8, 9: such as by determining with respect to particular start and end times of an event a bounding box, such as different feature matrices with particular feature size, and/or stride; and/or to segment windows of determined lengths into a determined number of sub-segments); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: determines event boundaries), the time span as an extent of the sound event in the audio stream (Wang: Col 13:20-13:27; Fig 6: such as by detecting a particular event of a particular type); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as particular sound event classes), and an overall confidence score indicating the probability of presence of the sound class in the sound event (Wang: Abstract, etc.: such as a likelihood score that individual, pluralities, etc. of frames correspond to a sound event); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as by determining sound class activity prediction scores, thresholding same). The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 4
Wang in view of Perl teaches or suggests:
The method of claim 3, further comprising controlling a machine based on the result of the event detection (Wang: Col 5:32-5:39, 9:13-9:17; Fig 1A, 1B, etc.: such as causing an action to be performed based thereon). While Wang in view of Perl does not explicitly discusses “controlling a machine,” merely causing an action, turning on data logging, or causing a command to be executed this is considered substantially similar as output commands for causing an action such as the change of state of a switch may additionally cause a processor, device, or machine to awaken, sleep, change power modes, change state, etc. further Examiner has taken official notice which Applicant has failed to timely and explicitly traverse and it is thus accepted as Admitted Prior Art (APA: please see MPEP 2144.03) that controlling a machine as a result of an output command would have comprised an obvious inclusion. Thus the claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 5
Wang in view of Perl teaches or suggests:
The method of claim 3, further comprising identifying a source of a sound associated with the sound event, based on the result of the event detection (Wang: Col 10:28-10:50, 10:58-11:15, 11:28-11:63; 14:63-15:2; Fig 8, 9: such as by determining with respect to particular start and end times of an event a bounding box, such as different feature matrices with particular feature size and stride and/or to segment window of particular lengths into a determined number of segments). The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 6
Wang in view of Perl teaches or suggests:
The method of claim 3, wherein making the hard decision on the time span comprises identifying frames in the audio stream that belong to the sound class (Wang: Col 9:19-9:25, 10:28-10:50, 10:58-11:15, 11:28-11:63; Fig 8, 9: such as by determining with respect to particular start and end times of an event a bounding box, such as different feature matrices with particular feature size and stride and/or to segment window of particular lengths into a determined number of segments). The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 7
Wang in view of Perl teaches or suggests:
The method of claim 6, wherein identifying the frames that belong to the sound class comprises:
computing a probability of presence of the sound class in each frame of the audio stream; and applying a class threshold to the probability of presence of the sound class in each frame to filter a plurality of frames whose probability of presence exceeds the class threshold (Wang: Col 8:23-8:67, 9:18-10:5, 10:28-10:50, 10:58-11:15, 11:28-11:63; Fig 8, 9: such as by determining with respect to particular start and end times of an event a bounding box, such as different feature matrices with particular feature size and stride and/or to segment window of particular lengths into a determined number of segments); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as by resolving a subset of classification scores that exceed a threshold for each class). The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 8
Wang in view of Perl teaches or suggests:
The method of claim 7, wherein making the hard decision on the time span further comprises determining a start time instance and an end time instance of the time span based on the plurality of filtered frames (Wang: Col 3:33-3:45, 9:18-10:7, 12:11-12:25, 12:55-12:67: filter manages overlapping front and/or back ends of time windows in such a way as to adjust start and end time instances such that the decisions are conducted on the best fit time windows);. The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 9
Wang in view of Perl teaches or suggests:
The method of claim 8, wherein a time of occurrence of a sequentially first frame of the plurality of filtered frames is selected as the start time instance of the time span and a time of occurrence of a sequentially last frame of the plurality of filtered frames is selected as the end time instance of the time span (Wang: Col 9:18-10:7, 12:11-12:25, 12:55-12:67: filter manages overlapping front and/or back ends of time windows in such a way as to adjust start and end time instances such that the decisions are conducted on the best fit time windows). The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 11
Wang in view of Perl teaches or suggests:
The method of claim 1, wherein the computing includes one or more of obtaining the average, obtaining the maximum, obtaining the minimum, or obtaining the median of the probability of presence of the sound class in each frame of the entire time span (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as by determining sound class activity prediction scores, thresholding and smoothing same, wherein the smoothing of a prediction score corresponds to a smoothing function such as a mean, average, max , etc.). While Wang in view of Perl does not explicitly discuss selections among machine learning processes of averaging, maximizing, minimizing, taking a median, etc. Examiner has taken official notice which Applicant has failed to timely and explicitly traverse and it is thus accepted as Admitted Prior Art (APA: please see MPEP 2144.03) that such processes were well-known in the art before the effective filing date of the instant application and would have comprised an obvious inclusion such as for applying user desired machine learning processes to parameters pre or post processed based thereon. Thus the claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 14—the claim is considered to recite substantially similar subject matter to claim 1 under 35 USC 103 as discussed supra and is similarly rejected
Regarding claim 15
Wang in view of Perl teaches or suggests:
The method of claim 14, wherein the time-series data comprises an audio stream (Wang: Abstract); (Perl: Abstract), and wherein the soft decision comprises a sound class of a detected sound event in the audio stream (Wang: Col 10:28-10:50, 10:58-11:15, 11:28-11:63; Fig 8, 9: such as by determining with respect to particular start and end times of an event a bounding box, such as different feature matrices with particular feature size and stride and/or to segment window of particular lengths into a determined number of segments); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: computes frame wise and composite class likelihoods) and the time span as an extent of the detected sound event in the audio stream (Wang: Col 13:20-13:27; Fig 6: such as by detecting a particular event of a particular type); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as over extents illustrated in the figures). The claim is considered obvious over Wang as modified by Perl as addressed in the base claim as it would have been obvious to apply the further teaching of Wang and/or Perl to the modified device of Wang and Perl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 16—the claim is considered to recite substantially similar subject matter to claim 4 under 35 USC 103 as discussed supra and is similarly rejected
Regarding claim 17—the claim is considered to recite substantially similar subject matter to claim 5 under 35 USC 103 as discussed supra and is similarly rejected
Regarding claim 18—the claim is considered to recite substantially similar subject matter to claim 6 under 35 USC 103 as discussed supra and is similarly rejected
Regarding claim 19—the claim is considered to recite substantially similar subject matter to claim 7 under 35 USC 103 as discussed supra and is similarly rejected
Regarding claim 20—the claim is considered to recite substantially similar subject matter to claim 8 under 35 USC 103 as discussed supra and is similarly rejected
Claims 12 rejected under 35 U.S.C. 103 as being unpatentable over Wang: 10418957 further in view of Perl: 20230215460 as applied to claims 1-9, 11, 14-20 supra and further in view of Ema: 20210405909.
Regarding claim 12
Wang in view of Perl teaches or suggests:
The method of claim 3, further comprising:
computing for each frame of the audio stream, a class presence confidence score as a probability of presence of the sound class in each frame of the audio stream (Wang: Col 13:20-13:27; Fig 6: such as by detecting a particular event of a particular type; such as by a classifier); (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as particular sound event classes)
determining tentative onset times and tentative offset times of tentative events; and processing the tentative onset times and tentative offset times of tentative events to obtain time spans of sound events (Wang: Col 10:28-10:50, 10:58-11:15, 11:28-11:63; Fig 8, 9: such as by determining with respect to particular start and end times of an event different feature matrices with particular feature size, and/or stride; and/or to segment windows of determined lengths into a determined number of sub-segments) ; (Perl: ¶ 2, 29-35, 57, 67, 74; Figs 3A, 3B: such as over extents illustrated in the figures).
.
Wang in view of Perl does not explicitly teach filtering the class presence confidence scores with an ideal step filter in continuous time to determine a delta score for each frame as a difference between the average of class presence confidence scores in a predefined time period after a respective frame and the average of class presence confidence scores in the same- length time period before the respective frame; utilizing delta scores to process decision outputs.
In a related field of endeavor Ema teaches a system and method for management of data by the generation of differential data comprising filtering data using an ideal step filter (Ema: ¶ 66: a step value of a moving average or boxcar filter is optimized); in continuous time (Ema: ¶ 37, etc.: data continuously acquired in a time series from sensors; such as utilizing an embodiment processed over a time series); to determine a delta or differential score(s) for each frame (Ema; ¶ 66-69: thereby generating differential data acquired from maximum and minimum values by subtracting the filtered average).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to utilize the delta scores taught or suggested by Ema to better optimize change point detection in by taking a difference among averages of class-wise confidence scores in the manner taught or suggested by Wang in view of Perl for at least the point of utilizing change point detection of hard decision onset and offset times to manage, process, and improve the detected time spans of sound classes; one of ordinary skill in the art would have expected only predictable results therefrom.
Allowable Subject Matter
Claim 13 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Response to Arguments
Applicant’s arguments with respect to claim amendments, see Remarks and Claims filed 3/13/26, with respect to the rejection(s) of claim(s) 1-12, 14-20 under 35 USC 103 over Wang in view of Cakir; and Wang in view of Cakir in view of Ema have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Wang in view of Perl; and Wang in view of Perl in view of Ema.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL C MCCORD whose telephone number is (571)270-3701. The examiner can normally be reached 730-630 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN EDWARDS can be reached at (571) 270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAUL C MCCORD/Primary Examiner, Art Unit 2692