Prosecution Insights
Last updated: October 01, 2026
Application No. 19/018,115

METHODS AND APPARATUS TO PERFORM SIGNATURE MATCHING USING NOISE CANCELLATION MODELS TO ACHIEVE CONSENSUS

Non-Final OA §103
Filed
Jan 13, 2025
Priority
Aug 14, 2020 — continuation of 11/501,793 +1 more
Examiner
HASSAN, ALI MOHAMAD
Art Unit
Tech Center
Assignee
The Nielsen Company (US) LLC
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
11 granted / 16 resolved
+8.8% vs TC avg
Strong +38% interview lift
Without
With
+37.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
17 currently pending
Career history
35
Total Applications
across all art units

Statute-Specific Performance

§101
28.7%
-11.3% vs TC avg
§103
50.3%
+10.3% vs TC avg
§102
15.6%
-24.4% vs TC avg
§112
3.0%
-37.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 16 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1-20 are pending. Claims 1, 8 and 15 are independent. This Application was published as US 20250149059. Apparent priority: 14 August 2020. This Application is a continuation of US 17/986586 issued as US 12198717 which is in turn a continuation of US 16/994095 issued as US 11501793. Obviousness Double Patenting is not currently found. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 5-8, 12-15 and 19-20 are rejected under 35 U.S.C. 103 as obvious over Ye (US 20160275588) in view of Kupryjanow (US 20200184987) Claim 1 8, and 15 Regarding Claim 1, Ye teaches: 1. A media meter device comprising: at least one microphone; (Figure 3, “[0060] FIG. 3 is a block diagram illustrating a representative client device 104 associated with a social networking user account in accordance with some embodiments. …Client device 104 also includes a user interface 310. … User interface 310 also includes one or more input devices 314, including user interface components that facilitate user input such as a keyboard, a mouse, a voice-command input unit or microphone, a touch screen display, a touch-sensitive input pad, a camera, a gesture capturing camera, or other input buttons or controls. Furthermore, some client devices 104 use a microphone and voice recognition or a camera and gesture recognition to supplement or replace the keyboard.”) a processor; (Figure 3, “[0060] … Client device 104, typically, includes one or more processing units (CPUs) 302, …”) memory having stored thereon computer readable instructions that, when executed by the processor, cause the media meter device to perform a set of operations including: (Figure 3, “[0060] … Client device 104, typically, includes one or more processing units (CPUs) 302, one or more network interfaces 304, memory 306, and one or more communication buses 308 for interconnecting these components (sometimes called a chipset)….”) PNG media_image1.png 613 489 media_image1.png Greyscale (“[0214] FIGS. 22A-22D are flowchart diagrams illustrating a method 2200 of obtaining information based on an audio input in accordance with some embodiments. In some embodiments, method 2200 is performed by electronic device 104 with one or more processors and memory. For example, in some embodiments, method 2200 is performed by electronic device 104 (FIGS. 1A and 3) or a component thereof. In some embodiments, method 2200 is governed by instructions that are stored in a non-transitory computer readable storage medium and the instructions are executed by one or more processors of the electronic device. Optional operations are indicated by dashed lines (e.g., boxes with dashed-line borders).”) collecting raw ambient audio data using the at least one microphone; PNG media_image2.png 522 664 media_image2.png Greyscale PNG media_image3.png 543 748 media_image3.png Greyscale (Figure 1B and Figure 20 showing the ambient audio: “[0045] Referring to FIG. 1B, a network environment includes an electronic device (e.g., electronic device 104, FIG. 1A) connected with a server (e.g., server system 108, FIG. 1A) through network (e.g., network 110, FIG. 1A). In some embodiments, the electronic device is within the vicinity of the television (TV), so that the electronic device can receive sound broadcast by the TV. …” Figure 4, 402, 404: “[0079] When a triggering event is detected, the electronic device collects an environment sound through an audio input collector of the electronic device, thereby collecting the program sound of the current channel broadcast by the television in real time. In some embodiments, collection of the sound of the current channel broadcast in real time in the environment may be started and timing is performed from a moment when a triggering event is detected. …”) (“[0207] When a user watches a television program of a channel at home and sees an advertisement that interests the user, the user immediately shakes a mobile phone. The mobile phone senses the motion of the mobile phone using a motion sensor of the mobile phone, and starts a microphone to record an environment sound to obtain PCM audio which lasts for 5-15 seconds, whose sampling frequency is 8 kHz, and which is quantized with 16 bits, where the recorded audio is audio data. Then, the mobile phone obtains audio feature information by performing feature extraction from the audio data; generates, according to the audio feature information, an audio fingerprint including a collection timestamp; and sends the audio fingerprint to the server in a communication manner such as Wifi (a technology enables wireless network access through a radio wave), 2G (a second-generation mobile communications technology), 3G (a third-generation mobile communications technology), or 4G (a fourth-generation mobile communications technology).”) processing the raw ambient audio data with a noise reduction model to thereby generate modified audio data; (Figure 21A and Figure 22A, 2202: “[0216] Alternatively, in some embodiments, by initiating a “detect sound nearby” function of the social networking application, the microphone of the electronic device starts to capture and generate an audio input as soon as sounds (e.g., TV sounds) are detected in the surrounding environment. In some embodiments, the “detect sound nearby” optionally filter out sounds that are unlikely to be interesting to the user in the captured recording. For example, the detect sound nearby will only keep speech and music detected in the surrounding environment, but will filter out ambient noise such as the air conditioning fan turning, dog barking, cars passing by, etc. In some implementations, the first audio input includes captured audio recording of at least a first advertisement broadcasted on a first TV or radio broadcast channel.”) PNG media_image4.png 641 571 media_image4.png Greyscale generating a stream of modified audio signatures based on the modified audio data, wherein each of the modified audio signatures is generated from a modified version of a corresponding portion of the raw ambient audio; (Figure 7, “[0020] FIG. 7 is a schematic diagram of a process from audio collection by an electronic device to audio fingerprint extraction by a server in accordance with some embodiments.” Figure 4, 406, “[0080] Method 400 further comprises sending (406) the audio data, audio feature information extracted according to the audio data, and/or an audio fingerprint generated according to the audio data to a server,…” “[0081] The audio fingerprint refers to a content-based compact digital signature which represents an important acoustic feature of a piece of audio data….” ) (“[0137] Through Step 11) and Step 12), in M peak feature point pair sequences, each peak feature point pair in each peak feature point pair sequence may be represented by an audio fingerprint, each peak feature point pair sequence corresponds to an audio fingerprint sequence, and M peak feature point pair sequences correspond to M audio fingerprint sequences. A set formed by M audio fingerprint sequences can embody an acoustic feature of audio data; therefore, audio recognition can be performed accordingly to determine the matched channel identity.”) determining, for respective ones of the modified audio signatures, timestamp values corresponding to times at which respective corresponding portions of the raw ambient audio data were collected using the at least one microphone; (Two types of timestamp are taught by Ye: see Figure 10, e.g., “[0023] FIG. 10 is a schematic histogram in which the number of timestamp pairs corresponds to a counted difference between a channel timestamp and a collection timestamp when a collected audio fingerprint does not match a channel audio fingerprint of a channel in accordance with some embodiments.” The “collection timestamp” of Ye teaches the “timestamp values corresponding to times at which respective corresponding portions of the raw ambient audio data were collected” of the claim.) (see figure 9-11 for timestamp relationship from the captured audio fingerprint. “[0162] Assume that an audio fingerprint sequence from an electronic device is: F={(τ.sub.1,hashcode.sub.1),(τ.sub.2,hashcode.sub.2),L,(τ.sub.L,hashcode.sub.L)} [0163] where τ is a collection timestamp and may be a time offset away from a start time point of sound recording, hashcode is a Hash value of an audio fingerprint, and L is the number of audio fingerprints in an audio fingerprint sequence. [0164] In some embodiments, a channel audio fingerprint with a same Hash value as that of an audio fingerprint is searched for in the channel audio fingerprint database one after another from channels, so as to obtain a timestamp pair (t.sup.y,τ) which corresponds to each channel identity and is formed by a collection timestamp of the audio fingerprint and a channel timestamp of the channel audio fingerprint, where the audio fingerprint and the channel audio fingerprint have a same Hash value, y represents a channel identity, and t.sup.y represents a channel timestamp of a channel audio fingerprint corresponding to the channel identity y in the channel audio fingerprint database.” “[0223] In some implementations, the first audio input is accompanied with the timestamp for when the audio recording of the advertisement was captured by the electronic device. This timestamp can be used by the server to match the broadcast times for different advertisements to identify the captured first advertisement. In some embodiments, the user may optionally enter the channel number for the TV broadcast at the time that the recording was captured, so that the identification of the advertisement can be more accurate. The channel number can be sent to the server with the audio recording/audio fingerprint and the timestamp.”) associating the determined timestamp values with the respective ones of the modified audio signatures; and (Collection timestamps are associated with the time of collection of audio: “[0023] FIG. 10 is a schematic histogram in which the number of timestamp pairs corresponds to a counted difference between a channel timestamp and a collection timestamp when a collected audio fingerprint does not match a channel audio fingerprint of a channel in accordance with some embodiments.” “[0024] FIG. 11 is a schematic histogram in which the number of timestamp pairs corresponds to a counted difference between a channel timestamp and a collection timestamp when a collected audio fingerprint matches a channel audio fingerprint of a channel in accordance with some embodiments.”) (See figures 9-11 which compares each collected audio fingerprint together with its collection timestamp “[0136] As described in the foregoing, the four-tuple (t.sub.k,f.sub.k,Δf.sub.k,Δt.sub.k). sub.n is used to represent any peak feature point pair in a peak feature point pair sequence of any phase channel. Parameters in the four-tuple may be understood as follows: (f.sub.k, Δf.sub.k,Δt.sub.k) represents a feature part of a peak feature point pair, and t.sub.k represents a time when (f.sub.k, Δf.sub.k, Δt.sub.k) appears and represents a collection timestamp. In this step, the Hash operation may be performed on (f.sub.k, Δf.sub.k, Δt.sub.k), (f.sub.k, Δf.sub.k,Δt.sub.k) is represented using a Hash code with a fixed bit quantity, as follows: hashcode.sub.k=H (f.sub.k, Δf.sub.k, Δt.sub.k). Through the calculation in this step, any peak feature point pair in a peak feature point pair sequence of any phase channel may be represented by (t.sub.k,hashcode.sub.k).sub.n, where n represents a sequence number of a phase channel or a sequence number of a time-frequency sub-diagram, t.sub.k represents a time when hashcode.sub.k appears and (t.sub.k, hashcode.sub.k).sub.n is an audio fingerprint and may represent a peak feature point pair. An audio fingerprint is represented by a collection timestamp and a Hash value.”) transmitting the stream of modified audio signatures with the associated determined timestamp values to a matching server for comparison with reference audio signatures. (Figure 20 and Figure 1B showing the components and Figure 4, 406 showing the matching: “[0080] Method 400 further comprises sending (406) the audio data, audio feature information extracted according to the audio data, and/or an audio fingerprint generated according to the audio data to a server, so that the server obtains an audio fingerprint according to the audio data, the audio feature information, or the audio fingerprint, and determines, according to a channel audio fingerprint database buffered in real time, a matched channel identity corresponding to a channel audio fingerprint matching the audio fingerprint.” Figure 9: “[0022] FIG. 9 is a schematic flowchart of steps of determining, according to a channel audio fingerprint database buffered in real time, a matched channel identity corresponding to a channel audio fingerprint matching with an audio fingerprint in accordance with some embodiments.”) (“[0207] …Then, the mobile phone obtains audio feature information by performing feature extraction from the audio data; generates, according to the audio feature information, an audio fingerprint including a collection timestamp; and sends the audio fingerprint to the server in a communication manner such as Wifi (a technology enables wireless network access through a radio wave), 2G (a second-generation mobile communications technology), 3G (a third-generation mobile communications technology), or 4G (a fourth-generation mobile communications technology).”) PNG media_image3.png 543 748 media_image3.png Greyscale Ye is arguably a 102 reference for this Claim. Ye teaches generating the fingerprint or signature from unmodified audio and it also has an alternative option of generating the fingerprint or signature from a noise reduced modified audio: “[0216] … In some embodiments, the “detect sound nearby” optionally filter out sounds that are unlikely to be interesting to the user in the captured recording. For example, the detect sound nearby will only keep speech and music detected in the surrounding environment, but will filter out ambient noise such as the air conditioning fan turning, dog barking, cars passing by, etc. ...” Accordingly, Ye teaches generating the “modified audio signatures” of Claim 1. Ye does not explicitly teach a noise reduction “model.” Kupryjanow: PNG media_image5.png 569 763 media_image5.png Greyscale However, Kupryjanow teaches processing the raw ambient audio data with a noise reduction model to thereby generate modified audio data; (see Figure 1. “[0024] In response to detecting a disturbance, the noise reduction model selector 108 can select an appropriate noise reduction model. For example, the noise reduction model selector 108 can load a specific disturbance model 110 that is optimized for separating speech from the specific type of sound in the disturbance. In various examples, the specific disturbance model 110 can then be incorporated by the noise suppressor 104. The noise suppressor 104 can attenuate the components related to the disturbing sound using the specific disturbance model 110.”) It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ye to incorporate the teachings of Kupryjanow to provide a noise reduction model that modifies the noise Doing so would improve the quality of the audio, as recognized by Kupryjanow. (paragraph 1). Claim 8 is a method claim with limitations similar to those found in claim 1 and is rejected under similar rationale. Claim 15 contains limitations similar to those found in claim 1 and is rejected under similar rationale. Additionally, Ye teaches: 15. A non-transitory computer readable medium having stored thereon instructions which, when executed by at least one processor, cause performance of: ([0010] In some embodiments, a non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which, when executed by an electronic device (e.g., electronic device 104, FIGS. 1 and 3), cause the electronic device to perform the operations of any of the methods described herein. In some embodiments, an electronic device (e.g., electronic device 104, FIGS. 1 and 3) includes means for performing, or controlling performance of, the operations of any of the methods described herein.) Claim 5, 12, and 19 Regarding Claim 5, Ye in view of Kupryjanow teach the limitations of claim 1. Further, Ye teaches: 5. The media meter device of claim 1, wherein the noise reduction model attenuates at a portion of the raw ambient audio data associated with sources of household noise. (“[0216] Alternatively, in some embodiments, by initiating a “detect sound nearby” function of the social networking application, the microphone of the electronic device starts to capture and generate an audio input as soon as sounds (e.g., TV sounds) are detected in the surrounding environment. In some embodiments, the “detect sound nearby” optionally filter out sounds that are unlikely to be interesting to the user in the captured recording. For example, the detect sound nearby will only keep speech and music detected in the surrounding environment, but will filter out ambient noise such as the air conditioning fan turning, dog barking, cars passing by, etc. In some implementations, the first audio input includes captured audio recording of at least a first advertisement broadcasted on a first TV or radio broadcast channel.”) Ye does not explicitly teach a noise reduction model. However, Kupryjanow teaches 5. The media meter device of claim 1, wherein the noise reduction model attenuates at a portion of the raw ambient audio data associated with sources of household noise. (“[0010] The present disclosure relates generally to techniques for reducing noise in audio. Specifically, the techniques described herein include an apparatus, method and system for reducing noise in audio using specific disturbance models. An example apparatus includes a preprocessor to receive audio input from a microphone and preprocess the audio input to generate preprocessed audio. The apparatus also includes an acoustic event detector to detect an acoustic event corresponding to a disturbance in the preprocessed audio. The apparatus further includes a noise reduction model selector to select a specific disturbance model based the detected acoustic event. The apparatus also further includes a noise suppressor to attenuate components related to the disturbance in the preprocessed audio using the selected specific disturbance model to generate an enhanced audio with suppressed noise” “[0016] In the example of FIG. 1, system 100 can receive audio input 114 and generate audio output 116. For example, the system 100 may be a speech or voice preprocessing pipeline. In some examples, the audio input 114 may be received from one or more microphones included in the system 100. For example, the system 100 may be implemented on an audio capture device. In some examples, the audio input 114 may be received from an audio capture device via a network or other connection. For example, the system 100 may be implemented on a playback device. In various examples, the audio input 114 may include any number of disturbances, such as dog barking, baby crying, wind noise, music playing, etc. Generally, as used herein, noise is one or more disturbances in the audio input. Noise may be an undesirable audio that distorts the desired audio in the audio input 116. For example, the desired audio may be speech. The audio output 116 can be preprocessed audio in the case that a portion of audio input 114 does not contain disturbance. If any portion of the audio input 114 contains a disturbance, then such disturbances can be detected and the audio output 116 may be an enhanced audio that has any disturbance suppressed.”) See claim 1 for rationale. Claims 12 and 19 contain limitations similar to those found in claim 5 and is rejected under similar rationale. Claim 6, 13, and 20 Regarding Claim 6, Ye in view of Kupryjanow teach the limitations of claim 1. Further, Ye teaches: 6. The media meter device of claim 1, wherein the determining, for respective ones of the modified audio signatures, timestamp values includes determining a local time of the media meter device during collection of the raw ambient audio data. (“[0091] The collection time information is used to indicate time information during audio fingerprint collection. The collection time information may include: a local time which is obtained when the electronic device detects a triggering event and is sent to the server;...”) Claims 13 and 20 contain limitations similar to those found in claim 6 and is rejected under similar rationale. Claim 7 and 14 Regarding Claim 7, Ye in view of Kupryjanow teach the limitations of claim 1. Further, Ye teaches: 7. The media meter device of claim 1, wherein the media meter device is associated with a media presentation device, and wherein the raw ambient audio data is indicative of media being presented by the media presentation device. (“[0045] Referring to FIG. 1B, a network environment includes an electronic device (e.g., electronic device 104, FIG. 1A) connected with a server (e.g., server system 108, FIG. 1A) through network (e.g., network 110, FIG. 1A). In some embodiments, the electronic device is within the vicinity of the television (TV), so that the electronic device can receive sound broadcast by the TV…” “[0078] Method 400 further comprises collecting (404) an audio input of a current channel broadcast in real time to obtain audio data. A television or an external loudspeaker connected to the television is located within an audio input sensing range of the electronic device. The current channel may be a channel currently selected by a user…” “[0079] When a triggering event is detected, the electronic device collects an environment sound through an audio input collector of the electronic device, thereby collecting the program sound of the current channel broadcast by the television in real time. In some embodiments, collection of the sound of the current channel broadcast in real time in the environment may be started and timing is performed from a moment when a triggering event is detected.”) Claim 14 contains limitations similar to those found in claim 7 and is rejected under similar rationale. Claims 2-3, 9-10, and 16-17 are rejected under 35 U.S.C. 103 as obvious over Ye in view of Kupryjanow in further view of Mont-Reynaud (US 20120029670 ) Claim 2, 9, and 16 Regarding Claim 2, Ye in view of Kupryjanow teaches the limitations of claim 1. Further, Ye teaches 2. The media meter device of claim 1, wherein the set of operations further includes: generating a stream of unmodified audio signatures based on the raw ambient audio data such that the stream of unmodified audio signatures and the stream of modified audio signatures each correspond to a common time interval of the collected raw ambient audio; and transmitting the stream of unmodified audio signatures to the matching server. (See Figures 9-10. “[0079] When a triggering event is detected, the electronic device collects an environment sound through an audio input collector of the electronic device, thereby collecting the program sound of the current channel broadcast by the television in real time. In some embodiments, collection of the sound of the current channel broadcast in real time in the environment may be started and timing is performed from a moment when a triggering event is detected. When timing reaches a preset time length, the collection ends, and then audio data within the preset time length is obtained. The audio data refers to collected audio data. The preset time length is preferably 5-15 seconds, and in this way, audio can be effectively recognized and an occupied storage space is relatively small. Certainly, the user may also set a customized time length. By using the preset time length, in subsequent processing, a server can conveniently perform accurate audio recognition to determine a matched channel identity. In an embodiment, the audio data is pulse-code modulation (PCM) audio data whose sampling frequency is 8 kHZ and which is quantized with 16 bits. [0080] Method 400 further comprises sending (406) the audio data, audio feature information extracted according to the audio data, and/or an audio fingerprint generated according to the audio data to a server, so that the server obtains an audio fingerprint according to the audio data, the audio feature information, or the audio fingerprint, and determines, according to a channel audio fingerprint database buffered in real time, a matched channel identity corresponding to a channel audio fingerprint matching the audio fingerprint.”) PNG media_image6.png 602 515 media_image6.png Greyscale Ye in view of Kupryjanow does not explicitly teach both an unmodified and a modified audio signature being correlated to the same time interval. However, Mont-Reynaud teaches The media meter device of claim 1, wherein the set of operations further includes: generating a stream of unmodified audio signatures based on the raw ambient audio data such that the stream of unmodified audio signatures and the stream of modified audio signatures each correspond to a common time interval of the collected raw ambient audio; (See Figure 3 and the two types of audio fingerprint generated at 314 and 322. “[0076] The dispatch module 312 provides the audio signal to one or both of a fingerprint module 314 and a compression module 318. The fingerprint module 314 extracts fingerprints from segments of the incoming audio signal. The extracted fingerprints can be stored in a user fingerprint (FP) cache 316. The extracted fingerprints in the cache 316 can then be provided to the processing section 330, or may be provided directly from the fingerprint module 314. [0077] The compression module 318 compresses the incoming audio signal. The compressed representation of the incoming audio is then stored to a local user audio cache 320. The compressed representation in the user audio cache 320 can subsequently be provided to a decompress & fingerprint module 322. The decompress & fingerprint module 322 compute fingerprints using the compressed representation of the audio signal. Note that fingerprints computed directly from uncompressed audio generally provide more accuracy for identification, as computing fingerprints after audio compression usually decreases quality.” [0078] In general, both the fingerprints and the compressed representation may be stored. In some cases, only the fingerprints may be needed if playback options are not needed. In other cases, only compressed data may be needed if fingerprint creation can be postponed. The fingerprinting of audio queries may be done server-side, or client-side, and in either case it may be delayed.”) PNG media_image7.png 500 662 media_image7.png Greyscale and transmitting the stream of unmodified audio signatures to the matching server. (See Figure 4, transmission between client 304 and server 430. “[0033] The identifying mode 110 includes receiving extracted fingerprints or audio features (submode 202), searching a database using the received input (submode 204), and updating tracking and writing caches based on the search results (submode 206). In submode 202, a portable device sends the server extracted fingerprint(s) or audio features from the segment. Alternatively, as described below, audio segments can be sent from the portable device to the server and features extracted there. Receiving submode 202 progresses to searching in submode 204, in which the server searches a database to identify an audio item or multiple candidate audio items within based on the segment. One slow but simple way to find the best match in a database of audio items is to use an exhaustive search across all possible items and time alignments, giving a similarity score to each. Additional techniques can then be used to decide when a match is good enough, and when ambiguous matches are present.”) PNG media_image8.png 528 658 media_image8.png Greyscale It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Ye in view of Kupryjanow to incorporate the teachings of Mont-Reynaud to provide a both an unmodified and a modified audio signature being correlated to the same time interval. Doing so would improve the accuracy of audio identification, as recognized by Mont-Reynaud. (paragraph 77). Claims 9 and 16 contain limitations similar to those found in Claim 2 and is rejected under similar rationale. Claim 3,10 and 17 Regarding Claim 3, Ye in view of Kupryjanow In further view of Mont-Reynaud teach the limitations of claim 2. As provided with respect to Claim 1, Ye teaches generating the fingerprint or signature from unmodified audio and it also has an alternative option of generating the fingerprint or signature from a noise reduced modified audio: “[0216] … In some embodiments, the “detect sound nearby” optionally filter out sounds that are unlikely to be interesting to the user in the captured recording. For example, the detect sound nearby will only keep speech and music detected in the surrounding environment, but will filter out ambient noise such as the air conditioning fan turning, dog barking, cars passing by, etc. ...” Accordingly, Ye teaches generating the “modified audio signatures” and the “unmodified audio signatures” of this claim. Ye does not teach generating both types of signatures together. However, this claim does not require the two types of signatures together. Further Ye teaches 3. The media meter device of claim 2, wherein the set of operations further includes receiving, from the matching server, an indication that at least one of the stream of modified audio signatures or the stream of unmodified audio signatures matched at least one reference media signature. (“[0169] Step 906: Determine a channel identity corresponding to the channel audio fingerprint corresponding to the maximum value of the similarity measurement values as a matched channel identity. If the maximum value of the similarity measurement values exceeds the preset threshold, it indicates that the channel audio fingerprint corresponding to the maximum value of the similarity measurement values matches with the audio fingerprint, and the channel identity corresponding to the channel audio fingerprint is determined as the matched channel identity.” “[0173] The event detection module 1202 is used to detect a triggering event. The audio collection module 1204 is used to collect an audio input of a current channel broadcast in real time in an environment to obtain audio data. The sending module 1206 is used to send the audio data, audio feature information extracted according to the audio data, or an audio fingerprint generated according to the audio data to a server, so that the server obtains an audio fingerprint according to the audio data, the audio feature information, or the audio fingerprint and determines, according to a channel audio fingerprint database buffered in real time, a matched channel identity corresponding to a channel audio fingerprint matching with the audio fingerprint. The preset information receiving module 1208 is used to receive preset information which is obtained by the server from a preset information database, is sent by the server, and corresponds to the matched channel identity. [0174] In some embodiment, audio fingerprint collection corresponds to collection time information, and the preset information receiving module 1208 is further used to receive preset information which is obtained by the server from the preset information database, sent by the server, corresponds to the matched channel identity, and has a time attribute matching with the collection time information.”) Claims 10 and 17 contain limitations similar to those found in Claim 3 and is rejected under similar rationale. Allowable Subject Matter Claims 4,11, and 18, if rewritten in independent form including all of the limitations of the base claim and all limitations of any intervening claims, would comprise a particular combination of elements, which is neither taught nor suggested by the prior art. Reference Cited The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. US 20140254807 to Fonseca, JR discloses “[0044] Once the audio analyzer 48 has identified the temporal location of any detected noise in the audio signal received by the one or more microphones 16, the audio analyzer 48 provides that information to the audio signature generator 50, which may use that information to nullify those portions of the spectrogram it generates that are corrupted by noise.”. US 20090012638 to Lou discloses “[0023] FIG. 2 illustrates an exemplary flow diagram according to the present invention. At first, the audio signal 210 is fed to the preprocessor 220. Certain audio signals, such as songs, have short period of silence at the start. The pre-processor detects the onset of the song and disregards the silence. In case that the noise is present in the input audio, the pre-processor also remove the white additive noise. The resulting signal has higher signal to noise ratio, which in turn increases identification accuracy. After pre-processor, the feature extractor 230 is used to extract characteristic audio features. The resulting features can be used as fingerprints to compare against known fingerprints database for audio identification. Part or all of the features can also be used for classification purpose.”. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALI M HASSAN whose telephone number is (571)272-5331. The examiner can normally be reached Monday - Friday 8:00am - 4:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras Shah can be reached at (571)270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALI M HASSAN/ Examiner, Art Unit 2653 /FARIBA SIRJANI/Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jan 13, 2025
Application Filed
Aug 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737548
METHOD, DEVICE, AND COMPUTER PROGRAM PRODUCT FOR GENERATING TABLE
2y 4m to grant Granted Sep 15, 2026
Patent 12718028
CONVERSATION TOPIC AND SUMMARY EXTRACTION FOR REAL-TIME MESSAGING PLATFORMS
3y 0m to grant Granted Aug 25, 2026
Patent 12701315
Speech Recognition System and Method for Providing Speech Recognition Service
3y 7m to grant Granted Aug 04, 2026
Patent 12675740
TRAINING APPARATUS, TRAINING METHOD, AND TRAINING PROGRAM
2y 7m to grant Granted Jul 07, 2026
Patent 12670329
METHOD, APPARATUS, DEVICE, AND STORAGE MEDIUM FOR CLUSTERING EXTRACTION OF ENTITY RELATIONSHIPS
2y 3m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
99%
With Interview (+37.5%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 16 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month