Prosecution Insights
Last updated: August 17, 2026
Application No. 19/195,317

AUTOMATED VOICEMAIL DETECTION

Non-Final OA §103§112
Filed
Apr 30, 2025
Priority
Oct 22, 2024 — provisional 63/710,342
Examiner
ESCALANTE, OVIDIO
Art Unit
3992
Tech Center
3900
Assignee
Advanis Inc.
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
82%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
167 granted / 222 resolved
+15.2% vs TC avg
Moderate +7% lift
Without
With
+7.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
43 currently pending
Career history
261
Total Applications
across all art units

Statute-Specific Performance

§101
3.4%
-36.6% vs TC avg
§103
27.8%
-12.2% vs TC avg
§102
9.5%
-30.5% vs TC avg
§112
22.5%
-17.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 222 resolved cases

Office Action

§103 §112
DETAILED ACTION This action is in response to the Applicant’s preliminary amendment filed on June 6, 2025. As set forth therein, claims 1-20 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on December 18, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 4-5 and 14-15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 4 and 14 recite the limitation "the two numerical arrays" in line 2. There is insufficient antecedent basis for this limitation in the claim. Claims 5 and 15 recite the limitation "the highest normalized correlation value" in lines 1-2. There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 7, 8, 11, 17 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Noble US Patent Pub. 2006/0256949 in view of Brown et al. US Patent Pub. 2003/0086541. Regarding claim 11: A method for automated detection of voicemail calls, the method comprising: See also paragraph [0042] of Noble which discloses “if the call results a connection as detected by step 308, then step 314 determines if the call resulted in a human answer or not. If not, then step 322 is invoked to take appropriate action depending on the type of connection that has occurred (e.g., answer machine, voicemail, answering services such as a privacy director, etc.)”. See also Figure 3 steps 304-314. initiating a call to a telephone number associated with a user device; See paragraph [0040] of Noble which discloses “[w]ith reference to FIG. 3, when the scheduled callback time arrives for any such scheduled callback, step 302 is invoked to retrieve the stored callback number and possibly other associated data from its stored location in the database. Next, step 304 initiates a callback by automatically dialing the retrieved callback number”. capturing at least one audio sample frame from the call; analyzing the audio sample frame to determine if the call is a voicemail call; and Noble discloses in paragraph [0008] that the system will monitor the outbound call to see whether it is answered by a person, machine or service tones. See also paragraphs [0010], [0024] and [0040] Noble does not specifically disclose a specific method used for determining if the call is a voicemail call. That is, Noble does not specifically disclose analyzing the audio sample frame to determine if the call is a voicemail call. Nonetheless, Brown is directed to a method for determining if a person has answered a telephone versus an answering machine. See paragraph [0036]. In addition, Brown discloses “using well known techniques for detecting the energy in audio samples, energy analysis block 206 is used for answering machine detection, silence detection, and voice activity detection. Energy analysis block 206 performs answering machine detection by looking for the cadence in energy being received back in the voice samples.” See paragraph [0036]. See also paragraphs [0020] and [0050]. See also paragraph [0047] which discloses Block 1001 receives 10 milliseconds of audio data from block 801. Block 1001 segments this audio data into frames. Block 1002 is responsive to the audio frames to compute the raw energy level, perform energy normalization, and autocorrelation operations all of which are well known to those skilled in the art. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to analyze at least one sample frame to determine if the call is a voicemail call. As explained above, Noble discloses of determining if the call is a voicemail call; however, Noble does not specifically disclose of a specific method used for this determination. As explained by Brown, it was known to segment the audio data into frames and to perform an analysis on the audio frames in order to determine whether a human or a machine answered a call. As further explained by Brown, every frame of data is analyzed to see whether an end-point is reached. See paragraph [0044]. The end-point signifies a change in energy for a significant period of time. This is used to determine the detection of tones and speech. Thus, one of ordinary skill in the art would have found it obvious to capture an analyze audio sample frames in order to determine whether tones are speech are present. As explained above, Noble discloses that it was known to detect whether an answering device or a human answered the call. Brown discloses that analyzing frames was a well-known method to make an accurate determination if a person has answered a telephone versus an answering machine. if the call is a voicemail call, dropping the call connection, otherwise connecting the call to an agent operator device. See paragraph [0042] and Figure 3 of Noble which discloses that if there is no human answer (step 314), an appropriate action is taken, and the call is disconnected at step 324. If a human answers at step 314 and an agent is available, the call is connected to an agent at step 328. Regarding claim 7: The method of claim 1, wherein prior to analyzing, the method comprises: applying ringtone detection to detect audio sample frames comprising a ringtone, and analyzing audio sample frames not comprising a ringtone. Noble discloses “if the monitoring of the outbound call results in the detection of a service signal (e.g., disconnected or not in service signal), the system may update the related callback information from the database or otherwise note that service tones were received.”. See paragraph [0010]. The Examiner notes that since the detection of a service tone can be either a disconnected or not in service signal, then the call has not be answered and thus, is prior to the analyzing step since the analyzing step is based on detecting either a human or a voicemail system which occurs when there isn’t a disconnect or not in service signal. With respect to analyzing audio sample frames, Brown discloses upon receiving audio information from a destination endpoint of a call, the automatic speech recognition unit processes the audio information for speech and tones by first determining if the audio information is speech (audio sample frames not comprising a ringtone) or tones. If the audio information is speech, the automatic speech recognition unit separately executes automatic speech recognition procedures to detect words and phrases using an automatic speech recognition grammar for speech. If the audio information is tones, the automatic speech recognition unit separately executes automatic speech recognition procedures to detect tones using an automatic speech recognition grammar for tones. See paragraphs [0006]-[0007]. See also paragraph [0035] which discloses various types of tones, including ring back, dial tone, busy tone, reorder tones, etc. See also paragraph [0036] which discloses audio sample frames as explained above. As set forth above, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention to analyze audio frames. In addition, as set forth in paragraph [0008], Noble discloses that its system will process the call to determine whether it is answered by a person, machine or service tones. See paragraph [0035]. Thus, it would have been obvious to a person of ordinary skill in the art to analyze frames in order to make an accurate determination if a person has answered a telephone versus an answering machine. Regarding claim 8: The method of claim 7, wherein the ringtone detection is performed using a Goertzel algorithm. As set forth above, Noble and Brown discloses of using ringtone detection. See paragraphs [0006]. [0007] and [0035] of Brown. In addition, Brown discloses “[f]or the frequency analysis, processor 502 advantageously utilizes Goertzel algorithm which is a type of Discrete Fourier transform. One skilled in the art readily knows how to implement the Goertzel algorithm on processor 502 and to implement other algorithms for the detection of frequency.” See paragraph [0035] Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to utilize a Goertzel algorithm. As explained in paragraph [0035] of Brown, audio samples are analyzed to determine various types of tones. Brown discloses that detecting tones was well known. The Examiner notes that Noble also discloses the detection of tones. Thus, it would have been obvious to a person of ordinary skill in the art to use a well-known method for the detection of tones in order to precisely detect the type of tones that was received. As explained by Brown this will help determine whether an answering device has answered the call. Regarding claim 11: A system for automated detection of voicemail calls, the system comprising: See also paragraph [0042] of Noble which discloses “if the call results a connection as detected by step 308, then step 314 determines if the call resulted in a human answer or not. If not, then step 322 is invoked to take appropriate action depending on the type of connection that has occurred (e.g., answer machine, voicemail, answering services such as a privacy director, etc.)” See Figure 1 and the abstract which discloses a system for the detection of voicemail calls. See also Figure 3 steps 304-314. a communication interface; and See Figure 1 which discloses a connection between various components of the system. at least one processor coupled to the communication interface, the at least one processor configured for: See paragraph [0024] which discloses the use of a digital signal processor, ASIC, programmable IC (PIC), along with the use of analog to digital converters and/or digital to analog converters can monitor the phone line to determine if the outbound call was answered by a person, an answering machine, or resulted in a busy signal, disconnected signal, not in service signal, etc. See also paragraph [0026] initiating, via the communication interface, a call to a telephone number associated with a user device; See paragraph [0040] of Noble which discloses “[w]ith reference to FIG. 3, when the scheduled callback time arrives for any such scheduled callback, step 302 is invoked to retrieve the stored callback number and possibly other associated data from its stored location in the database. Next, step 304 initiates a callback by automatically dialing the retrieved callback number”. capturing at least one audio sample frame from the call; analyzing the audio sample frame to determine if the call is a voicemail call; and Noble discloses in paragraph [0008] that the system will monitor the outbound call to see whether it is answered by a person, machine or service tones. See also paragraphs [0010], [0024] and [0040] Noble does not specifically disclose a specific method used for determining if the call is a voicemail call. That is, Noble does not specifically disclose analyzing the audio sample frame to determine if the call is a voicemail call. Nonetheless, Brown is directed to a method for determining if a person has answered a telephone versus an answering machine. See paragraph [0036]. In addition, Brown discloses “using well known techniques for detecting the energy in audio samples, energy analysis block 206 is used for answering machine detection, silence detection, and voice activity detection. Energy analysis block 206 performs answering machine detection by looking for the cadence in energy being received back in the voice samples.” See paragraph [0036]. See also paragraphs [0020] and [0050]. See also paragraph [0047] which discloses Block 1001 receives 10 milliseconds of audio data from block 801. Block 1001 segments this audio data into frames. Block 1002 is responsive to the audio frames to compute the raw energy level, perform energy normalization, and autocorrelation operations all of which are well known to those skilled in the art. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to analyze at least one sample frame to determine if the call is a voicemail call. As explained above, Noble discloses of determining if the call is a voicemail call; however, Noble does not specifically disclose of a specific method used for this determination. As explained by Brown, it was known to segment the audio data into frames and to perform an analysis on the audio frames in order to determine whether a human or a machine answered a call. As further explained by Brown, every frame of data is analyzed to see whether an end-point is reached. See paragraph [0044]. The end-point signifies a change in energy for a significant period of time. This is used to determine the detection of tones and speech. Thus, one of ordinary skill in the art would have found it obvious to capture an analyze audio sample frames in order to determine whether tones are speech are present. As explained above, Noble discloses that it was known to detect whether an answering device or a human answered the call. Brown discloses that analyzing frames was a well-known method to make an accurate determination if a person has answered a telephone versus an answering machine. if the call is a voicemail call, dropping the call connection, otherwise connecting the call to an agent operator device. See paragraph [0042] and Figure 3 of Noble which discloses that if there is no human answer (step 314), an appropriate action is taken, and the call is disconnected at step 324. If a human answers at step 314 and an agent is available, the call is connected to an agent at step 328. Regarding claim 17: The system of claim 11, wherein prior to analyzing, the at least one processor is further configured for: applying ringtone detection to detect audio sample frames comprising a ringtone, and analyzing audio sample frames not comprising a ringtone. Noble discloses “if the monitoring of the outbound call results in the detection of a service signal (e.g., disconnected or not in service signal), the system may update the related callback information from the database or otherwise note that service tones were received.”. See paragraph [0010]. The Examiner notes that since the detection of a service tone can be either a disconnected or not in service signal, then the call has not be answered and thus, is prior to the analyzing step since the analyzing step is based on detecting either a human or a voicemail system which occurs when there isn’t a disconnect or not in service signal. With respect to analyzing audio sample frames, Brown discloses upon receiving audio information from a destination endpoint of a call, the automatic speech recognition unit processes the audio information for speech and tones by first determining if the audio information is speech (audio sample frames not comprising a ringtone) or tones. If the audio information is speech, the automatic speech recognition unit separately executes automatic speech recognition procedures to detect words and phrases using an automatic speech recognition grammar for speech. If the audio information is tones, the automatic speech recognition unit separately executes automatic speech recognition procedures to detect tones using an automatic speech recognition grammar for tones. See paragraphs [0006]-[0007]. See also paragraph [0035] which discloses various types of tones, including ring back, dial tone, busy tone, reorder tones, etc. See also paragraph [0036] which discloses audio sample frames as explained above. As set forth above, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention to analyze audio frames. In addition, as set forth in paragraph [0008], Noble discloses that its system will process the call to determine whether it is answered by a person, machine or service tones. See paragraph [0035]. Thus, it would have been obvious to a person of ordinary skill in the art to analyze frames in order to make an accurate determination if a person has answered a telephone versus an answering machine. Regarding claim 18: The system of claim 17, wherein the ringtone detection is performed using a Goertzel algorithm. As set forth above, Noble and Brown discloses of using ringtone detection. See paragraphs [0006]. [0007] and [0035] of Brown. In addition, Brown discloses “[f]or the frequency analysis, processor 502 advantageously utilizes Goertzel algorithm which is a type of Discrete Fourier transform. One skilled in the art readily knows how to implement the Goertzel algorithm on processor 502 and to implement other algorithms for the detection of frequency.” See paragraph [0035] Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to utilize a Goertzel algorithm. As explained in paragraph [0035] of Brown, audio samples are analyzed to determine various types of tones. Brown discloses that detecting tones was well known. The Examiner notes that Noble also discloses the detection of tones. Thus, it would have been obvious to a person of ordinary skill in the art to use a well-known method for the detection of tones in order to precisely detect the type of tones that was received. As explained by Brown this will help determine whether an answering device has answered the call. Claim(s) 2, 3 ,9, 10, 12, 13, 19 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Noble US Patent Pub. 2006/0256949 in view of Brown et al. US Patent Pub. 2003/0086541 and further in view of Clark et al. US 9,596,578. Regarding claim 2: The method of claim 1, wherein analyzing the audio sample frame to determine if the call is a voicemail call comprises: comparing each audio sample frame to a reference voicemail sample associated with the telephone number. As set forth above, Noble and Brown discloses of analyzing the audio sample frame to determine if the call is a voicemail call. Nobel and Brown do not specifically disclose comparing each audio sample frame to a reference voicemail sample associated with the telephone number. Nonetheless, Clark is directed to a method to detect when a call has been answered by voicemail. See the abstract. In addition, Clark discloses by using voicemail fingerprints for known voicemail greetings uniquely associated individual forwarding telephone numbers, the disclosed techniques may enable highly accurate detection of voicemail greetings on a per-telephone number basis, thus improving accuracy over previous systems that relied on comparing greetings to characteristics that are shared by all voicemail greetings. See col. 4, lines 46-52. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to compare each audio sample frame to a reference voicemail sample associated with the telephone number As explained by Clark by comparing using previous voicemail samples associated with the telephone number, this would improve accuracy in the detection of whether a voicemail system answered the call. Regarding claim 3: The method of claim 2, wherein the comparison is performed by comparing an array of the audio sample to an array of the reference voicemail sample. Clark discloses comparing the current audio analysis stream to each one of the audio analysis streams in the voicemail fingerprint of the known voicemail greeting associated with the forwarding telephone number. See col. 2, lines 27-30. See also col. 4, lines 46-60. As shown in Figure 3 and col. 10, lines 17-45, each audio sample is an array. See also col. 13, lines 55-65 which discloses “an embodiment in which the audio characteristics in the streams of the voicemail fingerprints of known voicemail greetings are made up of a series of audio characteristic chunks, and in which the audio characteristics in the current audio analysis stream are also made up of a series of audio characteristic chunks, the comparison of the current audio analysis stream to each one of the audio analysis streams in the voicemail fingerprint of the known voicemail greeting in step 210 of FIG. 2 may be performed on a chunk by chunk basis for each one of the audio analysis streams in the voicemail fingerprint of the known voicemail greeting.” As explained above, it would have been obvious to a person of ordinary skill in the art to compare using a reference voicemail sample. Thus, by comparing an array of the audio sample to an array of a reference voicemail sample, Clark states that this would improve accuracy in the detection of whether a voicemail system answered the call. See col. 4, lines 46-53 of Clark. See also col. 14, lines 5-11. Regarding claim 9: The method of claim 2, further comprising, initially generating the reference voicemail sample by: As set forth above, it would have been obvious to a person of ordinary skill in the art to use a reference voicemail sample. In addition, Clark discloses the disclosed system may obtain the known voicemail greeting associated with a given forwarding telephone number by detecting that a voicemail fingerprint is not stored for the forwarding telephone number, and then performing a candidate voicemail fingerprint generation operation for the forwarding telephone number. See col. 3, lines 55-60. initiating an initial call to the telephone number; See col. 3, line 55 – col. 4, line 19 of Clark which discloses recording audio received beginning when a first call to the forwarding telephone number is answered. recording the initial call to generate an audio call recording; See col. 3, line 55 – col. 4, line 19 of Clark which discloses recording audio received beginning when a first call to the forwarding telephone number is answered. applying a recorded voicemail detection model to the audio call recording to determine if the audio call is a voicemail call; Clark discloses in response to detecting that a requested user input (e.g. keypad selection or voice input) was not received prior to expiration of a time out period following the first call to the forwarding telephone number being answered, generating the candidate voicemail fingerprint using the recording of the audio received beginning when the first call to the forwarding telephone number was answered. See col. line 55 – col. 4, line 19. if the call is a voicemail call, extracting a reference voicemail sample from the audio call recording; and Clark discloses in response to detecting that a requested user input (e.g. keypad selection or voice input) was not received prior to expiration of a time out period following the first call to the forwarding telephone number being answered, generating the candidate voicemail fingerprint using the recording of the audio received beginning when the first call to the forwarding telephone number was answered. See col. line 55 – col. 4, line 19. storing the reference voicemail sample in association with the telephone number. Clark discloses in response to the new voicemail fingerprint matching the candidate voicemail fingerprint, storing the candidate fingerprint as the voicemail fingerprint of a known voicemail greeting associated with the forwarding telephone number. See col. line 55 – col. 4, line 19. As explained above, it would have been obvious to a person of ordinary skill in the art to compare using a reference voicemail sample. Clark discloses that this would improve accuracy in the detection of whether a voicemail system answered the call. See col. 4, lines 46-53 of Clark. Regarding claim 10: The method of claim 9, further comprising applying a tone detection model to the audio call recording to determine if the audio call recording is a voicemail call. The Examiner notes that although Noble discloses detecting tones and determining if the call was answered by a voicemail system, Noble does not specifically disclose of applying a tone detection model to the audio call recording. Nonetheless, see page [0050] of Brown which discloses performing a HMM analysis utilizing a unified model for both speech and tones. See also paragraph [0044]. As explained by Brown, this method is for determining the presence of an answering machine or to assist in the determination of tone detection. See paragraphs [0026] and [0036] Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to apply a tone detection model. As explained by Brown models for analyzing for speech and tones were well known in the prior art. Brown in paragraph [0044] discloses of other types of models that can be used “[o]ne skilled in the art would readily realize that other alternatives to HMM could be used such as Neural Net analysis”. Thus, one of ordinary skill in the art would have used a model in order to make an accurate determination of whether a call has been completed to a person. Regarding claim 12: The system of claim 11, wherein analyzing the audio sample frame to determine if the call is a voicemail call comprises the at least one processor being configured for: comparing each audio sample frame to a reference voicemail sample associated with the telephone number. As set forth above, Noble and Brown discloses of analyzing the audio sample frame to determine if the call is a voicemail call. Nobel and Brown do not specifically disclose comparing each audio sample frame to a reference voicemail sample associated with the telephone number. Nonetheless, Clark is directed to a method to detect when a call has been answered by voicemail. See the abstract. In addition, Clark discloses by using voicemail fingerprints for known voicemail greetings uniquely associated individual forwarding telephone numbers, the disclosed techniques may enable highly accurate detection of voicemail greetings on a per-telephone number basis, thus improving accuracy over previous systems that relied on comparing greetings to characteristics that are shared by all voicemail greetings. See col. 4, lines 46-52. Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to compare each audio sample frame to a reference voicemail sample associated with the telephone number As explained by Clark by comparing using previous voicemail samples associated with the telephone number, this would improve accuracy in the detection of whether a voicemail system answered the call. Regarding claim 13: The system of claim 12, wherein the comparison is performed by comparing an array of the audio sample to an array of the reference voicemail sample. Clark discloses comparing the current audio analysis stream to each one of the audio analysis streams in the voicemail fingerprint of the known voicemail greeting associated with the forwarding telephone number. See col. 2, lines 27-30. See also col. 4, lines 46-60. As shown in Figure 3 and col. 10, lines 17-45, each audio sample is an array. See also col. 13, lines 55-65 which discloses “an embodiment in which the audio characteristics in the streams of the voicemail fingerprints of known voicemail greetings are made up of a series of audio characteristic chunks, and in which the audio characteristics in the current audio analysis stream are also made up of a series of audio characteristic chunks, the comparison of the current audio analysis stream to each one of the audio analysis streams in the voicemail fingerprint of the known voicemail greeting in step 210 of FIG. 2 may be performed on a chunk by chunk basis for each one of the audio analysis streams in the voicemail fingerprint of the known voicemail greeting.” As explained above, it would have been obvious to a person of ordinary skill in the art to compare using a reference voicemail sample. Thus, by comparing an array of the audio sample to an array of a reference voicemail sample, Clark states that this would improve accuracy in the detection of whether a voicemail system answered the call. See col. 4, lines 46-53 of Clark. See also col. 14, lines 5-11. Regarding claim 19: The system of claim 12, wherein the at least one processor is further configured for, initially generating the reference voicemail sample by: As set forth above, it would have been obvious to a person of ordinary skill in the art to use a reference voicemail sample. In addition, Clark discloses the disclosed system may obtain the known voicemail greeting associated with a given forwarding telephone number by detecting that a voicemail fingerprint is not stored for the forwarding telephone number, and then performing a candidate voicemail fingerprint generation operation for the forwarding telephone number. See col. 3, lines 55-60. initiating an initial call to the telephone number; See col. 3, line 55 – col. 4, line 19 of Clark which discloses recording audio received beginning when a first call to the forwarding telephone number is answered. recording the initial call to generate an audio call recording; See col. 3, line 55 – col. 4, line 19 of Clark which discloses recording audio received beginning when a first call to the forwarding telephone number is answered. applying a recorded voicemail detection model to the audio call recording to determine if the audio call is a voicemail call; Clark discloses in response to detecting that a requested user input (e.g. keypad selection or voice input) was not received prior to expiration of a time out period following the first call to the forwarding telephone number being answered, generating the candidate voicemail fingerprint using the recording of the audio received beginning when the first call to the forwarding telephone number was answered. See col. line 55 – col. 4, line 19. if the call is a voicemail call, extracting a reference voicemail sample from the audio call recording; and Clark discloses in response to detecting that a requested user input (e.g. keypad selection or voice input) was not received prior to expiration of a time out period following the first call to the forwarding telephone number being answered, generating the candidate voicemail fingerprint using the recording of the audio received beginning when the first call to the forwarding telephone number was answered. See col. line 55 – col. 4, line 19. storing the reference voicemail sample in association with the telephone number. Clark discloses in response to the new voicemail fingerprint matching the candidate voicemail fingerprint, storing the candidate fingerprint as the voicemail fingerprint of a known voicemail greeting associated with the forwarding telephone number. See col. line 55 – col. 4, line 19. As explained above, it would have been obvious to a person of ordinary skill in the art to compare using a reference voicemail sample. Clark discloses that this would improve accuracy in the detection of whether a voicemail system answered the call. See col. 4, lines 46-53 of Clark. Regarding claim 20: The system of claim 19, wherein the at least one processor is further configured for: applying a tone detection model to the audio call recording to determine if the audio call recording is a voicemail call. The Examiner notes that although Noble discloses detecting tones and determining if the call was answered by a voicemail system, Noble does not specifically disclose of applying a tone detection model to the audio call recording. Nonetheless, see page [0050] of Brown which discloses performing a HMM analysis utilizing a unified model for both speech and tones. See also paragraph [0044]. As explained by Brown, this method is for determining the presence of an answering machine or to assist in the determination of tone detection. See paragraphs [0026] and [0036] Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to apply a tone detection model. As explained by Brown models for analyzing for speech and tones were well known in the prior art. Brown in paragraph [0044] discloses of other types of models that can be used “[o]ne skilled in the art would readily realize that other alternatives to HMM could be used such as Neural Net analysis”. Thus, one of ordinary skill in the art would have used a model in order to make an accurate determination of whether a call has been completed to a person. Claim(s) 4, 5, 14 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Noble US Patent Pub. 2006/0256949 in view of Brown et al. US Patent Pub. 2003/0086541 and further in view of Clark et al. US 9,596,578 and further in view of Solbach US 10,650,840. Regarding claim 4: The method of claim 3, wherein the comparison is performed by determining a cross correlation between the two numerical arrays. Noble, Brown and Clark do not specifically disclose performing the comparing by determining a cross-correlation between two arrays. The Examiner notes however that Brown discloses that it was known to perform autocorrelation operations with respect to audio data received during a call to determine whether a human or a machine answered the call. See paragraph [0047] of Brown. Therefore, one of ordinary skill in the art would have considered using cross-correlation for performing the comparison. Nonetheless, Solbach discloses that it was known to compare two audio samples by performing cross-correlation. As explained in col. 9, lines 64-67, the cross-correlation data may indicate a measure of similarity between the microphone audio data and the subsampled playback audio data as a function of the displacement of one relative to the other. Solbach discloses “[t]he device 110 may determine the cross-correlation data using any techniques known to one of skill in the art (e.g., normalized cross-covariance function or the like) without departing from the disclosure.” Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to use a cross correlation between the two arrays. As set forth above, Brown already discloses that it was known to perform autocorrelaton operations with respect to the audio received during a call. See paragraph [0047] of Brown which discloses “Block 1001 segments this audio data into frames. Block 1002 is responsive to the audio frames to compute the raw energy level, perform energy normalization, and autocorrelation operations all of which are well known to those skilled in the art.”. Solbech discloses that it was known to perform cross-correlation in order to determine the similarity between two sets of audio data. As set forth above, Clark discloses it was known to compare each audio sample frame to a reference voicemail sample associated with the telephone number. As further explained by Clark by comparing using previous voicemail samples associated with the telephone number, this would improve accuracy in detecting whether a voicemail system answering the call. Thus, using the known method of performing a cross-correlation function to determine similarity between two audio samples would have been obvious to a person of ordinary skill in the art since it would help improve accuracy in detecting whether a voicemail system answered the call. Regarding claim 5: The method of claim 4, wherein the call is determined to be a voicemail call if the highest normalized correlation value exceeds a predetermined threshold. Solbach discloses “the device 110 may determine a maximum value of the cross-correlation data (e.g., magnitude of a highest peak, Peak A indicating a highest correlation between the signals) and determine a threshold value based on the maximum value.” See col. 11, lines 9-36. As set forth above, it would have been obvious to a person of ordinary skill in the art to perform a correlation function in order to determine the similarities between two audio samples which will help improve the detection accuracy. Regarding claim 14: The system of claim 13, wherein the comparison is performed by determining a cross correlation between the two numerical arrays. Noble, Brown and Clark do not specifically disclose performing the comparing by determining a cross-correlation between two arrays. The Examiner notes however that Brown discloses that it was known to perform autocorrelation operations with respect to audio data received during a call to determine whether a human or a machine answered the call. See paragraph [0047] of Brown. Therefore, one of ordinary skill in the art would have considered using cross-correlation for performing the comparison. Nonetheless, Solbach discloses that it was known to compare two audio samples by performing cross-correlation. As explained in col. 9, lines 64-67, the cross-correlation data may indicate a measure of similarity between the microphone audio data and the subsampled playback audio data as a function of the displacement of one relative to the other. Solbach discloses “[t]he device 110 may determine the cross-correlation data using any techniques known to one of skill in the art (e.g., normalized cross-covariance function or the like) without departing from the disclosure.” Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to use a cross correlation between the two arrays. As set forth above, Brown already discloses that it was known to perform autocorrelaton operations with respect to the audio received during a call. See paragraph [0047] of Brown which discloses “Block 1001 segments this audio data into frames. Block 1002 is responsive to the audio frames to compute the raw energy level, perform energy normalization, and autocorrelation operations all of which are well known to those skilled in the art.”. Solbech discloses that it was known to perform cross-correlation in order to determine the similarity between two sets of audio data. As set forth above, Clark discloses it was known to compare each audio sample frame to a reference voicemail sample associated with the telephone number. As further explained by Clark by comparing using previous voicemail samples associated with the telephone number, this would improve accuracy in detecting whether a voicemail system answering the call. Thus, using the known method of performing a cross-correlation function to determine similarity between two audio samples would have been obvious to a person of ordinary skill in the art since it would help improve accuracy in detecting whether a voicemail system answered the call. Regarding claim 15: The system of claim 14, wherein the call is determined to be a voicemail call if the highest normalized correlation value exceeds a predetermined threshold. Solbach discloses “the device 110 may determine a maximum value of the cross-correlation data (e.g., magnitude of a highest peak, Peak A indicating a highest correlation between the signals) and determine a threshold value based on the maximum value.” See col. 11, lines 9-36. As set forth above, it would have been obvious to a person of ordinary skill in the art to perform a correlation function in order to determine the similarities between two audio samples which will help improve the detection accuracy. Claim(s) 6 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Noble US Patent Pub. 2006/0256949 in view of Brown et al. US Patent Pub. 2003/0086541 and further in view of Hung et al. US Patent 2021/0304742. Regarding claim 6: The method of claim 1, wherein analyzing the audio sample frame to determine if the call is a voicemail call comprises: applying a live voicemail detection model to the audio sample frames, the live voicemail detection model comprising a trained machine learning model. Noble and Brown do not specifically disclose applying a live voicemail detection model to the audio sample frames, the live voicemail detection model comprising a trained machine learning model Nonetheless, Hung discloses detecting a voicemail system using a machine learning model. As set forth in paragraph [0017], Hung discloses a machine learning model is trained using two categories of audio files, specifically files including beeps, and files that include speech with no beeps. See also paragraph [0049] Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to apply a voicemail detection model as explained by Hung in order to increase the accuracy in determining whether a voicemail system or a person answered the call. See paragraphs [0017] and [0049]. Regarding claim 16: The system of claim 11, wherein analyzing the audio sample frame to determine if the call is a voicemail call comprises the at least one processor being configured for: applying a live voicemail detection model to the audio sample frames, the live voicemail detection model comprising a trained machine learning model. Noble and Brown do not specifically disclose applying a live voicemail detection model to the audio sample frames, the live voicemail detection model comprising a trained machine learning model Nonetheless, Hung discloses detecting a voicemail system using a machine learning model. As set forth in paragraph [0017], Hung discloses a machine learning model is trained using two categories of audio files, specifically files including beeps, and files that include speech with no beeps. See also paragraph [0049] Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to apply a voicemail detection model as explained by Hung in order to increase the accuracy in determining whether a voicemail system or a person answered the call. See paragraphs [0017] and [0049]. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ovidio Escalante whose telephone number is (571)272-7537. The examiner can normally be reached on Monday to Friday - 6:00 AM to 2:30 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Fuelling, can be reached at telephone number (571)272-7537. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from Patent Center. Status information for published applications may be obtained from Patent Center. Status information for unpublished applications is available through Patent Center for authorized users only. Should you have questions about access to Patent Center, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Ovidio Escalante/ Primary Examiner, Art Unit 3992 1 The Examiner finds that claim 1 recites a contingent limitation. In accordance with MPEP 2111.04(II), The broadest reasonable interpretation of a method (or process) claim having contingent limitations requires only those steps that must be performed and does not include steps that are not required to be performed because the condition(s) precedent are not met.  Thus, in this claim, the Examiner finds that the contingent limitation (if statements and “otherwise”) are not required in the claim because the condition precedent are not met. For example, either the dropping the call connection or connecting the call to an agent operator device is not performed if the condition precedent are not met.
Read full office action

Prosecution Timeline

Apr 30, 2025
Application Filed
Aug 05, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent RE50982
APPARATUS, METHOD AND COMPUTER PROGRAM FOR UPMIXING A DOWNMIX AUDIO SIGNAL
1y 6m to grant Granted Aug 04, 2026
Patent RE50932
APPARATUS AND METHOD FOR TRANSMITTING AND RECEIVING SIGNAL IN A MOBILE COMMUNICATION SYSTEM
2y 6m to grant Granted Jun 23, 2026
Patent RE50803
DEVICES FOR ENHANCING TRANSMISSIONS OF STIMULI IN AUDITORY PROSTHESES
4y 4m to grant Granted Feb 17, 2026
Patent RE50766
SEMICONDUCTOR MEMORY DEVICE
1y 4m to grant Granted Jan 27, 2026
Patent RE50738
APPARATUS AND METHOD FOR GENERATING A BANDWIDTH EXTENDED SIGNAL
2y 1m to grant Granted Jan 06, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
82%
With Interview (+7.2%)
2y 4m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 222 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month