Detailed Action
This communication is in response to the Request for Continued Examination filed on 05/01/2026.
Claims 1-20 are pending and have been examined.
Claims 1-20 are rejected.
Claims 1 and 11 are independent.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 4/07/2026, 5/06/2026, 08/11/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Arguments and Amendments
Applicant has amended the independent claims to include “speech” “speech” “having machine generated speech of the caller” “speaker, the expected response time generated using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller.”
Regarding the Rejections under 35 U.S.C. 101 Applicant notes Independent claims 1 and 11, as amended, satisfy at least Prong 1 or Prong 2 of Step 2A. The claims are patent-eligible under Step 2A, Prong 1, at least because the claims describe operations that cannot practicably be performed in a human mind.
Examiner notes the claims recite limitations that are largely high level and functional which can be characterized as mental like steps. The claim limitations of “obtaining…”, “detecting…”, “determining …”, “identifying …” can be characterized as mental or data processing concepts. The claim is directed to judicial exceptions (mathematical concepts and mental process grouping)
Applicant notes the claims describe a particular machine-implemented series of operations rather than any abstract idea. Moreover, the claims are patent-eligible under Step 2A, Prong 2, at least because they recite features that describe technical improvements to technological shortcomings described in the Specification and improve computing functioning of existing approaches.
Step 2A, Prong 1 The USPTO's 2024 Guidance Update and MPEP 2106.04(a)(2) explain that a claim does not recite a mental process when it contains features that cannot practically be performed in the human mind, for instance when the human mind is not equipped to perform the claim limitations. As amended, the claims recite operations that the human mind is incapable of performing. Specifically, the human mind cannot determine a response delay from a speech segment of inbound audio data in the manner claimed.
Examiner notes the claims recite mental-like processes that a human can perform. A human can determine a response delay from a speech segment of audio data by listening and counting for the delay.
Applicant notes the human mind cannot execute a voice activity detection engine on inbound audio data to detect instances of speech and segment call audio into agent and caller speech regions. The human auditory and cognitive systems cannot process inbound audio data in the manner claimed to computationally extract timestamps from speech segments, determine response delays between agent and caller speech segments with millisecond precision, and perform this analysis in real-time during an active call. These operations require calculating running statistical measures (e.g., variance, mean, interquartile range) of response delays and adjusting response delay calculations based on dialogue context, operations that require continuous computational processing of streaming audio data that the human mind cannot replicate. Unlike pen-and-paper calculations on pre-recorded data, the claimed method operates on live inbound audio data during an ongoing call, requiring the computer to execute a voice activity detection engine and determine response delays as speech segments occur.
Examiner notes a human can naturally detect audio and use natural language processing to detect speech. A human can extract timestamps using observation. A human can determine response delays using observation in real time.
Applicant notes Step 2A, Prong 2The USPTO's 2024 Guidance Update and MPEP 2106.04(d)(1) explain that, under Step 2A, Prong 2, a claim is patent eligible when any alleged judicial exception (e.g., abstract idea) is integrated into a practical application. A practical application is found where the Specification includes a technical solution to a technical problem, and where the claims contain features describing the technical solution. Id. The independent claims meet this requirement. The Specification describes a concrete technical problem arising in conventional call authentication and fraud detection systems, which are vulnerable to machine-generated speech (deepfakes) and other automated attacks that can defeat traditional speaker verification or call analytics. See, e.g., Specification at [0003]-[0005]. As noted, "[d]eepfake technology has made significant advancements in recent years, enabling the creation of highly realistic, but fake, still imagery, audio playback, and video playback" and "[w]hat is needed is improved means for detecting fraudulent uses of audio-based deepfake technology over telecommunications channels." Id. at [0005]. The Specification describes a technical solution including a multi-stage, computer- implemented process that executes discrete computational operations for: (i) applying a voice activity detection (VAD) engine to parse inbound audio data into speech and non-speech regions; (ii) detecting and diarizing speech segments corresponding to distinct speakers within identified speech regions; (iii) extracting temporal features including response delay metrics derived from timestamp differentials between sequential agent-caller speech segment pairs; and (iv) performing statistical analysis on response delay distributions to discriminate between human-generated and machine-synthesized speech patterns in real-time call center environments. See id. at [0006]- [0009], [0020], [0049]-[0050]. The amended claims directly reflect these technical improvements recited in the Specification. See id. at [0018]-[0020], [0049]-[0050].
Examiner notes the response delay metrics and the statistical analysis regarding speech patterns is absent from the claim language.
Applicant notes independent claim 1, the claims recite a particular sequence of computer-implemented technical operations, including: "obtaining, by a computer, inbound audio data for a call, including a plurality of speech segments corresponding to a dialogue between a caller and an agent by executing a voice activity detection engine for detecting instances of speech of the caller or the agent in the inbound audio data"; "detecting, by the computer executing the voice activity detection engine, from the plurality of speech segments of the inbound audio data, a speech region including a first speech segment corresponding to the agent and a second speech segment corresponding to the caller"; "determining, by the computer, a response delay between the first segment corresponding to the agent and the second segment corresponding to the caller"; and "identifying, by the computer, the caller as a deepfake during the call, in response to determining that the response delay fails to satisfy an expected response time for a human speaker, the expected response time generated using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller." These claim limitations describe the technical solution disclosed in the Specification for improving call center security against deepfake attacks by using response delay analysis to detect machine-generated speech in real-time during calls. Id. at [0008]-[0009], [0143], [0145]. The claims therefore describe a technological solution to technological shortcomings discussed in the Specification.
Examiner notes the response delay metrics and the statistical analysis regarding speech patterns is absent from the claim language. The clams lack specificity tying the components to unconventional architectures, constrained parameterizations, training/regimen steps or demonstrable improvements. the recited elements appear to be routine, conventional uses of generic software components, and therefore fail to supply an inventive concept.
Applicant notes The Examples of the Updated Guidance support this conclusion. The analysis of the pending Application is analogous to the eligible claims in USPTO Example 47, claim 3 (network anomaly detection), and Example 48, claim 2 (speech separation), where the claims were found to integrate the abstract idea into a practical application by reciting a specific improvement to a technical field. Here, the claims recite a particular technical solution to the problem of detecting machine-synthesized speech by implementing a response delay analysis architecture that leverages voice activity detection, temporal feature extraction, and statistical threshold comparison to distinguish between human and deepfake audio sources, which are computational operations that are beyond the capability of the human mind to perform on streaming audio data in real time. As such, the claims satisfy Step 2A, Prong 2, because they describe a technical solution or improvement and therefore integrate any alleged abstract idea into a practical application.
Examiner notes Example 47 was eligible because the claim included steps that changed system behavior in a concrete, real time way (dropping/blocking packets) and that improvement was reflected in claim steps beyond mere data analysis. Here, claim 1 culminates in identifying a deepfake; while useful, identifying a deepfake is a data analysis task and not per se equivalent to controlling or altering physical system behavior or computer resources in the way Example 47’s real time packet mitigation did.
Examiner notes Example 48 included more explicit problem solution language and concrete mathematical steps and demonstrated how the claim’s pipeline provided an improvement over prior art speech separation by reciting clustering, masking, and resynthesis steps tied to the problem explanation. Example 48’s claim recited particular data transformations and a pipeline that the is treated as integrating the exception into a practical application; the present claim is similar in domain but, as written, is less explicit about the particular technical steps (e.g., precise alignment algorithm, constrained architecture, or demonstrable technical effect) that would make the improvement evident from claim language alone.
Applicant’s arguments and amendments with respect to the 35 U.S.C. § 101 rejections have been fully considered, but are not persuasive.
Applicant’s arguments with respect to Claims 1, 8-11 and 18-20 have been considered but are moot because the new ground of rejection does not rely on the primary reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Hence, new grounds for rejection have been made over HAVDAN (US Patent Number US 20230206925 A1), in view of Wang (US Patent Number US 11848029 B2)
As to the rejections of Claims 2-4 and 12-14, and Claims 5, 6, 15 and 16, and Claims 7 and 17 under 35 U.S.C. § 103 Applicant notes
Applicant respectfully requests reconsideration and withdrawal of the rejection in view of the amendments submitted with this paper and the comments made during the Examiner Interview. The references fail to teach or suggest at least the features of "determining, by the computer, a response delay between the first speech segment corresponding to the agent and the second speech segment corresponding to the caller," and "identifying, by the computer, the caller as a deepfake having machine generated speech of the caller during the call, in response to determining that the response delay fails to satisfy an expected response time for a human speaker, the expected response time generated using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller" as recited in independent claims 1 and 11.As discussed and tentatively agreed upon in the Examiner Interview, the references do not suggest determining a response delay between a first speech segment (of an agent) and a second speech segment (of a caller), and then identifying the caller as a deepfake using the response delay compared against an expected response time for a human speaker, where the expected response time is generated using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller, as in the independent claims. As such, independent claims 1 and 11, and the claims depending therefrom, are patentable over the combination of Dopuljic and Phatak.
Examiner notes the independent claims were discussed during the interview and it was agreed that the proposed amendments overcame the previously applied prior art for the claims discussed.
Applicant’s arguments with respect to the rejection(s) of Claims 2-4 and 12-14 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made over HAVDAN (US Patent Number US 20230206925 A1), in view of Wang (US Patent Number US 11848029 B2) and further in view of Kolbegger (US Patent Number US 20140270114 A1).
Applicant’s arguments with respect to the rejection(s) of Claims 5, 15 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of HAVDAN (US Patent Number US 20230206925 A1), in view of Wang (US Patent Number US 11848029 B2) and further in view of Kolbegger (US Patent Number US 20140270114 A1), and further in view of Gross (US Patent Number US 20160328547 A1).
Applicant’s arguments with respect to the rejection(s) of Claims 6 and 16 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of HAVDAN (US Patent Number US 20230206925 A1), in view of Wang (US Patent Number US 11848029 B2) and further in view of Gross (US Patent Number US 20160328547 A1).
Applicant’s arguments with respect to the rejection(s) of Claims 7 and 17 under 35 U.S.C. § 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of HAVDAN (US Patent Number US 20230206925 A1), in view of Wang (US Patent Number US 11848029 B2) and further in view of Newstadt (US Patent Number US-10657971-B1).
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding independent Claim 1, the claim recites “1. A computer-implemented method for detecting machine-based speech in calls, comprising:
obtaining, by a computer, inbound audio data for a call, including a plurality of speech segments corresponding to dialogue between a caller and an agent by executing a voice activity detection engine for detecting instances of speech of the caller or the agent in the inbound audio data;
detecting, by the computer, executing the voice activity detection engine, from the plurality of speech segments of the inbound audio data, a speech region including a first speech segment corresponding to the agent and a second speech segment corresponding to the caller;
determining, by the computer, a response delay between the first speech segment corresponding to the agent and the second speech segment corresponding to the caller;
and identifying, by the computer, the caller as a deepfake having machine generated speech of the caller during the call, in response to determining that the response delay fails to satisfy an expected response time for a human speaker, the expected response time generated using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller.”
The limitations of “obtaining…”, “detecting…”, “determining …”, “identifying …” as drafted covers a human mental activity or process.
More specifically, a human is capable of obtaining, inbound audio data for a call, including a plurality of speech segments corresponding to a dialogue between a caller and an agent using the human auditory system.
A human is capable of detecting, from the plurality of speech segments of the inbound audio data, a speech region including a first speech segment corresponding to the agent and a second speech segment corresponding to the caller using the human auditory system and natural language processing within the human mind.
A human is capable of determining, a response delay between the first speech segment corresponding to the agent and the second speech segment corresponding to the caller using natural language processing and logic and reasoning within the human mind.
A human is capable of identifying, the caller as a deepfake, having machine generated speech of the caller during the call, in response to determining that the response delay fails to satisfy an expected response time for a human speaker using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller using logic and reasoning within the human mind.
Regarding independent Claim 11, Claim 11 is a System claim with limitations similar to that of claim 1 and is rejected under the same rationale.
This judicial exception is not integrated into a practical application. In particular, claims 1 and 11 recites the additional element of “processor” and “computer” as per the independent claims. For example, in [0339] of the as filed specification, there is description of a processor-executable software module which may reside on a computer-readable or processor-readable storage medium… instructions or data structures and that may be accessed by a computer or processor. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a processor and computer is noted as a general computer as noted. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Further, the additional limitation in the claims noted above are directed towards insignificant solution activity. The claims are not patent eligible.
With respect to claims 2 and 12, the claims relate to identifying based on the plurality of timestamps, the response delay corresponding to a time difference between the first speech segment corresponding to a question by the agent and the second speech segment corresponding to an answer by the caller. This relates to a human performing natural language understanding and logic and reasoning within the human mind to distinguish a delay corresponding to a question by the agent and the second speech segment corresponding to an answer by the caller. No additional limitations are present.
With respect to claims 3 and 13, the claims relate to detecting a plurality of timestamps defining the first speech segment and the second speech segment in the speech region. This relates to a human performing natural language understanding and logic and reasoning within the human mind or pen and paper to define speech segments with timestamps. No additional limitations are present.
With respect to claims 4 and 14, the claims relate to determining a plurality of statistical measures of response delays based on a plurality of timestamps derived from the plurality of speech segments of the inbound audio data. This relates to a human performing natural language understanding and logic and reasoning within the human mind or pen and paper to apply statistical measures with timestamps. Statistical measurements are a mathematical process. No additional limitations are present.
With respect to claims 5 and 15, the claims relate to at least one of: (i) a running variance, (ii) a running inter-quartile range, or (iii) a running mean. This is a mathematical process. No additional limitations present.
With respect to claims 6 and 16, the claims relate to determining the response delay includes adjusting the response delay based on a context of the dialogue between the caller and the agent. This relates to a human performing natural language understanding and logic and reasoning within the human mind to determine that a response delay includes an adjustment. No additional limitations present.
With respect to claim 7 and 17, the claims relate to extracting a transcription containing text in chronological sequence of each caller speech segment and each agent speech segment.
This relates to a human performing natural language understanding and pen and paper to extract a transcription. No additional limitations present.
With respect to claim 8 and 18, the claims relate to determining the expected response time for the human speaker based on historical data of audio data between callers and agents. This relates to a human using natural language understanding and logic and reasoning to determine an expected response. No additional limitations present.
With respect to claim 9 and 19, the claims relate to identifying the caller as human, in response to determining that the response delay satisfies an expected response time for a human speaker. This relates to a human performing natural language understanding and logic and reasoning to determine and identify, using the delay, that a caller is human. No additional limitations present.
With respect to claims 10 and 20, the claims relate to generating an indication indicating the caller as one of the deepfake or human. This relates to a human using pen and paper or the vocal system to indicate that a caller is a deepfake or human. No additional limitations present.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 8-11 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over HAVDAN (US Patent Number US 20230206925 A1) in view of Wang (US Patent Number US 11848029 B2).
Regarding claim 1, HAVDAN teaches 1. (Currently Amended) A computer-implemented method for detecting machine- based speech in calls, comprising: obtaining, by a computer, inbound audio data for a call, (see HAVDAN [0031] Embodiments of the invention may provide an automatic method for spoofing detection. Embodiments of the invention may train a neural network (NN) model, or a deep learning (DL) NN model, for classifying input voice samples into genuine voice samples or spoofed voice samples, based on an input voice sample of as little as, in some embodiments three seconds of net voice recording (other amounts may be used)...”) (See Havdan [0034] “ performed by for example a conventional computer or GPU…”) (see HAVDAN [0035] An interaction may be a conversation or exchange of information between two or more people, for example a speaker or caller, e.g., a customer, person or fraudster on the one hand, and a call center agent or another person on the other hand. An interaction may be for example a telephone call or a voice over Internet Protocol (VoIP) call. An interaction may describe both the event of the conversation and also the data associated with that conversation, such as an audio recording of the conversation and metadata. …(examiner interprets inbound audio data as “interaction…telephone call”) “) including a plurality of speech segments corresponding to a dialogue between a caller and an agent (see HAVDAN [0035] “An interaction may be a conversation (examiner interprets plurality of speech segments as “conversation”) or exchange of information between two or more people, for example a speaker or caller, e.g., a customer, person or fraudster on the one hand, and a call center agent or another person on the other hand.”) by executing a voice activity detection engine for detecting instances of speech of the caller or the agent in the inbound audio data; (see HAVDAN [0066] In operation 810, a processor (e.g., processor 105) may detect speech in the voice sample, e.g., using VAD. In operation 820,…”) detecting, by the computer executing the voice activity detection engine, from the plurality of speech segments of the inbound audio data, a speech region (see HAVDAN [0066] In operation 810, a processor (e.g., processor 105) may detect speech in the voice sample, e.g., using VAD. In operation 820, the processor may remove parts of the voice sample that do not include speech, to generate a sample that is long enough (e.g., above 3 seconds) and includes only voiced parts. In operation 830, the processor may divide the voice sample into partially overlapping frames of a predetermined size, e.g., 200-600 samples. In operation 840, the processor may extract frequency-based features from the frames, e.g., by performing fast Fourier transform (FFT). In operation 850, the processor may apply convolutional layers on the frequency-based features. In operation 860, the processor may unify the results of the convolutional layers to obtain the set of features, e.g., using recurrent layers.”) (see HAVDAN [0040] Spoofing detection engine 40 may check the audio recording of a speaker part of the interaction (e.g., the part of the interaction not attributed to the agent of call center 10) for spoofing as disclosed herein. For example, spoofing detection engine 40 may obtain a data sample such as a voice sample, e.g., taken from the audio recording of a speaker part of the interaction, extract a set of features (e.g., organized as a feature matrix or a feature vector) from the voice sample, and provide or input the set of features of the voice sample to a trained neural network model (or DL NN model) to obtain a classification of the voice sample to spoofing (e.g. fake, not the person it purports to be from, or not from a real person) or genuine (e.g. from a live person, or from the person it purports to be from).”) including a first speech segment corresponding to the agent and a second speech segment corresponding to the caller; (see HAVDAN [0035] An interaction may be a conversation or exchange of information between two or more people, for example a speaker or caller, e.g., a customer, person or fraudster on the one hand, and a call center agent or another person on the other hand. An interaction may be for example a telephone call or a voice over Internet Protocol (VoIP) call. An interaction may describe both the event of the conversation and also the data associated with that conversation, such as an audio recording of the conversation and metadata. Metadata describing an interaction may include information beyond the exchange itself: metadata may include for example the telephone number of the customer or fraudster (e.g. gathered via ANI, automatic number identification), an internet protocol (IP) address associated with equipment used by the speaker or fraudster, account information relevant to or discussed in the interaction, the purported name or identification (ID) of the person, audio information related to the interaction such as background noise, and other information. “) (see HAVDAN [0050] “In one embodiment, DL NN model 300 may obtain or have input to it a training dataset including labeled feature sets extracted from net-speech samples, e.g., each sample may include a set of features extracted by block 220 and a label. A label of set of features may indicate whether the sample is spoof or genuine. A spoof sample may either include voice recording (e.g., instead of live speech) or synthetic speech. [0051] According to embodiments of the invention, the loss function may include a regulation factor, for example, the regulation factor may be added to a standard loss function (e.g., a cross-entropy function or other standard loss function). According to embodiments of the invention, the regulation factor may be set to zero for voice samples labeled as genuine and may be proportional (e.g., directly proportional) to the prediction of DL NN model 300 for data samples labeled as spoofing.”)
Havdan does not specifically teach determining, by the computer, a response delay between the first speech segment corresponding to the agent and the second speech segment corresponding to the caller; and However, Wang does teach this limitation (see Wang (8:30-45) “(58) When performing the detection, the audio to be detected may be divided into the speech segment and the non-speech segment, and the audio feature of the speech segment is inputted into the first real sound model and the first attack sound model respectively to obtain a probability score indicating that the speech segment belongs to the real sound and a probability score indicating that the speech segment belongs to the attack sound. The audio feature of the non-speech segment is inputted into the second real sound model and the second attack sound model respectively to obtain a probability score indicating that the non-speech segment belongs to the real sound and a probability score indicating that the non-speech segment belongs to the attack sound. identifying, by the computer, the caller as a deepfake having machine generated speech of the caller during the call, in response to determining that the response delay fails to satisfy an expected response time for a human speaker, the expected response time generated using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller. (see Wang (8:50-65) (61) Similar to the face recognition, in-vivo detection is needed for the voiceprint recognition, so as to determine whether the sound is real human voice or fake sound. The voiceprint recognition system may be attacked by the following ways, a. waveform splicing attack, b. record and replay attack, c. speech synthesizing attack, d. speech transforming attack, e. speech simulating attack. The record attack is easy to implement without any professional knowledge or specific hardware and difficult to detect. The attacker merely needs to record a speech of a target speaker by using a phone and replay the speech to pose as the target speaker to pass the authentication of the voiceprint recognition system.”)(see Wang (4:15:25) “(24) In some embodiments, obtaining the speech segment and the non-speech segment of the audio to be detected includes recognizing a first silent segment in the audio using a first recognition way, recognizing an unvoiced sound segment and a second silent segment in the audio using a second recognition way, determining a union of the unvoiced sound segment, the first silent segment and the second silent segment as the non-speech segment, and determining an audio segment other than the non-speech segment in the audio as the speech segment.”)(see Wang (6:1-10) “(38) The first detection score and the second detection score may form a final score set, or weighted averaging may be performed based on the first detection score and the second detection score to obtain the final score of the audio to be detected. A determination criterion determined based on a precision requirement, an application scenario, historical data and so on may be used to determine whether the final score of the audio to be detected is a score of real sound or a score of attack sound.”)
HAVDAN and Wang are in the same field of endeavor of signal processing, therefore It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of HAVDAN to incorporate the teachings of Wang to include determining, by the computer, a response delay between the first speech segment corresponding to the agent and the second speech segment corresponding to the caller; and identifying, by the computer, the caller as a deepfake having machine generated speech of the caller during the call, in response to determining that the response delay fails to satisfy an expected response time for a human speaker, the expected response time generated using prior stored audio data including historical response delay data between a historical agent and at least one of a human caller or a machine generated caller. Doing so allows for improved audio accuracy of attack audio as recognized by Wang in (4:7-15).
Regarding independent Claim 11, Claim 11 is a System claim with limitations similar to that of claim 1 and is rejected under the same rationale. Furthermore, HAVDAN teaches A system for detecting machine-based speech in calls, comprising: a computer having one or more processors and configured to: (see HAVDAN [0070] Memory 120 may be or may include, for example, a Random Access Memory (RAM), a read only memory (ROM), a Dynamic RAM (DRAM), a Synchronous DRAM (SD-RAM), a double data rate (DDR) memory chip, a Flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory or storage units. Memory 120 may be or may include a plurality of, possibly different memory units. Memory 120 may be a computer or processor non-transitory readable medium, or a computer non-transitory storage medium, e.g., a RAM.”)
Regarding claim 8, HAVDAN IN VIEW OF WANG teaches 8. The method of claim 1,
Furthermore, HAVDAN teaches, further comprising determining, by the computer, the expected response time for the human speaker based on historical data of audio data between callers and agents. (see HAVDAN [0044] “Reference is now made to FIG. 2, which is a schematic illustration of spoofing detection engine 40, according to embodiments of the present invention. Spoofing detection engine 40 may include silence removal block 200, feature extractor 210 and classifier 220, which may be a NN. Silence removal block 200 may remove silent periods, or periods that do not contain speech, from the voice sample, for example using voice activity detector (VAD), as known in the art. Using embodiments of the invention, a voiced sample with as little as three seconds of voice, or other durations, may be sufficient for detecting spoofing calls. Silence removal block 200 may generate net-speech samples of three seconds or longer from the voice sample. Silence removal block 200 may provide an indication, e.g., to analyst terminal 14, in case the voice sample does not include enough speech, e.g., in case the net-speech sample is shorter than a threshold duration, e.g., three seconds or other durations. Other threshold durations may be used.”)
Regarding claim 9, HAVDAN IN VIEW OF WANG teaches 9. The method of claim 1,
Furthermore, HAVDAN teaches, further comprising identifying, by the computer, the caller as human, in response to determining that the response delay satisfies an expected response time for a human speaker. (see HAVDAN [0044] “Reference is now made to FIG. 2, which is a schematic illustration of spoofing detection engine 40, according to embodiments of the present invention. Spoofing detection engine 40 may include silence removal block 200, feature extractor 210 and classifier 220, which may be a NN. Silence removal block 200 may remove silent periods, or periods that do not contain speech, from the voice sample, for example using voice activity detector (VAD), as known in the art. Using embodiments of the invention, a voiced sample with as little as three seconds of voice, or other durations, may be sufficient for detecting spoofing calls. Silence removal block 200 may generate net-speech samples of three seconds or longer from the voice sample. Silence removal block 200 may provide an indication, e.g., to analyst terminal 14, in case the voice sample does not include enough speech, e.g., in case the net-speech sample is shorter than a threshold duration, e.g., three seconds or other durations. Other threshold durations may be used.”)
Regarding claim 10, HAVDAN IN VIEW OF WANG teaches 10. The method of claim 1,
Furthermore, HAVDAN teaches, further comprising generating, by the computer, an indication, for a user interface, indicating the caller as one of the deepfake or human. (see HAVDAN [0064] In operation 750, the processor may extract a set of features from the new voice sample. In operation 760, the processor may classify the new voice sample, using the neural network model trained in operation 730. For example, the processor may provide the set of features to the trained neural network model that may provide a score indicative of the chances of the voice sample being genuine or spoof. The score may be compared against a threshold to provide a classification. In operation 770, the processor may provide the classification to a human operator, e.g., via analyst terminal 14, or to another computerized process, as required.”)
As to claim 18, claim 18 is a system claim with limitations similar to that of Claim 8 and is rejected under the same rationale.
As to claim 19, claim 19 is a system claim with limitations similar to that of Claim 9 and is rejected under the same rationale.
As to claim 20, claim 20 is a system claim with limitations similar to that of Claim 10 and is rejected under the same rationale.
Claims 2-4 and 12-14 are rejected under 35 U.S.C. 103 as being unpatentable over HAVDAN (US Patent Number US 20230206925 A1), in view of Wang (US Patent Number US 11848029 B2) and further in view of Kolbegger (US Patent Number US 20140270114 A1).
As to Claim 2, HAVDAN IN VIEW OF WANG teaches 2. The method of claim 1,
HAVDAN IN VIEW OF WANG does not specifically teach wherein determining the response delay includes identifying, by the computer, based on the plurality of timestamps, the response delay corresponding to a time difference between the first speech segment corresponding to a question by the agent and the second speech segment corresponding to an answer by the caller. However, Kolbegger does teach this limitation, (see Kolbegger [0066] “In FIG. 4B, a waveform 410B represents the time-in-speech on the agent channel and a waveform 420B represents the time-in-speech on the client channel. The waveform 410B shows a longer time-in-speech, which tends to indicate to the facility that the agent is asking preliminary information identifying the client such as, "With whom am I speaking to," "Can you spell that?", and/or "May I have your order number?" Moreover, the short pauses between the agent's and caller's time-in-speech tend to indicate to the facility that either the client is searching for information and/or that the agent is pulling up, looking for or reviewing information and/or putting the client on hold. The waveform 420B shows a shorter time-in-speech, which tends to indicate to the facility that the client is responding to the agents' preliminary questions.”) (see Kolbegger [0046] “At block 340, the clustered-frame representation or manifests for each channel are correlated by the facility 140 such that the visual representation of the client channel and the visual representation of the agent channel are presented together in one representation that depicts both channels simultaneously. The two channels may be correlated by a variety of different means, including syncing a start time for the client channel with a start time for the agent channel, by relying on embedded or captured time stamps at various points in the audio, etc.”)
HAVDAN IN VIEW OF WANG and Kolbegger are in the same field of endeavor of signal processing, therefore It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of HAVDAN IN VIEW OF WANG to incorporate the teachings of Kolbegger to include determining the response delay includes identifying, by the computer, based on the plurality of timestamps, the response delay corresponding to a time difference between the first speech segment corresponding to a question by the agent and the second speech segment corresponding to an answer by the caller. Doing so allows for a particular speech or call feature to more easily be distinguished from another speech or call feature as recognized by Kolbegger in [0044].
As to Claim 3, HAVDAN IN VIEW OF WANG teaches 3. The method of claim 1,
HAVDAN IN VIEW OF WANG does not specifically teach wherein detecting the speech region further includes detecting, by the computer, a plurality of timestamps defining the first speech segment and the second speech segment in the speech region. However, Kolbegger does teach this limitation, (see Kolbegger [0066] “In FIG. 4B, a waveform 410B represents the time-in-speech on the agent channel and a waveform 420B represents the time-in-speech on the client channel. The waveform 410B shows a longer time-in-speech, which tends to indicate to the facility that the agent is asking preliminary information identifying the client such as, "With whom am I speaking to," "Can you spell that?", and/or "May I have your order number?" Moreover, the short pauses between the agent's and caller's time-in-speech tend to indicate to the facility that either the client is searching for information and/or that the agent is pulling up, looking for or reviewing information and/or putting the client on hold. The waveform 420B shows a shorter time-in-speech, which tends to indicate to the facility that the client is responding to the agents' preliminary questions.”) (see Kolbegger, [0046] “At block 340, the clustered-frame representation or manifests for each channel are correlated by the facility 140 such that the visual representation of the client channel and the visual representation of the agent channel are presented together in one representation that depicts both channels simultaneously. The two channels may be correlated by a variety of different means, including syncing a start time for the client channel with a start time for the agent channel, by relying on embedded or captured time stamps at various points in the audio, etc.”)
HAVDAN IN VIEW OF WANG and Kolbegger are in the same field of endeavor of signal processing, therefore It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of HAVDAN IN VIEW OF WANG to incorporate the teachings of Kolbegger to include wherein detecting the speech region further includes detecting, by the computer, a plurality of timestamps defining the first speech segment and the second speech segment in the speech region. Doing so allows for a particular speech or call feature to more easily be distinguished from another speech or call feature as recognized by Kolbegger in [0044].
As to Claim 4, HAVDAN IN VIEW OF WANG teaches 4. The method of claim 1,
HAVDAN IN VIEW OF WANG does not specifically teach wherein determining the response delay includes determining, by the computer, a plurality of statistical measures (see Kolbegger [0065] “FIG. 4B depicts a clustered-frame representation of an example channel interaction pattern 400B known as an "exchange of personal data." A measure of time 430B (e.g., in milliseconds, seconds, minutes, etc.) is represented along the x-axis. The "exchange of personal data" represented by the pattern 400B typically occurs at the beginning of a call and reflects a lot of back- and forth in time-of-speech between the agent-channel and the client-channel.”) of response delays based on a plurality of timestamps derived from the plurality of speech segments of the inbound audio data. However, Kolbegger does teach this limitation, (see Kolbegger [0066] “In FIG. 4B, a waveform 410B represents the time-in-speech on the agent channel and a waveform 420B represents the time-in-speech on the client channel. The waveform 410B shows a longer time-in-speech, which tends to indicate to the facility that the agent is asking preliminary information identifying the client such as, "With whom am I speaking to," "Can you spell that?", and/or "May I have your order number?" Moreover, the short pauses between the agent's and caller's time-in-speech tend to indicate to the facility that either the client is searching for information and/or that the agent is pulling up, looking for or reviewing information and/or putting the client on hold. The waveform 420B shows a shorter time-in-speech, which tends to indicate to the facility that the client is responding to the agents' preliminary questions.”) (see Kolbegger, [0046] “At block 340, the clustered-frame representation or manifests for each channel are correlated by the facility 140 such that the visual representation of the client channel and the visual representation of the agent channel are presented together in one representation that depicts both channels simultaneously. The two channels may be correlated by a variety of different means, including syncing a start time for the client channel with a start time for the agent channel, by relying on embedded or captured time stamps at various points in the audio, etc.”)
HAVDAN IN VIEW OF WANG and Kolbegger are in the same field of endeavor of signal processing, therefore It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of HAVDAN IN VIEW OF WANG to incorporate the teachings of Kolbegger to include determining the response delay includes determining, by the computer, a plurality of statistical measures. Doing so allows for a particular speech or call feature to more easily be distinguished from another speech or call feature as recognized by Kolbegger in [0044].
As to claim 12, claim 12 is a system claim with limitations similar to that of Claim 2 and is rejected under the same rationale.
As to claim 13, claim 13 is a system claim with limitations similar to that of Claim 3 and is rejected under the same rationale.
As to claim 14, claim 14 is a system claim with limitations similar to that of Claim 4 and is rejected under the same rationale.
Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over HAVDAN (US Patent Number US 20230206925 A1), in view of Wang (US Patent Number US 11848029 B2) and further in view of Kolbegger (US Patent Number US 20140270114 A1), and further in view of Gross (US Patent Number US 20160328547 A1).
As to Claim 5, HAVDAN IN VIEW OF WANG in view of Kolbegger teaches 5. The method of claim 4,
HAVDAN IN VIEW OF WANG in view of Kolbegger do not specifically teach wherein the plurality of statistical measures further comprises at least one of: (i) a running variance, (ii) a running inter-quartile range, or (iii) a running mean. However, GROSS does teach this limitation (see GROSS [0179] “Through sufficient training samples the system should be able to identify, with controllable confidence levels, an appropriate set of sentences that are likely to weed out a machine imposter. Moreover candidate sentences which are confusing or take too long for a human can be eliminated as well. Again it is preferable that the challenge sentences include primarily samples that are rapidly processed and articulated as measured against a human reference set. The respective times required by human and machines can also be measured and compiled to determine minimum, maximum, average, mean and threshold times. For example, it may be desirable to select challenge sentences in which the time difference between human and machine articulations is greatest.”)
HAVDAN IN VIEW OF WANG in view of Kolbegger and Gross are in the same field of endeavor of signal processing, therefore It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of combination of HAVDAN IN VIEW OF WANG, and Kolbegger to incorporate the teachings of Gross to include the plurality of statistical measures further comprises at least one of: (i) a running variance, (ii) a running inter-quartile range, or (iii) a running mean. Doing so allows for a better selection and optimization for discriminating against machines as recognized by Gross in [0006].
Claims 6, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over HAVDAN (US Patent Number US 20230206925 A1) in view of Wang (US Patent Number US 11848029 B2) and further in view of Gross (US Patent Number US 20160328547 A1).
Regarding claim 6, HAVDAN IN VIEW OF WANG teaches 6. The method of claim 1,
HAVDAN IN VIEW OF WANG does not specifically teach wherein determining the response delay includes adjusting, by the computer, the response delay based on a context of the dialogue between the caller and the agent. However, GROSS does teach this limitation (see Gross [0083] After understanding the sentence, the machine imposter may also have to annotate the output of the desired articulation with appropriate prosodic elements at step 530. Acoustical aspects of prosodic structure also include modulation of fundamental frequency (F0), energy, relative timing of phonetic segments and pauses, and phonetic reduction or modification. For example a phoneme may have a variable duration which is highly dependent on context, such as preceding and following phonemes, phrase boundaries, word stress, phrase boundaries, etc.”)
HAVDAN IN VIEW OF WANG and Gross are in the same field of endeavor of signal processing, therefore It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of HAVDAN IN VIEW OF WANG, to incorporate the teachings of Gross to include determining the response delay includes adjusting, by the computer, the response delay based on a context of the dialogue between the caller and the agent. Doing so allows for a better selection and optimization for discriminating against machines as recognized by Gross in [0006].
As to claim 15, claim 15 is a system claim with limitations similar to that of Claim 5 and is rejected under the same rationale.
As to claim 16, claim 16 is a system claim with limitations similar to that of Claim 6 and is rejected under the same rationale.
Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over HAVDAN (US Patent Number US 20230206925 A1) in view of Wang (US Patent Number US 11848029 B2) and further in view of Newstadt (US Patent Number US-10657971-B1).
Regarding claim 7, HAVDAN IN VIEW OF WANG teaches 7. The method of claim 1,
HAVDAN IN VIEW OF WANG does not specifically teach further comprising extracting, by the computer, a transcription containing text in chronological sequence of each caller speech segment and each agent speech segment. However, Newstadt does teach this limitation (see Newstadt (12:55-13:14) “(59) As explained in connection with FIGS. 3 and 5, the systems and methods described herein may identify fraudulent and/or nuisance phone calls by using data collected by client applications installed within a community of users. The systems and methods described herein may collect data that captures identifying characteristics of the call as well as the user community's response to the call. In some embodiments, the systems and methods described herein may combine data across the community into a call reputation rating and then provide the reputation rating back to users to identify suspicious calls as the calls are received. In some embodiments, users may install the systems described herein as a client application on their device. In various embodiments, this client application may be a mobile application on a phone or a desktop application on a laptop or desktop. In some embodiments, the application may listen in on all audio or video calls made to the user's device. In some examples, the application may listen for the quality of the call, amount of time between answering the call and the caller speaking, time lag between the user speaking and the caller responding, background noise, and the tone, qualities, and emotional content of the caller's voice. In some embodiments, using voice recognition, the application may also monitor the content of the call itself, such as the specific words the caller is using as they initiate the conversation. For example, does the caller say, “Hi John,” or does the caller say, “for a small fee?””)
HAVDAN IN VIEW OF WANG and Newstadt are in the same field of endeavor of signal processing, therefore It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of HAVDAN IN VIEW OF WANG to incorporate the teachings of Newstadt to include extracting, by the computer, a transcription containing text in chronological sequence of each caller speech segment and each agent speech segment. Doing so allows for improve the functioning of a computing device by detecting potentially malicious calls with increased accuracy and thus reducing the computing device's user's likelihood of victimization by malicious callers as recognized by Newstadt in (4:6-9).
As to claim 17, claim 17 is a system claim with limitations similar to that of Claim 7 and is rejected under the same rationale.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KRISTEN MICHELLE MASTERS whose telephone number is (703)756-1274. The examiner can normally be reached M-F 8:30 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KRISTEN MICHELLE MASTERS/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659