Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/22/26 has been entered.
This office action is in response to correspondence 05/22/26 and 06/24/26 regarding application 18/647,410, in which claims 1, 10, 14, 17, and 19 were amended. Claims 1-20 are pending in the application and have been considered.
Response to Arguments
Applicant’s arguments on pages 9-10 regarding the 35 U.S.C. 103 rejections of the claims based on Ramadas, Mossoba, and Huffman have been considered but are moot in view of the new grounds for rejection based in part on the newly discovered reference to Patino et al. (“Speaker anonymisation using the McAdams coefficient”. arXiv:2011.01130v2 [eess.AS] 1 Sep 2021) which discloses using a random value both to obfuscate voice by altering voice characteristics and to provide proof of origin, authenticity, or integrity of audio data of the user, similarly to the amended language of independent claims 1, 10, and 17.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6 and 8-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ramadas et al. (US 20230269291) in view of Mossoba et al. (US 20230252190), in further view of Patino et al. (“Speaker anonymisation using the McAdams coefficient”. arXiv:2011.01130v2 [eess.AS] 1 Sep 2021).
Consider claim 1, Ramadas discloses a method (a method, [0160]), comprising:
detecting, by a device, a voice involving a user (user utters a spoken request to visit a healthcare provider, [0026], which is detected by local voice assistant, [0025], [0026]);
identifying, by the device, a usage of user-specific language or vocabulary based on a usage pattern (identifying that the user spoke their SSN based on detecting the pattern “SSN is” preceding it, [0029-0030]);
generating, by the device and based on the user-specific language or vocabulary, one or more replacement words to replace words spoken by the user during the voice (generating replacement words for the SSN, [0088]-[0090], from words in replacement vocabulary, [0065]);
generating, by the device, a random value to be applied to the voice call to create voice obfuscation for the voice, wherein the random value is used to obfuscate the voice (using a timed volatile random subset (t-VRS) to generate replacement words including natural language words, numbers, phrases, and the like to obfuscate the sensitive-information utterance, [0065] from the set of global replacement values, [0089]; this is considered to generate a random value from the set of global values); and
communicating, by the device, encoded data associated with the call, wherein the encoded data is in accordance with the voice obfuscation (encoder uses replacement utterance to generate an encoded spectrogram, [0113], which decoder and vocoder transform into a time domain speech waveform, [0115], which is transmitted to server 114, [0091]).
Ramadas does not specifically mention: a voice call; obfuscate one or more voice characteristics of the voice call.
Mossoba discloses a voice call (audio call example, Fig. 1D, [0043], [0047]); obfuscate one or more voice characteristics of the voice call (obfuscating by altering an audio frequency or amplitude associated with the section of the outgoing communication that includes the sensitive information, [0047]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas by utilizing it during a voice call as in Mossoba, and obfuscating one or more voice characteristics of the voice call in order to thwart malicious actions (Mossoba, [0011]) while still allowing communication without the sensitive information in situations which are not clearly fraud or not fraud (Mossoba, [0013]). Doing so would have led to predictable results of increased flexibility, as suggested by Mossoba ([0015]). The references cited are analogous art in the same field of audio processing.
Ramadas and Mossoba do not specifically mention applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user.
Patino discloses applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user (random McAdams coefficients, which are used to manipulate the formant positions, i.e. voice characteristics, of speech signals, drawn from different ranges of the same uniform distribution are used to anonymize the speech, the randomly chosen McAdams coefficient speaker-dependent for each speaker, i.e. proving that that the anonymized speech originated with a specific speaker, Sections 3.1, 3.2, pages 2-3)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas and Mossoba by applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user in order to improve protection from undesired reversibility, as suggested by Patino (page 1, Section 1), predictably helping to protect personally identifiable information in speech, a goal identified by Patino (page 1, Section 1). The references cited are analogous art in the same field of audio processing.
Consider claim 10, Ramadas discloses a device (devices, [0181]), comprising:
one or more processors (one or more processors, [0184]) configured to:
detect a voice involving a user (user utters a spoken request to visit a healthcare provider, [0026], which is detected by local voice assistant, [0025], [0026]);
identify a usage of user-specific language or vocabulary based on a usage pattern (identifying that the user spoke their SSN based on detecting the pattern “SSN is” preceding it, [0029-0030]);
generate, based on the user-specific language or vocabulary, one or more replacement words to replace words spoken by the user during the voice (generating replacement words for the SSN, [0088]-[0090], from words in replacement vocabulary, [0065]);
generate a random value to be applied to the voice call to create voice obfuscation for the voice, wherein the random value is used to obfuscate the voice (using a timed volatile random subset (t-VRS) to generate replacement words including natural language words, numbers, phrases, and the like to obfuscate the sensitive-information utterance, [0065] from the set of global replacement values, [0089]; this is considered to generate a random value from the set of global values); and
communicate encoded data associated with the call, wherein the encoded data is in accordance with the voice obfuscation (encoder uses replacement utterance to generate an encoded spectrogram, [0113], which decoder and vocoder transform into a time domain speech waveform, [0115], which is transmitted to server 114, [0091]).
Ramadas does not specifically mention: a voice call; obfuscate one or more voice characteristics of the voice call.
Mossoba discloses a voice call (audio call example, Fig. 1D, [0043], [0047]); obfuscate one or more voice characteristics of the voice call (altering an audio frequency or amplitude associated with the section of the outgoing communication that includes the sensitive information, [0047]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas by utilizing it during a voice call as in Mossoba, and obfuscating one or more voice characteristics of the voice call for reasons similar to those for claim 1.
Ramadas and Mossoba do not specifically mention applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user.
Patino discloses applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user (random McAdams coefficients, which are used to manipulate the formant positions, i.e. voice characteristics, of speech signals, drawn from different ranges of the same uniform distribution are used to anonymize the speech, the randomly chosen McAdams coefficient speaker-dependent for each speaker, i.e. proving that that the anonymized speech originated with a specific speaker, Sections 3.1, 3.2, pages 2-3)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas and Mossoba by applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user for reasons similar to those for claim 1.
Consider claim 17, Ramadas discloses a non-transitory computer-readable medium storing a set of instructions (non-transitory media with instructions, [0182]), the set of instructions comprising: one or more instructions that, when executed by one or more processors of a device (executed by processing circuitry of a device, [0182]), cause the device to:
detect a voice involving a user (user utters a spoken request to visit a healthcare provider, [0026], which is detected by local voice assistant, [0025], [0026]);
identify a usage of user-specific language or vocabulary based on a usage pattern (identifying that the user spoke their SSN based on detecting the pattern “SSN is” preceding it, [0029-0030]);
generate, based on the user-specific language or vocabulary, one or more replacement words to replace words spoken by the user during the voice (generating replacement words for the SSN, [0088]-[0090], from words in replacement vocabulary, [0065]);
generate a random value to be applied to the voice call to create voice obfuscation for the voice, wherein the random value is used to obfuscate the voice (using a timed volatile random subset (t-VRS) to generate replacement words including natural language words, numbers, phrases, and the like to obfuscate the sensitive-information utterance, [0065] from the set of global replacement values, [0089]; this is considered to generate a random value from the set of global values); and
communicate encoded data associated with the call, wherein the encoded data is in accordance with the voice obfuscation (encoder uses replacement utterance to generated an encoded spectrogram, [0113], which decoder and vocoder transform into a time domain speech waveform, [0115], which is transmitted to server 114, [0091]).
Ramadas does not specifically mention: a voice call; obfuscate one or more voice characteristics of the voice call.
Mossoba discloses a voice call (audio call example, Fig. 1D, [0043], [0047]); obfuscate one or more voice characteristics of the voice call (altering an audio frequency or amplitude associated with the section of the outgoing communication that includes the sensitive information, [0047]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas by utilizing it during a voice call as in Mossoba, and obfuscating one or more voice characteristics of the voice call for reasons similar to those for claim 1.
Ramadas and Mossoba do not specifically mention applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user.
Patino discloses applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user (random McAdams coefficients, which are used to manipulate the formant positions, i.e. voice characteristics, of speech signals, drawn from different ranges of the same uniform distribution are used to anonymize the speech, the randomly chosen McAdams coefficient speaker-dependent for each speaker, i.e. proving that that the anonymized speech originated with a specific speaker, Sections 3.1, 3.2, pages 2-3)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas and Mossoba by applying the random value with the one or more voice characteristics, wherein the random value is used to provide proof of origin, authenticity, or integrity of audio data of the user for reasons similar to those for claim 1.
Consider claim 2, Ramadas does not, but Mossoba discloses the voice characteristics is associated with one or more of: a pitch, a tone, or a note associated with a voice of the user, and the voice obfuscation is associated with one or more of: a change in pitch, a change in tone, or a change in note of the voice of the user (altering “audio frequency” is, at the very least, “associated with” a change in pitch, tone, and note, since the frequency spectrum of the audio determines the pitch, tone, and note of the speech, [0047]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas such that the voice characteristics is associated with one or more of: a pitch, a tone, or a note associated with a voice of the user, and the voice obfuscation is associated with one or more of: a change in pitch, a change in tone, or a change in note of the voice of the user for reasons similar to those for claim 1..
Consider claim 3, Ramadas discloses: applying the random value to one or more of replaced words or non-replaced words (using a timed volatile random subset (t-VRS) to generate replacement words including natural language words, numbers, phrases, and the like to obfuscate the sensitive-information utterance, [0065]).
Consider claim 4, Ramadas does not, but Mossoba discloses identifying the usage of user-specific language is based on an artificial intelligence or machine learning (AI/ML) model running on the device (machine learning model identifies a risk score for the communication, e.g. that they are speaking their SSN, [0025], [0031]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas by identifying the usage of user-specific language is based on an artificial intelligence or machine learning (AI/ML) model running on the device for reasons similar to those for claim 1.
Consider claim 5, Ramadas discloses generating the one or more replacement words is based on an artificial intelligence or machine learning (AI/ML) model running on the device (pre-trained model used to generate replacement words using encoder/decoder Translatotron, [0108], [0116]).
Consider claim 6, Ramadas discloses: storing, by the device, the random value in a local memory (mapping data 420 in storage device 406 of server 114 stores words corresponding to hash values, [0157]); or transmitting, by the device, the random value for storage in an external memory, (transmitting the t-VRS’s for the current interactive voice session to guardian system 112, [0152]).
Ramadas and Mossoba do not specifically mention wherein the random value is useable for non-repudiation of the user.
Patino discloses wherein a random value is useable for non-repudiation of the user (the randomly chosen McAdams coefficient is speaker-dependent for each speaker, i.e. usable for non-repudiation for the specific speaker, Sections 3.1, 3.2, pages 2-3).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas and Mossoba such that the random value is useable for non-repudiation of the user for reasons similar to those for claim 1.
Consider claim 8, Ramadas does not, but Mossoba discloses the user is a callee of the voice call or the user is a caller of the voice call (outgoing or incoming voice call for the user of a device, [0017], [0020]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas such that the user is a callee of the voice call or the user is a caller of the voice call in order to improve communications security, as suggested by Mossoba ([0002]).
Consider claim 9, Ramadas discloses the device is a network device or a user equipment (UE) (client device may be a mobile phone, [0023]).
Consider claim 11, Ramadas does not, but Mossoba discloses the voice characteristics is associated with one or more of: a pitch, a tone, or a note associated with a voice of the user, and the voice obfuscation is associated with one or more of: a change in pitch, a change in tone, or a change in note of the voice of the user (altering “audio frequency” is, at the very least, “associated with” a change in pitch, tone, and note, since the frequency spectrum of the audio determines the pitch, tone, and note of the speech, [0047]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas such that the voice characteristics is associated with one or more of: a pitch, a tone, or a note associated with a voice of the user, and the voice obfuscation is associated with one or more of: a change in pitch, a change in tone, or a change in note of the voice of the user for reasons similar to those for claim 1.
Consider claim 12, Ramadas discloses: the one or more processors are further configured to: apply the random value to one or more of replaced words or non-replaced words (using a timed volatile random subset (t-VRS) to generate replacement words including natural language words, numbers, phrases, and the like to obfuscate the sensitive-information utterance, [0065]).
Consider claim 13, Ramadas does not, but Mossoba discloses one or more processors are configured to identify the usage of user-specific language is based on an artificial intelligence or machine learning (AI/ML) model running on the device (machine learning model identifies a risk score for the communication, e.g. that they are speaking their SSN, [0025], [0031]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas by identifying the usage of user-specific language is based on an artificial intelligence or machine learning (AI/ML) model running on the device for reasons similar to those for claim 1.
Consider claim 14, Ramadas discloses one or more processors are configured to generate the one or more replacement words is based on an artificial intelligence or machine learning (AI/ML) model running on the device (pre-trained model used to generate replacement words using encoder/decoder Translatotron, [0108], [0116]).
Consider claim 15, Ramadas discloses the one or more processors are further configured to: store the random value in a local memory (mapping data 420 in storage device 406 of server 114 stores words corresponding to hash values, [0157]); or transmit the random value for storage in an external memory, (transmitting the t-VRS’s for the current interactive voice session to guardian system 112, [0152]).
Ramadas and Mossoba do not specifically mention wherein the random value is useable for non-repudiation of the user.
Patino discloses wherein a random value is useable for non-repudiation of the user (the randomly chosen McAdams coefficient is speaker-dependent for each speaker, i.e. usable for non-repudiation for the specific speaker, Sections 3.1, 3.2, pages 2-3).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas and Mossoba such that the random value is useable for non-repudiation of the user for reasons similar to those for claim 1.
Consider claim 16, Ramadas discloses the device is a network device or a user equipment (UE) (client device may be a mobile phone, [0023]).
Consider claim 18, Ramadas discloses the one or more instructions, when executed by the one or more processors, further cause the device to: apply the random value to one or more of replaced words or non-replaced words (using a timed volatile random subset (t-VRS) to generate replacement words including natural language words, numbers, phrases, and the like to obfuscate the sensitive-information utterance, [0065]); store the random value in a local memory (mapping data 420 in storage device 406 of server 114 stores words corresponding to hash values, [0157]); or transmit the random value for storage in an external memory, (transmitting the t-VRS’s for the current interactive voice session to guardian system 112, [0152]).
Ramadas and Mossoba do not specifically mention wherein the random value is useable for non-repudiation of the user.
Patino discloses wherein a random value is useable for non-repudiation of the user (the randomly chosen McAdams coefficient is speaker-dependent for each speaker, i.e. usable for non-repudiation for the specific speaker, Sections 3.1, 3.2, pages 2-3).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas and Mossoba such that the random value is useable for non-repudiation of the user for reasons similar to those for claim 1.
Consider claim 19, Ramadas discloses the one or more instructions, when executed by the one or more processors, further cause the device to: generate the one or more replacement words based on an artificial intelligence or machine learning (AI/ML) model running on the device (pre-trained model used to generate replacement words using encoder/decoder Translatotron, [0108], [0116]).
Ramadas does not specifically mention identifying the usage of user-specific language or vocabulary based on an artificial intelligence or machine learning (AI/ML) model running on the device.
Mossoba discloses identifying the usage of user-specific language is based on an artificial intelligence or machine learning (AI/ML) model running on the device (machine learning model identifies a risk score for the communication, e.g. that they are speaking their SSN, [0025], [0031]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas by identifying the usage of user-specific language is based on an artificial intelligence or machine learning (AI/ML) model running on the device for reasons similar to those for claim 1.
Consider claim 20, Ramadas discloses the device is a network device or a user equipment (UE) (client device may be a mobile phone, [0023]).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Ramadas et al. (US 20230269291) in view of Mossoba et al. (US 20230252190), in further view of Patino et al. (“Speaker anonymisation using the McAdams coefficient”. arXiv:2011.01130v2 [eess.AS] 1 Sep 2021), in further view of Kim (EP 4156665 A1).
Consider claim 7, Ramadas does not, but Mossoba discloses voice obfuscation is in response to a caller (during an incoming voice call for the user of a device, obfuscating sensitive audio information [0017]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas by performing voice obfuscation in response to a caller for reasons similar to those for claim 1.
Ramadas, Mossoba, and Patino do not specifically mention a caller in the voice call not being included in a contact list of a callee in the voice call.
Kim discloses a caller in the voice call not being included in a contact list of a callee in the voice call (call is incoming and the caller’s number is not in the user’s contact list, [0055]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Ramadas, Mossoba, and Patino by performing voice obfuscation in response to a caller, as in Mossoba, when a caller in the voice call is not included in a contact list of a callee in the voice call as in Kim in order to provide efficient protection against voice phishing, as suggested by Kim ([0008]), predictably helping to protect the elderly against fraud, as suggested by Kim ([0005]). The cited references are analogous art in the field of audio processing.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Nourtel et al. (“Evaluation of Speaker Anonymization on Emotional Speech”. arXiv:2305.01759v1 [eess.AS] 15 Apr 2023) discloses speech anonymization by randomly warping the F0 contour (See Section 2.2.2, page 2).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Jesse S Pullias/
Primary Examiner, Art Unit 2655 09/15/26