DETAILED ACTION
This office action is a First Action on the Merits (FAOM) for the claim set submitted on 01/14/2025. Claims 1-20 are pending and have been considered. The examiner would like to note that the claims have been deemed to be containing eligible subject matter under 35 U.S.C. 101 due to the inclusion of the “generating modified fundamental frequency information depending on the feature information”, wherein those modified functional frequencies are to be synthesized into speech (see synthesis step of independent claims). The examiner asserts that a user will not be able to meaningfully create modified fundamental frequencies, wherein the modification is with respect to the extracted real fundamental frequency. A user does not have the mental power to identify fundamental frequencies from listening to audio, nor generate additional fundamental frequencies based on frequencies which were not detected in the heard audio. Further, as these elements are not recited with mathematical operations, it is unreasonable to assume that these steps are generic mathematical operations when the signals being operated upon are not generic and/or simple. Even if a user making sounds could be interpreted to be generation of modified fundamental frequencies, synthesis of said sounds into a meaningful output is not reasonably performed in the mind with or without the aid of pen and paper (Step 2A, Prong 1, NO).
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119
(a)-(d). The certified copy has been filed for the parent Application No. EP22189150.0, filed on 08/05/2022.
Information Disclosure Statement
The information disclosure statement(s) submitted on 04/07/2025 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) is/are being considered by the examiner.
Drawings
The drawings are objected to because Fig. 2 has parts 2, 2A, and 2B. MPEP 608.02, Section V, (u) discloses drawing standards for numbering of views: “Partial views intended to form one complete view, on one or several sheets, must be identified by the same number followed by a capital letter.” It is unclear to the examiner whether all the figures with “2” are intended to form one complete view, but duplication of Fig. 2 and then Figs. 2A/B appears to be improper.
Further, Fig. 10 is split into “Part 1” and “Part 2”. This does not satisfy the above requirement. Applicant is recommended to amend the titles of these figures to “Fig. 10A” and “Fig. 10B” respectively. The examiner also recommends that Applicant labels the axes of Figs. 10 for clarity.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The examiner would like to note that any changes made to the titles of figures should also be reflected in the brief description of said figures in the specification.
Figures 8 and 9 should be designated by a legend such as --Prior Art-- because only that which is old is illustrated. Page 8 of the specification defines Figs. 8 and 9 to be illustrations according to conventional technology as described in the brief descriptions given to these figures. See MPEP § 608.02(g). Corrected drawings in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. The replacement sheet(s) should be labeled “Replacement Sheet” in the page header (as per 37 CFR 1.84(c)) so as not to obstruct any portion of the drawing figures. If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 10, 12, 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 10 recites the limitation "wherein the first neural network" in line 21 of pg. 4 of the claim set entered on 01/14/2025. There is insufficient antecedent basis for this limitation in the claim. Claim 5 first defines a neural network, but claim 10 is not dependent upon claim 5. Claim 8 first defines “a first neural network”, but claim 10 is not dependent upon claim 8. Claim 12 is rejected as being dependent upon a rejected base claim.
Claim 20 recites the limitation “…thereon to perform the method…” in line 27 of pg. 9 of the claim set entered on 01/14/2025. There is insufficient antecedent basis for this limitation in the claim. There is no method defined in claim 20. Applicant is recommended to amend the underlined language to “a method”.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-2, 4, 19, 20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kleinberger et al. (US-20210050029-A1), hereinafter Kleinberger.
Regarding claim 1, Kleinberger discloses: a system for conducting voice modification on an audio input signal comprising speech to acquire an audio output signal (Abstract, A feedback system may play back, to a user, an altered version of the user's voice in real time, [A voice alteration tracks to a modification, wherein playback indicates a required acquiring in order to be sent to the user]), wherein the system comprises:
a feature extractor for extracting feature information of the speech from the audio input signal ([0077] In some cases, features of the user's voice are extracted from a live audio stream of the user's voice),
a fundamental frequencies generator for generating modified fundamental frequency information depending on the feature information, such that the modified fundamental frequency information comprises modified fundamental frequencies being different from real fundamental frequencies of the speech ([0045] one or more pitch-shifted audio streams of the user's voice that are shifted in frequency relative to the fundamental frequency of the user's voice, [Pitch-shifted audio streams, wherein the shifting is based on fundamental frequency indicates the shifted frequencies to be generated, modified fundamental frequencies. The examiner asserts that shifting “relative to the fundamental frequency” indicates a changing of fundamental frequency, i.e. a harmonic operation as disclosed in [0059] harmonic rules for following harmonic shifts]), and/or such that the modified fundamental frequency information indicates a modified fundamental frequency trajectory being different from a real fundamental frequency trajectory of the speech ([The examiner would like to note that, due to this claim element containing optionally performed operations, i.e. and/or, no mapping is required]), and
a synthesizer for generating the audio output signal depending on the modified fundamental frequency information and depending on the feature information ([0115] audibly outputs the transformed signals in such a way that the user hears them, [Outputting an audio signal based on an input voice which has been transformed into features for processing indicates a synthesizer for transforming the modified features back into a signal in order for the output to be heard by the speaker]).
Regarding claim 2, Kleinberger discloses: the system according to claim 1.
Kleinberger further discloses:
wherein the feature information comprises first feature information and second feature information ([0077] The extracted speech features may include, among other things: (a) repetitions of words or parts of words (e.g., “wh-wh-which”); (b) prolongations of words or parts of words (e.g., “baaaat”); (c) blockages in speech (e.g., pauses of more than a specified threshold between words or parts of words); (d) tempo of speech after disregarding repetitions; and (e) excessive effort in speaking (while the user is trying to pronounce a word or part of a word), [Wherein any of the features described above could be first or second feature information]),
wherein the system comprises a modifier for generating modified second feature information depending on the second feature information, such that the modified second feature information is different from the second feature information ([0078] The extracted features may be fed as input to the machine learning model, during training of the model and during operation of the trained model. In some cases, a dimensionality reduction algorithm (e.g., principal component analysis) is performed, to reduce the dimensionality of the feature set,, [The examiner asserts that dimensionality reduction of features is a means of feature modification]),
wherein the fundamental frequencies generator is configured to generate the modified fundamental frequency information using the first feature information and using the modified second feature information ([Fig. 5, Pitch Shift 510-512], [0071] one or more pitch-shifted versions of the user's voice change fundamental frequency repeatedly as the fundamental frequency of the user's actual voice changes frequency. For instance, in Harmony, Pop, Retune and Pitch-Shift modes, each particular pitch-shifted version of the user's voice may repeatedly change pitch in order to maintain a constant frequency interval between the fundamental frequency of the particular pitch-shifted version and the fundamental frequency of the user's voice, [Wherein determining the pitch shifting based on the actual voice changes indicates generating modified fundamental frequency using the dimensionality reduced features, i.e. the second features, which will inherently describe the first feature information, indicating using both first and second feature information is required for the modified fundamental frequency generation]),
wherein the synthesizer is configured to generate the audio output signal using the modified fundamental frequency information, using the first feature information and using the modified second feature information ([0164] the transforming causes the transformed sound, which is outputted by the one or more speakers and is audible to the user, to comprise, at each pseudobeat in a set of pseudobeats, a superposition of two or more pitch-shifted versions of the user's voice, which pitch-shifted versions are sounded simultaneously with each other, in such a way that the fundamental frequencies of the respective pitch-shifted versions together form a chord in a chromatic musical scale, which chord has a root note that is the fundamental frequency of one of the pitch-shifted versions and is the nearest note in the scale to the fundamental frequency of the user's voice, [As previously disclosed, pitch-shifted versions of a user’s voice depend upon the first and second feature information; therefore, generating a transformed sound as output which is comprised of multiple pitch-shifted versions of an original voice with respect to the original voice’s fundamental frequency indicates the output to satisfy all required inputs to the synthesizer]).
Regarding claim 4, Kleinberger discloses: the system according to claim 2.
Kleinberger further discloses:
wherein the fundamental frequencies generator is implemented as a machine-trained system and/or is implemented as an artificial intelligence system ([0078] The extracted features may be fed as input to the machine learning model, during training of the model and during operation of the trained model. In some cases, a dimensionality reduction algorithm (e.g., principal component analysis) is performed, to reduce the dimensionality of the feature set, before feeding outputs (e.g., principal components) of the reduced dimensionality algorithm into the machine learning model, [As the pitch-shifting occurs after dimensionality reduction, as previously discussed, this indicates the pitch-shifting to be occurring using the machine learning model using the extracted features]).
Regarding claim 19, Kleinberger discloses: a method for conducting voice modification on an audio input signal comprising speech to acquire an audio output signal (Abstract, A feedback system may play back, to a user, an altered version of the user's voice in real time, [A voice alteration tracks to a modification, wherein playback indicates a required acquiring in order to be sent to the user]), wherein the method comprises:
extracting feature information of the speech from the audio input signal ([0077] In some cases, features of the user's voice are extracted from a live audio stream of the user's voice),
generating modified fundamental frequency information depending on the feature information, such that the modified fundamental frequency information comprises modified fundamental frequencies being different from real fundamental frequencies of the speech ([0045] one or more pitch-shifted audio streams of the user's voice that are shifted in frequency relative to the fundamental frequency of the user's voice, [Pitch-shifted audio streams, wherein the shifting is based on fundamental frequency indicates the shifted frequencies to be generated, modified fundamental frequencies. The examiner asserts that shifting “relative to the fundamental frequency” indicates a changing of fundamental frequency, i.e. a harmonic operation as disclosed in [0059] harmonic rules for following harmonic shifts]), and/or such that the modified fundamental frequency information indicates a modified fundamental frequency trajectory being different from a real fundamental frequency trajectory of the speech ([The examiner would like to note that, due to this claim element containing optionally performed operations, i.e. and/or, no mapping is required]), and
generating the audio output signal depending on the modified fundamental frequency information and depending on the feature information ([0115] audibly outputs the transformed signals in such a way that the user hears them, [Outputting an audio signal based on an input voice which has been transformed into features for processing indicates a synthesizer for transforming the modified features back into a signal in order for the output to be heard by the speaker]).
Regarding claim 20, Kleinberger discloses: a non-transitory digital storage medium having a computer program stored thereon ([0122] one or more computers execute programs according to instructions encoded in one or more tangible, non-transitory computer-readable media) to perform a method for conducting voice modification on an audio input signal comprising speech to acquire an audio output signal (Abstract, A feedback system may play back, to a user, an altered version of the user's voice in real time, [A voice alteration tracks to a modification, wherein playback indicates a required acquiring in order to be sent to the user]), wherein the method comprises:
extracting feature information of the speech from the audio input signal ([0077] In some cases, features of the user's voice are extracted from a live audio stream of the user's voice),
generating modified fundamental frequency information depending on the feature information, such that the modified fundamental frequency information comprises modified fundamental frequencies being different from real fundamental frequencies of the speech ([0045] one or more pitch-shifted audio streams of the user's voice that are shifted in frequency relative to the fundamental frequency of the user's voice, [Pitch-shifted audio streams, wherein the shifting is based on fundamental frequency indicates the shifted frequencies to be generated, modified fundamental frequencies. The examiner asserts that shifting “relative to the fundamental frequency” indicates a changing of fundamental frequency, i.e. a harmonic operation as disclosed in [0059] harmonic rules for following harmonic shifts]), and/or such that the modified fundamental frequency information indicates a modified fundamental frequency trajectory being different from a real fundamental frequency trajectory of the speech ([The examiner would like to note that, due to this claim element containing optionally performed operations, i.e. and/or, no mapping is required]), and
generating the audio output signal depending on the modified fundamental frequency information and depending on the feature information, when said computer program is run by a computer ([0115] audibly outputs the transformed signals in such a way that the user hears them, [Outputting an audio signal based on an input voice which has been transformed into features for processing indicates a synthesizer for transforming the modified features back into a signal in order for the output to be heard by the speaker. Implementing the method on a computer to be executed is previously disclosed in [0122].]).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 3, 10, 13-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kleinberger in view of Fang et al. (“Speaker Anonymization Using X-vector and Neural Waveform Models”), hereinafter Fang.
Regarding claim 3, Kleinberger discloses: the system according to claim 2.
Kleinberger does not disclose:
wherein the first feature information comprises phonetic posteriorgrams or other bottleneck features,
wherein the fundamental frequencies generator is configured to generate the modified fundamental frequency information using the phonetic posteriorgrams or the other bottleneck features of the speech and using the modified second feature information, and
wherein the synthesizer is configured to generate the audio output signal using the modified fundamental frequency information, using the phonetic posteriorgrams or the other bottleneck features of the speech and using the modified second feature information.
Fang discloses:
wherein the first feature information comprises phonetic posteriorgrams or other bottleneck features ([Fig. 1, “PPG”], [Introduction, Par. 3] capture linguistic
information in the form of a phoneme posteriorgram (PPG), [Which could be captured using the feature extraction of Kleinberger]).
wherein the fundamental frequencies generator is configured to generate the modified fundamental frequency information using the phonetic posteriorgrams or the other bottleneck features of the speech and using the modified second feature information ([Fig. 1, Acoustic Model receiving PPG, F0, and x-vector], [pg. 2, Section 3, par. 2] it uses an acoustic model and a neural waveform model to synthesize the speech waveform from the anonymized x-vector and the original PPG and F0, [Wherein the modified second feature data tracks to the anonymized x-vector being sent into the acoustic model]), and
wherein the synthesizer is configured to generate the audio output signal using the modified fundamental frequency information, using the phonetic posteriorgrams or the other bottleneck features of the speech and using the modified second feature information ([Fig. 1, Neural Waveform model resulting in output speech basd on a mel-spectrogram which is based on modified fundamental frequency, i.e. anonymized x-vector using the pitch-shift techniques of Kleinberger, and posteriorgrams]).
Kleinberger and Fang are considered analogous art within speech anonymization/modification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Kleinberger to incorporate the teachings of Fang, because of the novel way to extract linguistic and speaker identity features from an utterance to be used with neural acoustic and waveform models to synthesize anonymized speech which exploit x-vector representations, resulting in an increased error rate of speaker verification systems while maintaining high quality anonymized speech (Fang, Abstract).
Regarding claim 10, Kleinberger in view of Fang discloses: the system according to claim 3.
Fang further discloses:
wherein the second feature information is an x-vector of the speech ([Fig. 1, X-vector extractor]);
wherein the modifier is configured to generate a modified x-vector as the modified second feature information by choosing, depending on the x-vector of the speech, an x-vector from a group of available x-Vectors, such the x- vector being chosen from the group of x-vectors is different from the x-vector of the speech ([Fig. 2(b) Range Selection], [pgs. 2-3, Section 3.2] We devised two simple anonymization methods for modifying the x-vector of the input speech waveform. One is to use the mean x-vector of a set of randomly selected x-vectors from an x-vector pool, which produces a different anonymized psuedo speaker each time. The other is to compose an x-vector for which the similarity score to the original x-vector is s, [Composing an x-vector having a similarity to an original x-vector indicates selecting an x-vector based on the similarity to the original, wherein that could be chosen from the sets of x-vectors clearly disclosed in Fang for the other anonymization methods. The selected range forms a set from which one x-vector, i.e. starred, is chosen to anonymize the original x-vector]);
wherein the first neural network of the fundamental frequencies generator is configured to receive the phonetic posteriorgrams or the other bottleneck features of the speech and is configured to receive the modified x-vector as the input values of the first neural network ([Fig. 1, Acoustic Model receiving PPG (phoneme posteriorgram) and anonymized x-vector], [The “Conclusion and future work” section of Fang describes the proposed method being based on a “neural acoustic…model[s]”, indicating the acoustic model to be a neural network]), and is configured to output its output values comprising the modified fundamental frequencies ([Fig. 1, Acoustic Model], [pg. 3, Section 3.3, par. 1] an acoustic model that generates a Mel-spectrogram given the three input features (PPG, F0, and anonymized xvector), and a neural source-filter (NSF) waveform model [20] that produces a speech waveform given the F0, anonymized xvector, and generated Mel-spectrogram., [Wherein the fundamental frequency would be modified in view of the pitch-shifting of Kleinberger for anonymization.]) and/or indicating the modified fundamental frequencies trajectory ([The examiner would like to note that, due to this claim element being optional, no mapping is required]);
wherein the synthesizer is configured to generate the audio output signal using the phonetic posteriorgrams or the other bottleneck features of the speech and using the modified x-vector and depending on the output values of the first neural network that comprise the modified fundamental frequencies ([Fig. 1, Neural Waveform Model which receives fundamental frequency f0, Mel-spectrogram dependent upon phonetic posteriorgrams, and an anonymized x-vector to result in output speech]) and/or that indicate the modified fundamental frequencies trajectory ([The examiner would like to note that, due to this claim element being optional, no mapping is required]).
Regarding claim 13, Kleinberger discloses: the system according to claim 4.
Kleinberger does not disclose:
wherein the synthesizer is implemented as a neural vocoder and/or is implemented as a machine-trained system and/or is implemented as an artificial intelligence system and/or is implemented as a neural network.
Fang discloses:
wherein the synthesizer is implemented as a neural vocoder and/or is implemented as a machine-trained system and/or is implemented as an artificial intelligence system and/or is implemented as a neural network ([Fig. 1, Neural Waveform Model], [Fig. 3, RNN&CNN]).
Kleinberger and Fang are considered analogous art within speech anonymization/modification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Kleinberger to incorporate the teachings of Fang, because of the novel way to extract linguistic and speaker identity features from an utterance to be used with neural acoustic and waveform models to synthesize anonymized speech which exploit x-vector representations, resulting in an increased error rate of speaker verification systems while maintaining high quality anonymized speech (Fang, Abstract).
Regarding claim 14, Kleinberger discloses: the system according to claim 2.
Kleinberger does not disclose:
wherein the system is a system for conducting voice anonymization,
wherein the speech in the audio input signal is speech that has not been anonymized,
wherein the modifier is an anonymizer for generating anonymized second feature information as the modified second feature information depending on the second feature information, such that the anonymized second feature information is different from the second feature information,
wherein the fundamental frequencies generator is configured to generate anonymized fundamental frequency information as the modified fundamental frequency information using the first feature information and using the anonymized second feature information, and
wherein the synthesizer is configured to generate the audio output signal using the anonymized fundamental frequency information, using the first feature information and using the anonymized second feature information.
Fang discloses:
wherein the system is a system for conducting voice anonymization ([Title, “Speaker Anonymization….”]),
wherein the speech in the audio input signal is speech that has not been anonymized ([pg. 2, Section 3, para. 2] This system first extracts an x-vector, a PPG, and the fundamental frequency (F0) from the input waveform. It then anonymizes the x-vector on the basis of information gleaned from the x-vectors of external speakers, [Extracting an x-vector from speech to be anonymized indicates the input speech to be speech that has not been anonymized]),
wherein the modifier is an anonymizer for generating anonymized second feature information as the modified second feature information depending on the second feature information, such that the anonymized second feature information is different from the second feature information ([Fig. 1, Anonymize block], [pgs. 2-3, Section 3.2, para. 1] We devised two simple anonymization methods for modifying the x-vector of the input speech waveform. One is to use the mean x-vector of a set of randomly selected x-vectors from an x-vector pool, which produces a different anonymized psuedo speaker each time. The other is to compose an x-vector for which the similarity score to the original x-vector is s, [The anonymize block tracks to the modifier, wherein modification of the x-vector using either of the methods disclosed in Fang tracks to generation of anonymized second feature information]).
Kleinberger and Fang are considered analogous art within speech anonymization/modification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Kleinberger to incorporate the teachings of Fang, because of the novel way to extract linguistic and speaker identity features from an utterance to be used with neural acoustic and waveform models to synthesize anonymized speech which exploit x-vector representations, resulting in an increased error rate of speaker verification systems while maintaining high quality anonymized speech (Fang, Abstract).
Kleinberger further discloses:
wherein the fundamental frequencies generator is configured to generate anonymized fundamental frequency information as the modified fundamental frequency information using the first feature information and using the anonymized second feature information ([0070] the system shifts the fundamental frequency of an audio stream of the user's voice to a musical note in a musical scale, which note is the nearest to the fundamental frequency of the user's voice, [0092] (b) the feedback combines the effect of pitch shifting and choir speech by blending an original version of the voice (that respects all of the original vocal parameters) with additional versions where fundamental frequency and the formats are transformed, [Wherein the pitch shifts according to musical notes track to anonymized second feature information, i.e. the musical note(s)/pitch-shift itself, and the original voice tracks to first feature information]), and
wherein the synthesizer is configured to generate the audio output signal using the anonymized fundamental frequency information, using the first feature information and using the anonymized second feature information ([0164] the transforming causes the transformed sound, which is outputted by the one or more speakers and is audible to the user, to comprise, at each pseudobeat in a set of pseudobeats, a superposition of two or more pitch-shifted versions of the user's voice, which pitch-shifted versions are sounded simultaneously with each other, in such a way that the fundamental frequencies of the respective pitch-shifted versions together form a chord in a chromatic musical scale, which chord has a root note that is the fundamental frequency of one of the pitch-shifted versions and is the nearest note in the scale to the fundamental frequency of the user's voice, [As previously disclosed, pitch-shifted versions of a user’s voice track to anonymized second feature information; therefore, generating a transformed sound as output which is comprised of multiple pitch-shifted versions of an original voice with respect to the original voice’s fundamental frequency indicates the output to be based on first and second, anonymized feature information]).
Regarding claim 15, Kleinberger discloses: the system according to claim 2.
Kleinberger does not disclose:
wherein the system is a system for conducting voice de-anonymization,
wherein the speech in the audio input signal is speech that has been anonymized,
wherein the modifier is a de-anonymizer for generating de-anonymized second feature information as the modified second feature information depending on the second feature information, such that the de-anonymized second feature information is different from the second feature information,
wherein the fundamental frequencies generator is configured to generate de- anonymized fundamental frequency information as the modified fundamental frequency information using the first feature information and using the de- anonymized second feature information, and
wherein the synthesizer is configured to generate the audio output signal using the de-anonymized fundamental frequency information, using the first feature information and using the de-anonymized second feature information.
Fang discloses:
wherein the system is a system for conducting voice de-anonymization ([Title, “Speaker Anonymization….”], [The examiner asserts that if the input was already anonymized, the process of Fang could be used to de-anonymize the input speech without a change in functionality to Fang using the same “anonymization” methods disclosed]),
wherein the speech in the audio input signal is speech that has been anonymized ([pg. 2, Section 3, para. 2] This system first extracts an x-vector, a PPG, and the fundamental frequency (F0) from the input waveform. It then anonymizes the x-vector on the basis of information gleaned from the x-vectors of external speakers, [Extracting an x-vector from speech to be anonymized indicates the input speech to be speech that has not been anonymized]),
wherein the modifier is a de-anonymizer for generating de-anonymized second feature information as the modified second feature information depending on the second feature information, such that the de-anonymized second feature information is different from the second feature information ([Fig. 1, Anonymize block], [pgs. 2-3, Section 3.2, para. 1] We devised two simple anonymization methods for modifying the x-vector of the input speech waveform. One is to use the mean x-vector of a set of randomly selected x-vectors from an x-vector pool, which produces a different anonymized psuedo speaker each time. The other is to compose an x-vector for which the similarity score to the original x-vector is s, [The anonymize block tracks to the modifier, wherein modification of the x-vector using either of the methods disclosed in Fang tracks to generation of anonymized second feature information]).
Kleinberger and Fang are considered analogous art within speech anonymization/modification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Kleinberger to incorporate the teachings of Fang, because of the novel way to extract linguistic and speaker identity features from an utterance to be used with neural acoustic and waveform models to synthesize anonymized speech which exploit x-vector representations, resulting in an increased error rate of speaker verification systems while maintaining high quality anonymized speech (Fang, Abstract).
Kleinberger further discloses:
wherein the fundamental frequencies generator is configured to generate de-anonymized fundamental frequency information as the modified fundamental frequency information using the first feature information and using the de-anonymized second feature information ([0070] the system shifts the fundamental frequency of an audio stream of the user's voice to a musical note in a musical scale, which note is the nearest to the fundamental frequency of the user's voice, [0092] (b) the feedback combines the effect of pitch shifting and choir speech by blending an original version of the voice (that respects all of the original vocal parameters) with additional versions where fundamental frequency and the formats are transformed, [Wherein the pitch shifts according to musical notes track to anonymized second feature information, i.e. the musical note(s)/pitch-shift itself, and the original voice tracks to first feature information]), and
wherein the synthesizer is configured to generate the audio output signal using the de-anonymized fundamental frequency information, using the first feature information and using the de-anonymized second feature information ([0164] the transforming causes the transformed sound, which is outputted by the one or more speakers and is audible to the user, to comprise, at each pseudobeat in a set of pseudobeats, a superposition of two or more pitch-shifted versions of the user's voice, which pitch-shifted versions are sounded simultaneously with each other, in such a way that the fundamental frequencies of the respective pitch-shifted versions together form a chord in a chromatic musical scale, which chord has a root note that is the fundamental frequency of one of the pitch-shifted versions and is the nearest note in the scale to the fundamental frequency of the user's voice, [As previously disclosed, pitch-shifted versions of a user’s voice track to anonymized second feature information; therefore, generating a transformed sound as output which is comprised of multiple pitch-shifted versions of an original voice with respect to the original voice’s fundamental frequency indicates the output to be based on first and second, anonymized feature information]).
Regarding claim 16, Kleinberger in view of Fang discloses: the system according to claim 15.
Fang further discloses:
wherein the speech in the audio input signal is speech that has been anonymized according to a first mapping rule ([pgs. 2-3, Section 3.2, para. 1] We devised two simple anonymization methods for modifying the x-vector of the input speech waveform. One is to use the mean x-vector of a set of randomly selected x-vectors from an x-vector pool, which produces a different anonymized psuedo speaker each time. The other is to compose an x-vector for which the similarity score to the original x-vector is s),
wherein the de-anonymizer is configured to generating de-anonymized second feature information depending on the second feature information using a second mapping rule that depends on the first mapping rule ([As previously disclosed, the decision to “de-anonymize” can be performed using the same method of anonymization disclosed in Fang without a change in functionality to Fang. Either of the two rules for anonymization described above could be extended to de-anonymization if the received input signal of Fang is already anonymized. Using a second mapping rule that depends on a first mapping rule indicates that the second mapping rule could be the first mapping rule itself. A change in one rule will necessarily affect a change in the other rule, wherein the rules are the same]).
Regarding claim 17, Kleinberger in view of Fang discloses: the system according to claim 16.
Fang further discloses:
wherein the system is configured to receive information on the second mapping rule by receiving a bitstream that comprises the information on the second mapping rule ([Section 5.1, para. 1] Anonymized speech was obtained by averaging the nearest M speakers’ x-vectors in the pool, where M = 100, 200, and 300, [Anonymizing speech by averaging a plurality of x-vectors, wherein Fig. 1 indicates the anonymization occurs on a computing device, further wherein Fang discloses encoding of received speech (Section 3), indicating that the encodings which form the x-vectors are encoded bitstreams to be read/stored by the computer; therefore, performing the averaging of bitstreams for a series of x-vectors indicates a received second mapping rule by receiving a bitstream, i.e. the other x-vectors, to be averaged by the anonymization of the computing system. The examiner asserts that performing anonymization using a computing device will necessarily indicate the rule for anonymization to be in the form a bitstream in order for the computing system to be able to understand which of the two (and/or both) of Fang’s previously disclosed anonymization methods to apply]); or
wherein the system is configured to receive information on the first mapping rule by receiving a bitstream that comprises the information on the first mapping rule, and wherein the system is configured to derive information on the second mapping rule from the information on the first mapping rule ([The examiner would like to note that, due to the disjunctive construction of this claim element, it does not require a mapping]).
Regarding claim 18, Kleinberger discloses: a system (Abstract, A feedback system may play back, to a user, an altered version of the user's voice in real time, [A voice alteration tracks to a modification, wherein playback indicates a required acquiring in order to be sent to the user]) comprising:
a system for conducting voice anonymization ([0165] superposition of two or more pitch-shifted versions of the user's voice, which pitch-shifted versions are sounded simultaneously with each other, [The examiner asserts that pitch-shifting voice will effectively anonymize the original content]).
Kleinberger does not disclose:wherein the speech in the audio input signal is speech that has not been anonymized,
wherein the modifier is an anonymizer for generating anonymized second feature information as the modified second feature information depending on the second feature information, such that the anonymized second feature information is different from the second feature information; and
a system according to claim 15 for conducting voice de-anonymization,
wherein the system for conducting voice anonymization is configured to generate an audio output signal comprising speech that is anonymized,
wherein the system for conducting voice de-anonymization is configured to receive the audio output signal that has been generated by the system for conducting voice anonymization as an audio input signal, and
wherein the system for conducting voice de-anonymization is configured to generate an audio output signal from the audio input signal such that the speech in the audio output signal is de-anonymized.
Fang discloses:
wherein the speech in the audio input signal is speech that has not been anonymized ([pg. 2, Section 3, para. 2] This system first extracts an x-vector, a PPG, and the fundamental frequency (F0) from the input waveform. It then anonymizes the x-vector on the basis of information gleaned from the x-vectors of external speakers, [Extracting an x-vector from speech to be anonymized indicates the input speech to be speech that has not been anonymized]),
wherein the modifier is an anonymizer for generating anonymized second feature information as the modified second feature information depending on the second feature information, such that the anonymized second feature information is different from the second feature information ([Fig. 1, Anonymize block], [pgs. 2-3, Section 3.2, para. 1] We devised two simple anonymization methods for modifying the x-vector of the input speech waveform. One is to use the mean x-vector of a set of randomly selected x-vectors from an x-vector pool, which produces a different anonymized psuedo speaker each time. The other is to compose an x-vector for which the similarity score to the original x-vector is s, [The anonymize block tracks to the modifier, wherein modification of the x-vector using either of the methods disclosed in Fang tracks to generation of anonymized second feature information]).
Kleinberger and Fang are considered analogous art within speech anonymization/modification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Kleinberger to incorporate the teachings of Fang, because of the novel way to extract linguistic and speaker identity features from an utterance to be used with neural acoustic and waveform models to synthesize anonymized speech which exploit x-vector representations, resulting in an increased error rate of speaker verification systems while maintaining high quality anonymized speech (Fang, Abstract).
Kleinberger further discloses:
wherein the fundamental frequencies generator is configured to generate anonymized fundamental frequency information as the modified fundamental frequency information using the first feature information and using the anonymized second feature information ([0070] the system shifts the fundamental frequency of an audio stream of the user's voice to a musical note in a musical scale, which note is the nearest to the fundamental frequency of the user's voice, [0092] (b) the feedback combines the effect of pitch shifting and choir speech by blending an original version of the voice (that respects all of the original vocal parameters) with additional versions where fundamental frequency and the formats are transformed, [Wherein the pitch shifts according to musical notes track to anonymized second feature information, i.e. the musical note(s)/pitch-shift itself, and the original voice tracks to first feature information]), and
wherein the synthesizer is configured to generate the audio output signal using the anonymized fundamental frequency information, using the first feature information and using the anonymized second feature information ([0164] the transforming causes the transformed sound, which is outputted by the one or more speakers and is audible to the user, to comprise, at each pseudobeat in a set of pseudobeats, a superposition of two or more pitch-shifted versions of the user's voice, which pitch-shifted versions are sounded simultaneously with each other, in such a way that the fundamental frequencies of the respective pitch-shifted versions together form a chord in a chromatic musical scale, which chord has a root note that is the fundamental frequency of one of the pitch-shifted versions and is the nearest note in the scale to the fundamental frequency of the user's voice, [As previously disclosed, pitch-shifted versions of a user’s voice track to anonymized second feature information; therefore, generating a transformed sound as output which is comprised of multiple pitch-shifted versions of an original voice with respect to the original voice’s fundamental frequency indicates the output to be based on first and second, anonymized feature information]).
Kleinberger in view of Fang further discloses:
a system according to claim 15 for conducting voice de-anonymization (see rejection of claim 15).
Fang further discloses:
wherein the system for conducting voice anonymization is configured to generate an audio output signal comprising speech that is anonymized ([pg. 3, Section 3.3, para 1] a neural source-filter (NSF) waveform model [20] that produces a speech waveform given the F0, anonymized xvector, and generated Mel-spectrogram.),
wherein the system for conducting voice de-anonymization is configured to receive the audio output signal that has been generated by the system for conducting voice anonymization as an audio input signal ([Fig. 1, Input Audio Signal to be anonymized], [The examiner asserts that the input signal of Fang could be an anonymized signal to be de-anonymized without a change in functionality to the system of Fang. The additional step to de-anonymize previously anonymized speech is taught within the system of Fang, even if not explicitly disclosed as there is no functional difference between anonymization and de-anonymization as currently claimed. There is nothing preventing the use of Fang for de-anonymization as Fang discloses anonymization and the operations are identical, though in “reverse” order]), and
wherein the system for conducting voice de-anonymization is configured to generate an audio output signal from the audio input signal such that the speech in the audio output signal is de-anonymized ([Fig. 1, Output from Neural Waveform model], [The examiner asserts that the output signal of Fang could be an de-anonymized signal without a change in functionality to the system of Fang. The additional step to generate de-anonymized output speech is taught within the system of Fang, even if not explicitly disclosed as there is no functional difference between anonymization and de-anonymization as currently claimed. There is nothing preventing the use of Fang for de-anonymization as Fang discloses anonymization and the operations are identical, though in “reverse” order]).
Claim(s) 5-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kleinberger in view of Berlin et al. (US-11308657-B1), hereinafter Berlin.
Regarding claim 5, Kleinberger discloses: the system according to claim 2.
Kleinberger further discloses:
wherein the fundamental frequencies generator is implemented as a machine-trained system and/or is implemented as an artificial intelligence system ([0078] The extracted features may be fed as input to the machine learning model, during training of the model and during operation of the trained model. In some cases, a dimensionality reduction algorithm (e.g., principal component analysis) is performed, to reduce the dimensionality of the feature set, before feeding outputs (e.g., principal components) of the reduced dimensionality algorithm into the machine learning model, [As the pitch-shifting occurs after dimensionality reduction, as previously discussed, this indicates the pitch-shifting to be occurring using the machine learning model using the extracted features]).
Kleinberger does not disclose:
wherein the fundamental frequencies generator is implemented as a neural network, being configured to receive the first feature information and the modified second feature information as input values of the neural network, and
wherein the output values of the neural network comprise the modified fundamental frequencies and/or indicate the modified fundamental frequencies trajectory.
Berlin discloses:
wherein the fundamental frequencies generator is implemented as a neural network ([Col. 24, Lines 1-5] a neural network may be utilized to perform a shift in pitch, frequency, and/or modulation, and/or synthesize audio with a similar pitch, frequency, and/or modulation to a source voice), being configured to receive the first feature information and the modified second feature information as input values of the neural network ([Fig. 11, Source Voice 1102 and Destination Voice 1104], [Wherein a source voice tracks to first feature information and a destination voice tracks to modified second feature information (Berlin discloses low-dimension embeddings which would be gathered using the dimensionality reduction of Kleinberger, [Col. 24, Lines 30-45]), resulting in second modified features]), and
wherein the output values of the neural network comprise the modified fundamental frequencies ([Fig. 12, S1204 Swap Destination Voice with Source Voice…to Generate Modified Voice Track resulting in S1206 Generate Modified Audio]) and/or indicate the modified fundamental frequencies trajectory ([The examiner would like to note that, due to this claim element being optional, no mapping is required]).
Kleinberger and Berlin are considered analogous art within speech anonymization/modification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Kleinberger to incorporate the teachings of Berlin, because of the novel way to perform voice swapping using an autoencoder which includes a neural network which is trained using frame error minimization criteria and adjusted to minimize or reduce the error, resulting in cleaner voice swaps (wherein a voice signal which has had its fundamental frequency altered is functionally “swapped” with the original voice as disclosed in Kleinberger) (Berlin, [Col. 22, Lines 60-67], [Col. 23, Lines 1-10]).
Regarding claim 6, Kleinberger in view of Berlin discloses: the system according to claim 5.
Berlin further discloses:
wherein the neural network of the fundamental frequencies generator comprises one or more fully connected layers such that each node of the one or more fully connected layers depends on all input values of the neural network ([Col. 10, Lines 20-25] one or more intermediate networks may be utilized within the autoencoder to help learn more abstract representations. A first example intermediate network may include one layer or multiple fully connected layers), such that each node of the fully connected layers depends on the first feature information and depends on the modified second feature information ([A fully connected autoencoder receiving the first and second required inputs, as previously disclosed, indicates each node of the fully connected layers to depend on both the first and second feature information as input layers must receive all input]).
Regarding claim 7, Kleinberger in view of Berlin discloses: the system according to claim 5.
Berlin further discloses:
wherein the neural network of the fundamental frequencies generator has been trained by conducting training of the neural network using fundamental frequencies ([Col. 20, Lines 1-15] A similar user interface (including a generic face training set of controls, an extract generic face training control, a specify number of processing units for generic face training control, an initiate generic face training control, a terminate generic face training control, an initiate training control, a terminate training control, a select model control, select destination audio control, a select source audio control, select frame size) may be utilized to manage voice processing operations, such as the generic face training, training, and output voice creation, [Wherein the voice creation would be trained using the pitch-shifting described in Kleinberger]) and/or fundamental frequency trajectories of speech signals ([The examiner would like to note that, due to this claim element being optional, no mapping is required]).
Claim(s) 8-9, 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kleinberger in view of Berlin, further in view of Perero-Codosero et al. (“X-vector anonymization using autoencoders and adversarial training for preserving speech privacy”), hereinafter Perero.
Regarding claim 8, Kleinberger in view of Berlin discloses: the system according to claim 5.
Berlin further discloses:
wherein the neural network of the fundamental frequencies generator is a first neural network ([Col. 24, Lines 1-5] a neural network may be utilized to perform a shift in pitch, frequency, and/or modulation, and/or synthesize audio with a similar pitch, frequency, and/or modulation to a source voice).
Kleinberger in view of Berlin does not disclose:
wherein the modifier is implemented as a second neural network,
wherein the second neural network is configured to receive input values from a plurality of frames of the audio input signal,
wherein the second neural network is configured to output the second feature information as its output values.
Perero discloses:
wherein the modifier is implemented as a second neural network ([Fig. 2, AAN], [pg. 4, par. 3] The specific contribution of our proposal is the x-vector anonymization component, which we call the AAN, as shown in Fig. 2. This neural network architecture, represented in Fig. 3, includes an encoder–decoder branch that attempts to reconstruct the input x-vector),
wherein the second neural network is configured to receive input values from a plurality of frames of the audio input signal ([Fig. 2, X-vector extracted from Test Utterance], [The examiner asserts that generating a vector from a test utterance indicates each element of that vector to be a value from a frame with respect to the other components of the vector]),
wherein the second neural network is configured to output the second feature information as its output values ([Fig. 2, Anonymized x-vector output from AAN]).
Kleinberger, Berlin, and Perero are considered analogous art within speech anonymization/modification. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Kleinberger in view of Berlin to incorporate the teachings of Perero, because of the novel way to transform original x-vector representations of speech using an autoencoder to produce a new x-vector where speaker, gender, and accent information are suppressed through adversarial training, improving the intelligibility of anonymized speech (Perero, Abstract).
Regarding claim 9, Kleinberger in view of Berlin, further in view of Perero discloses: the system according to claim 8.
Perero further discloses:
wherein the second feature information is an x-vector of the speech ([Fig. 2, x-vector/anonymized x-vector]).
Regarding claim 11, Kleinberger in view of Berlin, further in view of Perero discloses: the system according to claim 8.
Berlin further discloses:
wherein the system further comprises an output value modifier for modifying the output values of the first neural network of the fundamental frequencies generator to acquire amended values that comprise amended fundamental frequencies ([Fig. 12, S1204 Swap Destination Voice with Source Voice in Voice track to Generate Modified Voice Track], [The component responsible for “swapping” is modifying the source voice to become the destination voice, wherein that swap comprises a pitch-shift as previously disclosed in Berlin]) and/or that indicate an amended fundamental frequencies trajectory ([The examiner would like to note that, due to this claim element being optional, no mapping is required]).
Perero further discloses:
wherein the synthesizer is configured to generate the audio output signal using the phonetic posteriorgrams or the other bottleneck features of the speech, using the modified x-vector and using the amended values ([Fig. 2, Synthesis model which receives Bottleneck features and Anonymized x-vector], [Wherein the modified x-vector of Perero (which removes speaker, gender, and/or accent) will have amended values to reflect the removal of these features of the audio]).
Allowable Subject Matter
Claim 12 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is an examiner’s statement of reasons for allowance:
Considering claim 12, the closest prior art of record is Kleinberger, Fang, Perero, Champion et al. (“A Study of Modification for X-Vector Based Speech Pseudonymization Across Gender”), hereinafter Champion, Mawalim et al. (“Speaker Anonymization by Modifying Fundamental Frequency and X-vector Singular Value”), hereinafter Mawalim, Agarwal et al. (“Speaker Anonymization for Machines using Sinusoidal Model”), hereinafter Agarwal, Gaznepoglu et al. (“Exploring the Importance of F0 Trajectories for Speaker Anonymization using X-vectors and Neural Waveform Models”), hereinafter Gaznepoglu, and Gaznepoglu et al. (“VoicePrivacy 2022 System Description: Speaker Anonymization with Feature-matched F0 Trajectories”), hereinafter Gaznepoglu2.
Kleinberger in view of Fang discloses: the system according to claim 10.
Fang further discloses:
wherein the system further comprises a fundamental frequencies extractor for extracting the real fundamental frequencies of the speech ([Fig. 1, F0 Extractor]).
Kleinberger in view of Fang does not disclose:
wherein the system comprises a second fundamental frequencies generator for generating second fundamental frequency information using the phonetic posteriorgrams or the other bottleneck features of the speech and using the x-vector of the speech,
wherein the system further comprises a first combiner for generating, depending on the real fundamental frequencies of the speech and depending on the second fundamental frequency information, values indicating a fundamental frequencies residuum,
wherein the system comprises a second combiner for combining the output values of the first neural network of the fundamental frequencies generator and the values indicating the fundamental frequencies residuum to acquire combined values, and
wherein the synthesizer is configured to generate the audio output signal depending on the combined values, using the phonetic posteriorgrams or the other bottleneck features of the speech and using the modified x-vector (emphasis added to underlined portions).
The examiner asserts that generation of fundamental frequencies using a second fundamental frequency generator, wherein said second fundamental frequencies are not contained within the received features of speech, wherein said frequencies are combined using first and second combiners to form a residuum and combined again with output from a first neural network of a fundamental frequency generator to be synthesized is a novel, non-obvious improvement over existing speech anonymization methods.
Fang discloses a F0 extractor as seen in Fig. 1, but there is no additional step of generation of fundamental frequencies based on x-vectors or bottleneck features of speech. All of these elements are present in Fang, but they are not used in the way claimed to generate additional fundamental frequencies.
Perero discloses a structure similar to that of Fang, see Fig. 2. Similar to Fang, Perero faces similar shortcomings with regard to the generation of additional fundamental frequency information outside of that extracted from input speech. The synthesis model of Perero operates in the same way as the system of Fang with similar inputs.
Champion discloses a system similar to that of Fang and Perero with an additional linear transformation to be performed on extracted fundamental frequencies, wherein the transformation is based on F0 stats from pooled x-vector/F0 statistics, see Fig. 1. The examiner asserts that the linear transformation of Champion could theoretically be interpreted to be representative of the second fundamental frequency generator, though the examiner asserts that modification of existing fundamental frequency values does not indicate generation of new, previously non-existent frequencies (see “F0 modification” section of Champion). Further, should the linear transform be interpreted to be the second fundamental frequency generator, Champion still fails to disclose the combination of the linear transformed frequencies with newly generated frequencies to form a frequency residuum. Champion is silent with regard to a frequency residuum.
Mawalim discloses a system for speech anonymization similar to that of Fang and Perero which have already been discussed, see Fig. 2. There are no functional differences between the structures of Mawalim, Fang, and Perero; therefore, Mawalim has the same shortcomings with regard to the second fundamental frequency generator/combiners as Fang and Perero as previously discussed.
Agarwal discloses a method for speaker anonymization which does consider residual and complementary speech signals, see section III; however, Agarwal explicitly mentions not using the generated residual speech because of phase discontinuity at frame boundaries (see pg. 3, “B. Motivation to use complementary speech for speaker anonymization”). Agarwal explicitly teaches away from using the claimed method of residual speech to anonymize. Further, Agarwal makes no mention of residual speech in the context of fundamental frequencies.
Gaznepoglu discloses a system which includes F0 modification in the context of speaker anonymization, see Fig. 1. This art does not disclose a frequency residuum, and even if this was disclosed, this art would be an exception under 35 U.S.C. 102(b)(1)(A) as it is the same inventors’ work and falls within the grace period of one year before the filing date of the claimed invention.
Gaznepoglu2 discloses a system which does generate F0 trajectories, see “F0 Regressor” of Fig. 1; however, like Gaznepoglu, this art is commonly owned and does not beat the EFD of the claimed invention. This is not available prior art.
The examiner believes this to be the best prior art available before the effective filing date of the claimed invention, 08/05/2022. As can be seen with the cited art, there is nothing which explicitly discloses or implicitly suggests the concepts of generating fundamental frequencies which are not included in a received speech feature signal/set, combining said generated fundamental frequencies with previously extracted fundamental frequencies to form a frequency residuum, nor using said residuum to synthesize speech which is anonymized based on x-vectors and other bottleneck features of input speech. The invention claims a novel, non-obvious way to improve speaker anonymization methods.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Henrotte et al. (US-20230410825-A1) discloses “A method includes masking the voice of a speaker by intentionally altering the pitch and the timbre of their voice. An audio signal corresponding to an original recording of the voice of the speaker is divided (11) into a series of successive audio segments of a determined constant duration. A rising frequency alteration (12a) is applied to a timbre (A) extracted from each audio segment. A falling frequency alteration (12b) is applied to a pitch (B) extracted from each audio segment. The altered pitch and the altered timbre of the audio segment are combined (14) so as to form a single resulting altered audio segment. From one audio segment to another in the series of audio segments, a variation (13a) of the rising alteration and a variation (13b) of the falling alteration are applied. These variations fluctuate randomly from one audio segment to another in the series of audio segments.” (abstract). See entire document.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THEODORE JOHN WITHEY whose telephone number is (703)756-1754. The examiner can normally be reached Monday - Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571) 272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THEODORE WITHEY/Examiner, Art Unit 2655
/ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655