DETAILED ACTION
This office action is in response to Applicant’s Request for Continued Examination (RCE), received on 06/05/2026. Claims 1, 9, and 17 have been amended. Claims 1-20 are pending and have been considered.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/05/2026 has been entered.
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119
(a)-(d). The certified copy has been filed for the parent Application No. CN202211297843.3, filed on 10/21/2022.
Response to Arguments
Applicant’s arguments, see pgs. 10-12, filed 06/05/2026, with respect to the rejection(s) of claim(s) 1, 9, and 17 under 35 U.S.C. 103 (Wei in view of Tan) have been fully considered and are persuasive (with respect to point 1, “Wei does not teach of suggest the step of extracting an initial recognition voice…wherein the mixed voice and initial recognition voice are time-domain voice signals”). Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Xu et al. (“SpEx: Multi-Scale Time Domain Speaker Extraction Network”), hereinafter Xu. Xu discloses “we propose a time-domain speaker extraction network (SpEx) that converts the mixture speech into multi-scale embedding coefficients instead of decomposing the speech signal into magnitude and phase spectra. In this way, we avoid phase estimation. The SpEx network consists of four network components, namely speaker encoder, speech encoder, speaker extractor, and speech decoder. Specifically, the speech encoder converts the mixture speech into multi-scale embedding coefficients, the speaker encoder learns to represent the target speaker with a speaker embedding. The speaker extractor takes the multi-scale embedding coefficients and target speaker embedding as input and estimates a receptive mask. Finally, the speech decoder reconstructs the target speaker’s speech from the masked embedding coefficients” (abstract). See updated rejections below.
Applicant's arguments filed 06/05/2026, see pgs. 12-13, with respect to point 2, “Wei does not teach or suggest the step of determining…a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice” have been fully considered but they are not persuasive.
Applicant’s representative asserts, “
2. Wei does not teach or suggest the step of determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice.
In the office action (pp. 5-6), the Examiner stated as follows:
determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice ([0175] the first noisy voice signal is input to the second encoding network frame by frame for feature extraction, to obtain a voice feature vector of each frame, [0224] when the third distortion score is greater than the eighth threshold (for example, 12 dB) and the SNR of the third noise segment is not less than the seventh threshold, it indicates that a voiceprint feature of the current user matches a stored sound feature, [Wherein SNR and/or distortion score tracks to the similarity metric for comparison to thresholds of current user voices,) i.e. initial recognition voices,) to stored, i.e. registered, threshold sound features, i.e. SNR and/or distortion values. See storing of registered voice samples, [0241]. Further, analyzing audio in "segments", in view of the previously disclosed frame-level feature extraction, indicates the segment corresponds to a frame as feature vectors are used for the similarity comparison]).
For reference, paragraphs [0221] and [0224] of Wei are reproduced below.
[0221] In an embodiment, the distortion score may be a
signal-to-distortion ratio (SDR) value or a perceptual
evaluation of speech quality (PESQ) value.
[0224] ... the terminal device captures the voice signal to
obtain the third noise segment; performs noise reduction on the
third noise segment based on the stored voiceprint feature vector of
the historical user to obtain the third noise-reduced noise segment;
and performs distortion evaluation based on the third noise
segment and the third noise-reduced noise segment to obtain
the third distortion score. When the third distortion score is
greater than the sixth threshold (for example, 8 dB) and the SNR of
the third noise segment is less than the seventh threshold (for example, 10 dB), or when the third distortion score is greater than the eighth threshold (for example, 12 dB) and the SNR of the third noise segment is not less than the seventh threshold, it indicates that a voiceprint feature of the current user matches a stored sound feature, and the terminal device sends the third prompt information to the user, where the third prompt information prompts the current user whether to enable the terminal device to enter the PNR mode.
Emphasis Added.
As highlighted above, the distortion score in Wei is to measure the speech quality of a voice signal, i.e., the extent to which that a voice signal is distorted by noise. This is why Wei states that signal-to-distortion ratio in the form of dB is used for representing the distortion score.
In contrast, the voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice is to measure the similarity between two different time-domain voice signals, i.e., the registered voice and the initial recognition voice, one time segment after another segment. See, e.g., paragraphs [0086] and [0087] of the present application reproduced below.
As highlighted above, the term "voice similarity" is to measure the extent to which the voice segment is close to the registered voice and it is fundamentally different from the distortion score of Wei.
Neither Tan nor any other cited references salvage the deficiencies of Wei.”
In response, the examiner would like to refer to the broadest reasonable interpretation (BRI) of “determining, based on the registered voice feature, a voice similarity between the registered voice and the voice information comprised in each voice segment in the initial recognition voice” in view of Wei. Specifically, the cited portion of the instant application [0086] explicitly discloses “In an embodiment, the voice similarity…may be calculated by the following formula…”. This appears to the examiner to be a preferred embodiment, but the claim interpretation is not limited to preferred embodiments. See MPEP 2111.01, Section II: “It is improper to import claim limitations from the specification”, But c.f. In re Am. Acad. of Sci. Tech. Ctr., 367 F.3d 1359, 1369, 70 USPQ2d 1827, 1834 (Fed. Cir. 2004) ("We have cautioned against reading limitations into a claim from the preferred embodiment described in the specification, even if it is the only embodiment described, absent clear disclaimer in the specification."). In view of the cosine formula being a preferred embodiment, any function which determines a voice similarity between a registered voice and voice information in each segment of an initial recognition voice tracks to the similarity measure as currently claimed.
Returning to Wei, there is disclosed a comparison between frames of a mixed, i.e. noisy, signal to stored sound features as compared to various distortion thresholds for determining if “a voiceprint feature of the current user matches a stored sound feature”. This indicates a binary assignment for similarity based on the distortion and SNR scores. Based on a comparison of the distortion and SNR scores to certain threshold conditions, the frame is determined to be a match or not a match. The examiner asserts that matching necessarily indicates similarity between the items being compared. The frames of signals of Wei could be in the time domain in view of Xu and/or the processing of Fig. 4 which suggests processing in the time domain.
Applicant’s arguments with respect to claim(s) 3-5, 11-13, and 19 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 6-10, 14-18, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei et al. (US-20240096343-A1), hereinafter Wei, in view of Xu et al. (“SpEx: Multi-Scale Time Domain Speaker Extraction Network”), hereinafter Xu, further in view of Tan et al. (US-20220375475-A1), hereinafter Tan.
Regarding claim 1, Wei discloses: a voice extraction method ([0010] to extract the voice signal of the target user from the noisy voice signal) performed by a computer device ([0084] the auxiliary device may be a device with a microphone array, for example, a computer or a tablet computer), comprising:
obtaining a registered voice of a speaker ([0162] registered voice of the target user are input to a voice noise reduction model for processing, [Obtained by the voice noise reduction model]);
determining a registered voice feature of the registered voice ([0174] extracting a feature vector of the registered voice signal of the target user from the registered voice signal, [A feature vector indicates at least one registered voice feature]);
extracting an initial recognition voice of the speaker from a mixed voice based on the registered voice feature ([Fig. 4], [0174] extracting a feature vector of the noisy voice signal from the noisy voice signal by using a second encoding network; obtaining a first feature vector based on the feature vector of the registered voice signal and the feature vector of the noisy voice signal, for example, a mathematical operation such as dot multiplication is performed on the feature vector of the registered voice signal and the feature vector of the noisy voice signal to obtain the first feature vector, [In view of Fig. 4 which demonstrates the output from the dot product operation to be sent into a TCN to be an initial recognition voice of the speaker from a mixed, i.e. noisy, voice based on the registered voice feature vector, i.e. output from the encoding networks combined through dot product]),
wherein the mixed voice is a time domain signal ([Fig. 4, Noisy voice signal and Output from the Second encoding network to be processed by Temporal Convolutional Network], [Processing a noisy, i.e. mixed, voice signal to extract a feature vector, i.e. initial recognition voice, as previously cited, to be passed into a Temporal convolutional network after a dot product calculation indicates the mixed voice to be time domain vectors in order to be processed by the Temporal convolutional network. This is in contrast with the embodiment of Fig. 7 which explicitly features FFT blocks for converting input into the frequency domain]).
Wei does not disclose:
wherein the initial recognition voice is a time-domain voice signal.
Xu discloses:
wherein the initial recognition voice is a time-domain voice signal ([Fig. 3, Speaker Extractor comprised of Temporal Convolutional Networks (TCNs) resulting in output M1-M3 and S1-S3], [pg. 1373, right column] we encode the time-domain signal into three temporal resolutions in the embedding E, [pg. 1374, B. Multi-Scale Encoding and Decoding] The speaker extractor then estimates multi-scale masks M1,M2,M3, and generates the multiscale modulated responses S1, S2, S3, [The examiner asserts that the presence of TCN blocks within the Speaker Extractor indicates operations of the extractor to be in the time-domain, wherein any of the modulated responses tracks to an initial recognition voice as being extracted from the mixed signal and is a function of time, see pg. 1374, Eq. (3)]).
Wei and Xu are considered analogous art within target speaker extraction from mixed speech. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei to incorporate the teachings of Xu, because of the novel way to perform multi-scale encoding and decoding over multiple temporal resolutions, improving the resultant voice quality extracted from the mixed speech signal (Xu, [pg. 1371, right column, contribution 4]). It would be obvious to take the input signals of Wei (which the examiner asserts are in the time domain before FFT conversion, Fig. 7, also see Fig. 4 which includes no FFT) and keep them in the time domain to be applied to the multi-scale process of Xu to result in the improved target speaker extraction quality as Fig. 4 of Wei discloses a system which appears to contain two signals in the time domain (in order to be processed by the temporal convolutional network of Fig. 4), wherein the output from the second encoding network is an extracted feature vector, i.e. initial recognition voice vector, as previously cited.
Wei in view of Xu does not disclose:
identifying multiple voice segments in the initial recognition voice.
Tan discloses:
identifying multiple voice segments in the initial recognition voice ([0018] The computing system 100 may use a portion of the audio (e.g., the first 5 seconds of the audio, the first 20 seconds of the audio, etc.) to generate a signature vector… The computing system 100 may compare the voice signature with other portions of the audio the vector generated for the beginning portion may be compared with vectors generated for other portions of the audio), [Other portions in addition to a first portion indicates multiple voice segments, i.e. portions]).
Wei, Xu, and Tan are considered analogous art within voice quality enhancement. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu to incorporate the teachings of Tan, because of the novel way to create a voice biometric for a user with call audio including background noise through removal of the background noise via segmentation of a received conversation into user voice/not based on a similarity threshold comparison to a known voiced segment, allowing for the creation of more accurate biometric samples for user authentication (Tan, [0001]-[0002]).
Wei further discloses:
determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice ([0175] the first noisy voice signal is input to the second encoding network frame by frame for feature extraction, to obtain a voice feature vector of each frame, [0224] when the third distortion score is greater than the eighth threshold (for example, 12 dB) and the SNR of the third noise segment is not less than the seventh threshold, it indicates that a voiceprint feature of the current user matches a stored sound feature, [Wherein SNR and/or distortion score tracks to the similarity metric for comparison to thresholds of current user voices, i.e. initial recognition voices, to stored, i.e. registered, threshold sound features, i.e. SNR and/or distortion values. See storing of registered voice samples, [0241]. Further, analyzing audio in “segments”, in view of the previously disclosed frame-level feature extraction, indicates the segment corresponds to a frame as feature vectors are used for the similarity comparison]).
Tan further discloses:
selecting, from the multiple voice segments, a first subset of voice segments whose associated voice similarities are less than a preset similarity ([0018] The computing system 100 may remove portions from the audio, for example, if the similarity score of a portion does not satisfy a similarity threshold. For example, a portion of the audio may include traffic noises and a vector generated for the portion may not satisfy a similarity threshold when compared with the signature vector, [Wherein the signature vector of Tan represents an earlier portion of the same audio, but it would have been obvious to apply the stored sound feature of Wei as the signature vector of Tan as Tan discloses a first portion of audio with only one person speaking (see [0019] disclose removal of anything other than one individual) indicating the stored sound feature voiceprint representing a historical user of Wei could be substituted for the signature vector of Tan without a change in functionality to Tan to reach the claimed functionality]); and
removing the first subset of voice segments from the initial recognition voice to obtain a clean voice of the speaker ([0019] Portions of the audio that include one or more interfering noises may be removed so that a voice biometric for the user may be generated using the remaining portions, [A voice capable of being used for a voice biometric with interfering portions removed indicates the biometric to be a clean voice]).
Regarding claim 2, Wei in view of Xu, further in view of Tan discloses: the method according to claim 1.
Wei further discloses:
wherein the extracting an initial recognition voice of the speaker from a mixed voice based on the registered voice feature comprises:
determining a mixed voice feature of the mixed voice ([0174] extracting a feature vector of the noisy voice signal from the noisy voice signal);
fusing the mixed voice feature and the registered voice feature of the registered voice to obtain a voice fusion feature ([0181] performing fusion on the frequency domain signal of the VPU signal of the target user and the frequency domain signal of the first noisy voice signal to obtain a first fusion frequency domain signal, [0182] the first fusion frequency domain signal is input to the third encoding network for feature extraction to obtain a feature vector of the first fusion frequency domain signal, [In view of [0163] which defines “registered VPU signal[s]” indicating the registered signals and VPU signals are synonymous. A first fusion feature vector indicates at least one voice fusion feature comprising the vector]); and,
performing initial recognition on voice information of the speaker in the mixed voice based on the voice fusion feature to obtain the initial recognition voice of the speaker ([0181] separately processing the first fusion frequency domain signal by using a third encoding network, a GRU, and a second decoding network to obtain a mask of a frequency domain signal of the voice signal of the target user, [Generating a mask based on a voice fusion signal for determining a voice signal of a target user indicates the mask to be representative of an initial recognition voice of the speaker]).
Regarding claim 6, Wei in view of Xu, further in view of Tan discloses: the method according to claim 1.
Wei further discloses:
wherein the determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice comprises:
repeating the each voice segment in the initial recognition voice based on a time length of the registered voice separately to obtain a recombined voice having the time length ([0219] the terminal device obtains, based on a preset cycle (for example, every 10 minutes), a first noise segment (for example, a 6 s voice signal currently captured by the microphone) and a second noise segment (for example, a 10 s voice signal after the 6 s voice signal currently captured by the microphone) of an environment in which the terminal device is located, and obtains an SNR and an SPL of the first noise segment, [Wherein the “repeating” is performed through the SNR/SPL calculations of each segment of the noise, i.e. mixed, signals and/or the obtaining of the segments]);
obtaining a recombined voice feature extracted from the recombined voice ([0219] determines whether the SNR of the first noise segment is greater than 20 dB and whether the SPL is greater than 40 dB; if the SNR of the first noise segment is greater than the first threshold (for example, 20 dB) and the SPL is greater than the second threshold (for example, 40 dB), extracts a first temporary feature vector of the first noise segment; performs noise reduction on the second noise segment by using the first temporary feature vector to obtain a second noise-reduced noise segment, [Generating a feature vector for the second noise segment, wherein that vector is dependent upon a first noise feature vector, indicates the second feature vector to be a combination of the first and second, i.e. recombined]);
determining a segment voice feature corresponding to the each voice segment in the initial recognition voice based on the recombined voice feature ([0219] performs distortion evaluation based on the second noise-reduced noise segment and the second noise segment to obtain a first distortion score, where the first distortion score indicates a degree of distortion of the signal captured by the microphone of the terminal device, [Performing distortion evaluation, wherein the evaluation is performed based on a feature comparison, see [0221] defining the distortion score to be a signal-to-distortion ratio (SDR), i.e. SDR is a feature of audio, indicates a required segment voice feature determination of the second noise segment to be compared to the first distortion score, i.e. also representing a segment of audio, wherein the second noise-reduced segment represents a recombined voice feature as previously disclosed. Further, the second noise segment tracks to an initial recognition voice]); and,
determining a voice similarity between the registered voice and the each voice segment based on the segment voice feature corresponding to the each voice segment and the registered voice feature separately ([0222] performing noise reduction on the third noise segment based on the reference temporary voiceprint feature vector to obtain a third noise-reduced noise segment; performing distortion evaluation based on the third noise segment and the third noise-reduced noise segment to obtain a third distortion score; if the third distortion score is greater than a sixth threshold and an SNR of the third noise segment is less than a seventh threshold, or the third distortion score is greater than an eighth threshold and the SNR of the third noise segment is not less than the seventh threshold, sending third prompt information, [0241] The registered voice of the current user may be obtained by connecting a plurality of segments of signals in series, and total duration is not less than 6 s, [Applying a temporary feature vector which results in a noise-reduced segment to be compared to a reference threshold, i.e. registered, segment SNR/SDR indicates the similarity comparison to be between a registered voice having the threshold ratio levels and each voice segment based on the segment voice feature, i.e. SNR/SDR values. Further, consider generation of temporary feature vectors for registered voice signals, [0231], indicating a noise reduction performed based on a similarity comparison of a registered feature vector to a voice segment feature vector]).
Regarding claim 7, Wei in view of Xu, further in view of Tan discloses the method according to claim 1.
Wei further discloses: wherein the obtaining a registered voice of a speaker comprises:
determining a speaker in response to a call triggering operation ([0240] When it is detected that the terminal device is in a hands-free call state, the terminal device enters the PNR mode, and a terminal device owner whose voiceprint feature has been registered is a target user, [PNR tracks to “personalized noise reduction”, [0007]]); and,
determining the registered voice of the speaker from a prestored candidate registered voice ([0178] The voice signal of the target user is registered in advance, [0241] If registered voice of the current user or a voice feature of the current user is stored, the current user is determined as a target user, [Generating a registered voice, wherein the registration is performed in advance, indicates the registered voice is determined from a prestored candidate registered voice]).
Regarding claim 8, Wei in view of Xu, further in view of Tan discloses: the method according to claim 1.
Wei further discloses:
wherein the mixed voice is transmitted from a terminal of a speaker after a voice call is established with the terminal ([0227] when it is detected that the terminal device is used again for a call, the default microphone of the terminal device captures the second noisy voice signal, [0246] a call function of the terminal device is implemented through a “Phone” application, [A second noisy voice signal tracks to a mixed signal. Further, disclosing the terminal to be a phone indicates at least one speaker for playing the noisy sound to be captured by the microphone of the terminal device. See Fig. 15 “Volume” control on the phone terminal and Fig. 17 phone call setting to “Enable PNR” after the phone call is connected]).
Regarding claim 9, Wei discloses: a computer device ([0084] the auxiliary device may be a device with a microphone array, for example, a computer or a tablet computer), comprising a memory ([Fig. 26, Memory 2602]) and a processor ([Fig. 26, Processor 2601]), the memory having computer-readable instructions therein ([0387] random access memory (RAM) or another type of dynamic storage device capable of storing information and instructions), and the computer-readable instructions, when executed by the processor, cause the computer device to perform a voice extraction method ([0010] to extract the voice signal of the target user from the noisy voice signal) including:
obtaining a registered voice of a speaker ([0162] registered voice of the target user are input to a voice noise reduction model for processing, [Obtained by the voice noise reduction model]);
determining a registered voice feature of the registered voice ([0174] extracting a feature vector of the registered voice signal of the target user from the registered voice signal, [A feature vector indicates at least one registered voice feature]);
extracting an initial recognition voice of the speaker from a mixed voice based on the registered voice feature ([Fig. 4], [0174] extracting a feature vector of the noisy voice signal from the noisy voice signal by using a second encoding network; obtaining a first feature vector based on the feature vector of the registered voice signal and the feature vector of the noisy voice signal, for example, a mathematical operation such as dot multiplication is performed on the feature vector of the registered voice signal and the feature vector of the noisy voice signal to obtain the first feature vector, [In view of Fig. 4 which demonstrates the output from the dot product operation to be sent into a TCN to be an initial recognition voice of the speaker from a mixed, i.e. noisy, voice based on the registered voice feature vector, i.e. output from the encoding networks combined through dot product]),
wherein the mixed voice is a time domain signal ([Fig. 4, Noisy voice signal and Output from the Second encoding network to be processed by Temporal Convolutional Network], [Processing a noisy, i.e. mixed, voice signal to extract a feature vector, i.e. initial recognition voice, as previously cited, to be passed into a Temporal convolutional network after a dot product calculation indicates the mixed voice to be time domain vectors in order to be processed by the Temporal convolutional network. This is in contrast with the embodiment of Fig. 7 which explicitly features FFT blocks for converting input into the frequency domain]).
Wei does not disclose:
wherein the initial recognition voice is a time-domain voice signal.
Xu discloses:
wherein the initial recognition voice is a time-domain voice signal ([Fig. 3, Speaker Extractor comprised of Temporal Convolutional Networks (TCNs) resulting in output M1-M3 and S1-S3], [pg. 1373, right column] we encode the time-domain signal into three temporal resolutions in the embedding E, [pg. 1374, B. Multi-Scale Encoding and Decoding] The speaker extractor then estimates multi-scale masks M1,M2,M3, and generates the multiscale modulated responses S1, S2, S3, [The examiner asserts that the presence of TCN blocks within the Speaker Extractor indicates operations of the extractor to be in the time-domain, wherein any of the modulated responses tracks to an initial recognition voice as being extracted from the mixed signal and is a function of time, see pg. 1374, Eq. (3)]).
Wei and Xu are considered analogous art within target speaker extraction from mixed speech. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei to incorporate the teachings of Xu, because of the novel way to perform multi-scale encoding and decoding over multiple temporal resolutions, improving the resultant voice quality extracted from the mixed speech signal (Xu, [pg. 1371, right column, contribution 4]). It would be obvious to take the input signals of Wei (which the examiner asserts are in the time domain before FFT conversion, Fig. 7) and keep them in the time domain to be applied to the multi-scale process of Xu to result in the improved target speaker extraction quality as Fig. 4 of Wei discloses a system which appears to contain two signals in the time domain (in order to be processed by the temporal convolutional network of Fig. 4), wherein the output from the second encoding network is an extracted feature vector, i.e. initial recognition voice vector, as previously cited.
Wei in view of Xu does not disclose:
identifying multiple voice segments in the initial recognition voice.
Tan discloses:
identifying multiple voice segments in the initial recognition voice ([0018] The computing system 100 may use a portion of the audio (e.g., the first 5 seconds of the audio, the first 20 seconds of the audio, etc.) to generate a signature vector… The computing system 100 may compare the voice signature with other portions of the audio the vector generated for the beginning portion may be compared with vectors generated for other portions of the audio), [Other portions in addition to a first portion indicates multiple voice segments, i.e. portions]).
Wei, Xu, and Tan are considered analogous art within voice quality enhancement. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu to incorporate the teachings of Tan, because of the novel way to create a voice biometric for a user with call audio including background noise through removal of the background noise via segmentation of a received conversation into user voice/not based on a similarity threshold comparison to a known voiced segment, allowing for the creation of more accurate biometric samples for user authentication (Tan, [0001]-[0002]).
Wei further discloses:
determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice ([0175] the first noisy voice signal is input to the second encoding network frame by frame for feature extraction, to obtain a voice feature vector of each frame, [0224] when the third distortion score is greater than the eighth threshold (for example, 12 dB) and the SNR of the third noise segment is not less than the seventh threshold, it indicates that a voiceprint feature of the current user matches a stored sound feature, [Wherein SNR and/or distortion score tracks to the similarity metric for comparison to thresholds of current user voices, i.e. initial recognition voices, to stored, i.e. registered, threshold sound features, i.e. SNR and/or distortion values. See storing of registered voice samples, [0241]. Further, analyzing audio in “segments”, in view of the previously disclosed frame-level feature extraction, indicates the segment corresponds to a frame as feature vectors are used for the similarity comparison]).
Tan further discloses:
selecting, from the multiple voice segments, a first subset of voice segments whose associated voice similarities are less than a preset similarity ([0018] The computing system 100 may remove portions from the audio, for example, if the similarity score of a portion does not satisfy a similarity threshold. For example, a portion of the audio may include traffic noises and a vector generated for the portion may not satisfy a similarity threshold when compared with the signature vector, [Wherein the signature vector of Tan represents an earlier portion of the same audio, but it would have been obvious to apply the stored sound feature of Wei as the signature vector of Tan as Tan discloses a first portion of audio with only one person speaking (see [0019] disclose removal of anything other than one individual) indicating the stored sound feature voiceprint representing a historical user of Wei could be substituted for the signature vector of Tan without a change in functionality to Tan to come to reach the claimed functionality]); and
removing the first subset of voice segments from the initial recognition voice to obtain a clean voice of the speaker ([0019] Portions of the audio that include one or more interfering noises may be removed so that a voice biometric for the user may be generated using the remaining portions, [A voice capable of being used for a voice biometric with interfering portions removed indicates the biometric to be a clean voice]).
Regarding claim 10, Wei in view of Xu, further in view of Tan discloses: the computer device according to claim 9.
Wei further discloses:
wherein the extracting an initial recognition voice of the speaker from a mixed voice based on the registered voice feature comprises:
determining a mixed voice feature of the mixed voice ([0174] extracting a feature vector of the noisy voice signal from the noisy voice signal);
fusing the mixed voice feature and the registered voice feature of the registered voice to obtain a voice fusion feature ([0181] performing fusion on the frequency domain signal of the VPU signal of the target user and the frequency domain signal of the first noisy voice signal to obtain a first fusion frequency domain signal, [0182] the first fusion frequency domain signal is input to the third encoding network for feature extraction to obtain a feature vector of the first fusion frequency domain signal, [In view of [0163] which defines “registered VPU signal[s]” indicating the registered signals and VPU signals are synonymous. A first fusion feature vector indicates at least one voice fusion feature comprising the vector]); and,
performing initial recognition on voice information of the speaker in the mixed voice based on the voice fusion feature to obtain the initial recognition voice of the speaker ([0181] separately processing the first fusion frequency domain signal by using a third encoding network, a GRU, and a second decoding network to obtain a mask of a frequency domain signal of the voice signal of the target user, [Generating a mask based on a voice fusion signal for determining a voice signal of a target user indicates the mask to be representative of an initial recognition voice of the speaker]).
Regarding claim 14, Wei in view of Xu, further in view of Tan discloses: the computer device according to claim 9.
Wei further discloses:
wherein the determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice comprises:
repeating the each voice segment in the initial recognition voice based on a time length of the registered voice separately to obtain a recombined voice having the time length ([0219] the terminal device obtains, based on a preset cycle (for example, every 10 minutes), a first noise segment (for example, a 6 s voice signal currently captured by the microphone) and a second noise segment (for example, a 10 s voice signal after the 6 s voice signal currently captured by the microphone) of an environment in which the terminal device is located, and obtains an SNR and an SPL of the first noise segment, [Wherein the “repeating” is performed through the SNR/SPL calculations of each segment of the noise, i.e. mixed, signals and/or the obtaining of the segments]);
obtaining a recombined voice feature extracted from the recombined voice ([0219] determines whether the SNR of the first noise segment is greater than 20 dB and whether the SPL is greater than 40 dB; if the SNR of the first noise segment is greater than the first threshold (for example, 20 dB) and the SPL is greater than the second threshold (for example, 40 dB), extracts a first temporary feature vector of the first noise segment; performs noise reduction on the second noise segment by using the first temporary feature vector to obtain a second noise-reduced noise segment, [Generating a feature vector for the second noise segment, wherein that vector is dependent upon a first noise feature vector, indicates the second feature vector to be a combination of the first and second, i.e. recombined]);
determining a segment voice feature corresponding to the each voice segment in the initial recognition voice based on the recombined voice feature ([0219] performs distortion evaluation based on the second noise-reduced noise segment and the second noise segment to obtain a first distortion score, where the first distortion score indicates a degree of distortion of the signal captured by the microphone of the terminal device, [Performing distortion evaluation, wherein the evaluation is performed based on a feature comparison, see [0221] defining the distortion score to be a signal-to-distortion ratio (SDR), i.e. SDR is a feature of audio, indicates a required segment voice feature determination of the second noise segment to be compared to the first distortion score, i.e. also representing a segment of audio, wherein the second noise-reduced segment represents a recombined voice feature as previously disclosed. Further, the second noise segment tracks to an initial recognition voice]); and,
determining a voice similarity between the registered voice and the each voice segment based on the segment voice feature corresponding to the each voice segment and the registered voice feature separately ([0222] performing noise reduction on the third noise segment based on the reference temporary voiceprint feature vector to obtain a third noise-reduced noise segment; performing distortion evaluation based on the third noise segment and the third noise-reduced noise segment to obtain a third distortion score; if the third distortion score is greater than a sixth threshold and an SNR of the third noise segment is less than a seventh threshold, or the third distortion score is greater than an eighth threshold and the SNR of the third noise segment is not less than the seventh threshold, sending third prompt information, [0241] The registered voice of the current user may be obtained by connecting a plurality of segments of signals in series, and total duration is not less than 6 s, [Applying a temporary feature vector which results in a noise-reduced segment to be compared to a reference threshold, i.e. registered, segment SNR/SDR indicates the similarity comparison to be between a registered voice having the threshold ratio levels and each voice segment based on the segment voice feature, i.e. SNR/SDR values. Further, consider generation of temporary feature vectors for registered voice signals, [0231], indicating a noise reduction performed based on a similarity comparison of a registered feature vector to a voice segment feature vector]).
Regarding claim 15, Wei in view of Xu, further in view of Tan discloses the computer device according to claim 9.
Wei further discloses: wherein the obtaining a registered voice of a speaker comprises:
determining a speaker in response to a call triggering operation ([0240] When it is detected that the terminal device is in a hands-free call state, the terminal device enters the PNR mode, and a terminal device owner whose voiceprint feature has been registered is a target user, [PNR tracks to “personalized noise reduction”, [0007]]); and,
determining the registered voice of the speaker from a prestored candidate registered voice ([0178] The voice signal of the target user is registered in advance, [0241] If registered voice of the current user or a voice feature of the current user is stored, the current user is determined as a target user, [Generating a registered voice, wherein the registration is performed in advance, indicates the registered voice is determined from a prestored candidate registered voice]).
Regarding claim 16, Wei in view of Xu, further in view of Tan discloses: the computer device according to claim 9.
Wei further discloses:
wherein the mixed voice is transmitted from a terminal of a speaker after a voice call is established with the terminal ([0227] when it is detected that the terminal device is used again for a call, the default microphone of the terminal device captures the second noisy voice signal, [0246] a call function of the terminal device is implemented through a “Phone” application, [A second noisy voice signal tracks to a mixed signal. Further, disclosing the terminal to be a phone indicates at least one speaker for playing the noisy sound to be captured by the microphone of the terminal device. See Fig. 15 “Volume” control on the phone terminal and Fig. 17 phone call setting to “Enable PNR” after the phone call is connected]).
Regarding claim 17, Wei discloses: a non-transitory computer-readable storage medium, having computer-readable instructions stored thereon (Claim 20, A non-transitory machine readable storage medium having instructions stored therein, [0387] random access memory (RAM) or another type of dynamic storage device capable of storing information and instructions, [RAM is physical memory]), and the computer-readable instructions, when executed by a processor of a computing device ([Fig. 26, Processor 2601]), causing the computer device to perform a voice extraction method ([0010] to extract the voice signal of the target user from the noisy voice signal) including:
obtaining a registered voice of a speaker ([0162] registered voice of the target user are input to a voice noise reduction model for processing, [Obtained by the voice noise reduction model]);
determining a registered voice feature of the registered voice ([0174] extracting a feature vector of the registered voice signal of the target user from the registered voice signal, [A feature vector indicates at least one registered voice feature]);
extracting an initial recognition voice of the speaker from a mixed voice based on the registered voice feature ([Fig. 4], [0174] extracting a feature vector of the noisy voice signal from the noisy voice signal by using a second encoding network; obtaining a first feature vector based on the feature vector of the registered voice signal and the feature vector of the noisy voice signal, for example, a mathematical operation such as dot multiplication is performed on the feature vector of the registered voice signal and the feature vector of the noisy voice signal to obtain the first feature vector, [In view of Fig. 4 which demonstrates the output from the dot product operation to be sent into a TCN to be an initial recognition voice of the speaker from a mixed, i.e. noisy, voice based on the registered voice feature vector, i.e. output from the encoding networks combined through dot product]),
wherein the mixed voice is a time domain signal ([Fig. 4, Noisy voice signal and Output from the Second encoding network to be processed by Temporal Convolutional Network], [Processing a noisy, i.e. mixed, voice signal to extract a feature vector, i.e. initial recognition voice, as previously cited, to be passed into a Temporal convolutional network after a dot product calculation indicates the mixed voice to be time domain vectors in order to be processed by the Temporal convolutional network. This is in contrast with the embodiment of Fig. 7 which explicitly features FFT blocks for converting input into the frequency domain]).
Wei does not disclose:
wherein the initial recognition voice is a time-domain voice signal.
Xu discloses:
wherein the initial recognition voice is a time-domain voice signal ([Fig. 3, Speaker Extractor comprised of Temporal Convolutional Networks (TCNs) resulting in output M1-M3 and S1-S3], [pg. 1373, right column] we encode the time-domain signal into three temporal resolutions in the embedding E, [pg. 1374, B. Multi-Scale Encoding and Decoding] The speaker extractor then estimates multi-scale masks M1,M2,M3, and generates the multiscale modulated responses S1, S2, S3, [The examiner asserts that the presence of TCN blocks within the Speaker Extractor indicates operations of the extractor to be in the time-domain, wherein any of the modulated responses tracks to an initial recognition voice as being extracted from the mixed signal and is a function of time, see pg. 1374, Eq. (3)]).
Wei and Xu are considered analogous art within target speaker extraction from mixed speech. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei to incorporate the teachings of Xu, because of the novel way to perform multi-scale encoding and decoding over multiple temporal resolutions, improving the resultant voice quality extracted from the mixed speech signal (Xu, [pg. 1371, right column, contribution 4]). It would be obvious to take the input signals of Wei (which the examiner asserts are in the time domain before FFT conversion, Fig. 7) and keep them in the time domain to be applied to the multi-scale process of Xu to result in the improved target speaker extraction quality as Fig. 4 of Wei discloses a system which appears to contain two signals in the time domain (in order to be processed by the temporal convolutional network of Fig. 4), wherein the output from the second encoding network is an extracted feature vector, i.e. initial recognition voice vector, as previously cited.
Wei in view of Xu does not disclose:
identifying multiple voice segments in the initial recognition voice.
Tan discloses:
identifying multiple voice segments in the initial recognition voice ([0018] The computing system 100 may use a portion of the audio (e.g., the first 5 seconds of the audio, the first 20 seconds of the audio, etc.) to generate a signature vector… The computing system 100 may compare the voice signature with other portions of the audio the vector generated for the beginning portion may be compared with vectors generated for other portions of the audio), [Other portions in addition to a first portion indicates multiple voice segments, i.e. portions]).
Wei, Xu, and Tan are considered analogous art within voice quality enhancement. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu to incorporate the teachings of Tan, because of the novel way to create a voice biometric for a user with call audio including background noise through removal of the background noise via segmentation of a received conversation into user voice/not based on a similarity threshold comparison to a known voiced segment, allowing for the creation of more accurate biometric samples for user authentication (Tan, [0001]-[0002]).
Wei further discloses:
determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice ([0175] the first noisy voice signal is input to the second encoding network frame by frame for feature extraction, to obtain a voice feature vector of each frame, [0224] when the third distortion score is greater than the eighth threshold (for example, 12 dB) and the SNR of the third noise segment is not less than the seventh threshold, it indicates that a voiceprint feature of the current user matches a stored sound feature, [Wherein SNR and/or distortion score tracks to the similarity metric for comparison to thresholds of current user voices, i.e. initial recognition voices, to stored, i.e. registered, threshold sound features, i.e. SNR and/or distortion values. See storing of registered voice samples, [0241]. Further, analyzing audio in “segments”, in view of the previously disclosed frame-level feature extraction, indicates the segment corresponds to a frame as feature vectors are used for the similarity comparison]).
Tan further discloses:
selecting, from the multiple voice segments, a first subset of voice segments whose associated voice similarities are less than a preset similarity ([0018] The computing system 100 may remove portions from the audio, for example, if the similarity score of a portion does not satisfy a similarity threshold. For example, a portion of the audio may include traffic noises and a vector generated for the portion may not satisfy a similarity threshold when compared with the signature vector, [Wherein the signature vector of Tan represents an earlier portion of the same audio, but it would have been obvious to apply the stored sound feature of Wei as the signature vector of Tan as Tan discloses a first portion of audio with only one person speaking (see [0019] disclose removal of anything other than one individual) indicating the stored sound feature voiceprint representing a historical user of Wei could be substituted for the signature vector of Tan without a change in functionality to Tan to come to reach the claimed functionality]); and
removing the first subset of voice segments from the initial recognition voice to obtain a clean voice of the speaker ([0019] Portions of the audio that include one or more interfering noises may be removed so that a voice biometric for the user may be generated using the remaining portions, [A voice capable of being used for a voice biometric with interfering portions removed indicates the biometric to be a clean voice]).
Regarding claim 18, Wei in view of Xu, further in view of Tan discloses: the non-transitory computer-readable storage medium according to claim 17.
Wei further discloses:
wherein the extracting an initial recognition voice of the speaker from a mixed voice based on the registered voice feature comprises:
determining a mixed voice feature of the mixed voice ([0174] extracting a feature vector of the noisy voice signal from the noisy voice signal);
fusing the mixed voice feature and the registered voice feature of the registered voice to obtain a voice fusion feature ([0181] performing fusion on the frequency domain signal of the VPU signal of the target user and the frequency domain signal of the first noisy voice signal to obtain a first fusion frequency domain signal, [0182] the first fusion frequency domain signal is input to the third encoding network for feature extraction to obtain a feature vector of the first fusion frequency domain signal, [In view of [0163] which defines “registered VPU signal[s]” indicating the registered signals and VPU signals are synonymous. A first fusion feature vector indicates at least one voice fusion feature comprising the vector]); and,
performing initial recognition on voice information of the speaker in the mixed voice based on the voice fusion feature to obtain the initial recognition voice of the speaker ([0181] separately processing the first fusion frequency domain signal by using a third encoding network, a GRU, and a second decoding network to obtain a mask of a frequency domain signal of the voice signal of the target user, [Generating a mask based on a voice fusion signal for determining a voice signal of a target user indicates the mask to be representative of an initial recognition voice of the speaker]).
Regarding claim 20, Wei in view of Xu, further in view of Tan discloses: the non-transitory computer-readable storage medium according to claim 17.
Wei further discloses:
wherein the determining, based on the registered voice feature, a voice similarity between the registered voice and voice information comprised in each voice segment in the initial recognition voice comprises:
repeating the each voice segment in the initial recognition voice based on a time length of the registered voice separately to obtain a recombined voice having the time length ([0219] the terminal device obtains, based on a preset cycle (for example, every 10 minutes), a first noise segment (for example, a 6 s voice signal currently captured by the microphone) and a second noise segment (for example, a 10 s voice signal after the 6 s voice signal currently captured by the microphone) of an environment in which the terminal device is located, and obtains an SNR and an SPL of the first noise segment, [Wherein the “repeating” is performed through the SNR/SPL calculations of each segment of the noise, i.e. mixed, signals and/or the obtaining of the segments]);
obtaining a recombined voice feature extracted from the recombined voice ([0219] determines whether the SNR of the first noise segment is greater than 20 dB and whether the SPL is greater than 40 dB; if the SNR of the first noise segment is greater than the first threshold (for example, 20 dB) and the SPL is greater than the second threshold (for example, 40 dB), extracts a first temporary feature vector of the first noise segment; performs noise reduction on the second noise segment by using the first temporary feature vector to obtain a second noise-reduced noise segment, [Generating a feature vector for the second noise segment, wherein that vector is dependent upon a first noise feature vector, indicates the second feature vector to be a combination of the first and second, i.e. recombined]);
determining a segment voice feature corresponding to the each voice segment in the initial recognition voice based on the recombined voice feature ([0219] performs distortion evaluation based on the second noise-reduced noise segment and the second noise segment to obtain a first distortion score, where the first distortion score indicates a degree of distortion of the signal captured by the microphone of the terminal device, [Performing distortion evaluation, wherein the evaluation is performed based on a feature comparison, see [0221] defining the distortion score to be a signal-to-distortion ratio (SDR), i.e. SDR is a feature of audio, indicates a required segment voice feature determination of the second noise segment to be compared to the first distortion score, i.e. also representing a segment of audio, wherein the second noise-reduced segment represents a recombined voice feature as previously disclosed. Further, the second noise segment tracks to an initial recognition voice]); and,
determining a voice similarity between the registered voice and the each voice segment based on the segment voice feature corresponding to the each voice segment and the registered voice feature separately ([0222] performing noise reduction on the third noise segment based on the reference temporary voiceprint feature vector to obtain a third noise-reduced noise segment; performing distortion evaluation based on the third noise segment and the third noise-reduced noise segment to obtain a third distortion score; if the third distortion score is greater than a sixth threshold and an SNR of the third noise segment is less than a seventh threshold, or the third distortion score is greater than an eighth threshold and the SNR of the third noise segment is not less than the seventh threshold, sending third prompt information, [0241] The registered voice of the current user may be obtained by connecting a plurality of segments of signals in series, and total duration is not less than 6 s, [Applying a temporary feature vector which results in a noise-reduced segment to be compared to a reference threshold, i.e. registered, segment SNR/SDR indicates the similarity comparison to be between a registered voice having the threshold ratio levels and each voice segment based on the segment voice feature, i.e. SNR/SDR values. Further, consider generation of temporary feature vectors for registered voice signals, [0231], indicating a noise reduction performed based on a similarity comparison of a registered feature vector to a voice segment feature vector]).
Claim(s) 3, 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei in view of Xu, further in view of Tan, further in view of Yu (US-20170178666-A1).
Regarding claim 3, Wei in view of Xu, further in view of Tan discloses: the method according to claim 2.
Wei in view of Xu, further in view of Tan does not disclose:
wherein the determining a mixed voice feature of the mixed voice comprises:
extracting an amplitude spectrum of the mixed voice to obtain a first amplitude spectrum;
performing feature extraction on the first amplitude spectrum to obtain an amplitude spectrum feature; and,
performing feature extraction on the amplitude spectrum feature to obtain the mixed voice feature of the mixed voice.
Yu discloses:
wherein the determining a mixed voice feature of the mixed voice comprises:
extracting an amplitude spectrum of the mixed voice to obtain a first amplitude spectrum ([0045] an audio signal is processed using one or multiple types of feature extraction processes. The features can comprise multiple representations extracted during different methods. Exemplary methods include amplitude modulation spectrograph, [Wherein Yu discloses a multi-speaker speech input, see Abstract, indicating this operation to be applied to mixed voice in view of the previously disclosed mixed signal of Wei which could be substituted for the multi-speaker signal of Yu without a change in functionality to Yu]);
performing feature extraction on the first amplitude spectrum to obtain an amplitude spectrum feature ([0045] trained to receive features extracted using amplitude modulation spectrograms); and,
performing feature extraction on the amplitude spectrum feature to obtain the mixed voice feature of the mixed voice ([0087] wherein amplitude modulation spectrograph is used to generate the acoustic features, [The examiner would also like to note that it is unclear how one is able to extract additional features from previously extracted features and/or what the resultant of this additional extraction would represent. For analysis of this claim element, a similar extraction to that previously cited will be mapped here]).
Wei, Xu, Tan, and Yu are considered analogous art within multi-speaker speech enhancement. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu, further in view of Tan to incorporate the teachings of Yu, because of the novel way to implement a multiple-output layer RNN for processing acoustic speech signals comprising speech from multiple speakers, allowing for the tracking of each signal individually, improving the quality of masks generated for noise signals (Yu, [0004]).
Regarding claim 11, Wei in view of Xu, further in view of Tan discloses: the computer device according to claim 10.
Wei in view of Xu, further in view of Tan does not disclose:
wherein the determining a mixed voice feature of the mixed voice comprises:
extracting an amplitude spectrum of the mixed voice to obtain a first amplitude spectrum;
performing feature extraction on the first amplitude spectrum to obtain an amplitude spectrum feature; and,
performing feature extraction on the amplitude spectrum feature to obtain the mixed voice feature of the mixed voice.
Yu discloses:
wherein the determining a mixed voice feature of the mixed voice comprises:
extracting an amplitude spectrum of the mixed voice to obtain a first amplitude spectrum ([0045] an audio signal is processed using one or multiple types of feature extraction processes. The features can comprise multiple representations extracted during different methods. Exemplary methods include amplitude modulation spectrograph, [Wherein Yu discloses a multi-speaker speech input, see Abstract, indicating this operation to be applied to mixed voice in view of the previously disclosed mixed signal of Wei which could be substituted for the multi-speaker signal of Yu without a change in functionality to Yu]);
performing feature extraction on the first amplitude spectrum to obtain an amplitude spectrum feature ([0045] trained to receive features extracted using amplitude modulation spectrograms); and,
performing feature extraction on the amplitude spectrum feature to obtain the mixed voice feature of the mixed voice ([0087] wherein amplitude modulation spectrograph is used to generate the acoustic features, [The examiner would also like to note that it is unclear how one is able to extract additional features from previously extracted features and/or what the resultant of this additional extraction would represent. For analysis of this claim element, a similar extraction to that previously cited will be mapped here]).
Wei, Xu, Tan, and Yu are considered analogous art within multi-speaker speech enhancement. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu, further in view of Tan to incorporate the teachings of Yu, because of the novel way to implement a multiple-output layer RNN for processing acoustic speech signals comprising speech from multiple speakers, allowing for the tracking of each signal individually, improving the quality of masks generated for noise signals (Yu, [0004]).
Claim(s) 4, 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei in view of Xu, further in view of Tan, further in view of Mesgarani et al. (US-20230377595-A1), hereinafter Mesgarani.
Regarding claim 4, Wei in view of Xu, further in view of Tan discloses: the method according to claim 2.
Wei in view of Xu, further in view of Tan does not disclose:
wherein the performing initial recognition on voice information of the speaker in the mixed voice based on the voice fusion feature to obtain the initial recognition voice of the speaker comprises:
performing initial recognition on the voice information of the speaker in the mixed voice based on the voice fusion feature to obtain a voice feature of the speaker;
performing feature decoding on the voice feature of the speaker to obtain a second amplitude spectrum; and,
transforming the second amplitude spectrum based on a phase spectrum of the mixed voice to obtain the initial recognition voice of the speaker.
Mesgarani discloses:
wherein the performing initial recognition on voice information of the speaker in the mixed voice based on the voice fusion feature to obtain the initial recognition voice of the speaker comprises:
performing initial recognition on the voice information of the speaker in the mixed voice based on the voice fusion feature to obtain a voice feature of the speaker ([0100] additional separation techniques can, for example, include a hint fusion module to create a composite signal from the captured multi-source signal and the speaker-attending information, [0116] architecture 1400 includes a feature extraction section 1410 that includes multiple encoders (in the example of FIG. 14, two encoders 1412 and 1414 are depicted) which are shared by the mixture signals from both channels, and the encoder outputs for each channel are combined (e.g., concatenated, integrated, or fused in some manner)… encoders' output and passed to a separator section 1420 (also referred to as a mask estimation network), [The encoder output tracks to a series of voice features to be used for generating the mask. Further, taking the fusing of two encoder outputs in view of the fused registered and mixed signals of Wei indicates the fused feature vectors to be equivalent between the two pieces of art.]);
performing feature decoding on the voice feature of the speaker to obtain a second amplitude spectrum ([Fig. 15, Output from Encoders 1514 and 1512 being passed into decoders 1524 and 1522 respectively, ILD 1516], [0121] To enhance the extraction of the spatial features, the interaural phase difference (IPD) information and interaural level difference (ILD) information are explicitly added as additional information/features to the outputs of the encoders 1512 and 1514, [Adding amplitude information, i.e. interaural level difference (describing sound intensity or loudness between two receivers, tracking to amplitude), to an output encoding, wherein that encoding is directly passed into decoding indicates the amplitude information added to the output encoding will result in an output amplitude spectrum when decoded. Consider output spectrogram defined in [0084] used for generating a time-domain audio output signal, indicating a clear amplitude dependency on the spectrogram for this output signal generation.]); and,
transforming the second amplitude spectrum based on a phase spectrum of the mixed voice to obtain the initial recognition voice of the speaker ([Fig. 15, 1516 cosIPD/sinIPD], [0114] The TasNet pipeline incorporates cross-channel features into the single-channel model, where spatial features such as interaural phase difference (IPD) is concatenated with the mixture encoder output on a selected reference microphone for mask estimation, [As shown in Fig. 15, the phase and amplitude, i.e. level, differences are concatenated together with encoder output to be passed into the decoders which generate the second amplitude spectrum and time-domain audio output signal as previously disclosed. Considering this, it is apparent that the phase spectrum will have a role in transforming the second amplitude spectrum, i.e. output from the decoder (the second amplitude spectrum) will also be dependent upon the phase differences received as part of the decoder input, indicating the output from the decoder has a dependence in transforming the amplitude spectrum, e.g. decoder output, based on phase spectrum, i.e. decoder input]).
Wei, Xu, Tan, and Mesgarani are considered analogous art within multi-speaker speech separation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu, further in view of Tan to incorporate the teachings of Mesgarani, because of the novel way to jointly perform speech extraction and neural decoding guided in a robust single channel process, alleviating the need for a prior assumption of number of speakers in mixed audio, further improving speech separation capabilities (Mesgarani, [0004]).
Regarding claim 12, Wei in view of Xu, further in view of Tan discloses: the computer device according to claim 10.
Wei in view of Xu, further in view of Tan does not disclose:
wherein the performing initial recognition on voice information of the speaker in the mixed voice based on the voice fusion feature to obtain the initial recognition voice of the speaker comprises:
performing initial recognition on the voice information of the speaker in the mixed voice based on the voice fusion feature to obtain a voice feature of the speaker;
performing feature decoding on the voice feature of the speaker to obtain a second amplitude spectrum; and,
transforming the second amplitude spectrum based on a phase spectrum of the mixed voice to obtain the initial recognition voice of the speaker.
Mesgarani discloses:
wherein the performing initial recognition on voice information of the speaker in the mixed voice based on the voice fusion feature to obtain the initial recognition voice of the speaker comprises:
performing initial recognition on the voice information of the speaker in the mixed voice based on the voice fusion feature to obtain a voice feature of the speaker ([0100] additional separation techniques can, for example, include a hint fusion module to create a composite signal from the captured multi-source signal and the speaker-attending information, [0116] architecture 1400 includes a feature extraction section 1410 that includes multiple encoders (in the example of FIG. 14, two encoders 1412 and 1414 are depicted) which are shared by the mixture signals from both channels, and the encoder outputs for each channel are combined (e.g., concatenated, integrated, or fused in some manner)… encoders' output and passed to a separator section 1420 (also referred to as a mask estimation network), [The encoder output tracks to a series of voice features to be used for generating the mask. Further, taking the fusing of two encoder outputs in view of the fused registered and mixed signals of Wei indicates the fused feature vectors to be equivalent between the two pieces of art.]);
performing feature decoding on the voice feature of the speaker to obtain a second amplitude spectrum ([Fig. 15, Output from Encoders 1514 and 1512 being passed into decoders 1524 and 1522 respectively, ILD 1516], [0121] To enhance the extraction of the spatial features, the interaural phase difference (IPD) information and interaural level difference (ILD) information are explicitly added as additional information/features to the outputs of the encoders 1512 and 1514, [Adding amplitude information, i.e. interaural level difference (describing sound intensity or loudness between two receivers, tracking to amplitude), to an output encoding, wherein that encoding is directly passed into decoding indicates the amplitude information added to the output encoding will result in an output amplitude spectrum when decoded. Consider output spectrogram defined in [0084] used for generating a time-domain audio output signal, indicating a clear amplitude dependency on the spectrogram for this output signal generation.]); and,
transforming the second amplitude spectrum based on a phase spectrum of the mixed voice to obtain the initial recognition voice of the speaker ([Fig. 15, 1516 cosIPD/sinIPD], [0114] The TasNet pipeline incorporates cross-channel features into the single-channel model, where spatial features such as interaural phase difference (IPD) is concatenated with the mixture encoder output on a selected reference microphone for mask estimation, [As shown in Fig. 15, the phase and amplitude, i.e. level, differences are concatenated together with encoder output to be passed into the decoders which generate the second amplitude spectrum and time-domain audio output signal as previously disclosed. Considering this, it is apparent that the phase spectrum will have a role in transforming the second amplitude spectrum, i.e. output from the decoder (the second amplitude spectrum) will also be dependent upon the phase differences received as part of the decoder input, indicating the output from the decoder has a dependence in transforming the amplitude spectrum, e.g. decoder output, based on phase spectrum, i.e. decoder input]).
Wei, Xu, Tan, and Mesgarani are considered analogous art within multi-speaker speech separation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu, further in view of Tan to incorporate the teachings of Mesgarani, because of the novel way to jointly perform speech extraction and neural decoding guided in a robust single channel process, alleviating the need for a prior assumption of number of speakers in mixed audio, further improving speech separation capabilities (Mesgarani, [0004]).
Claim(s) 5, 13, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wei in view of Xu, further in view of Tan, further in view of Sivaraman (US-20220084509-A1), hereinafter Sivaraman.
Regarding claim 5, Wei in view of Xu, further in view of Tan discloses: the method according to claim 1.
Wei further discloses:
wherein the determining a registered voice feature of the registered voice comprises:
extracting a frequency spectrum of the registered voice ([0182] a spectrum of the frequency domain signal of the VPU signal of the target user, [Wherein a VPU signal reasonably tracks to a registered voice signal as previously disclosed, see [0163] “registered VPU signal”]). Wei in view of Xu, further in view of Tan does not disclose:
generating a mel-frequency spectrum of the registered voice based on the frequency spectrum; and,
performing feature extraction on the mel-frequency spectrum to obtain the registered voice feature of the registered voice.
Sivaraman discloses:
wherein the determining a registered voice feature of the registered voice comprises:
generating a mel-frequency spectrum of the registered voice based on the frequency spectrum ([0038] magnitude of a frequency spectrum of the clean signal, [Disclosing a frequency spectrum of a clean, i.e. registered, voice, wherein mel-frequency cepstrum coefficients (MFCCs) can be gathered from the spectrum, see below mapping, indicates the frequency spectrum must be a mel-frequency spectrum in order to have MFCCs extracted]); and,
performing feature extraction on the mel-frequency spectrum to obtain the registered voice feature of the registered voice ([0029] The speech separation engine extracts low-level spectral features, such as such as mel-frequency cepstrum coefficients (MFCCs), [In view of Fig. 1B which discloses a clear enrollment, i.e. registered, audio being sent through an embedding extraction (which also receives features), indicating the target voiceprint output from the embedding extraction engine 126 to be features of mel-frequency spectra. Further, the examiner would like to note that it is clear the MFCCs are extracted from the input mixed audio, not enrollment audio, but it would be reasonable to assume that the embeddings extracted from the enrollment audio are in the same form as those from the mixed signal in order for comparison to appropriately generate the speaker mask for speakers who aren’t the target speaker, requiring a comparison to the enrollment embeddings of the target speaker to other embeddings to properly determine the mask(s)]).
Wei, Xu, Tan, and Sivaraman are considered analogous art within multi-speaker speech separation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu, further in view of Tan to incorporate the teachings of Sivaraman, because of the novel way to jointly perform speaker mixture separation and background noise suppression tasks, enhancing the overall perceptual quality of target speech audio (Sivaraman, [0013]).
Regarding claim 13, Wei in view of Xu, further in view of Tan discloses: the computer device according to claim 9.
Wei further discloses:
wherein the determining a registered voice feature of the registered voice comprises:
extracting a frequency spectrum of the registered voice ([0182] a spectrum of the frequency domain signal of the VPU signal of the target user, [Wherein a VPU signal reasonably tracks to a registered voice signal as previously disclosed, see [0163] “registered VPU signal”]). Wei in view of Xu, further in view of Tan does not disclose:
generating a mel-frequency spectrum of the registered voice based on the frequency spectrum; and,
performing feature extraction on the mel-frequency spectrum to obtain the registered voice feature of the registered voice.
Sivaraman discloses:
wherein the determining a registered voice feature of the registered voice comprises:
generating a mel-frequency spectrum of the registered voice based on the frequency spectrum ([0038] magnitude of a frequency spectrum of the clean signal, [Disclosing a frequency spectrum of a clean, i.e. registered, voice, wherein mel-frequency cepstrum coefficients (MFCCs) can be gathered from the spectrum, see below mapping, indicates the frequency spectrum must be a mel-frequency spectrum in order to have MFCCs extracted]); and,
performing feature extraction on the mel-frequency spectrum to obtain the registered voice feature of the registered voice ([0029] The speech separation engine extracts low-level spectral features, such as such as mel-frequency cepstrum coefficients (MFCCs), [In view of Fig. 1B which discloses a clear enrollment, i.e. registered, audio being sent through an embedding extraction (which also receives features), indicating the target voiceprint output from the embedding extraction engine 126 to be features of mel-frequency spectra. Further, the examiner would like to note that it is clear the MFCCs are extracted from the input mixed audio, not enrollment audio, but it would be reasonable to assume that the embeddings extracted from the enrollment audio are in the same form as those from the mixed signal in order for comparison to appropriately generate the speaker mask for speakers who aren’t the target speaker, requiring a comparison to the enrollment embeddings of the target speaker to other embeddings to properly determine the mask(s)]).
Wei, Xu, Tan, and Sivaraman are considered analogous art within multi-speaker speech separation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu, further in view of Tan to incorporate the teachings of Sivaraman, because of the novel way to jointly perform speaker mixture separation and background noise suppression tasks, enhancing the overall perceptual quality of target speech audio (Sivaraman, [0013]).
Regarding claim 19, Wei in view of Xu, further in view of Tan discloses: the non-transitory computer-readable storage medium according to claim 17.
Wei further discloses:
wherein the determining a registered voice feature of the registered voice comprises:
extracting a frequency spectrum of the registered voice ([0182] a spectrum of the frequency domain signal of the VPU signal of the target user, [Wherein a VPU signal reasonably tracks to a registered voice signal as previously disclosed, see [0163] “registered VPU signal”]). Wei in view of Xu, further in view of Tan does not disclose:
generating a mel-frequency spectrum of the registered voice based on the frequency spectrum; and,
performing feature extraction on the mel-frequency spectrum to obtain the registered voice feature of the registered voice.
Sivaraman discloses:
wherein the determining a registered voice feature of the registered voice comprises:
generating a mel-frequency spectrum of the registered voice based on the frequency spectrum ([0038] magnitude of a frequency spectrum of the clean signal, [Disclosing a frequency spectrum of a clean, i.e. registered, voice, wherein mel-frequency cepstrum coefficients (MFCCs) can be gathered from the spectrum, see below mapping, indicates the frequency spectrum must be a mel-frequency spectrum in order to have MFCCs extracted]); and,
performing feature extraction on the mel-frequency spectrum to obtain the registered voice feature of the registered voice ([0029] The speech separation engine extracts low-level spectral features, such as such as mel-frequency cepstrum coefficients (MFCCs), [In view of Fig. 1B which discloses a clear enrollment, i.e. registered, audio being sent through an embedding extraction (which also receives features), indicating the target voiceprint output from the embedding extraction engine 126 to be features of mel-frequency spectra. Further, the examiner would like to note that it is clear the MFCCs are extracted from the input mixed audio, not enrollment audio, but it would be reasonable to assume that the embeddings extracted from the enrollment audio are in the same form as those from the mixed signal in order for comparison to appropriately generate the speaker mask for speakers who aren’t the target speaker, requiring a comparison to the enrollment embeddings of the target speaker to other embeddings to properly determine the mask(s)]).
Wei, Xu, Tan, and Sivaraman are considered analogous art within multi-speaker speech separation. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Wei in view of Xu, further in view of Tan to incorporate the teachings of Sivaraman, because of the novel way to jointly perform speaker mixture separation and background noise suppression tasks, enhancing the overall perceptual quality of target speech audio (Sivaraman, [0013]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Koshinaka et al. (US-20220238119-A1) discloses “A neural network input unit 81 inputs a neural network in which a first network having a layer for inputting an anchor signal belonging to a predetermined class and a mixed signal including a target signal belonging to the class and a layer for outputting, as an estimation result, a reconstruction mask indicating a time-frequency domain in which the target signal is present in the mixed signal, and a second network having a layer for inputting the target signal extracted by applying the mixed signal to the reconstruction mask and a layer for outputting a result obtained by classifying the input target signal into a predetermined class are combined. A reconstruction mask estimation unit 82 applies the anchor signal and mixed signal to the first network to estimate the reconstruction mask of the class to which the anchor signal belongs. A signal classification unit 83 applies the mixed signal to the estimated reconstruction mask to extract the target signal, and applies the extracted target signal to the second network to classify the target signal into the class.” (abstract). See entire document.
Peng et al. (US-20150325252-A1) discloses “A method and device for eliminating noise, and a mobile terminal. The method comprises: extracting, from the voice of a talker, an audio fingerprint of the talker voice in advance (101); and when the talker talks with an opposite listener, according to the audio fingerprint of the talker, extracting a voice which matches the audio fingerprint from the current talking voice, and sending to the opposite listener the voice which matches the audio fingerprint through a communication network (102)” (abstract). See entire document.
Dittmar et al. (US-10373623-B2) discloses “an apparatus described by a schematic block diagram for processing an audio signal to obtain a processed audio signal. The apparatus includes a phase calculator for calculating phase values for spectral values of a sequence of frequency-domain frames representing overlapping frames of the audio signal. Moreover, the phase calculator is configured to calculate the phase values based on information on a target time-domain envelope related to the processed audio signal, so that the processed audio signal has at least in an approximation the target time-domain envelope and a spectral envelope determined by the sequence of frequency-domain frames” (abstract). See entire document.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THEODORE JOHN WITHEY whose telephone number is (703)756-1754. The examiner can normally be reached Monday - Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571) 272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THEODORE WITHEY/Examiner, Art Unit 2655
/JESSE S PULLIAS/Primary Examiner, Art Unit 2655 08/18/26