DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to claim amendment filed on March 30, 2026 and wherein claims 1-5, 8-9, 11-12, 15-16 amended, claims 6-7, 10 canceled.
In virtue of this communication, claims 1-5, 8-9, 11-18 are currently pending in this Office Action.
With respect to the objection of claims 2-5 due to formality issue, as set forth in the previous Office Action, the claim amendment, and argument, see paragraph 3 of page 8 in Remarks filed on March 30, 2026, have been fully considered and the argument is persuasive. Therefore, the objection of claims 2-5 due to the formality issue, as set forth in the previous Office Action, has been withdrawn.
The Office appreciates the explanation of the amendment and analyses of the prior arts, and however, although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993) and MPEP 2145.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-5, 8-9, 11-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim recites “determining … at least one voice segment” “contains the reference voice” by matching “a voice feature corresponding to the at least one voice segment” to “(bottleneck) feature corresponding to the reference voice” from a “result” from “a voice separation model” by taking “the voice feature” and “the (bottleneck) feature” as inputs, etc., and wherein the limitation of “obtaining”, “generating”, “inputting”, “determining”, etc., as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation by mathematic concept and human mind. For example, “obtaining a voice feature corresponding to a voice” would be interpreted as using a well-known FFT or DFT to have time-frequency feature, or waveform extraction to have temporal feature, etc. Similarly, “generating a bottleneck feature” from “reference voice” by “performing a non-linear transformation on the reference voice” is typical math processing applied on “voice”, and wherein the claimed term “voice” herein would be nothing more than a math variable or symbol under its BRI. Claim further recites an intended purpose “to reduce a dimension of the reference voice” which would be interpreted as nothing more than modification of amount including size or number of math variable “reference voice” under its BRI. The claimed “a voice separation model” would be interpreted as feature comparator at a high level, and thus, covers human mind to perform the comparison of collected data (Example 40 of “101 Examples 37-42”). Claim broadly recites “voice to be processed” and “reference voice” with no recitation of what it is and how it is processed, a BRI would be given and interpreted as nothing more than math variables. Therefore, under their BRIs, the claimed “obtaining”, “generating”, “inputting”, “determining” applied on “features” of “voice to be processed” and “of reference voice”, covers performance of the limitations in mathematic concept and human mind. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim ends at the unprocessed “voice segments”, except a judgement that “at least one segment” may contain “similar” feature to “reference voice”, which does not counted as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to “a non-linear transformation” applied on “reference voice to reduce a dimension of the reference voice” as drafted in claim, which is math manipulation purported to achieve the modification of amount that may include size, number, or volume, etc., of math variable “reference voice”, which is not counted as additional element being sufficiently more than the judicial exception because it does not impose any meaningful limitation under its’ BRI. Therefore, the claim is not patent eligible, see 2019 Revised Patent Subject Matter Eligibility Guidance, “2019 PEG”, and “2025 Subject Matter Eligibility Updates”.
Claim 8 recited limitations of claim 1, and further recited additional elements “an electronic device comprising: a memory” “to store computer program instructions” and “a processor” “execute the computer program instructions” in a high level of generality to perform the abstract idea. A generic hardware to perform the abstract idea such that if it amounts no more than mere instructions to apply the exception using the generic computer component, and such generic hardware does not impose meaningful limits to precludes the steps from practically being performed mathematically as discussed in claim 1 above, Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
Claim 9 recited essential limitations as recited in claim 8 and thus, rejected according to at least reasons described in claim 8 above.
Claim 2 depends on claim 1 and further recites additional elements by reciting “a first neural network” and “a second neural network” comprised in the voice separation model, and the first neural networks is applied to “the voice feature” to output “a vector expression” and “the vector expression” and “bottleneck feature” are “spliced” to output “a fusion feature” that is an input to “a second neural network” for outputting “a matrix” and the “detection result” is based on “the matrix”, which is further math concept while broadly recited “neural network” would be interpreted as math algorithms by taking input(s) and outputting output(s), within the “voice separation model”, which does not count as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea and the additional elements are insufficient to amount to significantly more than the judicial exception. Accordingly, the claim does not rectify the 101 issue of parent claim 1 and is not patent eligible.
Claim 3 depends on claim 2 and further recites additional components by reciting “probability values“ related to “a first category” and “a second category” classified as each of “voice segments” and the classification is performed upon “element” in “the matrix” so that “first category” defined as matching “reference feature” and “second category” as unmatching the “reference feature” and the “voice detection result” is upon a maximum value of the probability, which does not count as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea and the additional components are insufficient to amount to significantly more than the judicial exception. Accordingly, the claim does not rectify the 101 issue of parent claim 2 and is not patent eligible.
Claim 4 depends on claim 1 and further added additional components by reciting “obtaining the voice separation model” is by “training” (nothing) upon “sample voice” having “voice feature”, ”bottleneck feature corresponding to the sample voice”, and “a labeled voice detection result of the sample voice, etc., which does not count as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea and the additional components are insufficient to amount to significantly more than the judicial exception. Accordingly, the claim does not rectify the 101 issue of parent claim 1 and is not patent eligible.
Claim 5 depends on claim 1 and further added additional components by reciting one of “FBank feature, a Mel frequency spectrum feature”, or “a pitch feature” comprised in “the voice feature”, which is merely definition of data and does not count as an integration of the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea and the additional components are insufficient to amount to significantly more than the judicial exception. Accordingly, the claim does not rectify the 101 issue of parent claim 1 and is not patent eligible.
Claim 11 depends on claim 8 and recited essential limitations as recited in claim 2 above and thus, rejected according to at least reasons described in claims 8, 2 above.
Claim 12 depends on claim 11 and recited essential limitations as recited in claim 3 above and thus, rejected according to at least reasons described in claims 11, 3 above.
Claim 13 depends on claim 8 and recited essential limitations as recited in claim 4 above and thus, rejected according to at least reasons described in claims 8, 4 above.
Claim 14 depends on claim 8 and recited essential limitations as recited in claim 5 above and thus, rejected according to at least reasons described in claims 8, 5 above.
Claim 15 depends on claim 9 and recited essential limitations as recited in claim 2 above and thus, rejected according to at least reasons described in claims 9, 2 above.
Claim 16 depends on claim 15 and recited essential limitations as recited in claim 3 above and thus, rejected according to at least reasons described in claims 9, 3 above.
Claim 17 depends on claim 9 and recited essential limitations as recited in claim 4 above and thus, rejected according to at least reasons described in claims 9, 4 above.
Claim 18 depends on claim 9 and recited essential limitations as recited in claim 5 above and thus, rejected according to at least reasons described in claims 9, 5 above.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(B) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 4, 13, 17 are rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention.
Claim 4 recited “the voice feature corresponding to the sample voice”, “the bottleneck feature corresponding to the sample voice”, wherein “the voice feature”, “the bottleneck feature”, and “the sample voice” have insufficient antecedent bases for the limitations and causes confusing because it is unclear what they are referred to and it is unclear what they are and thus, renders claim indefinite.
Claims 13, 17 rejected for the at least similar reasons described in claim 4 above since claims 13, 17 recited the similar deficient features as recited in claim 4.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 8, 9 are rejected under 35 U.S.C. 103 as being unpatentable over Xu et al. (CN 113113044 A, original and English translation are attached herein and the English translation applied while paragraphs cited under Xu below) and in view of reference Xu et al. (US 20220270627 A1, hereinafter Xu627) and Bocchieri et al (US 20150100312 A1, hereinafter Bocchieri).
Claim 1: Xu teaches a voice separation method (title and abstract, ln 1-6, a method in fig. 3), comprising:
obtaining a voice feature (including a second voiceprint feature of each pre-separated voice signal for effectively improving the performance, para 4, p.10) corresponding to a voice to be processed (extracted from each of pre-separated via mixed audio, last para of p.8, para 1 of p.9, including target audio and non-target audio, para 2-3, p.8 and the last para of p.8 and para 1-4, p.9), wherein the voice to be processed comprises a plurality of voice segments (the multiple voice signals obtained by a pre-separation module, para 3 from the last paragraph of p.2 and para 7, p.4);
generating a bottleneck feature corresponding to a reference voice (a first voiceprint feature of the target object is determined, the first voiceprint as the bottleneck feature and read specified text content as voice input from a target object or a person, S11 of p.6, para 5-7, p.7) by performing a feature extraction (by a predetermined voiceprint extraction network model to extract the first voiceprint feature of the target object, the last two para of p.4, para 4, p.14);
inputting, into a voice separation model (a portion of the predetermined voice separation network model), the voice feature corresponding to the voice to be processed and the bottleneck feature corresponding to the reference voice (through a third voiceprint feature generated by splicing the second and the first voiceprint features, and the third voiceprint feature is input into the predetermined voice separation network model, para 8-9, p.8 and para 5, p.9);
obtaining a voice detection result output by the voice separation model (the target audio in the mixed audio matches the target object from the person according to the first and the second voiceprints and discussed above, para 6-8, p.9);
determining, on a basis of the voice detection result (the matching is determined discussed above), at least one voice segment of the plurality of voice segments that contains the reference voice (a target audio in the mixed audio matches the target object according to the first voiceprint feature and the second voiceprint feature of each voice signal in the multiple voice signals, the last two para of p.2, and para 1-3 of p.3, and as discussed above, i.e., the audio of the target object is found as the target audio in the multiple voice signals).
However, Xu does not explicitly teach wherein a voice feature corresponding to the at least one voice segment matches the bottleneck feature corresponding to the reference voice and does not explicitly teach wherein performing the feature extraction by which the bottleneck feature corresponding to the reference voice is generated is to perform a non-linear transformation on the reference voice to reduce a dimension of the reference voice.
Xu627 teaches an analogous field of endeavor by disclosing a voice separation method (title and abstract, ln 1-8 and a method in figs. 1-2 and executed by one or more processors, para 6 and a hardware such as smartphones, tablet computers, etc., para 22 and the target audio is separated from the mixed audio) and wherein Xu627 further teaches obtaining a voice feature (obtaining audio features corresponding to audio frames, para 43) corresponding to a voice to be processed (corresponding to at least one of audio frames having the mixed audio during a voice call at step s201 in fig. 2, para 43), wherein the voice to be processed comprises a plurality of voice segments (audio frames corresponding to the mixed audio for audio features, para 43); generating a bottleneck feature corresponding to a reference voice (audio mixing feature of a target object is obtained and/or stored, para 23-24, e.g., a user reading specified text content to realize input voice) by processing the reference voice (sampling the target voice read by the user and discussed above, para 23); inputting, into a voice separation model, the voice feature corresponding to the voice to be processed and the bottleneck feature corresponding to the reference voice (inputting the audio features on the respective audio frames and the audio mixing feature of the target object into separation network model through respective sub-modules, para 43); obtaining a voice detection result output by the voice separation model (obtaining an output results from the respective sub-modules, para 43); and determining, on a basis of the voice detection result, at least one voice segment of the plurality of voice segments that contains the reference voice (the target audio matching with the target object in the mixed audio according to an overall output result of the output results of the respective sub-modules in series, para 43), wherein a voice feature corresponding to the at least one voice segment matches the bottleneck feature corresponding to the reference voice (the audio feature from mixed audio corresponding to audio frames and as an input to the separation network model, and the separation is performed by comparing the input audio feature according to the mixing feature of the audio mixing feature, i.e., checking whether it is matched or not, para 47 and the audio mixing feature is corresponding to audio mixing feature of the target object, para 23) for benefits of improving the sound quality (by improving the reliability of recognition in user recognition, para 18, and by separating voice from noise part of the audio signal, para 52, 55, and by including voiceprint feature and the pitch feature, the recognition rate is effectively improved, para 30 and increasing security of the voice communication by realizing the identification of the user, para 59).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied wherein the voice feature corresponding to the at least one voice segment matches the bottleneck feature corresponding to the reference voice, as taught by Xu627, to voice feature corresponding to the at least one voice segment and the bottleneck feature corresponding to the reference voice, as taught by Xu, for the benefits discussed above.
However, the combination of Xu and Xu627 does not explicitly teach wherein the feature extraction by which the bottleneck feature corresponding to the reference voice is generated is to perform a non-linear transformation on the reference voice to reduce a dimension of the reference voice.
Bocchieri teaches an analogous field of endeavor by disclosing a voice separation method (title and abstract, ln 1-11 and fig. 2-3 and executed on a system in fig. 1, separating the speech from noisy speech, para 13) and wherein the feature extraction is disclosed (fig. 3, from audio as input to have bottleneck features extracted in fig. 3, para 24-25) by which the bottleneck feature corresponding to the reference voice is generated (bottle-neck features extracted from decorrelation 322 based on the audio through 22 MFCCsx 11 frames 302 in fig. 3) and wherein the feature extraction is performed by performing a non-linear transformation on the reference voice (non-linear bottle-neck neural network or hybrid multi-layer perceptual MLP or tandem applied to measure the micro-modulation format-related audio features, as reference voice by using a bottle-neck neural network in fig. 3, para 14, 24-25) to reduce a dimension of the reference voice (from a dimension 242 raw feature supervector, as the dimension of the reference voice, to 60 bottle-neck features as output in fig. 3) and for benefits of improving the effectiveness (by adoption of the MLP-based transform to improve effectiveness with the formant frequencies, para 15, and improving the speech recognition accuracy by adding the micro-modulation feature, para 41, 46, and 50 and specifically in noisy environment, para 2).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied performing the non-linear transformation on the reference voice to reduce the dimension of the reference voice, by which the bottleneck feature corresponding to the reference voice is generated, as taught by Bocchieri, to generating the bottleneck feature corresponding to the reference voice by perform the feature extraction in the voice separation method, as taught by the combination of Xu and Xu627, for the benefits discussed above.
Claim 8 has been analyzed and rejected according to claim 1 above and the combination of Xu, Xu627, and Brocchieri further teaches, an electronic device (Xu, a terminal device in fig. 5, and Xu627, a terminal device in fig. 6, such as smartphone, para 22) comprising:
a memory and a processor (Xu, memory for storing instructions, claim 19, and Xu627, memory 602, para 88, and one or more processors 601 in fig. 6, para 87);
the memory is configured to store computer program instructions (Xu, processor, claim 19, and Xu627, the memory storing instructions, para 7, p.20); and
the processor is configured to execute the computer program instructions, so that the electronic device implements the voice separation method of claim 1 ((Xu, execution of the instructions by the processor, claim 19, and Xu627, the processors to execute instructions for completing method steps of the voice endpoint detection method, para 87).
Claim 9 has been analyzed and rejected according to claims 1, 8 above.
Claims 2, 4-5, 11, 13-15, 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Xu (above) and in view of references X627 (above), Brocchieri (above), and Qu et al. (CN 110648656 A, a copy of the original and an English translation version are attached here, hereinafter Qu, the English translation and original is attached herein and the English translation applied while paragraphs is cited under Qu below).
Claim 2: the combination of Xu, Xu627, and Brocchieri further teaches, according to claim 1 above, wherein the inputting, into a voice separation model, the voice feature corresponding to the voice to be processed and the bottleneck feature corresponding to the reference voice and obtaining a voice detection result output by the voice separation model (Xu and Xu627, the discussed in claim 1 above) comprises:
inputting the voice feature corresponding to the voice to be processed into a first neural network comprised in the voice separation model (Xu, other neural network model for voiceprint extraction by inputting the frequency spectrum of the audio signal to obtain the second voiceprint feature, para 6, p.11 and Xu627, a neural network model is used to obtain the voiceprint feature and obtain the pitch feature, para 36-38);
splicing the voice feature and the bottleneck feature to obtain a third feature (Xu, splicing the second voiceprint feature and the first voiceprint feature of target object to obtain a third voiceprint feature, para 7, p.8);
inputting the third feature into a second neural network comprised in the voice separation model (Xu, the third voiceprint feature is input into a predetermined voice separation network, para 9, p.13 and e.g., a recurrent neural network RNN, para 5, p.10 and Xu627, the target audio matching with the target object in the mixed audio is determined by a classification neural network, para 28), obtaining an output by the second neural network (Xu627, the target audio and the mixed audio are classified, para 28), and obtaining the voice detection result on a basis of the output (Xu627, the target audio is separated from the mixed audio, para 28).
However, the combination of Xu, Xu627, and Brocchieri does not explicitly teach the output is a matrix output by the second neural network and does not explicitly teach a vector expression corresponding to the voice feature output by the first neural network and the vector expression corresponding to the voice feature and the bottleneck feature corresponding to the reference voice is spliced to obtain a fusion feature and the fusion feature is input into the second neural network and obtaining the voice detection result on the basis of the matrix.
Qu teaches an analogous field of endeavor by disclosing a voice separation method (title and abstract, ln 1-15 and a method in figs. 1-5 and executed on an electronic device in fig. 6) and wherein Qu further teaches inputting, into a voice separation model (pre-trained voice detection model, para 4, p.12), the voice feature corresponding to the voice to be processed (frequency characteristics, energy characteristics, and zero-crossing rate characteristics of one or more to-be-detected sound frame, para 4, p.12) and the bottleneck feature corresponding to the reference voice (frequency characteristics, energy characteristics, and zero-crossing rate of other one or more to-be-detected sound frame, para 4, p.12) and obtaining a voice detection result output by the voice separation model (obtaining the detection results of each to-be-detected sound frame, including voice frames and non-speech frames, para 4, para 12) and wherein Qu further teaches:
inputting the voice feature corresponding to the voice to be processed into a first neural network comprised in the voice separation model (including a deep neural network model at the first classification layer, para 7, p.13), and obtaining a vector expression corresponding to the voice feature output by the first neural network (stitched feature matrix of each sound frame to be detected through second fusion unit as fusion feature and then input to a second classification layer as trained neural network model as the speech detection model, and including a delay neural network and LSTM, para 6-10, p.29);
splicing the vector expression corresponding to the voice feature and the bottleneck feature corresponding to the reference voice to obtain a fusion feature (obtaining a fusion feature of each sound frame by stitching features matrix of each sound frame to be detected is linearly mapped to obtain fusion feature of each sound frame to be detected, para 1, p.14); and
inputting the fusion feature into a second neural network comprised in the voice separation model (the fusion feature of each to-be-detected sound frame is input into the first classification layer that is TDNN+LSTM model with the 42-dimentional fusion feature in fig. 3, para 5-7, p.14 or combination of the first classification layer and a second classification layer comprising delay neural network, a long-term and short-term memory network, para 1, p.16), obtaining a matrix output by the second neural network (the matrix output from the first classification layer to the second classification layer, para 1, p.16 and the second classification layer can be delayed neural network TDNN+a long-term and short-term memory network LSTM, para 1, p.17), and obtaining a voice detection result on a basis of the matrix (detection result of each to-be-detected sound frame, speech frame or non-speech frame, para 4, p.14 and para 2, para 16, and including speech endpoint detected, para 2, p.17) for benefits of more accurately separating voice frames from others (the last paragraph of p.2, distinguishing sounds in different types by using different characteristics, para 1, p.10, and by using combined MFCC with FBank features, para 5, p.11, and with LSTER and HZCRR in fig. 3, and by using strong robust model of new technology applications including Deep Neural Network DNN plus LSTM, para 6, p.14);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied wherein the output is the matrix output by the second neural network and the vector expression corresponding to the voice feature output by the first neural network and the vector expression corresponding to the voice feature and the bottleneck feature corresponding to the reference voice is spliced to obtain the fusion feature and the fusion feature is then input into the second neural network and obtaining the voice detection result on the basis of the matrix, as taught by Qu, to inputting the voice feature corresponding to the voice to be processed into the first neural network and splicing the voice feature and the bottleneck feature to obtain the third feature and inputting the third feature into the second neural network, etc., in the voice separation method, as taught by the combination of Xu, Xu627, and Brocchieri for the benefits discussed above.
Claim 4: the combination of Xu, Xu627, Brocchieri, and Qu further teaches, according to claim 1 above, wherein the voice separation model is obtained by training (Xu, training the speech separation network model by using LSTM network and RNN, para 3-5, p.10, and Xu627, using a deep learning training, para 69-72, and Qu, training by using the concatenated or spliced or stitched MFCC feature, LSTER feature, and HZCRR feature, para 2, p.16) on a basis of the voice feature corresponding to the sample voice, the bottleneck feature corresponding to the sample voice (Qu, stitched MFCC feature, LSTER feature, and HZCRR feature above, para 2, p.16) and a labeled voice detection result of the sample voice (Qu, the class label of the frame to be trained and compared, para 2, p.16), and the sample voice comprises the reference voice (Qu, including voice frame and non-voice frame discussed above).
Claim 5: the combination of Xu, Xu627, Brocchieri, and Qu further teaches, according to claim 1 above, wherein the voice feature comprises one or more of a FBank feature, a Mel frequency spectrum feature or a pitch feature (applied Markush, see MPEP 2117, as Markush feature, Xu, including tone, timbre, intensity, sound wave wavelength, frequency, and rhythm of change, etc., para 4, p.7, Xu627, voiceprint feature includes tone, timbre, intensity, the wavelength of sound wave, frequency, and rhythm change, para 26, and Qu, MFCC or Mel Frequency Cepstrum Coefficient, Mel Frequency Cepstrum Coefficient, log spectral characteristics or Fbank feature).
Claim 11 has been analyzed and rejected according to claims 8, 2 above.
Claim 13 has been analyzed and rejected according to claims 8, 4 above.
Claim 14 has been analyzed and rejected according to claims 8, 5 above.
Claim 15 has been analyzed and rejected according to claims 9, 2 above.
Claim 17 has been analyzed and rejected according to claims 9, 4 above.
Claim 18 has been analyzed and rejected according to claims 9, 5 above.
Claims 3, 12, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Xu (above) and in view of references Xu627 (above), Brocchieri (above), Qu (above), and Bocklet et al. (US 20190043488 A1, hereinafter Bocklet).
Claim 3: the combination of Xu, Xu627, Brocchieri, and Qu further teaches, according to claim 2 above, wherein the obtaining the voice detection result on the basis of the matrix (discussion in claims 1-2 above), except the obtaining above comprises:
obtaining probability values matches of that each of the plurality of voice segments pertains to a first category and a second category respectively according to an element corresponding to the each voice segment comprised in the matrix; the voice feature corresponding to the voice segment comprised in the first category matches the bottleneck feature corresponding to the reference voice, and the voice feature corresponding to the voice segment comprised in the second category does not match the bottleneck feature corresponding to the reference voice; and determining a voice detection result corresponding to the each voice segment on a basis of a maximum value of the probability values that the each voice segment pertains to a first category and a second category respectively.
Bocklet teaches an analogous field of endeavor by disclosing a voice separation method (title and abstract, ln 1-5 and method steps in figs. 5, 15) and wherein Bocklet further teaches obtaining a voice detection result on a basis of a matrix (via keyphrase detection through a neural network in fig. 2 and based on an audio input from a user and an feature extraction module 202 that generated feature vectors 212 for acoustic scoring 203 in fig. 2) and comprising:
obtaining probability values (probabilities outputted from acoustic scores 214 and for each of feature vectors 212, para 49) matches of that each of the plurality of voice segments (audio frame of a frame sequence, para 55, and stored in a buffer having a length of 360ms, para 43) pertains to a first category (probabilities for spoken component from a phone, para 50) and a second category (probabilities for silence, background noise, etc., para 50) respectively according to an element corresponding to the each voice segment comprised in the matrix (based on the vector features outputted from feature extraction 202 in fig. 2, para 50); the voice feature corresponding to the voice segment comprised in the first category matches the bottleneck feature corresponding to the reference voice (predetermined key phrase as the bottleneck feature of reference voice stored in key phrase and rejection models 205, para 39, and performing a match of the input feature vector based on the scores or probabilities outputted from element 203 in fig. 2, para 39), and the voice feature corresponding to the voice segment comprised in the second category does not match the bottleneck feature corresponding to the reference voice (probabilities of feature vector of non-speech audio also evaluated and detected through the element 210 within the key phrase detection decoder 204 in fig. 2, para 54); and determining a voice detection result corresponding to the each voice segment on a basis of a maximum value of the probability values that the each voice segment pertains to a first category and a second category respectively (through a maximum pooling on multiple element state score vector, para 31, and pooling different speech and non-speech rejection categories scored by the element 203, and key phrase detection , and indicated by outputting key phrase score 215, para 54) for benefits of improving an efficiency (by accelerated implementation of neural network to reducing computational loads, para 38) and a performance (by implementing pre-processing of microphone signals, para 128) in the voice separation (increasing accurate detection of the voice by phoneme scores, other than word scores, para 28).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have applied obtaining the voice detection result on the basis of the matrix and the obtaining further comprising obtaining the probability values matches of that each of the plurality of voice segments pertains to the first category and the second category respectively according to the element corresponding to the each voice segment comprised in the matrix; the voice feature corresponding to the voice segment comprised in the first category matches the bottleneck feature corresponding to the reference voice, and the voice feature corresponding to the voice segment comprised in the second category does not match the bottleneck feature corresponding to the reference voice; and determining the voice detection result corresponding to the each voice segment on the basis of the maximum value of the probability values that the each voice segment pertains to the first category and the second category respectively, as taught by Bocklet, to obtaining the voice detection result on the basis of the matrix in the voice separation method, as taught by the combination of Xu, Xu627, Brocchieri, and Qu, for the benefits discussed above.
Claim 12 has been analyzed and rejected according to claims 11, 3 above.
Claim 16 has been analyzed and rejected according to claims 15, 3 above.
Response to Arguments
Applicant's arguments filed on March 30, 2026 have been fully considered and but are moot in view of the new ground(s) of rejection necessitated by the applicant amendment. The Office has thoroughly reviewed Applicants' arguments but firmly believes that the cited references to reasonably and properly meet the claimed limitations.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LESHUI ZHANG whose telephone number is (571)270-5589. The examiner can normally be reached Monday-Friday 6:30amp-4:00pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LESHUI ZHANG/
Primary Examiner,
Art Unit 2695