Prosecution Insights
Last updated: August 06, 2026
Application No. 18/830,160

PHONE RECOGNITION METHOD AND APPARATUS, ELECTRONIC DEVICE AND STORAGE MEDIUM

Non-Final OA §102§103§112
Filed
Sep 10, 2024
Priority
Nov 30, 2022 — CN 202211525113.4 +1 more
Examiner
SERRAGUARD, SEAN ERIN
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
1 (Non-Final)
70%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
107 granted / 154 resolved
+7.5% vs TC avg
Strong +35% interview lift
Without
With
+35.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
21 currently pending
Career history
186
Total Applications
across all art units

Statute-Specific Performance

§101
8.4%
-31.6% vs TC avg
§103
50.2%
+10.2% vs TC avg
§102
19.9%
-20.1% vs TC avg
§112
19.6%
-20.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 154 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 10 September 2024 is/are being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claim(s) 1-5 and 7-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the enablement requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to enable one skilled in the art to which it pertains, or with which it is most nearly connected, to make and/or use the invention. Regarding claim 1, and mutatis mutandis claims 14 and 18, the limitation “a trained phone recognition model to perform phone recognition” is not enabled commensurate with the scope of the claims. (See MPEP 2164.08). Claim 1 recites “a trained phone recognition model to perform phone recognition” at lines 4-5. This embodiment, as captured by claims 1, 14, and 18, is not enabled such that it can be practiced by one having ordinary skill in the art without undue experimentation. The courts have repeatedly held that “the specification must teach those skilled in the art how to make and use the full scope of the claimed invention without ‘undue experimentation’” or that any experimentation must be “reasonable”. See Amgen Inc. et al. v. Sanofi et al., 598 U.S. 594, 2023 USPQ2d 602 (2023); McRO, Inc. v. Bandai Namco Games Am. Inc., 959 F.3d 1091, 2020 USPQ2d 10550 (Fed. Cir. 2020); Wyeth & Cordis Corp. v. Abbott Laboratories, 720 F.3d 1380, 107 USPQ2d 1273 (Fed. Cir. 2013); Enzo Life Sciences, Inc. v. Roche Molecular Systems, Inc., 928 F.3d 1340 (Fed. Cir. 2019); and Idenix Pharmaceuticals LLC v. Gilead Sciences Inc., 941 F.3d 1149, 2019 USPQ2d 415844 (Fed. Cir. 2019). See also In re Wright, 999 F.2d 1557, 1561, 27 USPQ2d 1510, 1513 (Fed. Cir. 1993). In order to determine compliance with the enablement requirement of 35 U.S.C. 112(a), the Federal Circuit developed a framework of factors in In re Wands, 858 F.2d 731, 737, 8 USPQ2d 1400, 1404 (Fed. Cir. 1988), referred to as the Wands factors to assess whether any necessary experimentation required by the specification is “reasonable” or is “undue.” These factors include, but are not limited to: (A) The breadth of the claims; (B) The nature of the invention; (C) The state of the prior art; (D) The level of one of ordinary skill; (E) The level of predictability in the art; (F) The amount of direction provided by the inventor; (G) The existence of working examples; and (H) The quantity of experimentation needed to make or use the invention based on the content of the disclosure. (In re Wands, 858 F.2d at 737 (reversing the PTO’s determination that claims directed to methods for detection of hepatitis B surface antigens did not satisfy the enablement requirement)). All Wands factors are expressly considered here in light of the available evidence below. Regarding factor A, the claims broadly cover any phone recognition model. The broadest reasonable interpretation of “trained phone recognition model” is understood in light of the specification, including but not limited to, paragraphs [0089]-[0091]. The instant application at paragraph [0089] recites “the phone recognition model is a knowledge distillation model, and is composed of a teacher model (the base model) and a student model (the distillation model).” The instant application at paragraph [0091] recites “In this embodiment of this application, the neural network model respectively included in the teacher model and the student model may be… a wav2vec 3.0 model,” where the phone recognition model is trained, resulting in the trained phone recognition model. Thus, the trained phone recognition model, as provided for in claim 1, is understood as including embodiments resulting from the teacher model and/or the student model and can include “a wav2vec 3.0 model.” Because the specification fails to delineate the boundaries, structure, or minimum viable architecture of this model, by claiming a model with no known defining characteristics beyond a shared naming convention, the scope of the claims vastly exceeds the disclosure provided. Regarding factors B and H, we consider of the quantity of experimentation and the nature of the invention. It is noted that applicant has provided no direction which would lend itself to consideration regarding an expected level of experimentation. Though some level of experimentation would be considered expected in the field of neural networks, and network compositions and training schemes vary widely, applicant provides no direction regarding the wav2vec 3.0 model such that the use of such a model would vast experimentation with no expectation that the model derived would actually be the wav2vec 3.0 model which applicant envisions. As the wav2vec 3.0 model is an embodiment of a central component of the claims (e.g., the base model and/or the distillation model of the phone recognition model), the use of this embodiment by a PHOSITA would be tantamount to forcing the PHOSITA to invent the entire claimed embodiment, which is inherently undue experimentation Regarding factors C and E, while self-supervised learning frameworks such as “wav2vec 2.0” were known in the prior art, as shown in Baevski (U.S. Pat. No. 11,551,668, hereinafter Baevski), the art of deep neural networks and acoustic modeling is highly unpredictable. In the machine learning arts, minor alterations to a model’s depth, attention mechanisms, temporal masking strategies, or quantization modules can result in vast changes to model convergence and functionality. The recitation of a version “3.0” inherently implies a structural divergence from the 2.0 model. It is noted that a “wav2vec 3.0 model” does not exist currently and did not exist at the time of filing the instant application. Applicant provides no disclosure as to how the “wav2vec 3.0 model” would work or how it is integrated into the disclosed embodiments. Due to the above described unpredictability, a PHOSITA cannot accurately guess what specific structural modifications or training mechanisms that the applicant envisions for the “3.0” architecture. Regarding factor D, such a disclosure is not enabling to one of ordinary skill in the art. The field of neural networks is a highly skilled field of endeavor. However, the limits presented by such a disclosure, which relies only on naming conventions to name the non-existent model, would require a person having ordinary skill in the art (PHOSITA), to invent the model in its entirety. Regarding factors F and G, the specification merely names a wav2vec 3.0 model” as a black box component for processing speech features. The specification provides no architectural definitions, hyperparameter configurations, or mathematical loss functions required to construct/train it. Further, there are no working examples, architectural block diagrams, or pseudo-code disclosing how wav2vec 3.0 is structurally assembled or trained. Therefore, the full scope of claims 1, 14, and 18 is not enabled without undue experimentation, and the claims are rejected under 112(a). The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 2 recites the limitation “the phone recognition model” in line 1. There is insufficient antecedent basis for this limitation in the claim. Claim 1 does recite a trained phone recognition model. However, as the trained phone recognition model is distinguished from the phone recognition model in claim 2, the antecedent basis remains unclear. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1, 9-10, 13-14, 17-18, and 20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Fang (U.S. Pat. App. Pub. No. 2024/0395242, hereinafter Fang). Regarding claim 1, Fang discloses A phone recognition method, performed by a computer device (The systems and methods described with reference to the “speech recognition method,” where, under the BRI, the detection and classification of a spoken word fundamentally necessitates and encompasses the detection of articulatory phones constituting that word.; Fang, ¶ [0089]), the method comprising: obtaining a reference voiceprint feature of a target user and to-be-recognized audio (Discloses “obtaining... each speech frame in the target mixed speech” and “a speaker feature of a target speaker”; Fang, ¶ [0089], [0091]); inputting the to-be-recognized audio into a trained phone recognition model to perform phone recognition, to obtain a phone recognition result (The method further includes inputting the target mixed speech into a joint model, including “the speech recognition model” and “feature extraction model,” for “determining the obtained feature vector sequence as the speech feature of the target mixed speech”; Fang, ¶ [0091], 0110]), the trained phone recognition model being obtained through training based on first sample audio and second sample audio (“the speech recognition model may be obtained from a joint training with the feature extraction model.”; Fang, ¶ [0110]), the first sample audio comprising single-user utterance audio, and the second sample audio comprising multi-user utterance audio (“the training dataset S... includes a speech (clean speech) of a designated speaker and a training mixed speech including the speech of the designated speaker. The speech of the designated speaker includes an annotated text (where, the annotated text is the speech content of the speech of the designated speaker).”; Fang, ¶ [0114]-[0115]); extracting an audio feature of the to-be-recognized audio (“obtaining a feature vector (e.g., a spectrum feature) of each speech frame in the target mixed speech to obtain a feature vector sequence”; Fang, ¶ [0091]); performing denoising on the audio feature of the to-be-recognized audio based on the reference voiceprint feature, to obtain an acoustic voice feature of the target user (“the speech feature of the target mixed speech and the feature mask corresponding to the target speaker are multiplied frame by frame to obtain the extracted speech feature of the target speaker.”; Fang, ¶ [0156]); and performing the phone recognition on the acoustic voice feature, to obtain a phone recognition result corresponding to the target user (“After the extracted speech feature of the target speaker is obtained, the extracted speech feature of the target speaker and the registered speech feature of the target speaker are inputted into the speech recognition model to obtain the speech recognition result of the target speaker,” which is a frame level process {a phone recognition result}; Fang, ¶ [0157]). Regarding claim 9, Fang discloses wherein the performing denoising on the audio feature of the to-be-recognized audio based on the reference voiceprint feature, to obtain an acoustic voice feature of the target user comprises: performing masking on an audio feature in the audio feature of the to-be-recognized audio other than an audio feature of the target user (“the speech feature of the target mixed speech and the speaker feature of the target speaker are inputted into the feature extraction model to obtain a feature mask corresponding to the target speaker” where “the feature mask corresponding to the target speaker may represent a proportion of the speech feature of the target speaker in the speech feature of the target mixed speech,” by applying the feature mask to mixed speech via frame-by-frame multiplication, Fang performs masking on the audio features of non-target speakers in the mixture.; Fang, ¶ [0153]-[0154]) by using the trained phone recognition model based on the reference voiceprint feature, to obtain the acoustic voice feature of the target user (“the speech feature of the target mixed speech and the feature mask corresponding to the target speaker are multiplied frame by frame to obtain the extracted speech feature of the target speaker,” thus the model generates the mask based on the input speaker feature, and the result of applying the feature mask to the mixed speech is the extracted speech feature of the target speaker.; Fang, ¶ [0156]). Regarding claim 10, Fang discloses wherein the performing masking on an audio feature in the audio feature of the to-be-recognized audio other than an audio feature of the target user by using the trained phone recognition model based on the reference voiceprint feature, to obtain the acoustic voice feature of the target user comprises: splicing the reference voiceprint feature and the audio feature of the to-be-recognized audio, to obtain a spliced feature (“extracting the speech feature of the designated speaker from the target mixed speech feature, by using a pre-established feature extraction model, based on the target mixed speech feature” and “the feature mask corresponding to the target speaker, to obtain the extracted speech feature of the target speaker” and “the training speaker feature {reference voiceprint feature} may be concatenated {spliced} with each feature vector of the speech frames in the training mixed speech {the audio feature of the to-be-recognized audio}” where “after concatenating the feature vector of each speech frame in the training mixed speech with the short-term voiceprint feature and the long-term voiceprint feature, a concatenated feature vector of 120 dimensions can be obtained {to obtain a spliced feature}”; Fang, ¶ [0130], [0152], [0155]); performing nonlinear transformation on the spliced feature, to obtain a masked representation of the to-be-recognized audio (The “concatenated feature vector” is “then inputted into the feature extraction model 301” which, as a neural network model, applies a non-linear activation function {non-linear transformation} to calculate and output the feature mask (The application of a non-linear transformation is an inherent structural requirement of the feature mask disclosed by the prior art. The reference explicitly teaches that the feature mask represents a proportion where the values are “in a range of [0,1],” which mathematically necessitates the use of a non-linear activation function, as intermediate feature representations (logits) are unbounded and a linear activation function will result in ranges from negative to positive infinity. Therefore, the bounding of the unbounded internal logits to a bounded range of outputs (e.g., [0, 1]) requires the use of a non-linear activation function (such as a Sigmoid function or softmax). This fundamental principle of machine learning is further explained in Goodfellow (see at least at pages 179 and 180), which explain both unbounded logits and the necessary use of a non-linear activation function (e.g., sigmoid) to generate the “closed interval of valid probabilities [0, 1].”); Fang, ¶ [0130]-[0131]); and multiplying the masked representation of the to-be-recognized audio by the audio feature of the to-be-recognized audio, to obtain the acoustic voice feature of the target user (“the speech feature of the target mixed speech and the feature mask corresponding to the target speaker are multiplied frame by frame to obtain the extracted speech feature of the target speaker,” thus the model generates the mask based on the input speaker feature, and the result of applying the feature mask to the mixed speech is the extracted speech feature of the target speaker.; Fang, ¶ [0156]). Regarding claim 13, Fang discloses wherein the obtaining a reference voiceprint feature comprises: obtaining audio of the target user in an environment with a noise intensity less than a second preset value (Though not expressly recited as a “preset value” in Fang, the speech feature of the target speaker is obtained based on processing of “a registered speech of the target speaker” where “speech feature of the target speaker refers to a speech feature obtained based on a speech (clean speech) of the designated speaker,” thus the noise intensity value of the speech of the target speaker is less than noise intensity value which establishes the maximum level of noise that speech can have and still correspond to clean speech. As Fang differentiates between clean speech and not-clean speech, this differentiation necessarily has a boundary values such that clean and not-clean can be distinguished.; Fang, ¶ [0093], [0105]); and performing voiceprint feature recognition on the audio of the target user, to obtain the reference voiceprint feature (“the registered speech of the target speaker is obtained” and “a short-term voiceprint feature and a long-term voiceprint feature are extracted from the registered speech of the target speaker to obtain a multi-scale voiceprint feature as the speaker feature of the target speaker”; Fang, ¶ [0093], [0097]). Regarding claim 14, Fang discloses A phone recognition apparatus, comprising (The systems and methods described with reference to the “speech recognition method,” where, under the BRI, the detection and classification of a spoken word fundamentally necessitates and encompasses the detection of articulatory phones constituting that word.; Fang, ¶ [0089]): a memory storing a plurality of instructions; and a processor configured to execute the plurality of instructions, (Discloses the “speech recognition method” as “applied to a server” which may “include one or more central processing units and a memory” including instructions for performing “the speech recognition method provided in the present disclosure”; Fang, ¶ [0086]) wherein upon execution of the plurality of instructions, the processor is configured to: obtain a reference voiceprint feature of a target user and to-be-recognized audio (Discloses “obtaining... each speech frame in the target mixed speech” and “a speaker feature of a target speaker”; Fang, ¶ [0089], [0091]); input the to-be-recognized audio into a trained phone recognition model to perform phone recognition, to obtain a phone recognition result (The method further includes inputting the target mixed speech into a joint model, including “the speech recognition model” and “feature extraction model,” for “determining the obtained feature vector sequence as the speech feature of the target mixed speech”; Fang, ¶ [0091], 0110]), the trained phone recognition model being obtained through training based on first sample audio and second sample audio (“the speech recognition model may be obtained from a joint training with the feature extraction model.”; Fang, ¶ [0110]), the first sample audio comprising single-user utterance audio, and the second sample audio comprising multi-user utterance audio (“the training dataset S... includes a speech (clean speech) of a designated speaker and a training mixed speech including the speech of the designated speaker. The speech of the designated speaker includes an annotated text (where, the annotated text is the speech content of the speech of the designated speaker).”; Fang, ¶ [0114]-[0115]); extract an audio feature of the to-be-recognized audio (“obtaining a feature vector (e.g., a spectrum feature) of each speech frame in the target mixed speech to obtain a feature vector sequence”; Fang, ¶ [0091]); perform denoising on the audio feature of the to-be-recognized audio based on the reference voiceprint feature, to obtain an acoustic voice feature of the target user (“the speech feature of the target mixed speech and the feature mask corresponding to the target speaker are multiplied frame by frame to obtain the extracted speech feature of the target speaker.”; Fang, ¶ [0156]); and perform the phone recognition on the acoustic voice feature, to obtain a phone recognition result corresponding to the target user (“After the extracted speech feature of the target speaker is obtained, the extracted speech feature of the target speaker and the registered speech feature of the target speaker are inputted into the speech recognition model to obtain the speech recognition result of the target speaker,” which is a frame level process {a phone recognition result}; Fang, ¶ [0157]). Regarding claim 17, the rejection of claim 14 is incorporated. Claim 17 is substantially the same as claim 9 and is therefore rejected under the same rationale as above. Regarding claim 18, Fang discloses A non-transitory computer-readable storage medium storing a plurality of instructions executable by a processor, wherein upon being executed by the processor, the plurality of instructions is configured to cause the processor to (The systems and methods described with reference to the “speech recognition method,” as “applied to a server” which may “include one or more central processing units and a memory” including instructions for performing “the speech recognition method provided in the present disclosure,” where, under the BRI, the detection and classification of a spoken word fundamentally necessitates and encompasses the detection of articulatory phones constituting that word.; Fang, ¶ [0086], [0089]): obtain a reference voiceprint feature of a target user and to-be-recognized audio (Discloses “obtaining... each speech frame in the target mixed speech” and “a speaker feature of a target speaker”; Fang, ¶ [0089], [0091]); input the to-be-recognized audio into a trained phone recognition model to perform phone recognition, to obtain a phone recognition result (The method further includes inputting the target mixed speech into a joint model, including “the speech recognition model” and “feature extraction model,” for “determining the obtained feature vector sequence as the speech feature of the target mixed speech”; Fang, ¶ [0091], 0110]), the trained phone recognition model being obtained through training based on first sample audio and second sample audio (“the speech recognition model may be obtained from a joint training with the feature extraction model.”; Fang, ¶ [0110]), the first sample audio comprising single-user utterance audio, and the second sample audio comprising multi-user utterance audio (“the training dataset S... includes a speech (clean speech) of a designated speaker and a training mixed speech including the speech of the designated speaker. The speech of the designated speaker includes an annotated text (where, the annotated text is the speech content of the speech of the designated speaker).”; Fang, ¶ [0114]-[0115]); extract an audio feature of the to-be-recognized audio (“obtaining a feature vector (e.g., a spectrum feature) of each speech frame in the target mixed speech to obtain a feature vector sequence”; Fang, ¶ [0091]); perform denoising on the audio feature of the to-be-recognized audio based on the reference voiceprint feature, to obtain an acoustic voice feature of the target user (“the speech feature of the target mixed speech and the feature mask corresponding to the target speaker are multiplied frame by frame to obtain the extracted speech feature of the target speaker.”; Fang, ¶ [0156]); and perform the phone recognition on the acoustic voice feature, to obtain a phone recognition result corresponding to the target user (“After the extracted speech feature of the target speaker is obtained, the extracted speech feature of the target speaker and the registered speech feature of the target speaker are inputted into the speech recognition model to obtain the speech recognition result of the target speaker,” which is a frame level process {a phone recognition result}; Fang, ¶ [0157]). Regarding claim 20, the rejection of claim 18 is incorporated. Claim 20 is substantially the same as claim 13 and is therefore rejected under the same rationale as above. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2, 7, 11-12, 15 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fang as applied to claims 1, 14, and 18 above, and further in view of Qian (U.S. Pat. App. Pub. No. 2019/0304437, hereinafter Qian). Regarding claim 2, the rejection of claim 1 is incorporated. Fang discloses all of the elements of the current invention as stated above. However, Fang fails to expressly recite wherein the phone recognition model comprises a base model and a distillation model, a data dimension of the base model being greater than a data dimension of the distillation model, and the method further comprises: obtaining the first sample audio and the second sample audio; training the base model based on the first sample audio, to obtain a first loss value during the training of the base model, and training the distillation model based on the second sample audio, to obtain a second loss value during the training of the distillation model; and respectively adjusting a model parameter of the base model and a model parameter of the distillation model based on the first loss value and the second loss value, to obtain the trained phone recognition model. Qian teaches “adaptive permutation invariant training for multi-talker speech recognition.” (Qian, ¶ [0001]). Regarding claim 2, Qian teaches wherein the phone recognition model comprises a base model and a distillation model (Discloses a speech recognition system including a “single talker speech recognition model {base model}” and a “multi-talker speech recognition model {distillation model}”; Qian, ¶ [0070], [0073]), a data dimension of the base model being greater than a data dimension of the distillation model (Discloses the “single talker speech recognition model is a teacher model” and the “the multi-talker speech recognition model is a student model” where “knowledge is transferred from a large and complicated teacher network to a small student network,” thus the teacher model (base model) has a data dimension which is “large and complicated {a data dimension of the base model being greater than...}” as compared to the “small student network {...a data dimension of the distillation model}”; Qian, ¶ [0053], [0072]-[0073]), and the method further comprises: obtaining the first sample audio and the second sample audio (Discloses obtaining “original individual single-talker speech” also referred to as “single-talker clean speech” and “multi-talker mixed speech”; Qian, ¶ [0008], [0058]); training the base model based on the first sample audio, to obtain a first loss value during the training of the base model (“the processor may perform PIT model training on a single talker feature corresponding to one or more of the plurality of speakers and update a single talker speech recognition model” where “PIT aims to minimize the minimal average cross entropy (CE) {obtain a first loss value}” which occurs during the training of the “single talker speech recognition model {base model}”; Qian, ¶ [0050], [0072]), and training the distillation model based on the second sample audio, to obtain a second loss value during the training of the distillation model (“the processor may perform PIT model training on the multi-talker mixed speech signal based on the soft label input from the single talker speech recognition model and generate a plurality of estimated output segments” and “may minimize a minimal average cross entropy (CE) for utterances of all possible assignments between the plurality of estimated output segments and soft labels” to update the “multi-talker speech recognition model”; Qian, ¶ [0074]-[0076]); and respectively adjusting a model parameter of the base model and a model parameter of the distillation model based on the first loss value and the second loss value, to obtain the trained phone recognition model (Minimizing “a minimal average cross entropy (CE)” for each of the “single talker speech recognition model” and the “multi-talker speech recognition model” is updating or adjusting a model parameter of each model (i.e., “minimizing”) which is based on the average cross entropy for the single talker speech input {first loss value} and the multi-talker speech input {second loss value}, where the minimization results in the trained versions (i.e., at convergence, per Algorithm 1, step 7) of the “single talker speech recognition model” and the “multi-talker speech recognition model”; Qian, ¶ [0050], [0062]-[0063], [0075]-[0076]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang to incorporate the teachings of Qian to include wherein the phone recognition model comprises a base model and a distillation model, a data dimension of the base model being greater than a data dimension of the distillation model, and the method further comprises: obtaining the first sample audio and the second sample audio; training the base model based on the first sample audio, to obtain a first loss value during the training of the base model, and training the distillation model based on the second sample audio, to obtain a second loss value during the training of the distillation model; and respectively adjusting a model parameter of the base model and a model parameter of the distillation model based on the first loss value and the second loss value, to obtain the trained phone recognition model.. Qian teaches utilizing a multi-talker speech recognition model to separate overlapping speech into individual output posterior streams, which allows for each stream to be used for decoding as normal to obtain the final recognition result, thereby providing a predictable advantage of capturing speech recognition results for all participants in the mixed audio, rather than simply discarding the non-target speakers, as recognized by Qian. (Qian, ¶ [0043], [0051]). Regarding claim 7, the rejection of claim 2 is incorporated. Fang discloses all of the elements of the current invention as stated above. However, Fang fails to expressly recite wherein the performing the phone recognition on the acoustic voice feature, to obtain a phone recognition result corresponding to the target user comprises: calculating probabilities that the acoustic voice feature is classified as each phone by using a classification function in an output layer of a trained distillation model; and determining the phone recognition result corresponding to the target user based on the probabilities. The relevance of Qian is described above with relation to claim 2. Regarding claim 7, Qian teaches wherein the performing the phone recognition on the acoustic voice feature, to obtain a phone recognition result corresponding to the target user comprises: calculating probabilities that the acoustic voice feature is classified as each phone (Discloses that the model output is “fed through a linear layer for each speaker individually” and “softmax activations are separately computed on the outputs from the linear layer,” where the softmax activations, also referred to as the senone posteriori, are predictions for classification by the softmax layer.; Qian, ¶ [0043], [0046]) by using a classification function in an output layer of a trained distillation model (“the outputs from the linear layer 3, i.e., linear1 and linear2, are fed into softmax1 and softmax 2, respectively, of the softmax layer 4, which computes a softmax activation,” where the softmax function is a classification function at the output layer of neural networks to normalize outputs into a probability distribution; Qian, ¶ [0043]); and determining the phone recognition result corresponding to the target user based on the probabilities (“output of the softmax activation in softmax1 and softmax2 may be considered as predictions1 and predictions2 in prediction layer 5.”; Qian, ¶ [0043]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang to incorporate the teachings of Qian to include wherein the performing the phone recognition on the acoustic voice feature, to obtain a phone recognition result corresponding to the target user comprises: calculating probabilities that the acoustic voice feature is classified as each phone by using a classification function in an output layer of a trained distillation model; and determining the phone recognition result corresponding to the target user based on the probabilities. Qian teaches utilizing a multi-talker speech recognition model to separate overlapping speech into individual output posterior streams, which allows for each stream to be used for decoding as normal to obtain the final recognition result, thereby providing a predictable advantage of capturing speech recognition results for all participants in the mixed audio, rather than simply discarding the non-target speakers, as recognized by Qian. (Qian, ¶ [0043], [0051]). Regarding claim 11, the rejection of claim 10 is incorporated. Fang discloses all of the elements of the current invention as stated above. Fang further discloses wherein the performing nonlinear transformation on the spliced feature, to obtain a masked representation of the to-be-recognized audio comprises: performing the nonlinear transformation on the spliced feature by using a... [neural network] layer and an activation function, to obtain the masked representation of the to-be-recognized audio (The “concatenated feature vector” is “then inputted into the feature extraction model 301” which, as a neural network model, applies a non-linear activation function {non-linear transformation} to calculate and output the feature mask, where Fang teaches the model performing this transformation is a neural network (RNN, CNN, or DNN) and results in the feature mask M, where “where m1 to mK are in a range of [0,1]” (The application of a non-linear transformation is an inherent structural requirement of the feature mask disclosed by the prior art. The reference explicitly teaches that the feature mask represents a proportion where the values are “in a range of [0,1],” which mathematically necessitates the use of a non-linear activation function, as intermediate feature representations (logits) are unbounded and a linear activation function will result in ranges from negative to positive infinity. Therefore, the bounding of the unbounded internal logits to a bounded range of outputs (e.g., [0, 1]) requires the use of a non-linear activation function (such as a Sigmoid function or softmax). This fundamental principle of machine learning is further explained in Goodfellow (see at least at pages 179 and 180), which explain both unbounded logits and the necessary use of a non-linear activation function (e.g., sigmoid) to generate the “closed interval of valid probabilities [0, 1].”); Fang, ¶ [0130]-[0131]). However, Fang fails to expressly recite wherein the neural network layer is a fully connected layer. The relevance of Qian is described above with relation to claim 2. Regarding claim 11, Qian teaches wherein the performing nonlinear transformation on the spliced feature, to obtain a masked representation of the to-be-recognized audio comprises: performing the nonlinear transformation on the spliced feature by using a fully connected layer and an activation function in the trained phone recognition model (Discloses that “the input features are fed into a recurrent neural network (RNN) and RNN operations are performed” and “the output of the RNN operations [is] fed through a linear layer {fully connected layer} for each speaker individually. Thereafter, in operation S140, softmax activations {a non-linear activation function} are separately computed on the outputs from the linear layer,” which, in light of the disclosure of Fang, is included in the trained phone recognition model and the “softmax activations” results in a non-linear transformation on the “concatenated feature vectors” of Fang.; Qian, ¶ [0043], [0046]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang to incorporate the teachings of Qian to include wherein the neural network layer is a fully connected layer. Qian teaches utilizing a multi-talker speech recognition model to separate overlapping speech into individual output posterior streams, which allows for each stream to be used for decoding as normal to obtain the final recognition result, thereby providing a predictable advantage of capturing speech recognition results for all participants in the mixed audio, rather than simply discarding the non-target speakers, as recognized by Qian. (Qian, ¶ [0043], [0051]). Regarding claim 12, the rejection of claim 1 is incorporated. Fang discloses all of the elements of the current invention as stated above. Fang further discloses wherein the performing denoising on the audio feature of the to-be-recognized audio based on the reference voiceprint feature, to obtain an acoustic voice feature of the target user comprises: encoding the audio feature of the to-be-recognized audio by using... [an] encoder in the trained phone recognition model, to obtain audio features corresponding to different speakers (discloses “a speech feature of a target mixed speech” where “the target mixed speech refers to a speech of multiple speakers, including the speech of the target speaker as well as speeches of other speakers” and “obtaining the speech feature of the target mixed speech includes: obtaining a feature vector (e.g., a spectrum feature) of each speech frame in the target mixed speech to obtain a feature vector sequence” which “may be expressed as[x1, x2,.., xk,.., xK].”; Fang, ¶ [0089]-[0091]); and searching the audio features corresponding to the different speakers for an audio feature corresponding to the target user based on the reference voiceprint feature (Using the “speaker feature of the target speaker” as obtained based on processing of “a registered speech of the target speaker” where “the registered speech of the target speaker is obtained” and “a short-term voiceprint feature and a long-term voiceprint feature are extracted from the registered speech of the target speaker to obtain a multi-scale voiceprint feature as the speaker feature of the target speaker” and the “the speaker characteristic extraction model” uses the “speech feature sequence of the registered speech of the target speaker... so as to obtain a shallow feature and a deep feature” from the speech feature of the target mixed speech, where “the shallow feature has a smaller receptive field and therefore can better represent the short-term voiceprint. Hence, the shallow feature is regarded as the short-term voiceprint feature. The deep feature has a larger receptive field and therefore can better represent the long-term voiceprint. Hence, the deep feature is regarded as the long-term voiceprint feature.”; Fang, ¶ [0093], [0097]). However, Fang fails to expressly recite wherein the encoder is a multi-speaker encoder. The relevance of Qian is described above with relation to claim 2. Regarding claim 12, Qian teaches wherein the performing denoising on the audio feature of the to-be-recognized audio based on the reference voiceprint feature, to obtain an acoustic voice feature of the target user comprises: encoding the audio feature of the to-be-recognized audio by using a multi-speaker encoder in the trained phone recognition model, to obtain audio features corresponding to different speakers (Discloses a network architecture (“the output of the RNN layer 2 may go through linear1 and linear2 for each speaker individually”), which is designed to process multi-talker speech {a multi-speaker encoder} which is applied to encode the audio feature of the “speech feature of the target mixed speech” {to-be-recognized audio} into branched paths (“linear1 and linear2 for each speaker individually”), and “After the PIT model training, the individual output posterior stream can be used for decoding as normal to obtain the final recognition result.”; Qian, ¶ [0043], [0051]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang to incorporate the teachings of Qian to include wherein the encoder is a multi-speaker encoder. Qian teaches utilizing a multi-talker speech recognition model to separate overlapping speech into individual output posterior streams, which allows for each stream to be used for decoding as normal to obtain the final recognition result, thereby providing a predictable advantage of capturing speech recognition results for all participants in the mixed audio, rather than simply discarding the non-target speakers, as recognized by Qian. (Qian, ¶ [0043], [0051]). Regarding claim 15, the rejection of claim 14 is incorporated. Claim 15 is substantially the same as claim 2 and is therefore rejected under the same rationale as above. Regarding claim 19, the rejection of claim 18 is incorporated. Claim 19 is substantially the same as claim 12 and is therefore rejected under the same rationale as above. Claims 3-4, 8, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fang and Qian as applied to claims 1-2, 14-15, and 18 above, and further in view of Aslan (U.S. Pat. App. Pub. No. 2017/0132528, hereinafter Aslan). Regarding claim 3, the rejection of claim 2 is incorporated. Fang discloses all of the elements of the current invention as stated above. However, Fang fails to expressly recite wherein the respectively adjusting a model parameter of the base model and a model parameter of the distillation model based on the first loss value and the second loss value, to obtain the trained phone recognition model comprises: determining a target loss value based on the first loss value and the second loss value; respectively adjusting the model parameter of the base model and the model parameter of the distillation model based on the target loss value, so that the phone recognition model converges, to obtain the trained phone recognition mode. Aslan teaches “techniques and systems for jointly training multiple machine learning models.” (Aslan, ¶ [0004]). Regarding claim 3, Aslan teaches wherein the respectively adjusting a model parameter of the base model and a model parameter of the distillation model based on the first loss value and the second loss value, to obtain the trained phone recognition model comprises: determining a target loss value based on the first loss value and the second loss value (“joint training by optimizing an objective function jointly with respect to weight parameters of multiple models being trained in parallel, such as during joint training of the first model 100 and the second model 102” where “L_te and L_st” are the loss values “for the first (teacher) model 100 {first loss value} and the second (student) model 102 {second loss value}, respectively” and then formulates the “objective function for joint training of the first and second models 100 and 102” using Equation 2 which includes a weighted summation “L_te(Φ(te),Y)+γ1(L_st(Φ(st),Y)” of the first loss value and the second loss value, generating the target loss value.; Aslan, ¶ [0033], [0035]); respectively adjusting the model parameter of the base model and the model parameter of the distillation model based on the target loss value, so that the phone recognition model converges, to obtain the trained phone recognition mode (“by optimizing an objective function jointly with respect to weight parameters of multiple models” the first and second models 100 and 102 communicate “with each other via the objective function for purposes of joint training.”; Aslan, ¶ [0028], [0033]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang, as modified by the adaptive permutation invariant training of Qian, to incorporate the teachings of Aslan to include wherein the respectively adjusting a model parameter of the base model and a model parameter of the distillation model based on the first loss value and the second loss value, to obtain the trained phone recognition model comprises: determining a target loss value based on the first loss value and the second loss value; respectively adjusting the model parameter of the base model and the model parameter of the distillation model based on the target loss value, so that the phone recognition model converges, to obtain the trained phone recognition mode. The combination of Fang and Qian creates a highly robust system, which resolves permutation ambiguity and achieves high accuracy. However, multi-speaker separation models (like Qian’s PIT network) combined with voiceprint extraction layers are inherently large, structurally complex and computationally expensive. Aslan discloses the architectural framework of Teacher-Student knowledge distillation, wherein a large, complex model (the Teacher) is used to train smaller models, (the Student) by minimizing loss between their respective outputs, which provides the known benefit of reducing the size of the larger and more complex networks described in the combination of Fang and Qian, as recognized in light of the disclosure of Aslan. (Aslan, ¶ [0004], [0044]). Regarding claim 4, the rejection of claim 3 is incorporated. Fang discloses all of the elements of the current invention as stated above. However, Fang fails to expressly recite wherein the determining a target loss value based on the first loss value and the second loss value comprises: performing weighted summation on the first loss value and the second loss value, to obtain the target loss value; or selecting a larger one of the first loss value and the second loss value as the target loss value. The relevance of Aslan is described above with relation to claim 3. Regarding claim 4, Aslan teaches wherein the determining a target loss value based on the first loss value and the second loss value comprises: performing weighted summation on the first loss value and the second loss value, to obtain the target loss value; or selecting a larger one of the first loss value and the second loss value as the target loss value (as indicated above with relation to claim 3, “L_te and L_st” are the loss values “for the first (teacher) model 100 {first loss value} and the second (student) model 102 {second loss value}, respectively” and the training includes the objective function described by Equation 2 which includes the weighted summation “L_te(Φ(te),Y)+γ1(L_st(Φ(st),Y)” of the first loss value and the second loss value, to generate/obtain the target loss value.; Aslan, ¶ [0033], [0035]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang, as modified by the adaptive permutation invariant training of Qian, to incorporate the teachings of Aslan to include wherein the determining a target loss value based on the first loss value and the second loss value comprises: performing weighted summation on the first loss value and the second loss value, to obtain the target loss value; or selecting a larger one of the first loss value and the second loss value as the target loss value. The combination of Fang and Qian creates a highly robust system, which resolves permutation ambiguity and achieves high accuracy. However, multi-speaker separation models (like Qian’s PIT network) combined with voiceprint extraction layers are inherently large, structurally complex and computationally expensive. Aslan discloses the architectural framework of Teacher-Student knowledge distillation, wherein a large, complex model (the Teacher) is used to train smaller models, (the Student) by minimizing loss between their respective outputs, which provides the known benefit of reducing the size of the larger and more complex networks described in the combination of Fang and Qian, as recognized in light of the disclosure of Aslan. (Aslan, ¶ [0004], [0044]). Regarding claim 8, the rejection of claim 1 is incorporated. Fang discloses all of the elements of the current invention as stated above. However, Fang fails to expressly recite wherein the phone recognition model comprises a base model and a distillation model, a data dimension of the base model being greater than a data dimension of the distillation model, and the method further comprises: obtaining the first sample audio and the second sample audio; training the base model by using the first sample audio, to obtain a trained base model, inputting the second sample audio into the trained base model and the distillation model, to obtain a first output result of the trained base model and a second output result of the distillation model; obtaining a third loss value based on the first output result and a phone label of the second sample audio, and obtaining a fourth loss value based on the second output result and the phone label of the second sample audio; and adjusting a model parameter of the distillation model based on the third loss value and the fourth loss value, to obtain a trained distillation model. The relevance of Qian is described above with relation to claim 2. Regarding claim 8, Qian teaches wherein the phone recognition model comprises a base model and a distillation model (Discloses a speech recognition system including a “single talker speech recognition model {base model}” and a “multi-talker speech recognition model {distillation model}”; Qian, ¶ [0070], [0073]), a data dimension of the base model being greater than a data dimension of the distillation model (Discloses the “single talker speech recognition model is a teacher model” and the “the multi-talker speech recognition model is a student model” where “knowledge is transferred from a large and complicated teacher network to a small student network,” thus the teacher model (base model) has a data dimension which is “large and complicated {a data dimension of the base model being greater than...}” as compared to the “small student network {...a data dimension of the distillation model}”; Qian, ¶ [0053], [0072]-[0073]), and the method further comprises: obtaining the first sample audio and the second sample audio (Discloses obtaining “original individual single-talker speech” also referred to as “single-talker clean speech” and “multi-talker mixed speech”; Qian, ¶ [0008], [0058]); training the base model by using the first sample audio, to obtain a trained base model (“the processor may perform PIT model training on a single talker feature corresponding to one or more of the plurality of speakers and update a single talker speech recognition model” where “PIT aims to minimize the minimal average cross entropy (CE) {obtain a first loss value}” which occurs during the training of the “single talker speech recognition model {base model}”; Qian, ¶ [0050], [0072]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang to incorporate the teachings of Qian to include wherein the phone recognition model comprises a base model and a distillation model, a data dimension of the base model being greater than a data dimension of the distillation model, and the method further comprises: obtaining the first sample audio and the second sample audio; training the base model by using the first sample audio, to obtain a trained base model. Qian teaches utilizing a multi-talker speech recognition model to separate overlapping speech into individual output posterior streams, which allows for each stream to be used for decoding as normal to obtain the final recognition result, thereby providing a predictable advantage of capturing speech recognition results for all participants in the mixed audio, rather than simply discarding the non-target speakers, as recognized by Qian. (Qian, ¶ [0043], [0051]). However, Fang and Qian fail to expressly recite inputting the second sample audio into the trained base model and the distillation model, to obtain a first output result of the trained base model and a second output result of the distillation model; obtaining a third loss value based on the first output result and a phone label of the second sample audio, and obtaining a fourth loss value based on the second output result and the phone label of the second sample audio; and adjusting a model parameter of the distillation model based on the third loss value and the fourth loss value, to obtain a trained distillation model. The relevance of Aslan is described above with relation to claim 3. Regarding claim 8, Aslan teaches inputting the second sample audio into the trained base model and the distillation model, to obtain a first output result of the trained base model and a second output result of the distillation model (Discloses “The training data 104” as “speech data” which may be received by “the first model 100 and the second model 102 during joint training” where the “data can be processed by each model 100 and 102, and the objective function used for joint training of the models 100 and 102 can determine the degree to which the models 100 and 102 agree with each other” where each model predicts “a set of probabilities” for the same data; Aslan, ¶ [0024], [0031]); obtaining a third loss value based on the first output result and a phone label of the second sample audio, and obtaining a fourth loss value based on the second output result and the phone label of the second sample audio (“joint training by optimizing an objective function jointly with respect to weight parameters of multiple models being trained in parallel, such as during joint training of the first model 100 and the second model 102” where “L_te and L_st” are the loss values “for the first (teacher) model 100 {obtaining a third loss value} and the second (student) model 102 {obtaining a fourth loss value}, respectively” and then formulates the “objective function for joint training of the first and second models 100 and 102” using Equation 2 which includes a weighted summation “L_te(Φ(te),Y)+γ1(L_st(Φ(st),Y)” of the first loss value and the second loss value, where each of the first loss value and fourth loss value are compared against Y, where “Y represents the original labels from the training data 104 when the training data 104 comprises labeled training data 104”, generating the target loss value from the third loss value and the fourth loss value each of which being based on original label Y {the phone label of the second sample audio}; Aslan, ¶ [0033], [0035]-[0036]); and adjusting a model parameter of the distillation model based on the third loss value and the fourth loss value, to obtain a trained distillation model (“by optimizing an objective function jointly with respect to weight parameters of multiple models” the first and second models 100 and 102 communicate “with each other via the objective function for purposes of joint training.”; Aslan, ¶ [0028], [0033]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang, as modified by the adaptive permutation invariant training of Qian, to incorporate the teachings of Aslan to include inputting the second sample audio into the trained base model and the distillation model, to obtain a first output result of the trained base model and a second output result of the distillation model; obtaining a third loss value based on the first output result and a phone label of the second sample audio, and obtaining a fourth loss value based on the second output result and the phone label of the second sample audio; and adjusting a model parameter of the distillation model based on the third loss value and the fourth loss value, to obtain a trained distillation model. The combination of Fang and Qian creates a highly robust system, which resolves permutation ambiguity and achieves high accuracy. However, multi-speaker separation models (like Qian’s PIT network) combined with voiceprint extraction layers are inherently large, structurally complex and computationally expensive. Aslan discloses the architectural framework of Teacher-Student knowledge distillation, wherein a large, complex model (the Teacher) is used to train smaller models, (the Student) by minimizing loss between their respective outputs, which provides the known benefit of reducing the size of the larger and more complex networks described in the combination of Fang and Qian, as recognized in light of the disclosure of Aslan. (Aslan, ¶ [0004], [0044]). Regarding claim 16, the rejection of claim 14 is incorporated. Claim 16 is substantially the same as claim 8 and is therefore rejected under the same rationale as above. Claim 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fang and Qian as applied to claims 1-2 above, and further in view of Li (U.S. Pat. App. Pub. No. 2019/0287515, hereinafter Li). Regarding claim 5, the rejection of claim 2 is incorporated. Fang discloses all of the elements of the current invention as stated above. However, Fang fails to expressly recite wherein the obtaining the first sample audio comprises: obtaining single-user utterance audio in an environment with a noise intensity less than a first preset value as the first sample audio. The relevance of Qian is described above with relation to claim 2. Regarding claim 8, Qian teaches wherein the obtaining the first sample audio comprises: obtaining single-user utterance audio... with a noise intensity less than a first preset value as the first sample audio (Discloses the first sample audio is clean speech, where clean is with reference to the noise intensity in the audio, and “parallel data used as inputs to the model may be generated by varying the relative energy of the involved talkers without transcribing the source streams,” where relative energy refers to a difference in intensity, both with regards to noise and signal.; Qian, ¶ [0008], [0064]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang to incorporate the teachings of Qian to include wherein the obtaining the first sample audio comprises: obtaining single-user utterance audio... with a noise intensity less than a first preset value as the first sample audio. Qian teaches utilizing a multi-talker speech recognition model to separate overlapping speech into individual output posterior streams, which allows for each stream to be used for decoding as normal to obtain the final recognition result, thereby providing a predictable advantage of capturing speech recognition results for all participants in the mixed audio, rather than simply discarding the non-target speakers, as recognized by Qian. (Qian, ¶ [0043], [0051]). However, Fang and Qian fail to expressly recite wherein the obtaining the first sample audio comprises: obtaining single-user utterance audio in an environment with a noise intensity less than a first preset value as the first sample audio. Li teaches systems and methods for training “a student model for speech recognition based on a teacher model.. (Li, ¶ [0015]). Regarding claim 5, Li teaches wherein the obtaining the first sample audio comprises: obtaining single-user utterance audio in an environment with a noise intensity less than a first preset value as the first sample audio (The system gives express consideration to “the characteristics of the acoustic environment (e.g., level of noise, distance to the microphone)” in collecting the sample audio, where “to build an adaptation system, very clean data is used (e.g., clear speaker in a clear environment without noise), and a teacher model is built for this clean domain,” where phrases such as “very clean”, “clear environment”, and “without noise” are all with reference to a threshold noise intensity (i.e., the threshold is the noise level which no longer qualifies as “very clean”, “clear”, and/or “without noise”).; Li, ¶ [0026], [0033]). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang, as modified by the adaptive permutation invariant training of Qian, to incorporate the teachings of Li to include wherein the obtaining the first sample audio comprises: obtaining single-user utterance audio in an environment with a noise intensity less than a first preset value as the first sample audio. Li teaches that acoustic models suffer large performance degradation when presented in a new domain or varying acoustic environments. The AT/S described in Li can improve condition robustness, allowing the multi-speaker extraction system described by the combination of Fang and Qian to maintain high speech recognition accuracy across multiple noisy conditions and environments without the need for transcribed adaptation data, as recognized by Li. (Li, ¶ [0016]-[0018]). Claim 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fang and Qian as applied to claims 1 and 2 above, and further in view of Baevski (U.S. Pat. No. 11,551,668, hereinafter Baevski). Regarding claim 6, the rejection of claim 2 is incorporated. Fang discloses all of the elements of the current invention as stated above. However, Fang fails to expressly recite wherein the extracting an audio feature of the to-be-recognized audio comprises: inputting the to-be-recognized audio into a voice encoder comprised in a trained distillation model, and performing discrete quantization on the to-be-recognized audio by using a shallow feature extraction layer of the voice encoder, to obtain a plurality of frames of voices comprised in the to-be-recognized audio; and extracting an audio feature corresponding to each frame of voice in the to-be-recognized audio by using a deep feature extraction layer of the voice encoder. Baevski teaches systems and methods for “ learning representations of speech audio using self-supervised learning.” (Baevski, ¶ Col. 1, lines 48-51). Regarding claim 6, Baevski teaches wherein the extracting an audio feature of the to-be-recognized audio comprises: inputting the to-be-recognized audio into a voice encoder comprised in a trained distillation model (Discloses a “model for learning representations of speech audio using self-supervised learning”; Baevski, ¶ Col. 3, lines 12-15), and performing discrete quantization on the to-be-recognized audio by using a shallow feature extraction layer of the voice encoder, to obtain a plurality of frames of voices comprised in the to-be-recognized audio (“Our model may be composed of a multi-layer convolutional feature encoder ƒ: X→Z which takes as input raw audio X and outputs latent speech representations z1,.. zT for T time-steps” which are “then fed to a Transformer g:Z→C to build representations c1,.. cT capturing information from the entire sequence,” where the convolutional feature encoder is positioned at the bottom of the stack, as shown in FIG. 1, to extract the temporal representations {shallow feature extraction layer of the voice encoder} and “The output of the feature encoder may be discretized to qt with a quantization module Z→Q to represent the targets (see FIG. 1) in the self-supervised objective {discrete quantization}” and “encoder output z” can be mapped to codebook entries by the Gumbel softmax.; Baevski, ¶ Col. 5, line 64 - Col. 6, line 14; Col. 6, line 64 – Col. 7, line 15); and extracting an audio feature corresponding to each frame of voice in the to-be-recognized audio by using a deep feature extraction layer of the voice encoder (After the audio is quantized into discrete frames, those frames are fed into a “Transformer network 135 to build contextualized representations 105,” where the transformer network is the deep feature extraction layer.; Baevski, ¶ Col. 5, lines 14-26). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition systems for mixed speech of Fang, as modified by the adaptive permutation invariant training of Qian, to incorporate the teachings of Baevski to include wherein the extracting an audio feature of the to-be-recognized audio comprises: inputting the to-be-recognized audio into a voice encoder comprised in a trained distillation model, and performing discrete quantization on the to-be-recognized audio by using a shallow feature extraction layer of the voice encoder, to obtain a plurality of frames of voices comprised in the to-be-recognized audio; and extracting an audio feature corresponding to each frame of voice in the to-be-recognized audio by using a deep feature extraction layer of the voice encoder. Baevski teaches that conventional speech recognition systems suffer from the requirement of thousands of hours of scarce, transcribed (labeled) speech. A PHOSITA would have been motivated to the disclosed self-supervised learning technique within the acoustic encoders of the Fang-Qian system to achieve Baevski’s stated goal of learning powerful representations from raw speech alone, providing the predictable benefit of drastically reducing the amount of labeled training data required to achieve acceptable performance in downstream speech recognition tasks, as recognized by Baevski. (Baevski, ¶ Col. 4, lines 7-12, and 44-54). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sean E. Serraguard whose telephone number is (313)446-6627. The examiner can normally be reached 07:00-17:00 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C. Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Sean E Serraguard/Primary Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Sep 10, 2024
Application Filed
May 04, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699834
SYSTEM AND METHOD FOR INTERACTIVE DIALOGUE
4y 5m to grant Granted Aug 04, 2026
Patent 12700402
SYSTEM AND METHOD FOR COMMAND FULFILLMENT WITHOUT WAKE WORD
3y 10m to grant Granted Aug 04, 2026
Patent 12682919
SYSTEM AND METHOD FOR REAL-TIME IDENTIFICATION OF DISSATISFACTION DATA
3y 7m to grant Granted Jul 14, 2026
Patent 12675641
DETECTION OF INTERACTION EVENTS IN RECORDED AUDIO STREAMS
3y 5m to grant Granted Jul 07, 2026
Patent 12646505
CONVERSATIONAL RECOMMENDATION METHOD, METHOD OF TRAINING MODEL, DEVICE AND MEDIUM
3y 6m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
70%
Grant Probability
99%
With Interview (+35.1%)
3y 0m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 154 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month