Prosecution Insights
Last updated: October 02, 2026
Application No. 18/264,531

METHOD AND APPARATUS FOR DETERMINING SPEECH SIMILARITY, AND PROGRAM PRODUCT

Final Rejection §103
Filed
Aug 07, 2023
Priority
Feb 07, 2021 — CN 202110179824.X +1 more
Examiner
MEIS, JON CHRISTOPHER
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Lemon Inc.
OA Round
4 (Final)
33%
Grant Probability
At Risk
5-6
OA Rounds
0m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants only 33% of cases
33%
Career Allowance Rate
11 granted / 33 resolved
-28.7% vs TC avg
Strong +52% interview lift
Without
With
+52.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
17 currently pending
Career history
60
Total Applications
across all art units

Statute-Specific Performance

§101
21.7%
-18.3% vs TC avg
§103
55.9%
+15.9% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
9.2%
-30.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 33 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1, 4-14, 16, 19 and 22-26 are pending. Claims 1 and 12-13 are independent. This Application was published as US 20240096347. Apparent priority is 02/07/2021. The instant Application is directed to a method of scoring pronunciation similarity to an example audio. Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection that, if presented, were necessitated by the amendments to the Claims. This action is Final. Response to Arguments 35 USC 103 Applicant’s arguments regarding feature fusion with respect to 35 USC 103 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant’s arguments regarding preset sampling points are not persuasive. Li discloses in [0074] that a Hamming window is used for extracting features. A Hamming window uses a preset width, and therefore each sampling window reads on preset sampling points. Grover also discloses (pg. 3, section 3.2) that an FFT window with a hop length of 512 samples is used; therefore the combination using the encoder disclosed by Grover to extract features also reads on preset sampling points. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 4-6, 8-14, 16, 19, and 22-26 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (CN104732977A) in view of Grover et al. ("Multi-modal Automated Speech Scoring using Attention Fusion"), Bartosik (US 20010025240 A1), and Du et al. (US 20090204398 A1). Regarding claim 1, Li discloses: 1. A method for determining speech similarity based on speech interaction, ("[0002] The present invention relates to the field of speech recognition and evaluation technology, and in particular to an online spoken pronunciation quality evaluation method and system." ) applied to a user terminal, (“[0067] In the specific implementation, the mobile client is installed on the user's mobile phone or other mobile devices in the form of an application…”) comprising: playing exemplary audio, (Not explicitly disclosed – Li discloses “[0018] Get standard voice; [0019] Preprocessing the standard speech;” – the standard speech is exemplary audio.) and acquiring evaluation audio of a user, ("[0009] Receiving a test voice collected by a mobile client through a network;" ) wherein the exemplary audio is audio of specified content that is read by using a specified language; ("[0166] The online spoken English pronunciation quality evaluation method and system provided by the embodiments of the present invention can be applied to spoken English learning to detect the pronunciation quality of spoken English. It can also be applied to pronunciation quality assessment of other languages, such as Japanese and French." - Whichever language quality assessment is used would be the specified language.) acquiring a standard pronunciation feature corresponding to the exemplary audio, and ("[0020] The speech feature parameters are extracted from the preprocessed standard speech to obtain the feature parameters of the standard speech." ) extracting, from the evaluation audio, an evaluation pronunciation feature corresponding to the standard pronunciation feature based on an encoder of a speech recognition model, ("[0011] Extracting speech feature parameters from the preprocessed test speech to obtain feature parameters of the test speech;" ) wherein: the standard pronunciation feature is used to reflect a specific pronunciation of the specified content in the specified; ("[0025] Calculating the pronunciation duration of the test speech and obtaining a pronunciation duration feature parameter of the test speech;" - Note that any of the extracted features would reflect specific pronunciation. See also [0022]-[0029].) and the standard pronunciation feature corresponding to the exemplary audio is obtained by fusing a plurality of reference pronunciation features; ("[0024] Extracting the fundamental frequency feature, the short-time energy feature, and the formant feature of the test speech, and using the fundamental frequency feature, the short-time energy feature, and the formant feature to form the emotional feature parameters of the test speech;" - See also “[0117] 4.2) According to the emotional feature parameters of the test speech, based on the SVM (Support Vector Machine) emotion model, emotion recognition is performed on the test speech to obtain an emotion recognition result.” – See also: “[0111]…The specific steps for extracting the feature parameters of the standard speech are consistent with the feature parameter extraction process of the test speech, and will not be repeated here.”) the plurality of reference pronunciation features are obtained by using the encoder to respectively extract features from a plurality of pieces of reference audio, ("[0024] Extracting the fundamental frequency feature, the short-time energy feature, and the formant feature of the test speech, and using the fundamental frequency feature, the short-time energy feature, and the formant feature to form the emotional feature parameters of the test speech;" see also "[0026] Divide the test speech into stress units, extract the start frame position group and the end frame position group of the stress, and obtain stress position feature parameters of the test speech;" – These features read on extracting a plurality of features, but Li only discloses extracting from a single piece of reference audio. ) wherein, when extracting a reference pronunciation feature of each piece of reference audio, features of preset sampling points in each piece of reference audio are collected, and wherein the standard pronunciation feature comprises the features of preset sampling points in the plurality of pieces of reference audio; (“[0074] 2.3) Windowing: In order to emphasize the speech waveform near the sampling position in the test speech and weaken the rest of the waveform, the Hamming window is selected to window the test speech in this embodiment. Windowing after framing can reduce the Gibbs phenomenon caused by truncation, making the spectrum of the test speech smoother. In one possible implementation, the windowing calculation formula is as follows:” – see also [0073] which discloses framing. Both framing and windowing read on using preset sampling points.) the plurality of pieces of reference audio are audio of the specified content that are respectively read by using the specified language; (Li discloses in [0166] that the method is applied in the chosen language) and the exemplary audio is any piece of audio of the reference audio; ([0024] discloses that the features are extracted from the test speech which is the reference audio) determining a feature difference between the standard pronunciation feature and the evaluation pronunciation feature, ("[0012] evaluating the test speech according to the feature parameters of the test speech and the feature parameters of the standard speech to obtain an evaluation result;" ) and determining a similarity between the evaluation audio and the exemplary audio according to the feature difference; ("[0030]... similarity calculation is performed on the MFCC feature parameters of the test speech and the MFCC feature parameters of the standard speech to obtain an MFCC correlation coefficient; and according to the speech recognition result and the MFCC correlation coefficient, an accuracy score of the test speech is calculated;" ) and mapping the determined similarity to a score, (“[0037] The accuracy score, the emotion score, the speaking speed score, the stress score, the rhythm score and the intonation score are weighted and summed to obtain a comprehensive score;…”) and displaying the score (“[0049] The data display unit is used to display the evaluation result.”) Li does not explicitly disclose that the exemplary audio is played. Li discloses that pronunciation guidance opinions are fed back to the user (Li [0040]). Li also does not explicitly disclose an encoder or a plurality of reference audio. Bartosik discloses: the standard pronunciation feature corresponding to the exemplary audio is obtained by fusing a plurality of reference pronunciation features (“[0002]…By forming an average value of the feature vectors of all the reference speakers for the pronunciation of each phoneme of words of the text, the idiosyncrasies of the individual reference speakers are averaged and the thus determined reference information is suitable for a speaker-independent speech recognition device…”; see also: “[0024] Reference information RI, which is largely independent of the type of pronunciation of words by individual speakers and may also be referred to as speaker-independent reference information RI, is determined by the transformation matrix generator 1. For this purpose, a plurality of users speak a predefined text into the input devices 3, 5 and 7 in accordance with the reference determining method, to statistically average the differences of the individual speakers, as this is generally known…” – averaging features from different speakers reads on the feature fusion) Li and Bartosik are considered analogous art to the claimed invention because they disclose methods of analyzing speech. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Li to use an average of multiple reference audio sample features as taught by Bartosik. Doing so would have beneficial so that the standard speech is speaker-independent. (Bartosik [0003]). This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Bartosik does not disclose playing exemplary audio, or an encoder. Grover discloses: extracting, from the evaluation audio, an evaluation pronunciation feature corresponding to the standard pronunciation feature based on an encoder of a speech recognition model. (“The pipeline employs Bi-directional Recurrent Convolutional Neural Networks and Bi-directional Long Short-Term Memory Neural Networks to encode acoustic and lexical cues from spectrograms and transcriptions, respectively.”) Li, Bartosik, and Grover are considered analogous art to the claimed invention because they disclose devices for analyzing speech. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to encode the features with a RCNN as taught by Grover. Doing so would have been beneficial to extract local features and summarize long temporal information. (Grover Pg. 1 para 4) This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Grover does not disclose playing exemplary audio. Du discloses: playing exemplary audio, ("Play benchmark A/V 304" Fig. 3) Li, Bartosik, Grover, and Du are considered analogous art to the claimed invention because they disclose methods of speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to playback the exemplary audio as taught by Du. Doing so would have been beneficial so a student could hear the correct pronunciation. Regarding claim 4, Li discloses: 4. The method according to claim 1, wherein the determining the feature difference between the standard pronunciation feature and the evaluation pronunciation feature, and determining the similarity between the evaluation audio and the exemplary audio according to the feature difference comprises: determining a time wrapping function according to the standard pronunciation feature and the evaluation pronunciation feature; ("[0134] 4.6) According to the pitch characteristic parameters of the test speech and the pitch characteristic parameters of the standard speech, a DTW (Dynamic Time Warping) algorithm is utilized to obtain the pitch difference between the standard speech and the test speech, and based on the pitch difference, the intonation score of the test speech is calculated." – based on the specification of the instant application, Time Wrapping/Dynamic Time Wrapping are understood to mean the same thing as Time Warping/Dynamic Time Warping.) determining a plurality of combinations of alignment points according to the time wrapping function, the standard pronunciation feature and the evaluation pronunciation feature, wherein each combination of alignment points comprises a standard feature point in the standard pronunciation feature and an evaluation feature point in the evaluation pronunciation feature; ("[0132] Among them, d is the duration of the speech unit divided into sentences (such as: d <sub> k < /sub> is the duration of the kth speech unit), m=min(S <sub> snum </sub> , T <sub> snum < /sub> ), S <sub> snum </sub> is the number of speech units of standard speech, T <sub> snum </sub> is the number of speech units of test speech, and Len is the duration of standard speech." - Determining alignment points is implied by DTW. It would be obvious to one of ordinary skill in the art that the speech units to be compared are aligned by the DTW algorithm.) determining, according to the standard feature point and the evaluation feature point comprised in each combination of alignment points, the feature difference corresponding to each combination of alignment points; ("[0130] After extracting the speech unit duration characteristic parameters of the test speech, the syllable unit duration characteristic parameters of the test speech are compared with the speech unit duration characteristic parameters of the standard speech ..." ) determining the similarity between the evaluation audio and the exemplary audio according to the feature difference of each combination of alignment points. ("[0130]...and the dPVI parameters used as the basis for system scoring are converted. The calculation formula of the dPVI parameters is as follows:" ) Regarding claim 5, Li discloses: 5. The method according to claim 1, further comprising: acquiring a mapping function, and configuration information corresponding to the exemplary audio, wherein the configuration information is used to indicate a mapping relationship between the score and the similarity which is between the evaluation audio and the exemplary audio; mapping the similarity between the evaluation audio and the exemplary audio to the score according to the mapping function and the configuration information corresponding to the exemplary audio. ("[0037] The accuracy score, the emotion score, the speaking speed score, the stress score, the rhythm score and the intonation score are weighted and summed to obtain a comprehensive score; …" ; see also "[0138]... And according to the accuracy score, the emotion score, the speaking speed score, the stress score, the rhythm score, the intonation score and the comprehensive score, combined with the mapping relationship between each score and the grade evaluation, the accuracy grade evaluation, emotion grade evaluation, speaking speed grade evaluation, stress grade evaluation, rhythm grade evaluation, intonation grade evaluation and comprehensive grade evaluation of the test speech are obtained...." ; regarding the configuration information, see: "[0139] In the process of weighted summing the accuracy score, the emotion score, the speaking speed score, the stress score, the rhythm score and the intonation score, the weight of each indicator score may adopt different values according to different needs, and the weight combination that meets the user's needs may be selected according to the user's own characteristics...." ) Regarding claim 6, Li discloses: 6. The method according to claim 5, wherein the configuration information comprises a maximum score, similarity corresponding to the maximum score, a minimum score, and similarity corresponding to the minimum score. ("[0139]… For example, if the accuracy score is in the score range of 90 to 100, the accuracy grade is evaluated as A; if the accuracy score is in the score range of 70 to 90, the accuracy grade is evaluated as B; if the accuracy score is in the score range of 60 to 70, the accuracy grade is evaluated as C; if the accuracy score is in the score range of 0 to 60, the accuracy grade is evaluated as D.…" – “A” is the maximum score with a similarity range of 90-100, and “D” is the minimum score with a similarity range of 0-60.) Regarding claim 8, Li and Bartosik do not disclose the additional limitations. Grover discloses: 8. The method according to claim 6, wherein the similarity corresponding to the minimum score is an average value of multiple pieces of white noise similarity, each piece of white noise similarity is similarity between each white noise feature and the standard pronunciation feature, and each white noise feature is obtained by using the encoder to extract a feature from each piece of preset white noise audio. ("White noise with intact text: Replacing audio inputs to the model with white noise …" Pg. 4 para 3) Li, Bartosik, Grover, and Du are considered analogous art to the claimed invention because they disclose devices for speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination with white noise as the low score. Doing so would have been beneficial because white noise represents unintelligible speech. Regarding claim 9, Li discloses: 9. The method according to claim 1, before playing the exemplary audio, further comprising: transmitting, in response to a start instruction, a data request instruction to a server; ("[0055] The embodiment of the present invention is based on a C/S (Client/Server) architecture, constructs a mobile client and a server, collects a user's test voice signal through the mobile client and sends it to the server, the server evaluates the test voice and returns the voice evaluation result to the mobile client, and finally displays the evaluation result through the mobile client." - A data request instruction is implied in a Client/Server architecture.) receiving the encoder, ("[0113]… It should be noted that the segmented clustering probabilistic neural network integrated speech recognition model is trained in advance, stored in a database, and retrieved when needed.".) the exemplary audio, ("[0165]... The mobile client 100 collects the user's test voice signal and sends it to the server 200. …" ) the standard pronunciation feature corresponding to the exemplary audio. ("[0020] The speech feature parameters are extracted from the preprocessed standard speech to obtain the feature parameters of the standard speech." – It is noted that the claim does not limit which part of the system receives the encoder, exemplary audio, and standard pronunciation feature.) Regarding claim 10, Li discloses: 10. The method according to claim 1, wherein the speech recognition model is obtained by performing training on an initial model using speech recognition data; ("[0113]… It should be noted that the segmented clustering probabilistic neural network integrated speech recognition model is trained in advance, stored in a database, and retrieved when needed." ) the encoder for extracting the pronunciation feature is obtained by performing training on the encoder in the speech recognition model using audio data in plural categories of languages. ("[0166] The online spoken English pronunciation quality evaluation method and system provided by the embodiments of the present invention can be applied to spoken English learning to detect the pronunciation quality of spoken English. It can also be applied to pronunciation quality assessment of other languages, such as Japanese and French." ) Regarding claim 11, Li and Bartosik do not disclose the additional limitations. Grover discloses: 11. The method according to claim 2, wherein the encoder is a three-layer long short-term memory network. (Fig. 1 shows multiple LSTM layers for encoding the audio.) Li, Bartosik, Du, and Grover are considered analogous art to the claimed invention because they disclose devices for speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to include LSTM layers in the encoder. Doing so would have been beneficial “to capture the sequential structure of audio, taking into account both past and future speech production.” (Grover Pg. 1 para 5) This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. One of ordinary skill in the art would understand that the number of layers in a neural network can be varied to find the best results. Therefore, using 3 layers would have been obvious to try. See MPEP 2143. Regarding claim 12, Li discloses: 12. A method for processing a data request instruction, applied to a server, and comprises: ("[0050] Preferably, the system further comprises a webpage management terminal, which is connected to the server terminal via a network; the server terminal further comprises a database and a statistical analysis unit;" ) receiving the data request instruction; ("[0165]… The mobile client 100 collects the user's test voice signal and sends it to the server 200. The server 200 evaluates the test voice and returns the voice evaluation result to the mobile client 100 …" – This is a data request.) transmitting, according to the data request instruction, an encoder based on a speech recognition model, (see [003] and claim 2) exemplary audio (“[0019] … standard speech;”), and a standard pronunciation feature corresponding to the exemplary audio (“[0020] The speech feature parameters are extracted from the preprocessed standard speech to obtain the feature parameters of the standard speech.”) to a user terminal; wherein the exemplary audio is audio of specified content that is read by using a specified language, and (“[0166] The online spoken English pronunciation quality evaluation method and system provided by the embodiments of the present invention can be applied to spoken English learning to detect the pronunciation quality of spoken English. It can also be applied to pronunciation quality assessment of other languages, such as Japanese and French.” - Whichever language quality assessment is used would be the specified language.) the encoder of the speech recognition model is used to extract, from evaluation audio, an evaluation pronunciation feature corresponding to the standard pronunciation feature, wherein the standard pronunciation feature is used to reflect a specific pronunciation of the specified content in the specified language; (“[0020] The speech feature parameters are extracted from the preprocessed standard speech to obtain the feature parameters of the standard speech.” ; "[0025] Calculating the pronunciation duration of the test speech and obtaining a pronunciation duration feature parameter of the test speech;") the standard pronunciation feature corresponding to the exemplary audio is obtained by fusing a plurality of reference pronunciation features; the plurality of reference pronunciation features are obtained by using the encoder to respectively extract features from a plurality of pieces of reference audio, ("[0024] Extracting the fundamental frequency feature, the short-time energy feature, and the formant feature of the test speech, and using the fundamental frequency feature, the short-time energy feature, and the formant feature to form the emotional feature parameters of the test speech;" see also "[0026] Divide the test speech into stress units, extract the start frame position group and the end frame position group of the stress, and obtain stress position feature parameters of the test speech;" – These features read on extracting a plurality of features, but Li only discloses extracting from a single piece of reference audio. ) wherein, when extracting a reference pronunciation feature of each piece of reference audio, features of preset sampling points in each piece of reference audio are collected, and wherein the standard pronunciation feature comprises the features of preset sampling points in the plurality of pieces of reference audio; (“[0074] 2.3) Windowing: In order to emphasize the speech waveform near the sampling position in the test speech and weaken the rest of the waveform, the Hamming window is selected to window the test speech in this embodiment. Windowing after framing can reduce the Gibbs phenomenon caused by truncation, making the spectrum of the test speech smoother. In one possible implementation, the windowing calculation formula is as follows:” – see also [0073] which discloses framing. Both framing and windowing read on using preset sampling points.) the plurality of pieces of reference audio are audio of the specified content that are respectively read by using the specified language; (Li discloses in [0166] that the method is applied in the chosen language) and the exemplary audio is any piece of audio of the reference audio; ([0024] discloses that the features are extracted from the test speech which is the reference audio) wherein a feature difference between the standard pronunciation feature and the evaluation pronunciation feature is determined for determining a similarity between the evaluation audio and the exemplary audio; and ("[0024] Extracting the fundamental frequency feature, the short-time energy feature, and the formant feature of the test speech, and using the fundamental frequency feature, the short-time energy feature, and the formant feature to form the emotional feature parameters of the test speech;" - See also “[0117] 4.2) According to the emotional feature parameters of the test speech, based on the SVM (Support Vector Machine) emotion model, emotion recognition is performed on the test speech to obtain an emotion recognition result.” – See also: “[0111]…The specific steps for extracting the feature parameters of the standard speech are consistent with the feature parameter extraction process of the test speech, and will not be repeated here.”) wherein the determined similarity is mapped to a score for displaying the score. (“[0037] The accuracy score, the emotion score, the speaking speed score, the stress score, the rhythm score and the intonation score are weighted and summed to obtain a comprehensive score;…”; “[0049] The data display unit is used to display the evaluation result.”) Li does not disclose that encoder, exemplary audio, and standard pronunciation feature are transmitted to a user terminal. Li discloses that the similarity analysis is done by the server, not the user terminal. Li does not explicitly disclose an encoder or a plurality of reference audio. Bartosik discloses: the standard pronunciation feature corresponding to the exemplary audio is obtained by fusing a plurality of reference pronunciation features (“[0002]…By forming an average value of the feature vectors of all the reference speakers for the pronunciation of each phoneme of words of the text, the idiosyncrasies of the individual reference speakers are averaged and the thus determined reference information is suitable for a speaker-independent speech recognition device…”; see also: “[0024] Reference information RI, which is largely independent of the type of pronunciation of words by individual speakers and may also be referred to as speaker-independent reference information RI, is determined by the transformation matrix generator 1. For this purpose, a plurality of users speak a predefined text into the input devices 3, 5 and 7 in accordance with the reference determining method, to statistically average the differences of the individual speakers, as this is generally known…” – averaging features from different speakers reads on the feature fusion) Li and Bartosik are considered analogous art to the claimed invention because they disclose methods of analyzing speech. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Li to use an average of multiple reference audio sample features as taught by Bartosik. Doing so would have beneficial so that the standard speech is speaker-independent. (Bartosik [0003]). This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Bartosik does not disclose an encoder, or that encoder, exemplary audio, and standard pronunciation feature are transmitted to a user terminal. Grover discloses: extracting, from the evaluation audio, an evaluation pronunciation feature corresponding to the standard pronunciation feature based on an encoder of a speech recognition model, (“The pipeline employs Bi-directional Recurrent Convolutional Neural Networks and Bi-directional Long Short-Term Memory Neural Networks to encode acoustic and lexical cues from spectrograms and transcriptions, respectively.”) Li, Bartosik, and Grover are considered analogous art to the claimed invention because they disclose devices for speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to encode the features with a RCNN as taught by Grover. Doing so would have been beneficial to extract local features and summarize long temporal information. (Grover Pg. 1 para 4) This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Grover does not disclose that the encoder, exemplary audio, and standard pronunciation feature are transmitted to a user terminal. Du discloses that similarity analysis is performed by the user terminal (client). (“[0021] VLT online client 104 comprises client software that enables a student to obtain oral practice assignments assigned by the teacher, perform the oral practice assignments, and receive performance results or feedback and grading based on their performance of the oral practice assignments.” – Fig. 1 shows VLT Online Client 104 is part of the Client Side 102 of the system. See also “[0017] Virtual language tutor online server 112 comprises a virtual language tutor content management module 112 a, a homework management module 112 b, and a virtual language tutor learner information management module 112 c. VLT content management module 112 a comprises content modules that may be used for assignments, or to prepare assignments…” – the VLT Content Management Module 112a contains the information to be transmitted to the client terminal for use in a learning session. See also “[0019]… In this instance, the student may select a language course offered by VLT online server…”) Li, Bartosik, Grover, and Du are considered analogous art to the claimed invention because they disclose methods of speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to transmit the encoder, exemplary audio, and standard pronunciation feature to the client terminal as taught by Du. Doing so would have been beneficial so a student could complete lessons without a continuous connection to the server. Claim 13 is an apparatus claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally, a memory; a processor; and a computer program; wherein the computer program is stored in the memory, and configured to be executed by the processor are taught by Li. (“[0167] Through the description of the above implementation methods, technicians in the relevant field can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware. Of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. The technical solution of the present invention can essentially be embodied in the form of a software product or the part that contributes to the prior art. The software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.”) Claim 14 is an apparatus, claim with limitations corresponding to the limitations of Claim 12 and is rejected under similar rationale. Additionally, a memory; a processor; and a computer program; of the Claim are taught by Li. (See [0167] as mapped in claim 13. ) Claim 16 is a non-transitory computer-readable storage medium claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally, A non-transitory computer-readable storage medium having, stored thereon, a computer program of the Claim is taught by Li. (“[0167] The software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.. ) Claim 19 is a non-transitory computer-readable storage medium claim with limitations corresponding to the limitations of Claim 12 and is rejected under similar rationale. Additionally, A non-transitory computer-readable storage medium having, stored thereon, a computer program of the Claim is taught by Li. (“[0167] The software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.. ) Claim 22 is an apparatus claim with limitations corresponding to the limitations of Claim 4 and is rejected under similar rationale. Claim 23 is an apparatus claim with limitations corresponding to the limitations of Claim 5 and is rejected under similar rationale. Regarding claim 24, Li, Bartosik, and Grover do not disclose the additional limitations. Du discloses: 24. The method according to claim 1, wherein when the specified content is read in the specified language, the preset sampling points in each piece of reference audio are determined based on pronunciation positions characterized in the specified language. (“[0040] VLT online client may then analyze the student's accuracy, by assessing for example the pronunciation and intonation of each word or phoneme by comparing it with the pronunciation and intonation of the benchmark voice or in some other way (block 310). This may be accomplished in any of a variety of different ways including using forced alignment, speech analysis, and pattern recognition techniques. VLT online client may also analyze the student's speed by measuring the elapsed time or duration of the recorded utterance and comparing it to the duration of the benchmark voice. The speed measurement may be determined on a per word, per sentence, per passage or total utterance basis. Alternatively, one or more of these speed measures may be combined. The accuracy and speed may then be combined into a fluency score (block 311), using, for example any one or more of Equations 1, 2, or 3, described above.” – performing forced alignment and assessing by word or phoneme reads on preset sampling points because the word is specific to the language used. Measuring the speed on a per word/sentence/passage basis also reads on preset sampling points.) Li, Bartosik, Grover, and Du are considered analogous art to the claimed invention because they disclose methods of speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to analyze the accuracy of each phoneme by using forced alignment as taught by Du. Doing so would have been beneficial to aid the student in knowing which sentence, word, or phoneme needs improvement (Du [0041]). Regarding claim 25, Li, Bartosik, and Grover do not disclose the additional limitations. Du discloses: 25. The method according to claim 12, wherein when the specified content is read in the specified language, the preset sampling points in each piece of reference audio are determined based on pronunciation positions characterized in the specified language. (“[0040] VLT online client may then analyze the student's accuracy, by assessing for example the pronunciation and intonation of each word or phoneme by comparing it with the pronunciation and intonation of the benchmark voice or in some other way (block 310). This may be accomplished in any of a variety of different ways including using forced alignment, speech analysis, and pattern recognition techniques. VLT online client may also analyze the student's speed by measuring the elapsed time or duration of the recorded utterance and comparing it to the duration of the benchmark voice. The speed measurement may be determined on a per word, per sentence, per passage or total utterance basis. Alternatively, one or more of these speed measures may be combined. The accuracy and speed may then be combined into a fluency score (block 311), using, for example any one or more of Equations 1, 2, or 3, described above.” – performing forced alignment and assessing by word or phoneme reads on preset sampling points because the word is specific to the language used. Measuring the speed on a per word/sentence/passage basis also reads on preset sampling points.) Li, Bartosik, Grover, and Du are considered analogous art to the claimed invention because they disclose methods of speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to analyze the accuracy of each phoneme by using forced alignment as taught by Du. Doing so would have been beneficial to aid the student in knowing which sentence, word, or phoneme needs improvement (Du [0041]). Claim 26 is an apparatus claim with limitations corresponding to the limitations of Claim 24 and is rejected under similar rationale. Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Bartosik, Grover, and Du as applied to claim 6 above, and further in view of Nakamura (US 20160055763 A1). Regarding claim 7, Li discloses: 7. The method according to claim 6, wherein the similarity corresponding to the maximum score is an average value of multiple pieces of reference similarity, and each piece of reference similarity is similarity between each reference pronunciation feature and the standard pronunciation feature. ("[0037] The accuracy score, the emotion score, the speaking speed score, the stress score, the rhythm score and the intonation score are weighted and summed to obtain a comprehensive score;…" ) Li, Bartosik, Grover, and Du do not disclose an average value for the score. Li discloses a summed value. Nakamura discloses: 7. The method according to claim 6, wherein the similarity corresponding to the maximum score is an average value of multiple pieces of reference similarity, and each piece of reference similarity is similarity between each reference pronunciation feature and the standard pronunciation feature. ("[0044] The evaluation score data storage area 33 d stores data of evaluation scores for respective phonetic symbols and average scores of the evaluation scores. The evaluation scores are obtained in accordance with the degree of similarity by comparison between the speech data by the user's pronunciation, which is stored in the recorded speech data storage area 33 c, and the model speech data of the corresponding word or illustrative sentence, or the important word for practice, which is stored in the word database 32 b or illustrative sentence database 32 c, or the word-for-practice search table 32 d." ) Li, Bartosik, Grover, Du, and Nakamura are considered analogous art to the claimed invention because they disclose methods of speech analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination with an average score as taught by Nakamura. This combination falls under combining prior art elements according to known methods to yield predictable results or simple substitution of one known element for another to obtain predictable results. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Tian (CN 111554300 A). Tian discloses an audio data processing method wherein the selected audio features of candidate users are averaged to make the reference audio more representative and distinctive. Evanini (US 20140195239 A1). Evanini discloses a pronunciation assessment system which aligns words and phonemes before extracting measurement characteristics. See [0018] and Fig. 3. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JON C MEIS whose telephone number is (703)756-1566. The examiner can normally be reached Monday - Thursday, 8:30 am - 5:30 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JON CHRISTOPHER MEIS/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Show 2 earlier events
Oct 06, 2025
Response Filed
Nov 18, 2025
Final Rejection mailed — §103
Jan 20, 2026
Response after Non-Final Action
Apr 20, 2026
Request for Continued Examination
Apr 24, 2026
Response after Non-Final Action
May 15, 2026
Non-Final Rejection mailed — §103
Aug 13, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711961
SYSTEM AND METHOD FOR DIGITAL VOICE DATA PROCESSING AND AUTHENTICATION
3y 4m to grant Granted Aug 18, 2026
Patent 12603087
VOICE RECOGNITION USING ACCELEROMETERS FOR SENSING BONE CONDUCTION
3y 8m to grant Granted Apr 14, 2026
Patent 12579975
Detecting Unintended Memorization in Language-Model-Fused ASR Systems
2y 11m to grant Granted Mar 17, 2026
Patent 12482487
MULTI-SCALE SPEAKER DIARIZATION FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
3y 0m to grant Granted Nov 25, 2025
Patent 12475312
FOREIGN LANGUAGE PHRASES LEARNING SYSTEM BASED ON BASIC SENTENCE PATTERN UNIT DECOMPOSITION
2y 9m to grant Granted Nov 18, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
33%
Grant Probability
86%
With Interview (+52.4%)
2y 10m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 33 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month