Prosecution Insights
Last updated: August 18, 2026
Application No. 18/673,609

Speech Recognition Method, Speech Recognition Apparatus, and System

Final Rejection §103
Filed
May 24, 2024
Priority
Nov 25, 2021 — continuation of PCTCN2021133207
Examiner
BOGGS JR., JAMES
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Shenzhen Yinwang Intelligent Technology Co., Ltd.
OA Round
2 (Final)
63%
Grant Probability
Moderate
3-4
OA Rounds
11m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
75 granted / 119 resolved
+1.0% vs TC avg
Strong +34% interview lift
Without
With
+34.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
27 currently pending
Career history
142
Total Applications
across all art units

Statute-Specific Performance

§101
11.8%
-28.2% vs TC avg
§103
50.4%
+10.4% vs TC avg
§102
15.7%
-24.3% vs TC avg
§112
18.5%
-21.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 119 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed June 1, 2026, has been entered. Claims 1 – 2, 4 – 10 and 12 – 22 are pending in the application. Response to Arguments Applicant’s arguments, filed June 1, 2026, with respect to claims 1 – 2, 4 – 10 and 12 – 22 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference signs mentioned in the description: “400” in paragraphs 0115, 0130, 0136, 0165, and 0177 “404” in paragraphs 0130, 0131, and 0136 “500” in paragraphs 0160 and 0171. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 – 2, 6 – 10, 14 – 18 and 21 – 22 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US Patent No. 9,437,186), hereinafter Liu, in view of Maas et al. (Maas, Roland, Ariya Rastrow, Chengyuan Ma, Guitang Lan, Kyle Goehner, Gautam Tiwari, Shaun Joseph, and Björn Hoffmeister, "Combining Acoustic Embeddings and Decoding Features for End-of-Utterance Detection in Real-Time Far-Field Speech Recognition Systems", April 2018, 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2018), pp. 5544-5548.), hereinafter Maas. Regarding claim 1, Liu discloses a speech recognition method, comprising: obtaining first audio data comprising a plurality of audio frames (Column 8, lines 4-6, "The ASR module 314 may include an acoustic front end (AFE), not shown. The AFE transforms audio data into data for processing by the speech recognition engine."; Column 8, lines 11-15, "The AFE may reduce noise in the audio data and divide the digitized audio data into frames representing time intervals for which the AFE determines a set of values, called a feature vector, representing the features/qualities of the utterance portion within the frame."); extracting sound categories of the audio frames and semantics of the audio frames based on relationships between energies of the audio frames and preset energy thresholds (Column 4, lines 30-48, "Audio detection processing for endpoint determination may be performed by determining an energy level of the audio input. In some embodiments, the endpointing/audio detection may include a low-power digital signal processor (or other type of processor) configured to determine an energy level (such as a volume, intensity, amplitude, etc.) of an obtained audio input and for comparing the energy level of the audio input to an energy level threshold. The energy level threshold may be set according to user input, or may be set by a device. In some embodiments, the endpointing/audio detection may be further configured to determine that the audio input has an energy level satisfying a threshold for at least a threshold duration of time. In such embodiments, high-energy audio inputs of relatively short duration, which may correspond to sudden noises that are relatively unlikely to include speech, may be ignored. The endpointing/audio detection may compare the energy level to the energy level threshold (and optionally to the threshold duration) to determine whether the energy level threshold is met."; Column 4, lines 49-64, "If the endpointing/audio detection determines that the obtained audio input has an energy level satisfying an energy level threshold it may process audio input to determine whether the audio input includes speech. In some embodiments, the endpointing/audio detection works in conjunction with digital signal processing to implement one or more techniques to determine whether the audio input includes speech. Some embodiments may apply voice activity detection (VAD) techniques, such as harmonicity detection. Such techniques may determine whether speech is present in an audio input based on various quantitative aspects of the audio input, such as the spectral slope between one or more frames of the audio input; the energy levels of the audio input in one or more spectral bands; the signal-to-noise ratios of the audio input in one or more spectral bands; or other quantitative aspects."; Column 11, lines 27-44, "In a semantic interpretation process (which may be part of traditional NLU processing), which commonly takes places after ASR processing, semantic tagging is a process of recognizing and identifying specific important words of an ASR output and assigning a tag to those words, where the tag is a classification of the associated word. The tags may be called entities or named entities. Some words in a phrase may be considered less important, thus not considered for a named entity and may not receive a tag or may be given a catchall or default tag such as “Unknown” or “DontCare.” The tagging process may also be referred to as named entity recognition (NER). In this aspect, the semantic information and tags are built in to the models used to perform ASR processing such that the semantic information (which may be less comprehensive than tags available in a post-ASR semantic tagging process) is output with the text as a result of the ASR process."); Comparing the energy level of the audio input to an energy level threshold to determine whether speech is present in an audio input, and performing semantic tagging of the speech, reads on extracting sound categories of the audio frames and semantics of the audio frames based on relationships between energies of the audio frames and preset energy thresholds.); and obtaining a speech ending point of the first audio data [based on the integrated feature] (Column 3, lines 58-65, "In a further implementation according to the disclosure, semantic information in the user's speech may be used to help determine the end of an utterance. That is, an ASR processor may be configured so that semantic tags or other indicators may be included as part of the ASR output and used to determine, as the user's speech is recognized, whether the user's utterance has reached a logical stopping point."; Column 11, line 67 - Column 12, line 2, "Thus semantic information in the user's speech may be used to help determine the ending of an utterance, instead of basing it on non-speech audio frames only."; Determining the ending of an utterance using semantic information and non-speech audio frames reads on obtaining a speech ending point of the first audio data.). Liu does not specifically disclose: integrating the sound categories and the semantics to obtain an integrated feature of the audio frames; and obtaining a speech ending point of the first audio data based on the integrated feature. Maas teaches: integrating the sound categories and the semantics to obtain an integrated feature of the audio frames (Section 1, lines 29-38, "In this publication, we investigate a combination of LSTM-based classifiers, similarly to [25, 26], as follows: First, an LSTM is trained on acoustic features with frame-wise multi-task targets to predict both the utterance end-point as well as voice activity. Second, an LSTM is trained on embeddings of the 1-best ASR hypothesis. Third, a DNN is trained on frame-wise end-pointing targets combining three types of input features: the final layer representations of the acoustic and word LSTMs as well as pause duration estimates from the ASR decoder."; Section 2.1, lines 51-56, "Subsequently, we form joint feature vectors f t = [ a t , h t , d t ] at every frame by concatenating the three feature types, as shown in Fig. 1 c): i) the hidden representations a t of the last layer of the acoustic LSTM, ii) the hidden representation h t of the last layer the word LSTM, iii) the decoder feature d t ."; Joint feature vectors f t read on an integrated feature of the audio frames, acoustic LSTM representations a t read on sound categories, and word LSTM representations h t read on semantics.); and obtaining a speech ending point of the first audio data based on the integrated feature (Abstract, lines 1-9, "We present an end-of-utterance detector for real-time automatic speech recognition in far-field scenarios. The proposed system consists of three components: a long short-term memory (LSTM) neural network trained on acoustic features, an LSTM trained on 1-best recognition hypotheses of the automatic speech recognition (ASR) decoder, and a feed-forward deep neural network (DNN) combining embeddings derived from both LSTMs with pause duration features from the ASR decoder."; Section 2.1, lines 59-61, "Finally, a joint classification layer, i.e., a fully connected DNN depicted in Fig. 2, is trained on the joint feature vectors f t shown in Fig. 1 c)."; An end-of-utterance detector with a deep neural network classifier that performs end-of-utterance classification based on joint feature vectors reads on obtaining a speech ending point of the first audio data based on the integrated feature.). Maas is considered to be analogous to the claimed invention because it is in the same field of end-of-utterance detection. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu to incorporate the teachings of Maas to form joint feature vectors from acoustic LSTM representations and word LSTM representations, and perform end-of-utterance detection with a deep neural network classifier that performs end-of-utterance classification based on the joint feature vectors. Doing so would allow for implementing an end-of-utterance detection system for real-time speech recognition in far-field scenarios that allows for resource-efficient adaptation to new domains and across languages (Maas; Section 4, lines 1-7). Regarding claim 2, Liu in view of Maas discloses the speech recognition method as claimed in claim 1. Liu further discloses: wherein after obtaining the first audio data, the method further comprises responding to an instruction corresponding to second audio data that is prior to the speech ending point (Column 2, lines 27-43, "The downstream component may be any number of components or applications that operate on ASR output. Although many such downstream applications are envisioned for these techniques, for purposes of illustration this description will use an NLU process and application as the NLU process illustrates the benefits of early ASR output as described below. For example, the NLU process may take ASR output and determine, for example, the actions (sometimes referred to as an “application response” or “app response”) based on the recognized speech of the early ASR output. The app response based on the early ASR output may not be immediately activated but may be held pending successful comparison of the early ASR output to the final ASR output. If the comparison is successful, then the app response is immediately available for execution (rather than having to wait for NLU processing in a typically staged process), thus improving latency from a user's perspective."; Performing an application response based on the recognized speech of the early automatic speech recognition output reads on responding to an instruction corresponding to second audio data that is prior to the speech ending point.). Regarding claim 6, Liu in view of Maas discloses the speech recognition method as claimed in claim 1. Liu further discloses: wherein the audio frames comprise a first audio frame and a second audio frame, wherein the first audio frame includes the semantics, wherein the second audio frame is subsequent to the first audio frame in the audio frames (Column 11, line 65 - Column 12, line 9, "Early or final endpoints, or both, may be adjusted in a system according to the disclosure. Thus semantic information in the user's speech may be used to help determine the ending of an utterance, instead of basing it on non-speech audio frames only. The threshold of the number of non-speech frames may be dynamically changed based on the semantic meaning the speech that has been recognized so far. Thus, the ASR module 314 may determine a likelihood that an utterance includes a complete command and use that utterance to adjust the threshold of non-speech frames for determining the end of the utterance."; Semantic information in the user's speech reads on a first audio frame including the semantics, and non-speech audio frames following speech reads on a second audio frame subsequent to the first audio frame.). Maas further teaches: and wherein integrating the sound categories and the semantics to obtain the integrated feature comprises integrating the semantics and a first sound category of the second audio frame to obtain the integrated feature (Section 1, lines 29-38, "In this publication, we investigate a combination of LSTM-based classifiers, similarly to [25, 26], as follows: First, an LSTM is trained on acoustic features with frame-wise multi-task targets to predict both the utterance end-point as well as voice activity. Second, an LSTM is trained on embeddings of the 1-best ASR hypothesis. Third, a DNN is trained on frame-wise end-pointing targets combining three types of input features: the final layer representations of the acoustic and word LSTMs as well as pause duration estimates from the ASR decoder."; Section 2.1, lines 51-56, "Subsequently, we form joint feature vectors f t = [ a t , h t , d t ] at every frame by concatenating the three feature types, as shown in Fig. 1 c): i) the hidden representations a t of the last layer of the acoustic LSTM, ii) the hidden representation h t of the last layer the word LSTM, iii) the decoder feature d t ."; Joint feature vectors f t read on an integrated feature of the audio frames, acoustic LSTM representations a t read on sound categories, and word LSTM representations h t read on semantics.); Maas is considered to be analogous to the claimed invention because it is in the same field of end-of-utterance detection. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Maas to further incorporate the teachings of Maas to form joint feature vectors at every frame from acoustic LSTM representations and word LSTM representations. Doing so would allow for implementing an end-of-utterance detection system for real-time speech recognition in far-field scenarios that allows for resource-efficient adaptation to new domains and across languages (Maas; Section 4, lines 1-7). Regarding claim 7, Liu in view of Maas discloses the speech recognition method as claimed in claim 6. Liu further discloses: wherein speech endpoint categories comprise "speaking", "thinking", and "ending" (Column 12, lines 38-46, "The pauses, P1 and P2, may be tagged as “Non-Speech” and be associated with respective durations or numbers of non-speech frames, or time ranges, as a result of the phrase and semantics. Some pauses, such as P2 in FIG. 6A, may represent a tag for the end of a command. The end of a command may be inserted based on whether the ASR model recognizes that the preceding words have formed a complete command."; Column 12, lines 36-60, "During runtime the ASR module 314 in FIG. 3, receives the textual input, compares the input with its available models and as part of ASR processing, outputs tags along with the associated processed speech with each word or pause (i.e. non-speech frames). As described above with respect to FIG. 2, and further illustrated in FIGS. 6 and 7, this semantic information in the user's speech, and represented by the semantic tags output by the ASR, may be used to help determine the end of an utterance. The amount of time after non-speech is detected until the ASR is activated to produce speech results may be dynamically adjusted based on tags in the language model that provide the semantic information appropriate to more accurately determine the end of an utterance for a given user. Based on training data, the language model may be adjusted to reflect the likelihood that the ASR should await more speech to process a complete utterance. For example, if the ASR module 314 is processing the word “songs” and it knows that the word directly following is “songs” is “by” it may be more likely to continue processing in order to get to and apply the tag “ArtistName” to the words “Michael Jackson” as in FIG. 6B. The tag for pause P5 after “Jackson” may signal the end of the phrase/command and prompt generation of the ASR output without waiting for a traditional number of silent frames."; Detecting speech frames reads on a speech endpoint category of "speaking", detecting non-speech frames and determining that the automatic speech recognition should await more speech to process a complete utterance reads on a speech endpoint category of "thinking", and detecting non-speech frames and recognizing that the preceding words have formed a complete command reads on a speech endpoint category of "ending".), and wherein obtaining the speech ending point based on the integrated feature comprises: determining a first speech endpoint category of the first audio data [based on the semantics and the first sound category] (Column 11, line 65 - Column 12, line 9, "Early or final endpoints, or both, may be adjusted in a system according to the disclosure. Thus semantic information in the user's speech may be used to help determine the ending of an utterance, instead of basing it on non-speech audio frames only. The threshold of the number of non-speech frames may be dynamically changed based on the semantic meaning the speech that has been recognized so far. Thus, the ASR module 314 may determine a likelihood that an utterance includes a complete command and use that utterance to adjust the threshold of non-speech frames for determining the end of the utterance."; Determining the ending of an utterance using semantic information in speech and non-speech audio frames following speech reads on determining a first speech endpoint category of the first audio data.); and obtaining the speech ending point in response to the first speech endpoint category being "ending" (Column 11, line 65 - Column 12, line 9, "Early or final endpoints, or both, may be adjusted in a system according to the disclosure. Thus semantic information in the user's speech may be used to help determine the ending of an utterance, instead of basing it on non-speech audio frames only. The threshold of the number of non-speech frames may be dynamically changed based on the semantic meaning the speech that has been recognized so far. Thus, the ASR module 314 may determine a likelihood that an utterance includes a complete command and use that utterance to adjust the threshold of non-speech frames for determining the end of the utterance."; Determining the ending of an utterance using semantic information in speech and non-speech audio frames following speech reads on obtaining the speech ending point in response to the first speech endpoint category is "ending".). Maas further teaches: wherein obtaining the speech ending point based on the integrated feature comprises: determining a first speech endpoint category of the first audio data based on the semantics and the first sound category (Abstract, lines 1-9, "We present an end-of-utterance detector for real-time automatic speech recognition in far-field scenarios. The proposed system consists of three components: a long short-term memory (LSTM) neural network trained on acoustic features, an LSTM trained on 1-best recognition hypotheses of the automatic speech recognition (ASR) decoder, and a feed-forward deep neural network (DNN) combining embeddings derived from both LSTMs with pause duration features from the ASR decoder."; Section 2.1, lines 51-56, "Subsequently, we form joint feature vectors f t = [ a t , h t , d t ] at every frame by concatenating the three feature types, as shown in Fig. 1 c): i) the hidden representations a t of the last layer of the acoustic LSTM, ii) the hidden representation h t of the last layer the word LSTM, iii) the decoder feature d t ."; Section 2.1, lines 59-61, "Finally, a joint classification layer, i.e., a fully connected DNN depicted in Fig. 2, is trained on the joint feature vectors f t shown in Fig. 1 c)."; An end-of-utterance detector with a deep neural network classifier that performs end-of-utterance classification based on joint feature vectors, where the joint feature vectors are formed at every frame from acoustic LSTM representations and word LSTM representations, reads on determining a first speech endpoint category of the first audio data based on the semantics and the first sound category.). Maas is considered to be analogous to the claimed invention because it is in the same field of end-of-utterance detection. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Maas to further incorporate the teachings of Maas to perform end-of-utterance classification based on joint feature vectors, where the joint feature vectors are formed at every frame from acoustic LSTM representations and word LSTM representations. Doing so would allow for implementing an end-of-utterance detection system for real-time speech recognition in far-field scenarios that allows for resource-efficient adaptation to new domains and across languages (Maas; Section 4, lines 1-7). Regarding claim 8, Liu in view of Maas discloses the speech recognition method as claimed in claim 7. Liu further discloses: wherein the speech endpoint classification model is based on a speech sample and an endpoint category label of the speech sample (Column 12, lines 36-60, "During runtime the ASR module 314 in FIG. 3, receives the textual input, compares the input with its available models and as part of ASR processing, outputs tags along with the associated processed speech with each word or pause (i.e. non-speech frames). As described above with respect to FIG. 2, and further illustrated in FIGS. 6 and 7, this semantic information in the user's speech, and represented by the semantic tags output by the ASR, may be used to help determine the end of an utterance. The amount of time after non-speech is detected until the ASR is activated to produce speech results may be dynamically adjusted based on tags in the language model that provide the semantic information appropriate to more accurately determine the end of an utterance for a given user. Based on training data, the language model may be adjusted to reflect the likelihood that the ASR should await more speech to process a complete utterance. For example, if the ASR module 314 is processing the word “songs” and it knows that the word directly following is “songs” is “by” it may be more likely to continue processing in order to get to and apply the tag “ArtistName” to the words “Michael Jackson” as in FIG. 6B. The tag for pause P5 after “Jackson” may signal the end of the phrase/command and prompt generation of the ASR output without waiting for a traditional number of silent frames."; Detecting speech and using semantic tags to determine that processing of an utterance should continue reads on obtaining the first speech endpoint category, wherein the speech endpoint classification model is based on a speech sample and an endpoint category label of the speech sample.), and wherein an endpoint category in the endpoint category label corresponds to the first speech endpoint category (Column 12, lines 36-60, "During runtime the ASR module 314 in FIG. 3, receives the textual input, compares the input with its available models and as part of ASR processing, outputs tags along with the associated processed speech with each word or pause (i.e. non-speech frames). As described above with respect to FIG. 2, and further illustrated in FIGS. 6 and 7, this semantic information in the user's speech, and represented by the semantic tags output by the ASR, may be used to help determine the end of an utterance. The amount of time after non-speech is detected until the ASR is activated to produce speech results may be dynamically adjusted based on tags in the language model that provide the semantic information appropriate to more accurately determine the end of an utterance for a given user. Based on training data, the language model may be adjusted to reflect the likelihood that the ASR should await more speech to process a complete utterance. For example, if the ASR module 314 is processing the word “songs” and it knows that the word directly following is “songs” is “by” it may be more likely to continue processing in order to get to and apply the tag “ArtistName” to the words “Michael Jackson” as in FIG. 6B. The tag for pause P5 after “Jackson” may signal the end of the phrase/command and prompt generation of the ASR output without waiting for a traditional number of silent frames."; Detecting speech and using semantic tags to determine that processing of an utterance should continue reads on obtaining the first speech endpoint category, wherein an endpoint category in the endpoint category label corresponds to the first speech endpoint category.). Maas further teaches: wherein determining the first speech endpoint category comprises processing the integrated feature using a speech endpoint classification model to obtain the first speech endpoint category (Abstract, lines 1-9, "We present an end-of-utterance detector for real-time automatic speech recognition in far-field scenarios. The proposed system consists of three components: a long short-term memory (LSTM) neural network trained on acoustic features, an LSTM trained on 1-best recognition hypotheses of the automatic speech recognition (ASR) decoder, and a feed-forward deep neural network (DNN) combining embeddings derived from both LSTMs with pause duration features from the ASR decoder."; Section 2.1, lines 59-61, "Finally, a joint classification layer, i.e., a fully connected DNN depicted in Fig. 2, is trained on the joint feature vectors f t shown in Fig. 1 c)."; An end-of-utterance detector with a deep neural network classifier that performs end-of-utterance classification based on joint feature vectors reads on processing the integrated feature using a speech endpoint classification model to obtain the first speech endpoint category.); wherein a first format of the speech sample corresponds to a second format of the integrated feature (Section 1, lines 29-38, "In this publication, we investigate a combination of LSTM-based classifiers, similarly to [25, 26], as follows: First, an LSTM is trained on acoustic features with frame-wise multi-task targets to predict both the utterance end-point as well as voice activity. Second, an LSTM is trained on embeddings of the 1-best ASR hypothesis. Third, a DNN is trained on frame-wise end-pointing targets combining three types of input features: the final layer representations of the acoustic and word LSTMs as well as pause duration estimates from the ASR decoder."; Section 2.1, lines 51-56, "Subsequently, we form joint feature vectors f t = [ a t , h t , d t ] at every frame by concatenating the three feature types, as shown in Fig. 1 c): i) the hidden representations a t of the last layer of the acoustic LSTM, ii) the hidden representation h t of the last layer the word LSTM, iii) the decoder feature d t ."; Forming joint feature vectors at every audio frame from acoustic LSTM representations and word LSTM representations reads on a format of the speech sample corresponds to a format of the integrated feature.). Maas is considered to be analogous to the claimed invention because it is in the same field of end-of-utterance detection. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Maas to further incorporate the teachings of Maas to form joint feature vectors at every audio frame from acoustic LSTM representations and word LSTM representations, and perform end-of-utterance classification based on the joint feature vectors with a deep neural network classifier. Doing so would allow for implementing an end-of-utterance detection system for real-time speech recognition in far-field scenarios that allows for resource-efficient adaptation to new domains and across languages (Maas; Section 4, lines 1-7). Regarding claim 9, arguments analogous to claim 1 are applicable. In addition, Liu discloses a speech recognition apparatus (Column 2, lines 54-57, “As illustrated, a speech recognition process 100 may be implemented on a client or local device 102 such as a smart phone or other local device”), comprising: an obtainer circuit configured to obtain first audio data comprising a plurality of audio frames (Column 8, lines 4-6, "The ASR module 314 may include an acoustic front end (AFE), not shown. The AFE transforms audio data into data for processing by the speech recognition engine."; Column 8, lines 11-15, "The AFE may reduce noise in the audio data and divide the digitized audio data into frames representing time intervals for which the AFE determines a set of values, called a feature vector, representing the features/qualities of the utterance portion within the frame."); and a processor (Column 5, lines 23-26, “As discussed above, any or all of the modules may be embodied in one or more general-purpose microprocessors, or in one or more special-purpose digital signal processors or other dedicated microprocessing hardware.”) configured to perform the steps of claim 1. Regarding claim 10, arguments analogous to claim 2 are applicable. Regarding claim 14, arguments analogous to claim 6 are applicable. Regarding claim 15, arguments analogous to claim 7 are applicable. Regarding claim 16, arguments analogous to claim 8 are applicable. Regarding claim 17, arguments analogous to claim 1 are applicable. In addition, Liu discloses a computer program product comprising computer-executable instructions that are stored on a non-transitory computer-readable storage media (Column 5, lines 47-51, “The memory can further include computer program instructions that an application processing module and/or processing unit in the device 102 executes in order to implement one or more embodiments of a speech recognition system with distributed endpointing according to the disclosure.”) and that, when executed by a processor, cause a speech recognition apparatus (Column 2, lines 54-57, “As illustrated, a speech recognition process 100 may be implemented on a client or local device 102 such as a smart phone or other local device”) to perform the steps of claim 1. Regarding claim 18, arguments analogous to claim 2 are applicable. Regarding claim 21 arguments analogous to claim 6 are applicable. Regarding claim 22, arguments analogous to claim 7 are applicable. Claims 4 – 5, 12 -– 13 and 19 – 20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Maas, and further in view of Rajan et al. (US Patent Application Publication No. 2002/0198704), hereinafter Rajan. Regarding claim 4, Liu in view of Maas discloses the speech recognition method as claimed in claim 1, but does not specifically disclose: wherein the sound categories comprise "speech", "neutral", and "silence", wherein the preset energy thresholds comprise a first energy threshold and a second energy threshold, wherein the first energy threshold is greater than the second energy threshold, wherein a first sound category of the sound categories and of a first audio frame in the audio frames and with a first energy that is greater than or equal to the first energy threshold is "speech", wherein a second sound category of the sound categories and of a second audio frame in the audio frames and with a second energy that is less than the first energy threshold and is greater than the second energy threshold is "neutral", and wherein a third sound category of a third audio frame in the audio frames and with a third energy that is less than or equal to the second energy threshold is "silence". Rajan teaches: wherein the sound categories comprise "speech", "neutral", and "silence", wherein the preset energy thresholds comprise a first energy threshold and a second energy threshold, wherein the first energy threshold is greater than the second energy threshold, wherein a first sound category of the sound categories and of a first audio frame in the audio frames and with a first energy that is greater than or equal to the first energy threshold is "speech", wherein a second sound category of the sound categories and of a second audio frame in the audio frames and with a second energy that is less than the first energy threshold and is greater than the second energy threshold is "neutral", and wherein a third sound category of a third audio frame in the audio frames and with a third energy that is less than or equal to the second energy threshold is "silence" (Paragraph 0038, lines 1-9, "In this embodiment, two threshold values are actually determined and stored within the threshold memory 39--a coarse threshold value which is used to indicate the start of the signal which is clearly not background noise and a fine threshold value which is used to determine the start point of speech more accurately. In this embodiment, the fine threshold value is the 0.01 percentile energy value discussed above and the coarse threshold value is the 0.05 percentile level."; Paragraph 0039, lines 9-21, "The speech/noise decision unit 38 then compares the energy values calculated for each block of samples (as determined by the block energy determining unit 35) with the threshold energy levels stored in the threshold memory 39. If the residual energy value for the current block being processed is below the thresholds, then the decision unit 38 decides that the corresponding audio corresponds to background noise. However, once the speech/noise decision unit 38 determines that there are a number of consecutive blocks (e.g. five consecutive blocks) whose residual energy values exceed the coarse threshold, then the decision unit 38 determines that the corresponding audio is speech."; Paragraph 0046, lines 1-12, "In the above embodiment, the speech/noise decision unit used two threshold values in determining whether or not the incoming audio was speech or noise. As those skilled in the art will appreciate, other decision strategies may be used. For example, the decision unit may decide that the input audio corresponds to speech as soon as a predetermined threshold value has been exceeded, however, such an embodiment is not preferred because it is susceptible to false detection of speech due to spurious short sounds or noises. Similarly, when detecting the end of speech, both the fine threshold and the coarse threshold could be used rather than just the fine threshold."; A fine threshold value reads on a first energy threshold, a coarse threshold value reads on a second energy threshold, determining that audio corresponds to background noise when the residual energy value for the block being processed is below the thresholds reads on the sound category "silence", determining the start point of speech with the fine threshold value reads on the sound category "speech", and indicating the start of the signal which is clearly not background noise with the coarse threshold value reads on the sound category "neutral".). Rajan is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Maas to incorporate the teachings of Rajan to determine that audio corresponds to background noise when the residual energy value for the block being processed is below a coarse threshold value, indicating the start of the signal which is clearly not background noise with the coarse threshold value, and determine the start point of speech with the fine threshold value. Doing so would allow for detecting the boundary between speech and noise (Rajan; Paragraph 0006, lines 1-14). Regarding claim 5, Liu in view of Maas, and further in view of Rajan, discloses the speech recognition method as claimed in claim 4. Rajan further teaches: further comprising determining the first energy threshold and the second energy threshold based on a second energy of background sound of the first audio data (Paragraph 0036, lines 1-15, "In this embodiment, one second of background noise is used in the training algorithm which, with the 16 kHz sampling rate, means that approximately 16,000 background noise samples are processed in the maximum likelihood analysis unit 31. Further, in this embodiment, the block energy determining unit 35 divides the residual error values determined for these samples into non-overlapping blocks of approximately eighty samples. Therefore, the block energy determining unit determines approximately 200 energy values for the training background noise. During the training routine, the energy values determined by the block energy determining unit 35 are passed via the switch 36 to a histogram analysis unit 37 which analyses the energy values to determine appropriate threshold values for use in detecting speech."; Determining appropriate threshold values for use in detecting speech based on the energy values of background noise samples reads on determining the first energy threshold and the second energy threshold based on a second energy of background sound of the first audio data.). Rajan is considered to be analogous to the claimed invention because it is in the same field of speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Liu in view of Maas, and further in view of Rajan, to further incorporate the teachings of Rajan to determine appropriate threshold values for use in detecting speech based on the energy values of background noise samples. Doing so would allow for detecting the boundary between speech and noise (Rajan; Paragraph 0006, lines 1-14). Regarding claim 12, arguments analogous to claim 4 are applicable. Regarding claim 13, arguments analogous to claim 5 are applicable. Regarding claim 19, arguments analogous to claim 4 are applicable. Regarding claim 20, arguments analogous to claim 5 are applicable. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Hwang et al. (Hwang, Inyoung, and Joon-Hyuk Chang, "End-to-End Speech Endpoint Detection Utilizing Acoustic and Language Modeling Knowledge for Online Low-Latency Speech Recognition", September 2020, IEEE Access, Vol. 8, pp. 161109-161123.) teaches a method for speech endpoint detection using a language model (LM) based end-of-utterance (EOU) predictor that incorporates phonetic embedding (PE) based acoustic model knowledge. Masumura et al. (Masumura, Ryo, Taichi Asami, Hirokazu Masataki, Ryo Ishii, and Ryuichiro Higashinaka, "Online End-of-Turn Detection from Speech Based on Stacked Time-Asynchronous Sequential Networks", August 2017, Interspeech 2017, pp. 1661-1665.) teaches a method for online end-of-turn detection using long-range sequential information of multiple time-asynchronous sequential features, such as prosodic, phonetic, and lexical sequential features, to enhance online end-of-turn detection performance. Chang et al. (Chang, Shuo-Yiin, Bo Li, Tara N. Sainath, Gabor Simko, and Carolina Parada, "Endpoint Detection using Grid Long Short-Term Memory Networks for Streaming Speech Recognition", August 2017, Interspeech 2017, pp. 3812-3816.) teaches a method for endpoint detection using a grid long short-term memory deep neural network endpointer model that models both spectral and temporal variations through recurrent connections. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Boggs whose telephone number is (571)272-2968. The examiner can normally be reached M-F 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAMES BOGGS/Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

May 24, 2024
Application Filed
Jul 10, 2024
Response after Non-Final Action
Mar 03, 2026
Non-Final Rejection mailed — §103
Jun 01, 2026
Response Filed
Jul 02, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694880
METHOD AND DEVICE FOR AUDIO BAND-WIDTH DETECTION AND AUDIO BAND-WIDTH SWITCHING IN AN AUDIO CODEC
3y 3m to grant Granted Jul 28, 2026
Patent 12682911
AUTOMATIC DETECTION AND ATTENUATION OF SPEECH-ARTICULATION NOISE EVENTS
3y 5m to grant Granted Jul 14, 2026
Patent 12682181
MULTIMODAL DIALOGS USING LARGE LANGUAGE MODEL(S) AND VISUAL LANGUAGE MODEL(S)
3y 0m to grant Granted Jul 14, 2026
Patent 12670922
AUDIO PROCESSING METHOD AND APPARATUS
2y 3m to grant Granted Jun 30, 2026
Patent 12651112
INTELLIGENTLY IDENTIFYING FRESHNESS OF TERMS IN DOCUMENTATION
3y 9m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
63%
Grant Probability
97%
With Interview (+34.0%)
3y 2m (~11m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 119 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month