Prosecution Insights
Last updated: October 02, 2026
Application No. 18/992,038

SPEECH INTERACTION METHOD AND RELATED ELECTRONIC DEVICE

Non-Final OA §101§102§103§112§Other
Filed
Jan 07, 2025
Priority
Nov 04, 2022 — CN 202211376580.5 +1 more
Examiner
SMITH, SEAN THOMAS
Art Unit
Tech Center
Assignee
Honor Device Co., Ltd.
OA Round
1 (Non-Final)
72%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
13 granted / 18 resolved
+12.2% vs TC avg
Strong +28% interview lift
Without
With
+27.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
23 currently pending
Career history
48
Total Applications
across all art units

Statute-Specific Performance

§101
25.3%
-14.7% vs TC avg
§103
53.1%
+13.1% vs TC avg
§102
13.0%
-27.0% vs TC avg
§112
7.2%
-32.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 18 resolved cases

Office Action

§101 §102 §103 §112 §Other
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Accordingly, the present application is afforded the benefit of the earlier filing date of CN-202211376580.5, filed November 4th, 2022. Information Disclosure Statement The information disclosure statement (IDS) submitted on March 14th, 2025, is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections Claim 15 is objected to because of the following informalities: The terms "the user's a first speech signal" and "the user's a second speech signal" make it unclear whether each speech signal belongs to a particular user or are meant as indefinite speech signals. Appropriate correction is required. Claim 26 is objected to for reciting a function “ K = f m × W 1 + L m × W 2 + R m × W 3 ” but only defining variables K and R m . Claim 27 is objected to for reciting two functions, “ W 1 = 1 / a b s ( f m - f k ) ∑ k = 1 Q 1 / a b s ( f m - f k ) ” and “ W 2 = 1 / a b s ( L m - L k ) ∑ k = 1 Q 1 / a b s ( L m - L k ) ” but not defining variables f k or L k . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 15-22 and 24-2 rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. Claim 15 recites the limitation “the electronic device acquires the user's a first speech signal.” There is insufficient antecedent basis for “the user” in the claim. Claim 18 recites the limitation “the second value is less than the first threshold.” There is insufficient antecedent basis for “the first threshold” in the claim. Claim 21 recites the limitation “a difference function between the voiceprint feature information of the first speech signal and the voiceprint feature information of the pre-set voice is less than the second threshold.” There is insufficient antecedent basis for “the voiceprint feature information” or “the second threshold” in the claim. Claims 16-17, 19-20 and 22 depend from claim 15, and therefore are inherently rejected under 35 U.S.C. 112(b). Claim 24 recites the limitation “acquiring acceleration data of the electronic device based on the acceleration sensor.” There is insufficient antecedent basis for “the acceleration sensor” in the claim. Claims 25-26 depend from claim 24, and therefore are inherently rejected under 35 U.S.C. 112(b). The term “this time” in claim 27 is a relative term which renders the claim indefinite. The term “this time” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 15-28 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite a mental process that can be performed in the human mind or with the aid of pen and paper. This judicial exception is not integrated into a practical application because a computer is invoked merely as a tool to execute an abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because an abstract idea is merely applied on a generic computer without any element that would otherwise preclude performance of the abstrac. Regarding claim 15, the claim recites “A speech interaction method, applied to an electronic device, wherein the electronic device comprises a first microphone and a second microphone, and the method comprises:in a first time period, the electronic device acquires the user's a first speech signal based on the first microphone and the second microphone, the signal strength of the first speech signal acquired by the first microphone is a first signal strength value, and the signal strength of the first speech signal acquired by the second microphone is a second signal strength value, the difference between the first signal strength value and the second signal strength value is a first value, the first speech signal does not include wake-up words;the electronic device performs a first operation based on the first speech signal; in a second time period, the electronic device acquires the user's a second speech signal based on the first microphone and the second microphone, the signal strength of the second speech signal acquired by the first microphone is the third signal strength value, and the signal strength of the second speech signal acquired by the second microphone is the fourth signal strength value, the difference between the third signal strength value and the fourth signal strength value is the second value, the first value is greater than the second value, and the second speech signal does not include the wake-up words, the semantics of the first speech signal and the second speech signal are the same;the first time period is earlier than the second time period;the electronic device does not perform the first operation based on the second speech signal.” The limitations of acquire speech signals and performing or not performing an action in response cover mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole, these limitations describe acts which are equivalent to human mental work of listening and making decisions. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and no additional features in the claims would preclude them from being performed as such. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 16, the claim depends from claim 15, and thus recites the limitations of claim 15, wherein “the second value is less than the first value, the electronic device performs the first operation based on the first speech signal, and does not perform the first operation based on the second speech signal.” Taken individually, or as a whole with claim 15, these limitations describe acts which are equivalent to human mental work of making decisions or comparing values. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 17, the claim depends from claim 16, and thus recites the limitations of claims 15 and 16, wherein “the first value is greater than or equal to a first threshold ,the electronic device performs the first operation based on the first speech signal.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of making decisions or comparing values. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 18, the claim depends from claim 16, and thus recites the limitations of claims 15 and 16, wherein “the second value is less than the first threshold, the electronic device does not performs the first operation based on the second speech signal.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of making decisions or comparing values. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 19, the claim depends from claim 15, and thus recites the limitations of claim 15, wherein “in a fourth time period, the electronic device is in motion state, and the electronic device performs the first operation based on the first speech signal, wherein, based on the electronic device being in motion and the first speech signal, the electronic device performs the first operation;the fourth time period is earlier than a third time period.” Taken individually, or as a whole with claim 15, these limitations describe acts which are equivalent to human mental work of making decisions or comparing values. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 20, the claim depends from claim 19, and thus recites the limitations of claims 15 and 19, wherein “in the third time period, the electronic device is in a stationary state, and the electronic device acquire a third speech signal based on the first microphone and the second microphone, the signal strength of the third speech signal acquired by the first microphone is a fifth signal strength value, and the signal strength of the third speech signal acquired by the second microphone is a sixth signal strength value, the difference between the fifth signal strength value and the sixth signal strength value is a third value, the third value is greater than the first value, the semantics of the third speech signal are the same as those of the first speech signal, and the third speech signal does not include the wake-up words;the third time period is earlier than the first time period;the electronic device does not perform the first operation based on the third speech signal.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of making decisions or comparing values. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 21, the claim depends from claim 20, and thus recites the limitations of claims 15 and 19-20, wherein “a target user pre-set voice on the electronic device, and a difference function between the voiceprint feature information of the first speech signal and the voiceprint feature information of the pre-set voice is less than the second threshold, based on the first speech signal, perform the first operation;the difference function between the voiceprint feature information of the third speech signal and the voiceprint feature information of the preset speech is greater than the second threshold, and the first operation is not performed based on the third speech signal.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of making decisions or comparing values. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 22, the claim depends from claim 21, and thus recites the limitations of claims 15 and 19-21, wherein “display a setting interface, wherein the setting interface includes a wake-up free words component, and in response to clicking to activate the wake-up free words component, enabling the wake-up free words function of the electronic device.” Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of making decisions or comparing values, with only computer functions that are well-understood and commonplace to a person having ordinary skill in the art. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 23, the claim recites “A speech interaction method, applied to an electronic device, wherein the electronic device comprises a speech interaction application, and the method comprises:receiving a first speech signal;obtaining speech signal data based on the first speech signal when it is determined that speech detection is to be performed on the first speech signal;the speech signal data include mel frequency cepstral coefficients and signal strength differences;processing the speech signal data by using a speech detection model to obtain a first confidence level and speech data, wherein the first confidence level is used for representing a probability that the first speech signal is a speech instruction issued by a user to the electronic device.” The limitations of speech processing cover mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole, these limitations describe acts which are equivalent to human mental work of listening and evaluating. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the steps of the claimed invention can be performed mentally, and Mel frequency coefficients are a mathematical tool that is well-understood and commonplace for persons having skill in the art. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 24, the claim depends from claim 23, and thus recites the limitations of claim 23, “acquiring acceleration data of the electronic device based on the acceleration sensor, and obtaining pose information of the electronic device based on the acceleration data;processing the pose information by using a pose detection model to obtain a second confidence level and target pose information, wherein the second confidence level is used for representing a probability that the electronic device is in a hand-held raised state;processing the target pose information and the speech data by using an speech-pose detection fusion model to obtain a third confidence level, wherein the third confidence level is used for representing a probability that the electronic device is in a hand-held raised state and the first speech signal is a speech instruction sent by a user to the electronic device; anddetermining, based on the first confidence level, the second confidence level, and the third confidence level, whether to start the speech interaction application.” The limitation of obtaining and processing pose information cover mental activities which can be performed in the mind or with the aid of pen and paper, or otherwise describe technical steps that are well-understood and commonplace to persons having skill in the art. Taken individually, or as a whole with claim 23, these limitations describe acts which are equivalent to human mental work of listening and making decisions. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 25, the claim depends from claim 24, and thus recites the limitations of claims 23 and 24, “wherein the detection, the electronic device further comprises a voiceprint detection module, based on the first confidence level, the second confidence level, and the third confidence level, whether to start the speech interaction application specifically comprises:setting a first confidence identifier to 1 when the first confidence level is greater than or equal to a first confidence threshold;setting the first confidence identifier to 0 when the first confidence level is less than the first confidence threshold;setting a second confidence identifier to 1 when the second confidence level is greater than or equal to a second confidence threshold;setting the second confidence identifier to 0 when the second confidence level is less than the second confidence threshold;setting a third confidence identifier to 1 when the third confidence level is greater than or equal to a third confidence threshold;setting the third confidence identifier to 0 when the third confidence level is less than the third confidence threshold;performing an AND logical operation on the first confidence identifier, the second confidence identifier, and the third confidence identifier to obtain a determining result; anddetermining, based on the determining result, whether to start the speech interaction application; therein, skipping starting the speech interaction application when the determining result is 0; or when the determining result is 1, detecting whether the first speech signal is a voice of a target user by using the voiceprint detection module, wherein the target user is a user of the electronic device;starting the speech interaction application if the first speech signal is the speech of the target user; orskipping starting the speech interaction application if the first speech signal is not the speech of the target user. ” These limitation as drafted covers logical operations and mental activities which can be performed in the mind or with the aid of pen and paper. Taken individually, or as a whole with the preceding claims, these limitations describe acts which are equivalent to human mental work of comparing values and making decisions. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 26, the claim depends from claim 25, and thus recites the limitations of claims 23-25, “wherein the determining, based on the first confidence level, the second confidence level, and the third confidence level, whether to start the speech interaction application specifically comprises: calculating a first weight value of the first confidence level, a second weight value of the second confidence level, and a third weight value of the third confidence level; performing calculation, based on the first confidence level, the first weight value, the second confidence level, the second weight value, the third confidence level, and the third weight value, to obtain a fused confidence level; calculating the fused confidence level according to a formula K = f m × W 1 + L m × W 2 + R m × W 3 , wherein K is the fused confidence level, and Rm is the third confidence level, and determining, based on the fused confidence level, whether to start the speech interaction application.” These limitations as drafted covers mental activities which can be performed in the mind or with the aid of pen and paper such as decision making in the form of mathematical formulae. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 27, the claim depends from claim 26, and thus recites the limitations of claims 23-26, “wherein the calculating a first weight value of the first confidence level, a second weight value of the second confidence level, and a third weight value of the third confidence level specifically comprises: calculating the first weight value according to a formula W 1 = 1 / a b s ( f m - f k ) ∑ k = 1 Q 1 / a b s ( f m - f k ) , wherein W1 is the first weight value, abs is an absolute value function, fm is a first confidence level output by the speech detection model this time, and k is a number of first Q first confidence levels closest to the first confidence level output this time; calculating the second weight value according to a formula W 2 = 1 / a b s ( L m - L k ) ∑ k = 1 Q 1 / a b s ( L m - L k ) wherein W2 is the second weight value, Lm is a second confidence level output by the pose detection model this time, and k is a number of first Q second confidence levels closest to the second confidence level output this time; and calculating the third weight value according to a formula W 3 = 1 - W 1 - W 2 , wherein W3 is the third weight value.” These limitations as drafted covers mathematical calculations which can be performed in the mind or with the aid of pen and paper. Accordingly, the claim is directed to an abstract idea without significantly more. The claim is not patent eligible. Regarding claim 28, the claim recites “An electronic device, comprising:a memory, a processor, and a touch screen, wherein the touch screen is configured to display content;the memory is configured to store a computer program, and the computer program comprises program instructions;the microphone is configured to collect speech signals, noise reduction, recognizing speech sources, and directional recording;the acceleration sensor is configured to detect the magnitude and direction of gravity, recognize the posture of electronic devices, switch between horizontal and vertical screens, or for applications such as pedometers; andthe processor is configured to invoke the program instructions to enable the electronic device to perform the method according to claim 23.” These limitations describe an electronic device with features that are well-understood and commonplace to persons having skill in the art. The performance of claim 23 describes acts which are equivalent to human mental work of listening and decision making, and thus claim 28 is directed to a mental process without significantly more. The claim is not patent eligible. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 23-24 and 28 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by U.S. Patent 10,515,623 to Grizzel (hereinafter, “Grizzel”). Regarding claim 23, Grizzel teaches a speech interaction method, applied to an electronic device, wherein the electronic device comprises a speech interaction application, and the method comprises: receiving a first speech signal (column 6, line 31, " An audio capture component, such as the microphone 103 of the voice input device 110 (or other device), captures input audio 11 corresponding to a spoken utterance."); obtaining speech signal data based on the first speech signal when it is determined that speech detection is to be performed on the first speech signal (column 6, line 34, "The device 110, using a wake command detection component 220, then processes audio data corresponding to the input audio 11 to determine if a keyword (such as a wakeword) is detected in the audio data."); the speech signal data include mel frequency cepstral coefficients and signal strength differences (column 6, line 55, "The voice input device 110 may use various techniques to determine whether audio data includes speech. Some embodiments may apply voice activity detection (VAD) techniques. Such techniques may determine whether speech is present in input audio based on various quantitative aspects of the input audio, such as a spectral slope between one or more frames of the input audio; energy levels of the input audio in one or more spectral bands; signal-to-noise ratios of the input audio in one or more spectral bands; or other quantitative aspects."); processing the speech signal data by using a speech detection model to obtain a first confidence level and speech data, wherein the first confidence level is used for representing a probability that the first speech signal is a speech instruction issued by a user to the electronic device (column 8, line 22, "The different ways a spoken utterance may be interpreted (i.e., the different hypotheses) may each be assigned a probability or a confidence score representing a likelihood that a particular set of words matches those spoken in the spoken utterance. The confidence score may be based on a number of factors including, for example, a similarity of the sound in the spoken utterance to models for language sounds (e.g., an acoustic model 253 stored in the ASR model storage 252), and a likelihood that a particular word that matches the sound would be included in the sentence at the specific location (e.g., using a language model 254 stored in the ASR model storage 252). Thus, each potential textual interpretation of the spoken utterance (i.e., hypothesis) is associated with a confidence score. Based on the considered factors and the assigned confidence score, the ASR component 250 outputs the most likely text recognized in the audio data 111."). Regarding claim 24, Grizzel further teaches a speech interaction method to include acquiring acceleration data of the electronic device based on the acceleration sensor, and obtaining pose information of the electronic device based on the acceleration data (column 31, line 38, "As shown in FIG. 13, the device may receive sensor data 302 corresponding to device movement."); processing the pose information by using a pose detection model to obtain a second confidence level and target pose information, wherein the second confidence level is used for representing a probability that the electronic device is in a hand-held raised state (column 31, line 38, "As shown in FIG. 13, the device may receive sensor data 302 corresponding to device movement. The device 110 may then process the sensor data 302 using a wake command detection component 220, gesture detection component 620, or the like, to determine (1302) a first confidence that a wake gesture was detected."); processing the target pose information and the speech data by using an speech-pose detection fusion model to obtain a third confidence level, wherein the third confidence level is used for representing a probability that the electronic device is in a hand-held raised state and the first speech signal is a speech instruction sent by a user to the electronic device (column 31, line 43, "The device 110 may also receive (1304) input audio and determine input audio data corresponding to the input audio. The device 110 may then process the input audio data using a wake command detection component 220 or the like to determine (1306) a second confidence that a wakeword is represented in the input audio data."); and determining, based on the first confidence level, the second confidence level, and the third confidence level, whether to start the speech interaction application (column 31, line 49, "Using both the first confidence and the second confidence the device 110 may then determine (1308) if a wake command was detected, and if so, send audio data to the server(s) 120. To determine if a wake command was detected the device 110 may weight the first confidence by a first weight and the second confidence by a second weight where the weights are determined based on operating conditions."). Regarding claim 28, Grizzel teaches an electronic device, comprising: a memory, a processor, and a touch screen, wherein the touch screen is configured to display content (column 35, line 47, "Each of these devices (110/120) may include one or more controllers/processors (1904/2004), that may each include a central processing unit (CPU) for processing data and computer-readable instructions, and a memory (1906/2006) for storing data and instructions of the respective device," and column 36, line 18, "Referring to the device 110 of FIG. 19, the device 110 may include a display, which may comprise a touch interface configured to receive limited touch inputs."); the memory is configured to store a computer program, and the computer program comprises program instructions (column 35, line 47, " Each of these devices (110/120) may include one or more controllers/processors (1904/2004), that may each include a central processing unit (CPU) for processing data and computer-readable instructions, and a memory (1906/2006) for storing data and instructions of the respective device."); the microphone is configured to collect speech signals, noise reduction, recognizing speech sources, and directional recording (column 36, line 31, "The device 110 may also include an audio capture component. The audio capture component may be, for example, a microphone 103 or array of microphones included in a headset or wireless headset. The microphone 103 may be configured to capture audio. If an array of microphones is included, approximate distance to a sound's point of origin may be determined by acoustic localization based on time and amplitude differences between sounds captured by different microphones of the array."); the acceleration sensor is configured to detect the magnitude and direction of gravity, recognize the posture of electronic devices, switch between horizontal and vertical screens, or for applications such as pedometers (column 36, line 59, "The device 110 may include one or more motion sensors 630. As discussed above, the device 110 may include one or more motion sensors 630. The sensors 630 may be any appropriate motion sensor(s) capable of providing information about rotations and/or translations of the device, and may include electronic accelerometer(s) that may measure linear acceleration about three dimensions (such as, x-, y-, and z-axis), electronic gyroscope(s) that may measure rotational acceleration about three dimensions (e.g., roll, pitch, and yaw), inertial sensor(s), barometer(s), gravity sensor(s), electronic compass(es), inclinometer(s), magnetometer(s), proximity sensor(s), distance sensor(s), depth sensor(s), range finder(s), ultrasonic transceiver(s), global position system (GPS) or other location determining sensor(s) and/or the like. The device can be configured to monitor for a change in position and/or orientation of the device using these motion sensor(s) 630."); and the processor is configured to invoke the program instructions to enable the electronic device to perform the method according to claim 23 (column 35, line 47, " Each of these devices (110/120) may include one or more controllers/processors (1904/2004), that may each include a central processing unit (CPU) for processing data and computer-readable instructions, and a memory (1906/2006) for storing data and instructions of the respective device."). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 15-21 are rejected under 35 U.S.C. 103 as being unpatentable over Grizzel in view of U.S. Patent 11,393,491 to Han et al. (hereinafter, "Han"). Regarding claim 15, Grizzel teaches a speech interaction method, applied to an electronic device, wherein the electronic device comprises a first microphone and a second microphone (column 36, line 31, "The device 110 may also include an audio capture component. The audio capture component may be, for example, a microphone 103 or array of microphones included in a headset or wireless headset. The microphone 103 may be configured to capture audio. If an array of microphones is included, approximate distance to a sound's point of origin may be determined by acoustic localization based on time and amplitude differences between sounds captured by different microphones of the array."), and the method comprises: in a first time period, the electronic device acquires the user's a first speech signal based on the first microphone and the second microphone, the signal strength of the first speech signal acquired by the first microphone is a first signal strength value, and the signal strength of the first speech signal acquired by the second microphone is a second signal strength value, the difference between the first signal strength value and the second signal strength value is a first value, the first speech signal does not include wake-up words; the electronic device performs a first operation based on the first speech signal (column 28, line 9, "The method may include receiving (1002) first audio data from a microphone connected to the wearable device. A signal quality metric, such as signal-to-noise ratio (SNR) of the first audio data may be determined (1004) for comparison to a threshold. If the signal quality metric is determined to be at or above a threshold, the first audio data may be sent (1006) to a remote device or further component of the wearable device for processing."). Grizzel does not explicitly teach acquiring a user’s speech signal at a second time period, however Han teaches a method for speech interaction wherein in a second time period, the electronic device acquires the user's a second speech signal based on the first microphone and the second microphone, the signal strength of the second speech signal acquired by the first microphone is the third signal strength value, and the signal strength of the second speech signal acquired by the second microphone is the fourth signal strength value, the difference between the third signal strength value and the fourth signal strength value is the second value, the first value is greater than the second value, and the second speech signal does not include the wake-up words, the semantics of the first speech signal and the second speech signal are the same (column 25, line 35, "Meanwhile, the microphone 122 of the artificial intelligence device 100-1 receives a second operation command (S1815). The second operation command may be a command continuously uttered after the user has uttered the first operation command," and column 25, line 66, "For example, in the case where the speech quality level is a volume, the processor 180 may determine that the speech quality level is changed when a difference between the first volume of the first operation command and the second volume of the second operation command is equal to or greater than a predetermined volume range," and column 26, line 19, "The processor 180 of the artificial intelligence device 100-1 transmits a second control command for performing operation suiting the second intention to the determined second external artificial intelligence device 100-3 through the short-range communication module 114 (S1823)." See also Grizzel column 24, line 59, "The system may then use the first confidence and the second confidence to determine if a wake command confidence threshold is satisfied. The individual confidences may be weighted and/or combined in various ways depending on system configuration and operating conditions."); and the first time period is earlier than the second time period (column 25, line 41, "The second operation command may be a command received one second after the first operation command is received. Here, one second is merely an example."). Grizzel also teaches a speech interaction method wherein the electronic device does not perform the first operation based on the second speech signal (column 30, line 44, "If a wakeword is not detected (1104:No) the device 110 may determine (1106) a signal quality metric corresponding to audio data of the spoken audio. In an alternate embodiment the device 110 may send the audio data to the server 120 and the server may determine (1106) the signal quality metric. If the signal quality metric is not determined to be below a threshold (1108:No), the system may determine that the audio was of sufficient quality and a wakeword was not detected. Thus the system may continue to receive new audio and attempting to detect a wakeword."). Han also describes a method wherein a second device is called to execute the second command; therefore, the first device does not execute the second command based on the comparison of the first and second signal values. Grizzel and Han are considered analogous because they are each concerned with speech-based device activation. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Grizzel with the teachings of Han for the purpose of improving activation response accuracy. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. The omission of an explicit wakeword is an obvious prompted variation on the combination, in view of common usage habits among consumers, and would be implemented to improve user experience with predictable results. Regarding claim 16, Han goes on to teach the method according to claim 15, and the method further comprises: the second value is less than the first value, the electronic device performs the first operation based on the first speech signal, and does not perform the first operation based on the second speech signal (column 25, line 35, "Meanwhile, the microphone 122 of the artificial intelligence device 100-1 receives a second operation command (S1815). The second operation command may be a command continuously uttered after the user has uttered the first operation command," and column 25, line 66, "For example, in the case where the speech quality level is a volume, the processor 180 may determine that the speech quality level is changed when a difference between the first volume of the first operation command and the second volume of the second operation command is equal to or greater than a predetermined volume range," and column 26, line 19, "The processor 180 of the artificial intelligence device 100-1 transmits a second control command for performing operation suiting the second intention to the determined second external artificial intelligence device 100-3 through the short-range communication module 114 (S1823)."). As taught in regards to claim 15, Han teaches a method wherein a first and second value are calculated, the first value is identified as greater than the second (logically identical to the second being less than the fist), and then the second command is sent away from the first device. Grizzel and Han are considered analogous because they are each concerned with speech-based device activation. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Grizzel with the teachings of Han for the purpose of improving activation response accuracy. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 17, Grizzel further teaches a method wherein the first value is greater than or equal to a first threshold, the electronic device performs the first operation based on the first speech signal (column 24, line 12, "Input audio data 111 may be processed to determine a first confidence that the audio data 111 includes a representation of a wakeword. The processor may determine that the first confidence is above or below a wakeword confidence threshold. If the first confidence is above the threshold the device 110 may determine that a wake command has been received."). Regarding claim 18, Grizzel further teaches a method wherein the second value is less than the first threshold, the electronic device does not performs the first operation based on the second speech signal (column 31, line 45, "The device 110 may then process the input audio data using a wake command detection component 220 or the like to determine (1306) a second confidence that a wakeword is represented in the input audio data. Using both the first confidence and the second confidence the device 110 may then determine (1308) if a wake command was detected, and if so, send audio data to the server(s) 120."). Regarding claim 19, Grizzel further teaches a method wherein in a fourth time period, the electronic device is in motion state, and the electronic device performs the first operation based on the first speech signal, wherein, based on the electronic device being in motion and the first speech signal, the electronic device performs the first operation (column 30, line 8, "The device may operate in a motion detection mode where the wearable device only responds to a wakeword when the wakeword is detected in conjunction with a particular movement detected by motion sensors in the wearable device."); the fourth time period is earlier than a third time period (column 21, line 39, "Although the figures and discussion illustrate certain operational steps of the system 100 in a particular order, the steps described may be performed in a different order (as well as certain steps removed or added) without departing from the intent of the disclosure."). Regarding claim 20, Han is referenced to teach a method wherein in the third time period, the electronic device is in a stationary state, and the electronic device acquire a third speech signal based on the first microphone and the second microphone, the signal strength of the third speech signal acquired by the first microphone is a fifth signal strength value, and the signal strength of the third speech signal acquired by the second microphone is a sixth signal strength value, the difference between the fifth signal strength value and the sixth signal strength value is a third value, the third value is greater than the first value, the semantics of the third speech signal are the same as those of the first speech signal, and the third speech signal does not include the wake-up words (column 4, line 9, "However, the artificial intelligence device 100 described in this specification is applicable to stationary artificial intelligence devices such as smart TVs, desktop computers or digital signages," and column 27, line 55, "Meanwhile, referring to FIG. 20 again, the AI hub device 1900-1 may receive the third operation command 2050 and acquire the third volume and third intention of the third operation command 2050," and column 27, line 64, "For example, assume that the second volume is 50, the third volume is 20, and a difference between the second volume and the third volume is 30 and is not within the predetermined volume range," and column 28, line 10, "In another example, when the difference between the third volume and the first volume is within the predetermined volume range, the AI hub device 1900-1 may recognize trigger for changing the device to be controlled to an existing device to be controlled."). Grizzel teaches that the electronic device does not perform the first operation based on the third speech signal (column 30, line 44, "If a wakeword is not detected (1104:No) the device 110 may determine (1106) a signal quality metric corresponding to audio data of the spoken audio. In an alternate embodiment the device 110 may send the audio data to the server 120 and the server may determine (1106) the signal quality metric. If the signal quality metric is not determined to be below a threshold (1108:No), the system may determine that the audio was of sufficient quality and a wakeword was not detected. Thus the system may continue to receive new audio and attempting to detect a wakeword.") and the third time period is earlier than the first time period (column 21, line 39, "Although the figures and discussion illustrate certain operational steps of the system 100 in a particular order, the steps described may be performed in a different order (as well as certain steps removed or added) without departing from the intent of the disclosure."). Save for the recitation of a stationary device, this is a restatement of the method of claim 15; the fifth and sixth signal strength values are analogous to the earlier recited third and fourth signal strength values, first and second values are compared, and a determination is made from that comparison. Additionally, Han teaches a method wherein a second device is called to perform an action based on the comparison of a first and second signal value, thus a first device does not perform the action. Grizzel and Han are considered analogous because they are each concerned with speech-based device activation. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Grizzel with the teachings of Han for the purpose of improving activation response accuracy. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Regarding claim 21, Grizzel teaches a speech interaction method wherein a target user pre-set voice on the electronic device, and a difference function between the voiceprint feature information of the first speech signal and the voiceprint feature information of the pre-set voice is less than the second threshold, based on the first speech signal, perform the first operation (column 7, line 16, "Specifically, keyword detection is typically performed without performing linguistic analysis, textual analysis, or semantic analysis. Instead, incoming audio (or audio data) is analyzed to determine if specific characteristics of the audio match preconfigured acoustic waveforms, audio signatures, or other data to determine if the incoming audio 'matches' stored audio data corresponding to a keyword."); the difference function between the voiceprint feature information of the third speech signal and the voiceprint feature information of the preset speech is greater than the second threshold, and the first operation is not performed based on the third speech signal (column 4, line 5, "In traditional speech-controlled systems the wake command may be a wakeword which is spoken to, and recognized by, a local device, which then captures the audio for an utterance and either processes it or forwards audio data of the utterance to another device for processing. The local device may continually listen for the wakeword and may disregard any audio detected that does not include the wakeword or is not preceded by the wakeword."). Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over Grizzel and Han as applied to claims 15-21 above, and further in view of U.S. Patent Application Publication 2020/0005789 to Chae (hereinafter, "Chae"). Regarding claim 22, the combination of Grizzel and Han does not teach a speech interaction method that includes “display a setting interface, wherein the setting interface includes a wake-up free words component, and in response to clicking to activate the wake-up free words component, enabling the wake-up free words function of the electronic device,” and thus, Chae is introduced. Chae teaches a method for speech interaction that comprises display a setting interface, wherein the setting interface includes a wake-up free words component, and in response to clicking to activate the wake-up free words component, enabling the wake-up free words function of the electronic device (paragraph [0086], "For example, the operator 120 may further include a text input unit (for example, a keyboard) to which an additional wake-up word text is to be inputted to set an additional wake-up word. When an additional wake-up word setting menu is displayed on the display 170 by the operation of a contact switch, the user may call the text input unit to set the additional wake-up word, and the controller 190 may display the text input unit on the display 170 in response to a call signal."). Grizzel, Han and Chae are considered analogous because they are each concerned with speech-based device activation. Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Grizzel and Han with the setting menu of Chae for the purpose of improving user experience. Given that all the claimed elements were known in the prior art, one skilled in the art could have combined the elements by known methods with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Allowable Subject Matter Claims 2 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form, overcoming all eligibility rejections and including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: China Patent Application CN 110033758 to Zhao teaches a method for voice-activated devices that uses combined probability calculations. China Patent Application CN 111048089 to Chen et al. teaches a method for voice-activated devices including the use of motion and acceleration measurements. China Patent Application CN 115223561 to Huang et al. teaches a method for voice-activated devices including the use of motion and acceleration measurements. China Patent CN 111933112 to Jin et al. teaches a method for improving voice activation accuracy based on multi-stage recognition. U.S. Patent Application Publication 2014/0358552 to Xu teaches method for improving voice activation in the presence of noise with multi-stage speech detection. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN T SMITH whose telephone number is (571)272-6643. The examiner can normally be reached Monday - Friday 8:00am - 5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SEAN THOMAS SMITH/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jan 07, 2025
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12750536
SYSTEMS AND METHODS FOR JUST IN TIME TRANSCODING OF VIDEO ON DEMAND
3y 8m to grant Granted Sep 29, 2026
Patent 12737549
METHOD AND APPARATUS FOR SENTIMENT ANALYSIS, ELECTRONIC DEVICE AND COMPUTER-READABLE STORAGE MEDIUM
2y 3m to grant Granted Sep 15, 2026
Patent 12688358
SYSTEMS AND METHODS FOR DYNAMICALLY PROVIDING A CORRECT PRONUNCIATION FOR A USER NAME BASED ON USER LOCATION
2y 6m to grant Granted Jul 21, 2026
Patent 12626056
GENERATING NATURAL LANGUAGE MODEL INSIGHTS FOR DATA CHARTS USING LIGHT LANGUAGE MODELS DISTILLED FROM LARGE LANGUAGE MODELS
2y 10m to grant Granted May 12, 2026
Patent 12602540
LEVERAGING A LARGE LANGUAGE MODEL ENCODER TO EVALUATE PREDICTIVE MODELS
2y 3m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
72%
Grant Probability
99%
With Interview (+27.5%)
2y 9m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 18 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month