Prosecution Insights
Last updated: October 02, 2026
Application No. 19/022,814

VOICE CONTROL METHOD AND APPARATUS FOR DEVICE, STORAGE MEDIUM, AND ELECTRONIC APPARATUS

Non-Final OA §101§103§DOUBLEPATENT
Filed
Jan 15, 2025
Priority
Mar 14, 2022 — CN 202210248231.9 +2 more
Examiner
KIM, JONATHAN C
Art Unit
Tech Center
Assignee
Dreame Innovation Technology(Suzhou) Co. Ltd.
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
271 granted / 368 resolved
+13.6% vs TC avg
Strong +39% interview lift
Without
With
+38.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
17 currently pending
Career history
392
Total Applications
across all art units

Statute-Specific Performance

§101
19.9%
-20.1% vs TC avg
§103
50.8%
+10.8% vs TC avg
§102
11.5%
-28.5% vs TC avg
§112
10.4%
-29.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 368 resolved cases

Office Action

§101 §103 §DOUBLEPATENT
DETAILED ACTION This Office Action is in response to the correspondence filed by the applicant on 1/15/2025. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file. Information Disclosure Statement The Information Statements (IDS) filed on 1/15/2025 and 4/13/2026 have been accepted and considered in this office action and are in compliance with the provisions of 37 CFR 1.97. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the claims at issue are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the reference application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO internet Web site contains terminal disclaimer forms which may be used. Please visit http://www.uspto.gov/forms/. The filing date of the application will determine what form should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to http://www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 1-18 are rejected on the ground of nonstatutory double patenting as being unpatentable over Claims 1-13 of US PAT 12,230,273. Although the claims, at issue are not identical, they are not patentably distinct from each other because the claims of the instant application are rejected as being unpatentable over the claims of the US PAT. Please see below for the mapping in the table, where the bolded limitations indicate the corresponding limitations between the US PAT and instant application. Instant application: 19/022,814 US PAT 12,230,273 1. A voice control method for a device, comprising: acquiring a first voice feature of first voice data collected by a cleaning device, wherein the first voice data are voice data corresponding to a first wake-up instruction sent by a use object, and the first wake-up instruction is configured to wake up at least one of the cleaning device and a base station; acquiring a second voice feature of second voice data collected by the base station, wherein the second voice data are voice data corresponding to the first wake-up instruction; and selecting a first device to be woken up from the cleaning device and the base station according to the first voice feature and the second voice feature, and waking up the first device, wherein the first device in a wake-up state is configured to respond to a voice instruction sent by the use object, wherein after waking up the first device, the method further comprises: performing voice collection through a voice collection part of the cleaning device to obtain third voice data; recognizing a second wake-up instruction from the third voice data, wherein the second wake-up instruction is configured to wake up at least one of the cleaning device and the base station; and waking up the cleaning device in a case that no voice data matching the third voice data are received from the base station within a first time period, or no voice feature matching a voice feature of the third voice data is received from the base station within the first time period; or after waking up the first device, the method further comprises: receiving fourth voice data from the base station; searching for voice data matching the fourth voice data collected by the cleaning device in a case that a third wake-up instruction is recognized from the fourth voice data, wherein the third wake-up instruction is configured to wake up at least one of the cleaning device and the base station; and sending first indication information to the base station in a case that no voice data matching the fourth voice data collected by the cleaning device are found, wherein the first indication information is configured to indicate determining to wake up the base station; or receiving a third voice feature from the base station, wherein the third voice feature is obtained by performing voice feature extraction on fifth voice data collected by the base station in a case that a fourth wake-up instruction is recognized from the fifth voice data, and the fourth wake-up instruction is configured to wake up at least one of the cleaning device and the base station; searching for a voice feature matching the third voice feature in voice features of voice data collected by the cleaning device; and sending second indication information to the base station in a case that no voice feature matching the third voice feature is found, wherein the second indication information is configured to indicate determining to wake up the base station. 1. A voice control method for a device, comprising: acquiring a first voice feature of first voice data collected by a cleaning device, wherein the first voice data are voice data corresponding to a first wake-up instruction sent by a use object, and the first wake-up instruction is configured to wake up at least one of the cleaning device and a base station; acquiring a second voice feature of second voice data collected by the base station, wherein the second voice data are voice data corresponding to the first wake-up instruction; and selecting a first device to be woken up from the cleaning device and the base station according to the first voice feature and the second voice feature, and waking up the first device, wherein the first device in a wake-up state is configured to respond to a voice instruction sent by the use object, wherein the first voice feature comprises a first voice intensity of the first voice data, the second voice feature comprises a second voice intensity of the second voice data, and the step of selecting the first device to be woken up from the cleaning device and the base station according to the first voice feature and the second voice feature comprises: determining the cleaning device as the first device to be woken up in a case that the first voice intensity is greater than the second voice intensity, and a difference between the first voice intensity and the second voice intensity is greater than or equal to a target intensity threshold; determining the base station as the first device to be woken up in a case that the first voice intensity is less than the second voice intensity, and the difference between the first voice intensity and the second voice intensity is greater than or equal to the target intensity threshold; and determining a preset device in the cleaning device and the base station as the first device to be woken up in a case that the difference between the first voice intensity and the second voice intensity is less than the target intensity threshold. 4. The method according to claim 1, wherein after waking up the first device, the method further comprises: performing voice collection through a voice collection part of the cleaning device to obtain third voice data; recognizing a second wake-up instruction from the third voice data, wherein the second wake-up instruction is configured to wake up at least one of the cleaning device and the base station; and waking up the cleaning device in a case that no voice data matching the third voice data are received from the base station within a first time period, or no voice feature matching a voice feature of the third voice data is received from the base station within the first time period. 5. The method according to claim 1, wherein after waking up the first device, the method further comprises: receiving fourth voice data from the base station; searching for voice data matching the fourth voice data collected by the cleaning device in a case that a third wake-up instruction is recognized from the fourth voice data, wherein the third wake-up instruction is configured to wake up at least one of the cleaning device and the base station; and sending first indication information to the base station in a case that no voice data matching the fourth voice data collected by the cleaning device are found, wherein the first indication information is configured to indicate determining to wake up the base station; or receiving a third voice feature from the base station, wherein the third voice feature is obtained by performing voice feature extraction on fifth voice data collected by the base station in a case that a fourth wake-up instruction is recognized from the fifth voice data, and the fourth wake-up instruction is configured to wake up at least one of the cleaning device and the base station; searching for a voice feature matching the third voice feature in voice features of voice data collected by the cleaning device; and sending second indication information to the base station in a case that no voice feature matching the third voice feature is found, wherein the second indication information is configured to indicate determining to wake up the base station. Other independent claim 10 is also similar to the independent claims 1 of the US PAT. With respect to the dependent claims, each of the claims maps to a corresponding dependent claim of the US PAT or are found within the scope of the independent claim. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 8 and 17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. Claims 8 and 17 recite, “A computer-readable storage medium ...." In [0218], the specification describes the computer-readable medium, which is not specifically limited to non-transitory propagating signals. Examiner suggests adding the limitation "non-transitory" to the claim to avoid a rejection under 35 U.S.C. 101. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6, 8-15, and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over PASKO (US 2019/0311720 A1), and in further view of CHEUVRONT (US 2017/0361468 A1). REGARDING CLAIM 1, PASKO discloses a voice control method for a device, comprising: acquiring a first voice feature of first voice data (Par 34 – “An example is “energy data”, which may correspond to a signal strength value associated with a detected utterance. Such a signal strength value may be expressed as a signal to noise ratio (SNR) to indicate a comparison of the signal level (e.g., the power of the utterance) to the background noise level (e.g., the power of background noise in the environment). This signal strength value can be expressed in units of decibel (dB)”; Par 38 – “As sound travels through space, the sound gradually dissipates, decreasing in (perceived) energy at a rate (approximately) proportional to the square of the distance traveled (e.g., the sound pressure/amplitude decreases proportionally to the distance). Perceived volume decreases by about 6 dB for every doubling of the distance from the sound source.”; Par 61 – “The first device 102(1) may then generate a first score for itself, and a second score for the second device 102(2). The first score assigned to the first device 102(1) may be based on the first device's 102(1) time-based data (e.g., a first wakeword occurrence time, WT1), and the first device's 102(1) energy data (e.g., a first signal strength value (e.g., SNR) associated with audio data generated by, or received at, the first device 102(1) based on the utterance). … In some embodiments, sub-scores can be computed for the time-based data and the energy data, those sub-scores can be translated to the same domain, and then the translated sub-scores can be used as a weighted average for a total “device score.””) collected by a [cleaning] smart home device (Par 21 – “… smart home devices like thermostats, lights, refrigerators, ovens, etc.) that may be controllable by speech interface devices, such as the devices 102.”), wherein the first voice data are voice data corresponding to a first wake-up instruction sent by a use object (Fig. 1A – “[Wakeword], What time is it? 114”; Par 25 – “Now, to describe the device arbitration technique shown in FIG. 1A, consider an example where the user 112 utters the wakeword followed by the expression “What time is it?” The first device 102(1) may, at 116, detect this utterance 114 at a first time.”), and the first wake-up instruction is configured to wake up at least one of the [cleaning] smart home device and a base station (Par 29 – “Having received the second notification 119(2) from the second device 102(2) within this first timeout period (e.g., MIND), the first device 102(1) may, at 120, perform time-based device arbitration to determine whether to designate the first device 102(1) or the second device 102(2) as a designated device to field the utterance 114.”; Par 21 – “… smart home devices like thermostats, lights, refrigerators, ovens, etc.) that may be controllable by speech interface devices, such as the devices 102. The devices 102 may be configured as a “hubs” in order to connect a plurality of devices in the environment and control communications among them, thereby serving as a place of convergence where data arrives from one or more devices, and from which data is sent to one or more devices.”); acquiring a second voice feature of second voice data collected by the base station (Par 61 – “The first device 102(1) may then generate a first score for itself, and a second score for the second device 102(2). The first score assigned to the first device 102(1) may be based on the first device's 102(1) time-based data (e.g., a first wakeword occurrence time, WT1), and the first device's 102(1) energy data (e.g., a first signal strength value (e.g., SNR) associated with audio data generated by, or received at, the first device 102(1) based on the utterance). … In some embodiments, sub-scores can be computed for the time-based data and the energy data, those sub-scores can be translated to the same domain, and then the translated sub-scores can be used as a weighted average for a total “device score.””; Par 21 – “… smart home devices like thermostats, lights, refrigerators, ovens, etc.) that may be controllable by speech interface devices, such as the devices 102. The devices 102 may be configured as a “hubs” in order to connect a plurality of devices in the environment and control communications among them, thereby serving as a place of convergence where data arrives from one or more devices, and from which data is sent to one or more devices.”), wherein the second voice data are voice data corresponding to the first wake-up instruction (Par 22 – “If a user 112 is positioned M meters from the first device 102(1) of FIG. 1A, and N meters from the second device 102(2), and assuming N is a different value than M, then the devices 102 will notice a wakeword in an utterance 114 spoken by the user 112 at different times, and the difference between the perceived wakeword time, T, can be calculated as: T=|M−N|×2.91, in units of milliseconds (ms).”); and selecting a first device to be woken up from [the cleaning] smart home device and the base station according to the first voice feature and the second voice feature (Par 29 – “For example, the first device 102(1) may, at 120, designate the device with the earliest wakeword occurrence time as the designated device to perform an action with respect to the user speech. Thus, if the first device 102(1) determines, at 120, that the first wakeword occurrence time, WT1, is earlier than (or precedes) the second wakeword occurrence time, WT2, the first device 102(1) may designate itself to perform an action 121 with respect to the user speech.”; Par 38 – “As sound travels through space, the sound gradually dissipates, decreasing in (perceived) energy at a rate (approximately) proportional to the square of the distance traveled (e.g., the sound pressure/amplitude decreases proportionally to the distance). Perceived volume decreases by about 6 dB for every doubling of the distance from the sound source. Assuming similar characteristics of speech interface devices that record sound to generate audio data, and assuming a lack of other objects in the sound path that would affect the measurement (e.g., a speech interface device may cause the locally perceived volume to be higher than in an open space), one may use the relative energy difference to determine the device that's closest to the source of the sound. The precision of such a measurement may be affected by self-noise of the microphone, and/or the microphone's response characteristics, and/or, other sound sources in the environment. In ideal conditions, energy-based measurement (energy data) can be more precise than a time-based approach (time-based data), and energy-based measurements are also not dependent on a common clock source.”; Par 61 – “The first device 102(1) may then generate a first score for itself, and a second score for the second device 102(2). The first score assigned to the first device 102(1) may be based on the first device's 102(1) time-based data (e.g., a first wakeword occurrence time, WT1), and the first device's 102(1) energy data (e.g., a first signal strength value (e.g., SNR) associated with audio data generated by, or received at, the first device 102(1) based on the utterance). … In some embodiments, sub-scores can be computed for the time-based data and the energy data, those sub-scores can be translated to the same domain, and then the translated sub-scores can be used as a weighted average for a total “device score.””), and waking up the first device, wherein the first device in a wake-up state is configured to respond to a voice instruction sent by the use object (Par 29 – “The action 121 may include continuing to capture the user speech corresponding to the utterance 114 via a microphone of the designated device. In other words, the device arbitration logic may determine a most appropriate device to “listen” for sound representing user speech in the environment. For instance, a duration of the utterance may be longer than the time it takes to perform device arbitration, and, as such, a designated device can be determined for continuing to “listen” to the utterance 114.”; Par 30 – “In other words, the device arbitration logic may determine a most appropriate device to “respond” to the utterance 114. In order to determine a responsive action 121 that is to be performed, a local speech processing component of the first device 102(1) may be used to process the first audio data generated by the first device 102(1) (e.g., by performing automatic speech recognition (ASR) on the first audio data, and by perform natural language understanding (NLU) on the ASR text, etc.) to generate directive data, which tells the first device 102(1) how to respond to the user speech. Accordingly, the action 121 performed by the designated device (which, in this example, is the first device 102(1)) may be based on locally-generated directive data that tells the first device 102(1) to, for instance, output an audible response with the current time (e.g., a text-to-speech (TTS) response saying “It's 12:30 PM”).”), wherein after waking up the first device, the method further comprises: performing voice collection through a voice collection part of the [cleaning] smart home device to obtain third voice data (Fig. 7 – “Microphone 710”; Par 118 – “In the example of FIG. 1A, FIG. 2, and FIG. 3, the user 112 is shown as uttering the expression “What time is it?” Whether this utterance is captured by the microphone(s) 710 of the device 102 or captured by another speech interface device in the environment, the audio data representing this user's speech is ultimately received by a speech interaction manager (SIM) 758 of a voice services component 760 executing on the device 102.”); recognizing a second wake-up instruction from the third voice data (Par 25 – “The first device 102(1) may also detect the wakeword in the utterance 114 at block 116. As mentioned, the wakeword indicates to the first device 102(1) that the first audio data it generated is to be processed using speech processing techniques to determine an intent of the user 112.”), wherein the second wake- up instruction is configured to wake up at least one of the cleaning device and the base station (Par 29 – “Having received the second notification 119(2) from the second device 102(2) within this first timeout period (e.g., MIND), the first device 102(1) may, at 120, perform time-based device arbitration to determine whether to designate the first device 102(1) or the second device 102(2) as a designated device to field the utterance 114.”); and waking up [cleaning] smart home in a case that no voice data matching the third voice data are received from the base station within a first time period, or no voice feature matching a voice feature of the third voice data is received from the base station within the first time period (Fig. 1A – “Detect Utterance 116 -> Wait (First Timeout) and Advertise 118 -> Arbitrate 120 -> ….”; Par 27 – “At 118 of the process 100, the first device 102(1) may wait a period of time—starting from a first time at which the utterance 114 was first detected at the first device 102(1)—for data (e.g., audio data or notification data) to arrive at the first device 102(1) from other speech interface devices in the environment. This period of time may be a first timeout period that is sometimes referred to herein as “MIND” to indicate that the first timeout period represents the minimum amount of time that the first device 102(1) is configured to wait for data from other devices to arrive at the first device 102(1) before the first device 102(1) continues to respond to the user speech.”; Par 59 – “For example, if the device 102 waits at block 218 for other notifications or audio data to arrive at the device 102, proceeds to perform device arbitration at block 220 after a lapse of the first timeout period, and then receives a notification and/or audio data from another speech interface device prior to a lapse of a threshold time period corresponding to a second timeout period (e.g., “MAXD”), the device 102 may infer that the received audio data and/or notification(s) corresponds to the same utterance 114, and may de-duplicate by deleting the audio data, and/or ignoring the notification so that two actions are not output based on a single utterance 114.”); or after waking up the first device, the method further comprises: receiving fourth voice data from the base station (Par 115 – “In some embodiments, the remote system 352 may be configured to receive audio data from the device 102, to recognize speech in the received audio data using the remote speech processing system 354, and to perform functions in response to the recognized speech… Furthermore, the remote system 352 may perform device arbitration to designate a speech interface device in an environment to perform an action with respect to user speech.”); searching for voice data matching the fourth voice data collected by the [cleaning] smart home device (Fig. 1A – “Detect Utterance 116 -> Wait (First Timeout) and Advertise 118 -> Arbitrate 120 -> ….”; Par 14 – “To perform time-based local device arbitration, the device, upon detecting a wakeword in an utterance, can wait a period of time for data to arrive at the device, which, if received, indicates to the device that another speech interface device in the environment detected an utterance.”; Par 27 – “At 118 of the process 100, the first device 102(1) may wait a period of time—starting from a first time at which the utterance 114 was first detected at the first device 102(1)—for data (e.g., audio data or notification data) to arrive at the first device 102(1) from other speech interface devices in the environment. This period of time may be a first timeout period that is sometimes referred to herein as “MIND” to indicate that the first timeout period represents the minimum amount of time that the first device 102(1) is configured to wait for data from other devices to arrive at the first device 102(1) before the first device 102(1) continues to respond to the user speech.”) in a case that a third wake-up instruction is recognized from the fourth voice data (Par 25 – “The first device 102(1) may also detect the wakeword in the utterance 114 at block 116. As mentioned, the wakeword indicates to the first device 102(1) that the first audio data it generated is to be processed using speech processing techniques to determine an intent of the user 112.”), wherein the third wake-up instruction is configured to wake up at least one of the [cleaning] smart home device and the base station (Par 29 – “Having received the second notification 119(2) from the second device 102(2) within this first timeout period (e.g., MIND), the first device 102(1) may, at 120, perform time-based device arbitration to determine whether to designate the first device 102(1) or the second device 102(2) as a designated device to field the utterance 114.”); and sending first indication information to the base station (Par 115 – “In some embodiments, the remote system 352 may be configured to receive audio data from the device 102, to recognize speech in the received audio data using the remote speech processing system 354, and to perform functions in response to the recognized speech. In some embodiments, these functions involve sending directives, from the remote system 352, to the device 102 to cause the device 102 to perform an action, such as output an audible response to the user speech via a speaker(s) (i.e., an output device(s) 712), and/or control second devices in the environment by sending a control command via the wireless unit 730 and/or the antenna 732. Furthermore, the remote system 352 may perform device arbitration to designate a speech interface device in an environment to perform an action with respect to user speech. Thus, under normal conditions, when the device 102 is able to communicate with the remote system 352 over a wide area network 356 (e.g., the Internet), some or all of the functions capable of being performed by the remote system 352 may be performed by designating a device to field the utterance, and sending a directive(s) over the wide area network 356 to the designated device (e.g., the device 102), which, in turn, may process the directive(s), or send the directive(s) to the designated device (if the device 102 is not designated by the remote system 352), for performing an action(s).”) in a case that no voice data matching the fourth voice data collected by the [cleaning] smart home device are found (Fig. 1A – “Detect Utterance 116 -> Wait (First Timeout) and Advertise 118 -> Arbitrate 120 -> ….”; Par 27 – “At 118 of the process 100, the first device 102(1) may wait a period of time—starting from a first time at which the utterance 114 was first detected at the first device 102(1)—for data (e.g., audio data or notification data) to arrive at the first device 102(1) from other speech interface devices in the environment. This period of time may be a first timeout period that is sometimes referred to herein as “MIND” to indicate that the first timeout period represents the minimum amount of time that the first device 102(1) is configured to wait for data from other devices to arrive at the first device 102(1) before the first device 102(1) continues to respond to the user speech.”; Par 59 – “For example, if the device 102 waits at block 218 for other notifications or audio data to arrive at the device 102, proceeds to perform device arbitration at block 220 after a lapse of the first timeout period, and then receives a notification and/or audio data from another speech interface device prior to a lapse of a threshold time period corresponding to a second timeout period (e.g., “MAXD”), the device 102 may infer that the received audio data and/or notification(s) corresponds to the same utterance 114, and may de-duplicate by deleting the audio data, and/or ignoring the notification so that two actions are not output based on a single utterance 114.”), wherein the first indication information is configured to indicate determining to wake up the base station (Par 115 – “In some embodiments, these functions involve sending directives, from the remote system 352, to the device 102 to cause the device 102 to perform an action, such as output an audible response to the user speech via a speaker(s) (i.e., an output device(s) 712), and/or control second devices in the environment by sending a control command via the wireless unit 730 and/or the antenna 732.”); or receiving a third voice feature from the base station, wherein the third voice feature is obtained by performing voice feature extraction on fifth voice data collected by the base station in a case that a fourth wake-up instruction is recognized from the fifth voice data, and the fourth wake-up instruction is configured to wake up at least one of the cleaning device and the base station; searching for a voice feature matching the third voice feature in voice features of voice data collected by the cleaning device; and sending second indication information to the base station in a case that no voice feature matching the third voice feature is found, wherein the second indication information is configured to indicate determining to wake up the base station. PASKO does not explicitly teach the [square-bracketed] limitation and teaches the underlined feature instead. In other words, PASKO teaches voice interactions with smart home devices and/or a base station (i.e., hub device), but does not explicitly teach the smart home device is a [cleaning] device. CHEUVRONT disclose the [square-bracketed] limitation. CHEUVRONT discloses a method/system for controlling multiple devices comprising: interacting with smart home devices and a base station (Par 101 – “In another example depicted in FIG. 1, the user 100 utters an audible request for information regarding the second mobile robot 301 (“Robot #2”) to be emitted by the audio media device 400 (“AMID”): “AMID. Where is robot #2?” The audio media device 400 receives the audible request and transmits a wireless signal indicative of the audible request to the remote computing system 200.”), wherein the smart home device is a [cleaning] device (Par 58 – “In some examples, the mobile robot is a vacuum cleaning robot. The controller can be configured to control movement of the vacuum cleaning robot to a user-specified room in an environment in response to receiving the wireless command signal, wherein the audible user command and the wireless command signal are indicative of the user-specified room.”; Par 191 – “For instance, if the user 100 provides (702) an utterance commanding a vacuum cleaning robot to clean the room 20D, the remote computing system 200 determines that the second mobile robot 301 is closer to the room 20D and consequently transmits the command signal to the second mobile robot 301 to cause the second mobile robot 301 to clean the room 20D.”; Par 251 – “If the user 100 gives a voice command, but the AMD device 400 had problems hearing, the AMD-equipped mobile robot 300 could reorient and/or move to a position closer to the user 100 or away from another source of noise to hear better and ask the user 100 to repeat the voice command.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of PASKO to include an autonomous cleaning device as a smart home device, as taught by CHEUVRONT. One of ordinary skill would have been motivated to include an autonomous cleaning device as a smart home device, in order to enable a user to control a variety of home devices with a variety of functions needed for effective home management. REGARDING CLAIM 2, PASKO in view of CHEUVRONT discloses the method according to claim 1, wherein the step of acquiring the first voice feature of the first voice data collected by the cleaning device comprises: performing voice collection through a voice collection part of the cleaning device to obtain the first voice data (Fig. 7 – “Microphone 710”; Par 118 – “In the example of FIG. 1A, FIG. 2, and FIG. 3, the user 112 is shown as uttering the expression “What time is it?” Whether this utterance is captured by the microphone(s) 710 of the device 102 or captured by another speech interface device in the environment, the audio data representing this user's speech is ultimately received by a speech interaction manager (SIM) 758 of a voice services component 760 executing on the device 102.”); and performing voice feature extraction on the first voice data to obtain the first voice feature (Par 61 – “The first device 102(1) may then generate a first score for itself, and a second score for the second device 102(2). The first score assigned to the first device 102(1) may be based on the first device's 102(1) time-based data (e.g., a first wakeword occurrence time, WT1), and the first device's 102(1) energy data (e.g., a first signal strength value (e.g., SNR) associated with audio data generated by, or received at, the first device 102(1) based on the utterance). … In some embodiments, sub-scores can be computed for the time-based data and the energy data, those sub-scores can be translated to the same domain, and then the translated sub-scores can be used as a weighted average for a total “device score.””) in a case that the first wake-up instruction is recognized from the first voice data (Fig. 1A – “Detect Utterance 116 -> Wait (first Timeout) and Advertise 118 -> Arbitrate 120 ”; Par 25 – “The first device 102(1) may also detect the wakeword in the utterance 114 at block 116. As mentioned, the wakeword indicates to the first device 102(1) that the first audio data it generated is to be processed using speech processing techniques to determine an intent of the user 112.”). REGARDING CLAIM 3, PASKO in view of CHEUVRONT discloses the method according to claim 1, wherein the step of acquiring the second voice feature of the second voice data collected by the base station comprises: receiving the second voice data matching the first voice data from the base station; and performing voice feature extraction on the second voice data to obtain the second voice feature; or receiving the second voice feature matching the first voice feature from the base station, wherein the second voice feature is obtained by performing voice feature extraction on the second voice data (Fig. 2 – “Audio Data 204 Wakeword Occurrence Time (W12)”; Par 61 – “In the example of FIG. 2, the device 102 may detect the utterance (a first speech recognition event), and may receive a second speech recognition event from the device 202 in the form of the audio data 204. This speech recognition event (e.g., the received audio data 204) may include a second wakeword occurrence time, WT2, (which constitutes time-based data), and may include an additional type(s) of data in the form of energy data, for example. This energy data may correspond to a second signal strength value (e.g., SNR) associated with audio data generated by the device 202 based on the utterance. The device 102 may then generate a first score for itself, and a second score for the device 202.”). REGARDING CLAIM 4, PASKO in view of CHEUVRONT discloses the method according to claim 1, wherein the step of acquiring the second voice feature of the second voice data collected by the base station comprises: performing voice collection through a voice collection part of the base station to obtain the second voice data (Fig. 7 – “Microphone 710”; Par 118 – “In the example of FIG. 1A, FIG. 2, and FIG. 3, the user 112 is shown as uttering the expression “What time is it?” Whether this utterance is captured by the microphone(s) 710 of the device 102 or captured by another speech interface device in the environment, the audio data representing this user's speech is ultimately received by a speech interaction manager (SIM) 758 of a voice services component 760 executing on the device 102.”); and performing voice feature extraction on the second voice data to obtain the second voice feature (Par 61 – “The first device 102(1) may then generate a first score for itself, and a second score for the second device 102(2). The first score assigned to the first device 102(1) may be based on the first device's 102(1) time-based data (e.g., a first wakeword occurrence time, WT1), and the first device's 102(1) energy data (e.g., a first signal strength value (e.g., SNR) associated with audio data generated by, or received at, the first device 102(1) based on the utterance). … In some embodiments, sub-scores can be computed for the time-based data and the energy data, those sub-scores can be translated to the same domain, and then the translated sub-scores can be used as a weighted average for a total “device score.””) in a case that the first wake-up instruction is recognized from the second voice data (Fig. 1A – “Detect Utterance 116 -> Wait (first Timeout) and Advertise 118 -> Arbitrate 120 ”; Par 25 – “The first device 102(1) may also detect the wakeword in the utterance 114 at block 116. As mentioned, the wakeword indicates to the first device 102(1) that the first audio data it generated is to be processed using speech processing techniques to determine an intent of the user 112.”). REGARDING CLAIM 5, PASKO in view of CHEUVRONT discloses the method according to claim 1, wherein the step of acquiring the first voice feature of the first voice data collected by the cleaning device comprises: receiving the first voice data matching the second voice data from the cleaning device; and performing voice feature extraction on the first voice data to obtain the first voice feature; or receiving the first voice feature matching the second voice feature from the cleaning device, wherein the first voice feature is obtained by performing voice feature extraction on the first voice data (Fig. 2 – “Audio Data 204 Wakeword Occurrence Time (W12)”; Par 61 – “In the example of FIG. 2, the device 102 may detect the utterance (a first speech recognition event), and may receive a second speech recognition event from the device 202 in the form of the audio data 204. This speech recognition event (e.g., the received audio data 204) may include a second wakeword occurrence time, WT2, (which constitutes time-based data), and may include an additional type(s) of data in the form of energy data, for example. This energy data may correspond to a second signal strength value (e.g., SNR) associated with audio data generated by the device 202 based on the utterance. The device 102 may then generate a first score for itself, and a second score for the device 202.”). REGARDING CLAIM 6, PASKO in view of CHEUVRONT discloses the method according to claim 1, wherein the step of waking up the first device comprises: sending fifth indication information to the first device, wherein the fifth indication information is configured to indicate determining to wake up the first device (Par 115 – “In some embodiments, the remote system 352 may be configured to receive audio data from the device 102, to recognize speech in the received audio data using the remote speech processing system 354, and to perform functions in response to the recognized speech. In some embodiments, these functions involve sending directives, from the remote system 352, to the device 102 to cause the device 102 to perform an action, such as output an audible response to the user speech via a speaker(s) (i.e., an output device(s) 712), and/or control second devices in the environment by sending a control command via the wireless unit 730 and/or the antenna 732. Furthermore, the remote system 352 may perform device arbitration to designate a speech interface device in an environment to perform an action with respect to user speech. Thus, under normal conditions, when the device 102 is able to communicate with the remote system 352 over a wide area network 356 (e.g., the Internet), some or all of the functions capable of being performed by the remote system 352 may be performed by designating a device to field the utterance, and sending a directive(s) over the wide area network 356 to the designated device (e.g., the device 102), which, in turn, may process the directive(s), or send the directive(s) to the designated device (if the device 102 is not designated by the remote system 352), for performing an action(s).”). REGARDING CLAIM 8, PASKO in view of CHEUVRONT discloses a computer-readable storage medium, wherein the computer-readable storage medium comprises a stored program, wherein when the stored program runs (Par 103 – “In the illustrated implementation, the device 102 includes one or more processors 702 and computer-readable media 704.”), the method according to claim 1 is executed; thus, the claim is rejected under the same rationale explained in the rejection of claim 1. REGARDING CLAIM 9, PASKO in view of CHEUVRONT discloses an electronic apparatus, comprising a memory and a processor, wherein the memory stores a computer program, and the processor (Par 103 – “In the illustrated implementation, the device 102 includes one or more processors 702 and computer-readable media 704.”) is configured to execute the method according to claim 1 through the computer program; thus, the claim is rejected under the same rationale explained in the rejection of claim 1. CLAIM 10 is similar to the method of claim 1; thus, it is rejected under the same rationale. CLAIM 11 is similar to the method of claim 2; thus, it is rejected under the same rationale. CLAIM 12 is similar to the method of claim 3; thus, it is rejected under the same rationale. CLAIM 13 is similar to the method of claim 4; thus, it is rejected under the same rationale. CLAIM 14 is similar to the method of claim 5; thus, it is rejected under the same rationale. CLAIM 15 is similar to the method of claim 6; thus, it is rejected under the same rationale. CLAIM 17 is similar to the method of claim 8; thus, it is rejected under the same rationale. CLAIM 18 is similar to the method of claim 8; thus, it is rejected under the same rationale. Claim 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over PASKO (US 2019/0311720 A1) in view of CHEUVRONT (US 2017/0361468 A1), and in further view of KHAN (US 2016/0155443 A1). REGARDING CLAIM 7, PASKO in view of CHEUVRONT discloses the method according to claim 1. PASKO in view of CHEUVRONT does not explicitly teach retargeting a device to perform a second operation based on a voice data received by the first device that is currently awake. KHAN disclose a method/system for arbitration for listening devices, wherein after waking up the first device, the method further comprises: receiving ninth voice data sent by the first device, wherein the ninth voice data are voice data collected by the first device in the wake-up state (KHAN Fig. 8 – “Full wake up, play prompt, await voice command 860”; Par 55 – “In the case where no device is specified (e.g., “Send email to Bob”), if the initial device chosen by the system is incorrect (e.g., a desktop machine), a corrective utterance (e.g., “No, on my laptop,” “can we do this on my laptop?” or the like) can explicitly transfer the task to the specified device, to which the context is transferred (e.g., the user can continue typing the email to Bob). Such an utterance can be treated as an explicit utterance for purposes of machine learning or the like.”; Par 325 – “If the device receives information that it is the right device at 840, it can proceed to a full wake up, play an audio prompt and await a voice command at 860. If not, it can standby for an incoming handoff at 850 (e.g., in case a handoff comes in).”); in a case that a target voice instruction sent by the use object is recognized from the ninth voice data, determining a second device controlled by the target voice instruction in the cleaning device and the base station (KHAN Par 326 – “At 870 it can be determined whether the command can be carried out, or if a handoff is warranted. At 880, if the command can be carried out, it is. Otherwise, an error process can be invoked.”; Note that CHEUVRONT already teaches controlling a cleaning device and/or a base station.); and controlling the second device to execute a target device operation corresponding to the target voice instruction (KHAN Par 327 – “If the command cannot be carried out by the processing device, it can handoff at 890.”; Par 55 – “In the case where no device is specified (e.g., “Send email to Bob”), if the initial device chosen by the system is incorrect (e.g., a desktop machine), a corrective utterance (e.g., “No, on my laptop,” “can we do this on my laptop?” or the like) can explicitly transfer the task to the specified device, to which the context is transferred (e.g., the user can continue typing the email to Bob). Such an utterance can be treated as an explicit utterance for purposes of machine learning or the like.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of PASKO in view of CHEUVRONT to include retargeting a device to perform a second operation based on a voice data received by the first device that is currently awake, as taught by KHAN. One of ordinary skill would have been motivated to include retargeting a device to perform a second operation based on a voice data received by the first device that is currently awake, in order to effectively transfer the task to a specific device (Par 55). CLAIM 16 is similar to the method of claim 7; thus, it is rejected under the same rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN C KIM whose telephone number is (571)272-3327. The examiner can normally be reached Monday to Friday 8:00 AM thru 4:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JONATHAN C KIM/Primary Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Jan 15, 2025
Application Filed
Aug 05, 2026
Non-Final Rejection mailed — §101, §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744034
ELECTRONIC DEVICE AND CONTROL METHOD THEREOF
2y 8m to grant Granted Sep 22, 2026
Patent 12744038
SYSTEMS AND TECHNIQUES FOR USING A DIGITAL ASSISTANT WITH AN ENHANCED ENDPOINTER
2y 6m to grant Granted Sep 22, 2026
Patent 12711982
AUDIO PROCESSING
3y 3m to grant Granted Aug 18, 2026
Patent 12700403
METHOD, APPARATUS, ELECTRONIC DEVICE AND STORAGE MEDIUM FOR TEXT CONTENT MATCHING
2y 5m to grant Granted Aug 04, 2026
Patent 12688850
ELECTRONIC DEVICE AND CONTROL METHOD THEREFOR
2y 5m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+38.7%)
2y 5m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 368 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month