Prosecution Insights
Last updated: August 06, 2026
Application No. 18/373,859

METHODS, SYSTEMS, AND APPARATUS FOR AUTOMATON NETWORKS HAVING MULTIPLE VOICE AGENTS FOR SPEECH RECOGNITION

Non-Final OA §103§112
Filed
Sep 27, 2023
Priority
Sep 27, 2022 — provisional 63/410,286
Examiner
WASHBURN, DANIEL C
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Snap One LLC
OA Round
3 (Non-Final)
49%
Grant Probability
Moderate
3-4
OA Rounds
1y 3m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 49% of resolved cases
49%
Career Allowance Rate
79 granted / 160 resolved
-12.6% vs TC avg
Strong +29% interview lift
Without
With
+28.7%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
7 currently pending
Career history
172
Total Applications
across all art units

Statute-Specific Performance

§101
12.3%
-27.7% vs TC avg
§103
52.9%
+12.9% vs TC avg
§102
15.5%
-24.5% vs TC avg
§112
11.3%
-28.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 160 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 1/13/26 has been entered. Response to Arguments Applicant’s arguments with respect to claim(s) 1-11 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-6 and 12 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 1 recites the limitation "any of the voice targets" in line 10. There is insufficient antecedent basis for this limitation in the claim. The examiner recommends an amendment along the lines of “any number of voice targets”. Claims 2-6 and 12 are rejected due to their dependency on claim 1, as they inherit and do not correct the insufficient antecedent basis. The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claims 4, 5, 10, and 11 are rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 4 describes, “the voice coordinator determines mapping information for mapping any of the voice input devices to any of the voice targets.” However, claim 1, at lines 9-11, describes, “wherein the voice coordinator determines mapping information for mapping any of the voice input devices to any of the voice targets according to any of device registration information and device polling information”. Thus, claim 4 fails to further limit claim 1. Claim 5 describes, “the mapping information is determined according to any of device registration information and device polling information.” However, claim 1, at lines 9-11, describes, “wherein the voice coordinator determines mapping information for mapping any of the voice input devices to any of the voice targets according to any of device registration information and device polling information”. Thus, claim 5 fails to further limit claims 1 and 4. Claim 10 describes, “the mapping information is determined by mapping any one or more of the voice input devices to any one or more of the voice targets.” However, claim 7, at lines 7-10, describes, “determine mapping information mapping any number of voice input devices to any number of voice targets, the mapping information is determined by mapping any one or more of the voice input devices to any one or more of the voice targets according to any of device registration information and device polling information”. Thus, claim 10 fails to further limit claim 7. Claim 11 describes, “the mapping information is determined according to any of device registration information and device polling information.” However, claim 7, at lines 7-10, describes, “determine mapping information mapping any number of voice input devices to any number of voice targets, the mapping information is determined by mapping any one or more of the voice input devices to any one or more of the voice targets according to any of device registration information and device polling information”. Thus, claim 11 fails to further limit claims 7 and 10. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-12 are rejected under 35 U.S.C. 103 as being unpatentable over Wilberding (US 10,115,400, hereinafter “Wilberding”) in view of Mo et al. (US 11,700,141, hereinafter “Mo”) in view of Lang et al. (US 10,499,146, hereinafter “Lang”) and further in view of Lang et al. (US 10,475,449, hereinafter Lang ‘449) RE claim 1, Wilberding describes a whole home voice (WHV) system of an automation network connecting an automation server (FIG. 5 and col. 12 lns. 11-23: “the computing devices 504, 506, and 508 may be part of a cloud network 502. The cloud network 502 may include additional computing devices. In one example, the computing devices 504, 506, and 508 may be different servers.”), an automation controller (FIG. 5, col. 11 ln. 65-col. 12 ln. 2 “controller device (CR) 522”), and any number of automation devices (FIG. 5 and col. 11 ln. 65-col. 12 ln. 2 “network microphone devices (NMDs) 512, 514, and 516; playback devices (PBDs) 532, 534, 536, and 538”), the WHV system comprising: any number of voice input devices for receiving voice inputs vocalized by a user of the automation network (col. 12 lns. 24-48 describes network microphone devices (NMDs). Further, col. 16 lns. 45-55 describes that NMDs receive voice data indicating a voice input); and an automation controller separate from the voice input devices (col. 8 lns 15-16: “In one example, the control device 300 may be a dedicated controller for the media playback system 100.” Also see Col. 8 lns. 60-63: “the control device 300 may sometimes be referred to as a controller, whether the control device 300 is a dedicated controller or a network device on which media playback system controller application software is installed.” Further, see col. 16 lns. 46-55: “At block 702, implementation 700 involves receiving voice data indicating a voice input. For instance, a NMD, such as NMD 600, may receive, via a microphone, voice data indicating a voice input. As further examples, any of playback devices 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, and 124 or control devices 126 and 128 of FIG. 1 may be a NMD and may receive voice data indicating a voice input. Yet further examples NMDs include NMDs 512, 514, and 516, PBDs 532, 534, 536, and 538, and CR 522 of FIG. 5.”) for executing: (1) a voice coordinator for coordinating WHV system operations (col. 8 lns. 8-22 “In one example, the control device 300 may be a dedicated controller for the media playback system 100. In another example, the control device 300 may be a network device on which media playback system controller application software may be installed, such as for example, an iPhone™, iPad™ or any other smart phone, tablet or network device (e.g., a networked computer such as a PC or Mac™)”), the voice coordinator configures the voice-enabled components for runtime use (col. 8 lns 52-59: “As suggested above, changes to configurations of the media playback system 100 may also be performed by a user using the control device 300. The configuration changes may include adding/removing one or more playback devices to/from a zone, adding/removing one or more zones to/from a zone group, forming a bonded or consolidated player, separating one or more playback devices from a bonded or consolidated player, among others.”), wherein the voice coordinator determines mapping information for mapping any of the voice input devices to any of the voice targets according to any of device registration information and device polling information (col. 25 lns. 10-24 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).”), and wherein the voice coordinator broadcasts the mapping information to devices communicatively connected to the automation controller via the automation network (col. 25 lns. 34-43: “At block 906, implementation 900 involves causing registration of at least one of the detected voice services to be registered on the one or more second devices. For instance, the NMD may cause at least one of the detected voice services to be registered with a media playback system that includes one or more playback devices (e.g., media playback system 100 of FIG. 1). Causing the a voice service to be registered may involve transmitting, via a network interface, a message indicating credentials for that voice service to the media playback system (i.e., at least one device thereof).”), (2) any number of voice input proxies for proxying for the voice input devices (col 18 lns. 3-51: “Available voice services may include voice services registered with the NMD. Registration of a given voice service with the NMD may involve providing user credentials (e.g., user name and password) of the voice service to the NMD and/or providing an identifier of the NMD to the voice service. Such registration may configure the NMD to receive voice inputs on behalf of the voice service and perhaps configure the voice service to accept voice inputs from the NMD for processing.”), each voice input proxy maintains many voice target configurations per single input device (col 24 ln. 65 – col. 25 ln. 8: “At block 902, implementation 900 involves receiving input data indicating a command to register one or more voice services on one or more second devices. For instance, a first device (e.g., a NMD) may receive, via a user interface (e.g., a touchscreen), input data indicating a command to register one or more voice services with a media playback system that includes one or more playback devices. In one example, the NMD receives the input as part of a procedure to set-up the media playback system using any of the example techniques described above in connection with block 702 of implementation 700”, Also see col. 2 lns. 41-47: “Where two or more voice services are configured for a NMD, a particular voice service can be invoked by utterance of a wake-word corresponding to the particular voice service. For instance, in querying AMAZON®, a user might speak the wake-word “Alexa” followed by a voice input. Other examples include “Ok, Google” for querying GOOGLE® and “Hey, Siri” for querying APPLE®.”), , wherein each voice input proxy supports a current voice target and a current room for the respective voice input device (col. 25 lns. 11-19: “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD). For instance, a NMD that is a smartphone or tablet may have installed one or more applications (“apps”) that interface with voice services. The NMD may detect these applications using any suitable technique.” Also see col. 25 lns. 46-51: “In such manner, a user's media playback system may have registered one or more of the same voice services as registered on the user's NMD (e.g., smartphone) utilizing the same credentials as the user's NMD, which may hasten registration.” Further, regarding each proxy supporting a current room of a voice input device, see col. 7 lns. 2-11: “Referring back to the media playback system 100 of FIG. 1, the environment may have one or more playback zones, each with one or more playback devices. The media playback system 100 may be established with one or more playback zones, after which one or more zones may be added, or removed to arrive at the example configuration shown in FIG. 1. Each zone may be given a name according to a different room or space such as an office, bathroom, master bedroom, bedroom, kitchen, dining room, living room, and/or balcony.” And col. 7 lns. 47-59: “the zone configurations of the media playback system 100 may be dynamically modified, and in some embodiments, the media playback system 100 supports numerous configurations. For instance, if a user physically moves one or more playback devices to or from a zone, the media playback system 100 may be reconfigured to accommodate the change(s). For instance, if the user physically moves the playback device 102 from the balcony zone to the office zone, the office zone may now include both the playback device 118 and the playback device 102. The playback device 102 may be paired or grouped with the office zone and/or renamed if so desired via a control device such as the control devices 126 and 128.”), and wherein each voice input proxy is instantiated for a respective one of the voice input devices (col. 25 lns. 10-24 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).” Also see col. 25 lns. 34-43: “At block 906, implementation 900 involves causing registration of at least one of the detected voice services to be registered on the one or more second devices. For instance, the NMD may cause at least one of the detected voice services to be registered with a media playback system that includes one or more playback devices (e.g., media playback system 100 of FIG. 1). Causing the a voice service to be registered may involve transmitting, via a network interface, a message indicating credentials for that voice service to the media playback system (i.e., at least one device thereof).”); and (3) a voice daemon executing any number of voice daemon instances each respectively associated with one of any number of voice targets (col. 19 lns. 5-34: “a voice service may provide an application programming interface that the NMD can invoke to determine that whether the voice data includes the wake-word or phrase corresponding to that voice service. The NMD may invoke the API by transmitting a particular query of the voice service to the voice service along with data representing the wake-word portion of the received voice data. Alternatively, the NMD may invoke the API on the NMD itself. Registration of a voice service with the NMD or with the media playback system may integrate the API or other architecture of the voice service with the NMD.”), wherein each voice daemon instance is mapped to an individual voice target (col. 19 lns. 20-26: “Where multiple voice services are available to the NMD, the NMD might query wake-word detection algorithms corresponding to each voice service of the multiple voice services. As noted above, querying such detection algorithms may involve invoking respective APIs of the multiple voice services, either locally on the NMD or remotely using a network interface.”), , and wherein the voice daemon routes the audio data to the voice targets (col. 22 lns. 31-41: “At block 706, implementation 700 involves causing the identified voice service(s) to process the voice input. For instance, the NMD may transmit, via a network interface to one or more servers of the identified voice service(s), data representing the voice input and a command or query to process the data presenting the voice input. The command or query may cause the identified voice service(s) to process the voice command. The command or query may vary according to the identified voice service so as to conform the command or query to the identified voice service (e.g., to an API of the voice service).”), wherein the automation controller is communicatively connected to the voice targets, the voice targets receive voice audio commands and perform operations in response to the received voice audio commands (col. 22 lns. 31-41: “At block 706, implementation 700 involves causing the identified voice service(s) to process the voice input. For instance, the NMD may transmit, via a network interface to one or more servers of the identified voice service(s), data representing the voice input and a command or query to process the data presenting the voice input. The command or query may cause the identified voice service(s) to process the voice command. The command or query may vary according to the identified voice service so as to conform the command or query to the identified voice service (e.g., to an API of the voice service).” Also see col. 22 lns. 54-64: “After causing the identified voice service to process the voice input, the NMD may receive results of the processing. For instance, if the voice input represented a search query, the NMD may receive search results. As another example, if the voice input represented a command to a device (e.g., a media playback command to a playback device), the NMD may receive the command and perhaps additional data associated with the command (e.g., a source of media associated with the command). The NMD may output these results as appropriate to the type of command and the received results.”). Wilberding doesn’t describe a system or method wherein the voice coordinator discovers voice-enabled components of the automation network system, wherein each voice input proxy manages the voice input device's microphone, wherein the voice daemon spawns and manages separate processes for various types of services, wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol, wherein the voice daemon transcodes audio data from the voice input devices, and wherein the voice daemon routes the transcoded audio data to the voice targets. However, Mo describes a system and method wherein the voice coordinator discovers voice-enabled components of the automation network system (col. 8 lns. 46-67: “Various implementations described herein additionally or alternatively relate to utilizing local assistant client devices in discovering, provisioning, and/or registering smart devices for an account of a user. In some implementations, a smart device that is not yet registered can be discovered by causing each assistant client device, of an ecosystem of assistant client devices, to scan one or more communications channels (e.g., Wi-Fi, Bluetooth, and/or other) for smart device(s) that aren't registered. […] The discovery can occur at periodic or non-periodic intervals, or in response to a user request (e.g., a voice request, a request initiated via a smart phone app for the assistant, etc.).”), [and] wherein the voice daemon spawns and manages separate processes for various types of services (col. 9 lns. 1-17: “Once discovered, a smart device can be provisioned by causing the smart device to pair with at least one of the assistant client devices (e.g., in the case of Bluetooth) and/or to connect to a secured Wi-Fi network to which the assistant client devices are already connected. After provisioning, data transmitted by the smart device can be received, at an assistant client device, and processed using a local adapter of the assistant client device, to generate registration data in a schema of the automated assistant. For example, the data transmitted by the smart device can be in a protocol suite for a corresponding 3P, and can be interpreted by the 3P adapter into a schema for registration with the automated assistant. The registration data can be utilized by the automated assistant to add the smart device to a home graph or other structure that defines smart devices, assistant client devices, and various properties (e.g., name(s), location(s), etc.) of the smart devices and assistant client devices.” (emphasis added). The generated registration data maps to spawning and managing separate processes for various types of services). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Wilberding a system and method wherein the voice coordinator discovers voice-enabled components of the automation network system and wherein the voice daemon spawns and manages separate processes for various types of services, as suggested by Mo, in order to enable the system to quickly and easily incorporate new smart devices, while also being compatible with the communication protocol of the new smart devices (see Mo at col. 2 lns. 25-39). Wilberding in view of Mo doesn’t describe a system or method wherein each voice input proxy manages the voice input device's microphone, and wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol. However, Lang describes a system and method wherein each voice input proxy manages the voice input device's microphone (col. 82 lns. 24-33: “if the computing device configured to control the media playback system (e.g., CR 522) does not have a microphone (or if the microphone is in use by some other application running on CR 522, e.g., if CR 522 is engaged in a telephone call), then one or both of the networked microphone system and the media playback system, individually or in combination, may select some other device on the network with a microphone to use that microphone as a fallback microphone for receiving voice commands for the media playback system.”), and wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol (col. 82 lns. 24-33: “if the computing device configured to control the media playback system (e.g., CR 522) does not have a microphone (or if the microphone is in use by some other application running on CR 522, e.g., if CR 522 is engaged in a telephone call), then one or both of the networked microphone system and the media playback system, individually or in combination, may select some other device on the network with a microphone to use that microphone as a fallback microphone for receiving voice commands for the media playback system.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Wilberding in view of Mo a system and method wherein each voice input proxy manages the voice input device's microphone, and wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol, as suggested by Lang, in order to enable the system to continue receiving input commands when the primary microphone on the controller in unavailable, which improves the responsiveness of the system. Wilberding in view of Mo and further in view of Lang doesn’t describe a system or method wherein the voice daemon transcodes audio data from the voice input devices, and wherein the voice daemon routes the transcoded audio data to the voice targets. However, Lang ’449 describes a system and method wherein the voice daemon transcodes audio data from the voice input devices (col. 13 ln. 61 – col. 14 ln. 3: “In some instances, processing system 500 may convert the received audio content into a format suitable for wake word detection. For instance, if the audio content is provided to the audio/input component 502 via an analog line-in interface, the processing system 500 may digitize the analog audio (e.g., using a software or hardware-based analog-to-digital converter). As another example, if the received audio content is received in a digital form that is unsuitable for analysis, the processing system 500 may transcode the recording into a suitable format.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to include in Wilberding in view of Mo in view of Lang a system and method wherein the voice daemon transcodes audio data from the voice input devices, as suggested by Lang ‘449, and further wherein the voice daemon routes the transcoded audio data to the voice targets, as suggested by the proposed modification of Wilberding, in order to convert the format of the audio recording to a format that is compatible with the associated voice target, such that the voice target can successfully determine the user’s intent and respond appropriately. RE claim 2, Wilberding describes the WHV system of claim 1, wherein the voice targets are associated with respective voice assistant services (col 18 lns. 52-65: “For instance, the particular wake-word may be “Hey, Siri” to invoke APPLE®'s voice service, “Ok, Google” to invoke GOOGLE®'s voice service, “Alexa” to invoke AMAZON®'s voice service, or “Hey, Cortana” to invoke Microsoft's voice service.”). RE claim 3, Wilberding describes the WHV system of claim 1, wherein the voice input devices are any of: a touchscreen, a home speaker, a stand-alone microphone, a remote control, a television, a home appliance, an intercom device, a cell phone, a computer, and a personal electronic device (col 1 ln. 61-col. 2 ln 2, NMD may be a Sonos playback device, Amazon Echo, or Apple iPhone.). RE claim 4, Wilberding describes the WHV system of claim 1, wherein the voice coordinator determines mapping information for mapping any of the voice input devices to any of the voice targets (col. 25 lns. 10-24 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).”). RE claim 5, Wilberding describes the WHV system of claim 4, wherein the mapping information is determined according to any of device registration information and device polling information (col. 25 lns. 10-24 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).”). RE claim 6, Wilberding describes the WHV system of claim 1, wherein the received voice inputs are recordings of sounds or speech vocalized by a user of the automation network (col. 17 lns. 32-37 “In some cases, the NMD may receive voice data indicating the voice input via a network interface, perhaps from another NMD within a household. The NMD may receive this recording in addition to receiving voice data indicating the voice input via a microphone (e.g., if the two NMDs are both within detection range of the voice input).”). RE claim 7, Wilberding describes an automation controller separate from any number of input devices, the automation controller having a processor, a memory, and a communication interface (FIG. 3 – control device 300, col. 8 lns 15-16: “In one example, the control device 300 may be a dedicated controller for the media playback system 100.” Also see Col. 8 lns. 60-63: “the control device 300 may sometimes be referred to as a controller, whether the control device 300 is a dedicated controller or a network device on which media playback system controller application software is installed.” Further, see col. 16 lns. 46-55: “At block 702, implementation 700 involves receiving voice data indicating a voice input. For instance, a NMD, such as NMD 600, may receive, via a microphone, voice data indicating a voice input. As further examples, any of playback devices 102, 104, 106, 108, 110, 112, 114, 116, 118, 120, 122, and 124 or control devices 126 and 128 of FIG. 1 may be a NMD and may receive voice data indicating a voice input. Yet further examples NMDs include NMDs 512, 514, and 516, PBDs 532, 534, 536, and 538, and CR 522 of FIG. 5.”), the automation controller configured to: execute a voice coordinator, the voice coordinator configured to configure the voice-enabled components for runtime use (col. 8 lns 52-59: “As suggested above, changes to configurations of the media playback system 100 may also be performed by a user using the control device 300. The configuration changes may include adding/removing one or more playback devices to/from a zone, adding/removing one or more zones to/from a zone group, forming a bonded or consolidated player, separating one or more playback devices from a bonded or consolidated player, among others.”); determine mapping information mapping any number of voice input devices to any number of voice targets (col. 25 lns. 10-31 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).”), the mapping information is determined by mapping any one or more of the voice input devices to any one or more of the voice targets according to any of device registration information and device polling information (col. 25 lns. 10-24 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).”); broadcast the mapping information to any number of devices communicatively connected to the automation controller via an automation network (col. 25 lns. 34-51 “the NMD may cause at least one of the detected voice services to be registered with a media playback system that includes one or more playback devices (e.g., media playback system 100 of FIG. 1). Causing the a voice service to be registered may involve transmitting, via a network interface, a message indicating credentials for that voice service to the media playback system (i.e., at least one device thereof).”); transmit a message commanding a voice daemon to instantiate a voice daemon instance for each of the voice targets included in the mapping information (col. 25 lns. 34-51 “Causing the a voice service to be registered may involve transmitting, via a network interface, a message indicating credentials for that voice service to the media playback system (i.e., at least one device thereof). The message may also include a command, request, or other query to cause the media playback system to register with the voice service using the credentials from the NMD.” Also see col. 19 lns. 17-19: “Registration of a voice service with the NMD or with the media playback system may integrate the API or other architecture of the voice service with the NMD.”), and wherein each voice daemon instance is respectively associated with one of the voice targets (col. 19 lns. 20-26: “Where multiple voice services are available to the NMD, the NMD might query wake-word detection algorithms corresponding to each voice service of the multiple voice services. As noted above, querying such detection algorithms may involve invoking respective APIs of the multiple voice services, either locally on the NMD or remotely using a network interface.”); execute a voice input proxy for each of the voice input devices (col 18 lns. 3-51: “Available voice services may include voice services registered with the NMD. Registration of a given voice service with the NMD may involve providing user credentials (e.g., user name and password) of the voice service to the NMD and/or providing an identifier of the NMD to the voice service. Such registration may configure the NMD to receive voice inputs on behalf of the voice service and perhaps configure the voice service to accept voice inputs from the NMD for processing.”), wherein each voice input proxy is configured to: maintain many voice target configurations per single input device (col 24 ln. 65 – col. 25 ln. 8: “At block 902, implementation 900 involves receiving input data indicating a command to register one or more voice services on one or more second devices. For instance, a first device (e.g., a NMD) may receive, via a user interface (e.g., a touchscreen), input data indicating a command to register one or more voice services with a media playback system that includes one or more playback devices. In one example, the NMD receives the input as part of a procedure to set-up the media playback system using any of the example techniques described above in connection with block 702 of implementation 700”, Also see col. 2 lns. 41-47: “Where two or more voice services are configured for a NMD, a particular voice service can be invoked by utterance of a wake-word corresponding to the particular voice service. For instance, in querying AMAZON®, a user might speak the wake-word “Alexa” followed by a voice input. Other examples include “Ok, Google” for querying GOOGLE® and “Hey, Siri” for querying APPLE®.”); support a current voice target and a current room for the respective voice input device (col. 25 lns. 11-19: “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD). For instance, a NMD that is a smartphone or tablet may have installed one or more applications (“apps”) that interface with voice services. The NMD may detect these applications using any suitable technique.” Also see col. 25 lns. 46-51: “In such manner, a user's media playback system may have registered one or more of the same voice services as registered on the user's NMD (e.g., smartphone) utilizing the same credentials as the user's NMD, which may hasten registration.” Further, regarding each proxy supporting a current room of a voice input device, see col. 7 lns. 2-11: “Referring back to the media playback system 100 of FIG. 1, the environment may have one or more playback zones, each with one or more playback devices. The media playback system 100 may be established with one or more playback zones, after which one or more zones may be added, or removed to arrive at the example configuration shown in FIG. 1. Each zone may be given a name according to a different room or space such as an office, bathroom, master bedroom, bedroom, kitchen, dining room, living room, and/or balcony.” And col. 7 lns. 47-59: “the zone configurations of the media playback system 100 may be dynamically modified, and in some embodiments, the media playback system 100 supports numerous configurations. For instance, if a user physically moves one or more playback devices to or from a zone, the media playback system 100 may be reconfigured to accommodate the change(s). For instance, if the user physically moves the playback device 102 from the balcony zone to the office zone, the office zone may now include both the playback device 118 and the playback device 102. The playback device 102 may be paired or grouped with the office zone and/or renamed if so desired via a control device such as the control devices 126 and 128.”); and receive voice audio and provide the voice audio to the voice daemon (col 18 lns. 3-51: “Available voice services may include voice services registered with the NMD. Registration of a given voice service with the NMD may involve providing user credentials (e.g., user name and password) of the voice service to the NMD and/or providing an identifier of the NMD to the voice service. Such registration may configure the NMD to receive voice inputs on behalf of the voice service and perhaps configure the voice service to accept voice inputs from the NMD for processing.” Also see col. 19 lns. 5-34: “a voice service may provide an application programming interface that the NMD can invoke to determine that whether the voice data includes the wake-word or phrase corresponding to that voice service. The NMD may invoke the API by transmitting a particular query of the voice service to the voice service along with data representing the wake-word portion of the received voice data. Alternatively, the NMD may invoke the API on the NMD itself. Registration of a voice service with the NMD or with the media playback system may integrate the API or other architecture of the voice service with the NMD.”); execute the voice daemon (col. 19 lns. 5-34: “a voice service may provide an application programming interface that the NMD can invoke to determine that whether the voice data includes the wake-word or phrase corresponding to that voice service. The NMD may invoke the API by transmitting a particular query of the voice service to the voice service along with data representing the wake-word portion of the received voice data. Alternatively, the NMD may invoke the API on the NMD itself. Registration of a voice service with the NMD or with the media playback system may integrate the API or other architecture of the voice service with the NMD.”), wherein the voice daemon is configured to: route the audio data to the voice targets according to the mapping information (col. 22 lns. 31-41: “At block 706, implementation 700 involves causing the identified voice service(s) to process the voice input. For instance, the NMD may transmit, via a network interface to one or more servers of the identified voice service(s), data representing the voice input and a command or query to process the data presenting the voice input. The command or query may cause the identified voice service(s) to process the voice command. The command or query may vary according to the identified voice service so as to conform the command or query to the identified voice service (e.g., to an API of the voice service).”); and instruct the voice daemon to transmit user voice inputs according to the mapping information (col. 18 lns. 52-65 “Identification of a particular voice service to process the voice input may be based on a wake-word or phrase in the voice input. For instance, after receiving voice data indicating a voice input, the NMD may determine that a portion of the voice data represents a particular wake-word. Further, the NMD may determine that the particular wake-word corresponds to a specific voice service. In other words, the NMD may determine that the particular wake-word or phrase is used to invoke a specific voice service.”). Wilberding doesn’t describe a system or method wherein the voice coordinator discovers voice-enabled components of the automation network system, wherein the voice daemon spawns and manages separate processes for various types of services, wherein each voice input proxy manages the voice input device's microphone, wherein the voice input proxy receives voice audio data from the voice input devices, wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol, wherein the voice daemon transcodes audio data from the voice input devices, and wherein the voice daemon routes the transcoded audio data to the voice targets. However, Mo describes a system and method wherein the voice coordinator discovers voice-enabled components of the automation network system (col. 8 lns. 46-67: “Various implementations described herein additionally or alternatively relate to utilizing local assistant client devices in discovering, provisioning, and/or registering smart devices for an account of a user. In some implementations, a smart device that is not yet registered can be discovered by causing each assistant client device, of an ecosystem of assistant client devices, to scan one or more communications channels (e.g., Wi-Fi, Bluetooth, and/or other) for smart device(s) that aren't registered. […] The discovery can occur at periodic or non-periodic intervals, or in response to a user request (e.g., a voice request, a request initiated via a smart phone app for the assistant, etc.).”), [and] wherein the voice daemon spawns and manages separate processes for various types of services (col. 9 lns. 1-17: “Once discovered, a smart device can be provisioned by causing the smart device to pair with at least one of the assistant client devices (e.g., in the case of Bluetooth) and/or to connect to a secured Wi-Fi network to which the assistant client devices are already connected. After provisioning, data transmitted by the smart device can be received, at an assistant client device, and processed using a local adapter of the assistant client device, to generate registration data in a schema of the automated assistant. For example, the data transmitted by the smart device can be in a protocol suite for a corresponding 3P, and can be interpreted by the 3P adapter into a schema for registration with the automated assistant. The registration data can be utilized by the automated assistant to add the smart device to a home graph or other structure that defines smart devices, assistant client devices, and various properties (e.g., name(s), location(s), etc.) of the smart devices and assistant client devices.” (emphasis added). The generated registration data maps to spawning and managing separate processes for various types of services). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Wilberding a system and method wherein the voice coordinator discovers voice-enabled components of the automation network system and wherein the voice daemon spawns and manages separate processes for various types of services, as suggested by Mo, in order to enable the system to quickly and easily incorporate new smart devices, while also being compatible with the communication protocol of the new smart devices (see Mo at col. 2 lns. 25-39). Wilberding in view of Mo doesn’t describe a system or method wherein each voice input proxy manages the voice input device's microphone, wherein the voice input proxy receives voice audio data from the voice input devices and provides the voice audio to the voice daemon, and wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol. However, Lang describes a system and method wherein each voice input proxy manages the voice input device's microphone (col. 82 lns. 24-33: “if the computing device configured to control the media playback system (e.g., CR 522) does not have a microphone (or if the microphone is in use by some other application running on CR 522, e.g., if CR 522 is engaged in a telephone call), then one or both of the networked microphone system and the media playback system, individually or in combination, may select some other device on the network with a microphone to use that microphone as a fallback microphone for receiving voice commands for the media playback system.”), wherein the voice input proxy receives voice audio data from the voice input devices (col. 82 lns. 24-33: “if the computing device configured to control the media playback system (e.g., CR 522) does not have a microphone (or if the microphone is in use by some other application running on CR 522, e.g., if CR 522 is engaged in a telephone call), then one or both of the networked microphone system and the media playback system, individually or in combination, may select some other device on the network with a microphone to use that microphone as a fallback microphone for receiving voice commands for the media playback system.”), and wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol (col. 82 lns. 24-33: “if the computing device configured to control the media playback system (e.g., CR 522) does not have a microphone (or if the microphone is in use by some other application running on CR 522, e.g., if CR 522 is engaged in a telephone call), then one or both of the networked microphone system and the media playback system, individually or in combination, may select some other device on the network with a microphone to use that microphone as a fallback microphone for receiving voice commands for the media playback system.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Wilberding in view of Mo a system and method wherein each voice input proxy manages the voice input device's microphone, wherein the voice input proxy receives voice audio data from the voice input devices, and wherein the voice daemon receives voice audio from the voice input devices via a voice control protocol, as suggested by Lang, in order to enable the system to continue receiving input commands when the primary microphone on the controller in unavailable, which improves the responsiveness of the system. Wilberding in view of Mo and further in view of Lang doesn’t describe a system or method wherein the voice daemon transcodes audio data from the voice input devices, and wherein the voice daemon routes the transcoded audio data to the voice targets. However, Lang ’449 describes a system and method wherein the voice daemon transcodes audio data from the voice input devices (col. 13 ln. 61 – col. 14 ln. 3: “In some instances, processing system 500 may convert the received audio content into a format suitable for wake word detection. For instance, if the audio content is provided to the audio/input component 502 via an analog line-in interface, the processing system 500 may digitize the analog audio (e.g., using a software or hardware-based analog-to-digital converter). As another example, if the received audio content is received in a digital form that is unsuitable for analysis, the processing system 500 may transcode the recording into a suitable format.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to include in Wilberding in view of Mo in view of Lang a system and method wherein the voice daemon transcodes audio data from the voice input devices, as suggested by Lang ‘449, and further wherein the voice daemon routes the transcoded audio data to the voice targets, as suggested by the proposed modification of Wilberding, in order to convert the format of the audio recording to a format that is compatible with the associated voice target, such that the voice target can successfully determine the user’s intent and respond appropriately. RE claim 8, Wilberding describes the automation controller of claim 7, wherein the voice targets are associated with respective voice assistant services (col 18 lns. 52-65: “For instance, the particular wake-word may be “Hey, Siri” to invoke APPLE®'s voice service, “Ok, Google” to invoke GOOGLE®'s voice service, “Alexa” to invoke AMAZON®'s voice service, or “Hey, Cortana” to invoke Microsoft's voice service.”). RE claim 9, Wilberding describes the automation controller of claim 7, wherein the voice input devices are any of: a touchscreen, a home speaker, a stand-alone microphone, a remote control, a television, a home appliance, an intercom device, a cell phone, a computer, and a personal electronic device (col 1 ln. 61-col. 2 ln 2, NMD may be a Sonos playback device, Amazon Echo, or Apple iPhone.). RE claim 10, Wilberding describes the automation controller of claim 7, wherein the mapping information is determined by mapping any one or more of the voice input devices to any one or more of the voice targets (col. 25 lns. 10-24 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).”). RE claim 11, Wilberding describes the automation controller of claim 10, wherein the mapping information is determined according to any of device registration information and device polling information (col. 25 lns. 10-24 “At block 904, implementation 900 involves detecting one or more voice services that are registered to the first device (e.g., the NMD). Such voice services may include voice services that are installed on the NMD or that are native to the NMD (e.g., part of the operating system of the NMD).”). RE claim 12, Wilberding describes the WHV of claim 1, further comprising the mapping information comprising a context (col. 19 lns. 47-56: “the NMD may identify a default voice service to process the voice input based on context. A default voice service may be pre-determined (e.g., configured during a set-up procedure, such as the example procedures described above). Then, when the NMD determines that the received voice data excludes any wake-word corresponding to a specific voice service (e.g., the NMD does not detect a wake-word corresponding to the specific voice service in the voice data), the NMD may select the default voice service to process the voice input.” Also see col. 20 lns. 36-53: “Alternatively, the NMD may identify a particular voice service to process the voice input based on context. For instance, the NMD may identify a particular voice service based on the type of command. An NMD (e.g., a NMD that is associated with a media playback system) may recognize certain commands (e.g., play, pause, skip forward, etc.) as being a particular type of command (e.g., media playback commands). In such cases, when the NMD determines that the voice input includes a particular type of command (e.g., a media playback command), the NMD may identify, as the voice service to process that voice input, a particular voice service configured to process that type of command. To further illustrate, search queries may be another example type of command (e.g., “what's the weather today?” or “where was David Bowie born?”). When the NMD determines that a voice input includes a search query, the NMD may identify a particular voice service (e.g., “GOOGLE”) to process that voice inputs that includes the search.” Finally, see col. 20 ln. 54 – col. 21 ln. 3: “In some cases, the NMD may determine that the voice input includes a voice command that is directed to a particular type of device. In such cases, the NMD may identify a particular voice service that is configured to process voice inputs directed to that type of device to process the voice input. For example, the NMD may determine that a given voice input is directed to one or more wireless illumination devices (e.g., that “Turn on the lights in here” is directed to the “smart” lightbulbs in the same room as the NMD) and identify, as the voice service to process the voice input, a particular voice service that is configured to process voice inputs directed to wireless illumination devices. As another example, the NMD may determine that a given voice input is directed to a playback device and identify, as the voice service to process the voice input, a particular voice service that is configured to process voice inputs directed to playback devices.”). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: So et al. (US 12,279,008) describes transcoding audio. Rotschield et al. (US 2017/0242557), Kozura et al. (US 10,075,334), and Pognant (US 2020/0044884) all describe discovery methods for automatically discovering new smart devices in home automation systems. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniel C Washburn whose telephone number is (571)272-5551. The examiner can normally be reached Monday-Friday 9:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Sep 27, 2023
Application Filed
Jun 11, 2025
Non-Final Rejection mailed — §103, §112
Sep 04, 2025
Response Filed
Oct 24, 2025
Final Rejection mailed — §103, §112
Jan 13, 2026
Request for Continued Examination
Jan 26, 2026
Response after Non-Final Action
May 26, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12602555
METHOD FOR SEARCHING FOR TEXTS IN DIFFERENT LANGUAGES BASED ON PRONUNCIATION AND ELECTRONIC DEVICE APPLYING THE SAME
3y 3m to grant Granted Apr 14, 2026
Patent 12603084
METHOD, APPARATUS, AND COMPUTER-READABLE RECORDING MEDIUM FOR CONTROLLING RESPONSE UTTERANCE BEING REPRODUCED AND PREDICTING USER INTENTION
2y 7m to grant Granted Apr 14, 2026
Patent 12511480
Pattern Recognition Using NLP-Based Tokenizing and Clustering Models
2y 4m to grant Granted Dec 30, 2025
Patent 9614588
Smart Appliances
2y 2m to grant Granted Apr 04, 2017
Patent 8373711
IMAGE PROCESSING APPARATUS, IMAGE PROCESSING METHOD, AND COMPUTER-READABLE STORAGE MEDIUM
5y 2m to grant Granted Feb 12, 2013
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
49%
Grant Probability
78%
With Interview (+28.7%)
4y 1m (~1y 3m remaining)
Median Time to Grant
High
PTA Risk
Based on 160 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month