DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to because:
In Figure 3, element “316”, “Speeker” should read “Speaker”.
In Figures 3-4, the labels are not black, in Figures 3-6 the labels are placed upon hatched or shaded surfaces, and in Figure 6, the text is not legible. All drawings must be made by a process which will give them satisfactory reproduction characteristics. Every line, number, and letter must be durable, clean, black (except for color drawings), sufficiently dense and dark, and uniformly thick and well-defined. Numbers, letters, and reference characters should not be placed upon hatched or shaded surfaces (see 37 CFR 1.84(l) and 37 CFR 1.84(p)).
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character not mentioned in the description:
“100” in Figure 1.
Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
On page 8, line 16, “include a sample audio signals” should read “include a sample audio signal” or “include sample audio signals”.
On page 12, line 1, “system 102” should read “system 100”.
On page 12, line 2, “device 104” should read “device 108”.
On page 12, line 2, “component 106” should read “component 104”.
On page 15, line 27, “the product may reacts” should read “the product may react”.
Appropriate correction is required.
Claim Objections
Claim 12 is objected to because of the following informalities:
In claim 12, lines 3-4, “comparing to features to the at least one identifier of the device” should read “comparing the features to the at least one identifier of the device”.
Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1 – 9, 12 – 13, 18, and 20 – 21 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Smith (US Patent No. 11,200,900).
Regarding claim 1, Smith discloses a component for providing offline voice recognition capabilities for operating a device (Column 2, lines 38-40, "Example techniques described herein involve offline voice control using a networked microphone device (“NMD”)."; Column 38, lines 6-7, "As noted above, the NMD 703 may perform local (“offline”) voice input processing."), comprising:
a memory configured to store a text file including at least one command for operating the device (Column 18, lines 31-34, "The memory 413 of the controller device 104 may be configured to store controller application software and other data associated with the MPS 100 and/or a user of the system 100."; Column 29, lines 54-63, "As noted above, in some example implementations, the NMD 703 is configured to perform natural language processing, which may be carried out using an onboard natural language processor, referred to herein as a natural language unit (NLU) 776. The local NLU 776 is configured to analyze text output of the ASR 775 to spot (i.e., detect or identify) keywords in the voice input 780. In FIG. 7A, this output is illustrated as the signal SASR. The local NLU 776 includes a keyword library 778 (i.e., words and phrases) corresponding to respective commands and/or parameters."; A keyword library reads on a text file including at least one command for operating the device.);
a microphone (Column 11, lines 6-9, "The microphones 222 are configured to detect sound (i.e., acoustic waves) in the environment of the playback device 102, which is then provided to the voice processing components 220.");
and circuitry in communication with the memory and the microphone (Column 9, lines 23-26, "As shown, the playback device 102 includes at least one processor 212, which may be a clock-driven computing component configured to process input data according to instructions stored in memory 213."), the circuitry configured for:
extracting features from audio signals generated by the microphone (Column 32, lines 37-45, "The VAD 765 may utilize any suitable voice activity detection algorithms. Example voice detection algorithms involve determining whether a given frame includes one or more features or qualities that correspond to voice activity, and further determining whether those features or qualities diverge from noise to a given extent (e.g., if a value exceeds a threshold for a given frame). Some example voice detection algorithms involve filtering or otherwise reducing noise in the frames prior to identifying the features or qualities."; Column 26, lines 4-6, "In any event, the detected-sound data forms a digital representation (i.e., sound-data stream), SDS, of the sound detected by the microphones 720."; Column 29, lines 37-48, "When a local wake-word event is generated, the NMD 703 can employ an automatic speech recognizer 775. The ASR 775 is configured to output phonetic or phenomic representations, such as text corresponding to words, based on sound in the sound-data stream SDS to text. For instance, the ASR 775 may transcribe spoken words represented in the sound-data stream SDS to one or more strings representing the voice input 780 as text. The ASR 775 can feed ASR output (labeled as SASR) to a local natural language unit (NLU) 776 that identifies particular keywords as being local keywords for invoking local-keyword events, as described below."; Forming a digital representation of the sound detected by the microphones reads on extracting features from audio signals generated by the microphone.);
comparing the features to the at least one command of the text file stored by the memory (Column 29, lines 37-48, "When a local wake-word event is generated, the NMD 703 can employ an automatic speech recognizer 775. The ASR 775 is configured to output phonetic or phenomic representations, such as text corresponding to words, based on sound in the sound-data stream SDS to text. For instance, the ASR 775 may transcribe spoken words represented in the sound-data stream SDS to one or more strings representing the voice input 780 as text. The ASR 775 can feed ASR output (labeled as SASR) to a local natural language unit (NLU) 776 that identifies particular keywords as being local keywords for invoking local-keyword events, as described below."; Column 29, lines 54-63, "As noted above, in some example implementations, the NMD 703 is configured to perform natural language processing, which may be carried out using an onboard natural language processor, referred to herein as a natural language unit (NLU) 776. The local NLU 776 is configured to analyze text output of the ASR 775 to spot (i.e., detect or identify) keywords in the voice input 780. In FIG. 7A, this output is illustrated as the signal SASR. The local NLU 776 includes a keyword library 778 (i.e., words and phrases) corresponding to respective commands and/or parameters."; Detecting keywords in the voice input with a keyword library corresponding to respective commands reads on comparing the features to the at least one command of the text file stored by the memory.);
and in response to a match between the features and the at least one command, generating instructions for operating the device according to the at least one command (Column 28, lines 22-31, "Local keywords may also take the form of command keywords. In contrast to the nonce words typically as utilized as VAS wake words, command keywords function as both the activation word and the command itself. For instance, example command keywords may correspond to playback commands (e.g., “play,” “pause,” “skip,” etc.) as well as control commands (“turn on”), among other examples. Under appropriate conditions, based on detecting one of these command keywords, the NMD 703a performs the corresponding command."; Detecting a command keyword and performing the corresponding command reads on generating instructions for operating the device according to the at least one command in response to a match between the features and the at least one command.).
Regarding claim 2, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein comparing comprises measuring similarity between the features and the at least one command, and the match comprises the measured similarity is greater than a threshold and/or according to a requirement (Column 31, lines 6-21, "Similarly, some error in performing keyword matching is expected. Within examples, the local NLU 776 may generate a confidence score when determining an intent, which indicates how closely the transcribed words in the signal SASR match the corresponding keywords in the library 778 of the local NLU 776. In some implementations, performing an operation according to a determined intent is based on the confidence score for keywords matched in the signal SASR. For instance, the NMD 703 may perform an operation according to a determined intent when the confidence score for a given sound exceeds a given threshold value (e.g., 0.5 on a scale of 0-1, indicating that the given sound is more likely than not the command keyword). Conversely, when the confidence score for a given intent is at or below the given threshold value, the NMD 703 does not perform the operation according to the determined intent."; Generating a confidence score indicating how closely a transcribed word in the signal matches the corresponding keyword in the library reads on measuring similarity between the features and the command, and the confidence score exceeding a given threshold value reads on the measured similarity being greater than a threshold.).
Regarding claim 3, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein the component excludes a network interface (Column 22, lines 53-59, "Many of these components are similar to the playback device 102 of FIG. 2A. In some examples, the NMD 703 may be implemented in a playback device 102. In such cases, the NMD 703 might not include duplicate components (e.g., a network interface 224 and a network 724), but may instead share several components to carry out both playback and voice control functions."; Column 49, lines 49-56, "FIG. 12 is a flow diagram showing an example method 1200 to perform offline voice processing. The method 1200 may be performed by a networked microphone device, such as the NMD 703 (FIG. 7A). Alternatively, the method 1200 may be performed by any suitable device or by a system of devices, such as the playback devices 103, NMDs 103, control devices 104, computing devices 105, computing devices 106, and/or NMD 703."; A playback device that does not include a network interface reads on the component excluding a network interface.).
Regarding claim 4, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein the component includes a data interface for point to point communication between the component and a controller of the device (Column 12, lines 5-8, "The playback device 102 further includes a user interface 240 that may facilitate user interactions independent of or in conjunction with user interactions facilitated by one or more of the controller devices 104."; A user interface that facilitates user interactions with a controller device reads on a data interface for point to point communication between the component and a controller of the device.).
Regarding claim 5, Smith discloses the component as claimed in claim 4.
Smith further discloses:
wherein the data interface is physically wired to the controller of the device (Column 6, lines 16-22, "With reference still to FIG. 1B, the various playback, network microphone, and controller devices 102, 103, and 104 and/or other network devices of the MPS 100 may be coupled to one another via point-to-point connections and/or over other connections, which may be wired and/or wireless, via a network 111, such as a LAN including a network router 109.").
Regarding claim 6, Smith discloses the component as claimed in claim 4.
Smith further discloses:
wherein the data interface is a serial interface (Column 10, lines 39-55, "As shown, the at least one network interface 224, may take the form of one or more wireless interfaces 225 and/or one or more wired interfaces 226. A wireless interface may provide network interface functions for the playback device 102 to wirelessly communicate with other devices (e.g., other playback device(s), NMD(s), and/or controller device(s)) in accordance with a communication protocol (e.g., any wireless standard including IEEE 802.11a, 802.11b, 802.11g, 802.11n, 802.11ac, 802.15, 4G mobile communication standard, and so on). A wired interface may provide network interface functions for the playback device 102 to communicate over a wired connection with other devices in accordance with a communication protocol (e.g., IEEE 802.3). While the network interface 224 shown in FIG. 2A include both wired and wireless interfaces, the playback device 102 may in some implementations include only wireless interface(s) or only wired interface(s)."; Column 19, lines 5-13, "The user interface as shown is just one example of an interface that may be provided on a network device, such as the controller device shown in FIG. 4, and accessed by users to control a media playback system, such as the MPS 100. Other user interfaces of varying formats, styles, and interactive sequences may alternatively be implemented on one or more network devices to provide comparable control access to a media playback system."; An IEEE 802.3 communication protocol reads on a serial interface.).
Regarding claim 7, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein the device comprises at least one of: a consumer product and a home device (Column 2, lines 38-47, "Example techniques described herein involve offline voice control using a networked microphone device (“NMD”). An NMD is a networked computing device that typically includes an arrangement of microphones, such as a microphone array, that is configured to detect sound present in the NMD's environment. NMDs may facilitate voice control of smart home devices, such as wireless audio playback devices, illumination devices, appliances, and home-automation devices (e.g., thermostats, door locks, etc.).").
Regarding claim 8, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein the device is selected from: coffee maker, heater, fan, vacuum, light, diffuser, radio, oven, speaker, audio device, video device, television, personal care devices, and air conditioners (Column 2, lines 38-47, "Example techniques described herein involve offline voice control using a networked microphone device (“NMD”). An NMD is a networked computing device that typically includes an arrangement of microphones, such as a microphone array, that is configured to detect sound present in the NMD's environment. NMDs may facilitate voice control of smart home devices, such as wireless audio playback devices, illumination devices, appliances, and home-automation devices (e.g., thermostats, door locks, etc.)."; An audio playback device reads on a speaker.).
Regarding claim 9, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein the component is integrated with a controller of the device as an embedded system installed within the device (Column 7, lines 1-5, "In various implementations, one or more of the playback devices 102 may take the form of or include an on-board (e.g., integrated) network microphone device. For example, the playback devices 102a-e include or are otherwise equipped with corresponding NMDs 103a-e, respectively.").
Regarding claim 12, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein the memory is further configured to store a second text file including at least one identifier of the device, wherein circuitry is further configured for comparing the features to the at least one identifier of the device (Column 16, lines 32-40, "In some embodiments, the memory 213 of the playback device 102 may store instances of various variable types associated with the states. Variables instances may be stored with identifiers (e.g., tags) corresponding to type. For example, certain identifiers may be a first type “a1” to identify playback device(s) of a zone, a second type “b1” to identify playback device(s) that may be bonded in the zone, and a third type “cl” to identify a zone group to which the zone may belong."; Column 44, lines 21-38, "Within example implementations, the NMD 703 may populate the library 778 of the local NLU 776 locally within the network 111 (FIG. 1B). As noted above, the NMD 703 may maintain or have access to state variables indicating the respective states of devices connected to the network 111 (e.g., the playback devices 104). These state variables may include names of the various devices. For instance, the kitchen 101h may include the playback device 101b, which are assigned the zone name “Kitchen.” The NMD 703 may read these names from the state variables and include them in the library 778 of the local NLU 776 by training the local NLU 776 to recognize them as keywords. The keyword entry for a given name may then be associated with the corresponding device in an associated parameter (e.g., by an identifier of the device, such as a MAC address or IP address). The NMD 703a can then use the parameters to customize control commands and direct the commands to a particular device."; A keyword entry for a given name being associated with the corresponding device in an associated parameter reads on storing a second text file including at least one identifier of the device, and using the parameters to customize control commands and direct the commands to a particular device reads on comparing the features to the at least one identifier of the device.);
and in response to a match between the features and the at least one identifier, performing the comparison between the features and the at least one command (Column 35, lines 15-28, "In a second example, another example voice input 780 may be “play my favorites in the Kitchen” with “play” again being the command keyword portion (corresponding to a playback command) and “my favorites in the Kitchen” as the voice utterance portion. When analyzing this voice input 780, the NLU 776 may recognize that “favorites” and “Kitchen” match keywords in its library 778. In particular, “favorites” corresponds to a first parameter representing particular audio content (i.e., a particular playlist that includes a user's favorite audio tracks) while “Kitchen” corresponds to a second parameter representing a target for the playback command (i.e., the kitchen 101h zone. Accordingly, the NLU 776 may determine an intent to play this particular playlist in the kitchen 101h zone."; Column 44, lines 32-38, "The keyword entry for a given name may then be associated with the corresponding device in an associated parameter (e.g., by an identifier of the device, such as a MAC address or IP address). The NMD 703a can then use the parameters to customize control commands and direct the commands to a particular device."; Determining a keyword for a given name is associated with a corresponding device, then directing the commands to the particular device, reads on performing the comparison between the features and the at least one command in response to a match between the features and the at least one identifier.).
Regarding claim 13, Smith discloses the component as claimed in claim 1.
Smith further discloses:
further comprising circuitry for extracting features from the text file, wherein the comparison is done by comparing the features extracted from the audio signals to features extracted from the text file (Column 29, lines 37-48, "When a local wake-word event is generated, the NMD 703 can employ an automatic speech recognizer 775. The ASR 775 is configured to output phonetic or phenomic representations, such as text corresponding to words, based on sound in the sound-data stream SDS to text. For instance, the ASR 775 may transcribe spoken words represented in the sound-data stream SDS to one or more strings representing the voice input 780 as text. The ASR 775 can feed ASR output (labeled as SASR) to a local natural language unit (NLU) 776 that identifies particular keywords as being local keywords for invoking local-keyword events, as described below."; Column 29, lines 54-63, "As noted above, in some example implementations, the NMD 703 is configured to perform natural language processing, which may be carried out using an onboard natural language processor, referred to herein as a natural language unit (NLU) 776. The local NLU 776 is configured to analyze text output of the ASR 775 to spot (i.e., detect or identify) keywords in the voice input 780. In FIG. 7A, this output is illustrated as the signal SASR. The local NLU 776 includes a keyword library 778 (i.e., words and phrases) corresponding to respective commands and/or parameters."; Outputting phonetic or phenomic representations based on sound in the sound-data stream and performing natural language processing to identify keywords from a keyword library in the voice input reads on comparing the features extracted from the audio signals to features extracted from the text file.).
Regarding claim 18, Smith discloses the component as claimed in claim 1.
Smith further discloses:
wherein the text file includes at least 5 different commands for operating the device (Column 29, lines 22-29, “Local keywords may also take the form of command keywords. In contrast to the nonce words typically as utilized as VAS wake words, command keywords function as both the activation word and the command itself. For instance, example command keywords may correspond to playback commands (e.g., “play,” “pause,” “skip,” etc.) as well as control commands (“turn on”), among other examples."; The example commands “play”, “pause”, “skip”, and “turn on”, with other examples, reads on at least 5 different commands for operating the device.).
Regarding claim 20, Smith discloses a method of providing offline voice recognition capabilities for operating a device (Column 2, lines 38-40, "Example techniques described herein involve offline voice control using a networked microphone device (“NMD”)."; Column 38, lines 6-7, "As noted above, the NMD 703 may perform local (“offline”) voice input processing."), comprising:
extracting features from audio signals generated by a microphone (Column 32, lines 37-45, "The VAD 765 may utilize any suitable voice activity detection algorithms. Example voice detection algorithms involve determining whether a given frame includes one or more features or qualities that correspond to voice activity, and further determining whether those features or qualities diverge from noise to a given extent (e.g., if a value exceeds a threshold for a given frame). Some example voice detection algorithms involve filtering or otherwise reducing noise in the frames prior to identifying the features or qualities."; Column 26, lines 4-6, "In any event, the detected-sound data forms a digital representation (i.e., sound-data stream), SDS, of the sound detected by the microphones 720."; Column 29, lines 37-48, "When a local wake-word event is generated, the NMD 703 can employ an automatic speech recognizer 775. The ASR 775 is configured to output phonetic or phenomic representations, such as text corresponding to words, based on sound in the sound-data stream SDS to text. For instance, the ASR 775 may transcribe spoken words represented in the sound-data stream SDS to one or more strings representing the voice input 780 as text. The ASR 775 can feed ASR output (labeled as SASR) to a local natural language unit (NLU) 776 that identifies particular keywords as being local keywords for invoking local-keyword events, as described below."; Forming a digital representation of the sound detected by the microphones reads on extracting features from audio signals generated by the microphone.);
comparing the features to at least one command of a text file stored by a memory physically connected to the device (Column 18, lines 31-34, "The memory 413 of the controller device 104 may be configured to store controller application software and other data associated with the MPS 100 and/or a user of the system 100."; Column 29, lines 37-48, "When a local wake-word event is generated, the NMD 703 can employ an automatic speech recognizer 775. The ASR 775 is configured to output phonetic or phenomic representations, such as text corresponding to words, based on sound in the sound-data stream SDS to text. For instance, the ASR 775 may transcribe spoken words represented in the sound-data stream SDS to one or more strings representing the voice input 780 as text. The ASR 775 can feed ASR output (labeled as SASR) to a local natural language unit (NLU) 776 that identifies particular keywords as being local keywords for invoking local-keyword events, as described below."; Column 29, lines 54-63, "As noted above, in some example implementations, the NMD 703 is configured to perform natural language processing, which may be carried out using an onboard natural language processor, referred to herein as a natural language unit (NLU) 776. The local NLU 776 is configured to analyze text output of the ASR 775 to spot (i.e., detect or identify) keywords in the voice input 780. In FIG. 7A, this output is illustrated as the signal SASR. The local NLU 776 includes a keyword library 778 (i.e., words and phrases) corresponding to respective commands and/or parameters."; Detecting keywords in the voice input with a keyword library corresponding to respective commands reads on comparing the features to the at least one command of the text file stored by the memory.);
and in response to a match between the features and the at least one command, generating instructions for operating a controller of the device according to the at least one command (Column 28, lines 22-31, "Local keywords may also take the form of command keywords. In contrast to the nonce words typically as utilized as VAS wake words, command keywords function as both the activation word and the command itself. For instance, example command keywords may correspond to playback commands (e.g., “play,” “pause,” “skip,” etc.) as well as control commands (“turn on”), among other examples. Under appropriate conditions, based on detecting one of these command keywords, the NMD 703a performs the corresponding command."; Detecting a command keyword and performing the corresponding command reads on generating instructions for operating a controller of the device according to the at least one command in response to a match between the features and the at least one command.).
Regarding claim 21, Smith discloses a non-transitory medium storing program instructions (Column 18, lines 31-34, "The memory 413 of the controller device 104 may be configured to store controller application software and other data associated with the MPS 100 and/or a user of the system 100.") for providing offline voice recognition capabilities for operating a device (Column 2, lines 38-40, "Example techniques described herein involve offline voice control using a networked microphone device (“NMD”)."; Column 38, lines 6-7, "As noted above, the NMD 703 may perform local (“offline”) voice input processing."), which when executed by at least one processor (Column 9, lines 23-26, "As shown, the playback device 102 includes at least one processor 212, which may be a clock-driven computing component configured to process input data according to instructions stored in memory 213."), cause the at least one processor to:
extract features from audio signals generated by a microphone (Column 32, lines 37-45, "The VAD 765 may utilize any suitable voice activity detection algorithms. Example voice detection algorithms involve determining whether a given frame includes one or more features or qualities that correspond to voice activity, and further determining whether those features or qualities diverge from noise to a given extent (e.g., if a value exceeds a threshold for a given frame). Some example voice detection algorithms involve filtering or otherwise reducing noise in the frames prior to identifying the features or qualities."; Column 26, lines 4-6, "In any event, the detected-sound data forms a digital representation (i.e., sound-data stream), SDS, of the sound detected by the microphones 720."; Column 29, lines 37-48, "When a local wake-word event is generated, the NMD 703 can employ an automatic speech recognizer 775. The ASR 775 is configured to output phonetic or phenomic representations, such as text corresponding to words, based on sound in the sound-data stream SDS to text. For instance, the ASR 775 may transcribe spoken words represented in the sound-data stream SDS to one or more strings representing the voice input 780 as text. The ASR 775 can feed ASR output (labeled as SASR) to a local natural language unit (NLU) 776 that identifies particular keywords as being local keywords for invoking local-keyword events, as described below."; Forming a digital representation of the sound detected by the microphones reads on extracting features from audio signals generated by the microphone.);
compare the features to at least one command of a text file stored by a memory physically connected to the device (Column 18, lines 31-34, "The memory 413 of the controller device 104 may be configured to store controller application software and other data associated with the MPS 100 and/or a user of the system 100."; Column 29, lines 37-48, "When a local wake-word event is generated, the NMD 703 can employ an automatic speech recognizer 775. The ASR 775 is configured to output phonetic or phenomic representations, such as text corresponding to words, based on sound in the sound-data stream SDS to text. For instance, the ASR 775 may transcribe spoken words represented in the sound-data stream SDS to one or more strings representing the voice input 780 as text. The ASR 775 can feed ASR output (labeled as SASR) to a local natural language unit (NLU) 776 that identifies particular keywords as being local keywords for invoking local-keyword events, as described below."; Column 29, lines 54-63, "As noted above, in some example implementations, the NMD 703 is configured to perform natural language processing, which may be carried out using an onboard natural language processor, referred to herein as a natural language unit (NLU) 776. The local NLU 776 is configured to analyze text output of the ASR 775 to spot (i.e., detect or identify) keywords in the voice input 780. In FIG. 7A, this output is illustrated as the signal SASR. The local NLU 776 includes a keyword library 778 (i.e., words and phrases) corresponding to respective commands and/or parameters."; Detecting keywords in the voice input with a keyword library corresponding to respective commands reads on comparing the features to the at least one command of the text file stored by the memory.);
and in response to a match between the features and the at least one command, generate instructions for operating a controller of the device according to the at least one command (Column 28, lines 22-31, "Local keywords may also take the form of command keywords. In contrast to the nonce words typically as utilized as VAS wake words, command keywords function as both the activation word and the command itself. For instance, example command keywords may correspond to playback commands (e.g., “play,” “pause,” “skip,” etc.) as well as control commands (“turn on”), among other examples. Under appropriate conditions, based on detecting one of these command keywords, the NMD 703a performs the corresponding command."; Detecting a command keyword and performing the corresponding command reads on generating instructions for operating a controller of the device according to the at least one command in response to a match between the features and the at least one command.).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 10 – 11 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Smith in view of Mao et al. (US Patent No. 11,282,495), hereinafter Mao.
Regarding claim 10, Smith discloses the component as claimed in claim 1, but does not specifically disclose: wherein the circuitry is further configured for implementing a neural network that generates an embedding in response to an input of the audio signals, the extracted features include the embedding.
Mao teaches:
wherein the circuitry is further configured for implementing a neural network that generates an embedding in response to an input of the audio signals, the extracted features include the embedding (Column 3, lines 19-31, "The user device 110a may process (132) the audio data using a first neural network to determine first embedding data representing characteristics of the utterance. This first embedding data may include a data vector that represents vocal characteristics of the voice of the user 10. This embedding data compresses or “embeds” audio characteristics of a particular utterance over time. As described in greater detail below, the embedding data may be a vector of (for example) 100-200 floating-point values that represent an embedding of audio of duration between 0.5-2.0 seconds. The values may each represent one or more audio characteristics such as pitch, tone, speech rate, cadence, or other such characteristics."; A neural network determining embedding data representing characteristics of an utterance reads a on a neural network generating an embedding in response to an input of the audio signals.).
Mao is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith to incorporate the teachings of Mao to implement a neural network determining embedding data representing characteristics of an utterance. Doing so would allow for recognizing words spoken by a user based on the various qualities of received audio, understanding the intent of the user given the recognized words, and performing tasks based on the user's spoken commands (Mao; Column 1, lines 49-60).
Regarding claim 11, Smith in view of Mao discloses the component as claimed in claim 10.
Mao further teaches:
wherein the neural network is trained on a training dataset of multiple sample audio signals from a plurality of individuals referring to a plurality of different identifiers of devices and/or a plurality of different commands for operating the device (Column 8, lines 15-22, "Each of the feature extraction component 204, feature conversion component 208, and/or TTS component 216 may be initially trained by a system other than the user device 110, such as the remote system 120, using a corpus of training data including, for example, audio data representing one or more users uttering one or more words (and corresponding annotation data including text representations of those words).”; A corpus of training data including audio data representing one or more users uttering one or more words reads on a training dataset of multiple sample audio signals from a plurality of individuals referring to a plurality of different identifiers of devices.).
Mao is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith in view of Mao to further incorporate the teachings of Mao to train a neural network with a corpus of training data including audio data representing one or more users uttering one or more words. Doing so would allow for recognizing words spoken by a user based on the various qualities of received audio, understanding the intent of the user given the recognized words, and performing tasks based on the user's spoken commands (Mao; Column 1, lines 49-60).
Regarding claim 19, Smith discloses the component as claimed in claim 1, but does not specifically disclose: wherein the circuitry is further configured for digitizing an analogue signal obtained from the microphone, wherein the features are extracted from the digitized analogue signal.
Mao teaches:
wherein the circuitry is further configured for digitizing an analogue signal obtained from the microphone, wherein the features are extracted from the digitized analogue signal (Column 3, lines 6-21, "Referring first to FIG. 1A, in various embodiments, the user device 110a determines (130) audio data corresponding to an utterance. The user device may receive audio 12 from a user 10 using, for example, a microphone or array of microphones, and determine a digital representation of the analog audio signal, in which numbers in the audio data represent samples of an amplitude of the analog audio signal over time. The audio data may instead or in addition be processed to determine frequency-domain audio data (using, e.g., a Fourier transform); numbers in this audio data may represent frequencies of the audio data. The audio data may further be divided into frames, as described in greater detail below with reference to FIG. 6. The user device 110a may process (132) the audio data using a first neural network to determine first embedding data representing characteristics of the utterance."; Receiving audio from a user using a microphone, determining a digital representation of the analog audio signal, and processing the audio data using a neural network to determine embedding data representing characteristics of the utterance, reads on digitizing an analogue signal obtained from the microphone and extracting features from the digitized analogue signal.).
Mao is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith to incorporate the teachings of Mao to receive audio from a user using a microphone, determine a digital representation of the analog audio signal, and process the audio data using a neural network to determine embedding data representing characteristics of the utterance. Doing so would allow for recognizing words spoken by a user based on the various qualities of received audio, understanding the intent of the user given the recognized words, and performing tasks based on the user's spoken commands (Mao; Column 1, lines 49-60).
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Smith in view of Fidler et al. (US Patent No. 12,482,465), hereinafter Fidler.
Regarding claim 14, Smith discloses the component as claimed in claim 13, but does not specifically disclose: wherein features are extracted from the audio signals as embeddings arranged into a first vector outputted by a neural network, and features are extracted from the text file as a second vector, and the comparison is performed by computing a distance between the first vector and the second vector, and evaluating the distance according to a threshold or requirement.
Fidler teaches:
wherein features are extracted from the audio signals as embeddings arranged into a first vector outputted by a neural network, and features are extracted from the text file as a second vector, and the comparison is performed by computing a distance between the first vector and the second vector, and evaluating the distance according to a threshold or requirement (Column 3, lines 1-4, "As such, in examples, speech processing can be implemented with customized embeddings. Audio data may be input into component(s) and customized embeddings representing the audio data are output."; Column 3, line 56 - Column 4, line 1, "The storage manager may be configured to receive the reference interpretation and facilitate the generation and/or storage of a reference embedding that represents the reference interpretation. For example, the storage manager may parse the reference interpretation for the textual representation of the audio data or other ASR output data and, in examples, for the context data associated with the voice command. The textual representation of the audio data or other ASR output data and/or the context data may be sent to a context encoder of the device. The context encoder may include a number of components, including a text encoder, configured to generate a reference embedding that represents that voice command."; Column 4, lines 44-67, "Thereafter, the device may receive additional voice commands over time. The device may generate audio data from the received audio and may utilize an audio encoder to generate a runtime embedding of the audio data. As with the context encoder and the text encoder, the audio encoder may be configured to utilize RNNs and feed forward layers to generate a runtime embedding. The runtime embedding may then be analyzed with respect to the reference embeddings stored in the embedding space to determine whether the runtime embedding is similar to one or more of the reference embeddings to at least a threshold degree. For example, the runtime embedding may include a vector value indicating its location in the embedding space. A nearest reference embedding in the embedding space may be determined and a distance between the two locations may be determined. When the distance is within a threshold distance, the device may determine that the runtime embedding is sufficiently similar to the reference embedding. When this occurs, the intent associated with the reference embedding may be determined and corresponding intent data or other NLU output data and/or other data associated with the voice command may be sent to one or more applications to determine an action to be performed responsive to the voice command."; Generating a runtime embedding of received audio using a recurrent neural network reads on extracting features from the audio signals as embeddings arranged into a first vector outputted by a neural network, generating a reference embedding that represents a reference interpretation for a textual representation reads on extracting features from the text file as a second vector, and analyzing the runtime embedding with respect to the reference embeddings to determine whether the runtime embedding is similar to one or more of the reference embeddings to at least a threshold degree, where the runtime embedding includes a vector value indicating its location in the embedding space, a nearest reference embedding in the embedding space is determined, a distance between the two locations is determined, and when the distance is within a threshold distance, the device determines that the runtime embedding is sufficiently similar to the reference embedding, reads on computing a distance between the first vector and the second vector and evaluating the distance according to a threshold or requirement.).
Fidler is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith to incorporate the teachings of Fidler to generate a runtime embedding of received audio using a recurrent neural network, generate a reference embedding that represents a reference interpretation for a textual representation, and analyze the runtime embedding with respect to the reference embeddings to determine whether the runtime embedding is similar to one or more of the reference embeddings to at least a threshold degree, where the runtime embedding includes a vector value indicating its location in the embedding space, a nearest reference embedding in the embedding space is determined, a distance between the two locations is determined, and when the distance is within a threshold distance, the device determines that the runtime embedding is sufficiently similar to the reference embedding. Doing so would allow for determining the intent of a voice command and determining an action to be performed responsive to the voice command (Fidler; Column 3, lines 1-24).
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Smith in view of Herbig et al. (US Patent No. 8,346,551), hereinafter Herbig, and Shen et al. (US Patent No. 11,295,748), hereinafter Shen.
Regarding claim 15, Smith discloses the component as claimed in claim 1, but does not specifically disclose: wherein at least one of the extracted features includes Mel Frequency Cepstral Coefficients (MFCCs) computed by applying a Mel filterbank to a power spectrum representation of the audio signal, computing a logarithm of the outcome of applying the Mel filterbank, and applying a discrete cosine transform (DCT) to obtain the MFCCs,
Herbig teaches:
wherein at least one of the extracted features includes Mel Frequency Cepstral Coefficients (MFCCs) computed by applying a Mel filterbank to a power spectrum representation of the audio signal, computing a logarithm of the outcome of applying the Mel filterbank, and applying a discrete cosine transform (DCT) to obtain the MFCCs (Column 2, lines 30-40, "A feature vector is obtained by performing feature extraction on the speech input or utterance. Typically, a feature vector may be determined via a parameterization into the Mel-Cepstrum. For example, the power spectrum of an acoustic signal may be transformed into the Mel-Scale; for this purpose, for example, a Mel filterbank may be used. Then, the logarithm of the Mel frequency bands is taken, followed by a discrete cosine transform resulting in Mel frequency cepstral coefficients (MFCCs) constituting a feature vector. In general, there are also other possibilities to perform a feature extraction than by determining MFCCs.").
Herbig is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith to incorporate the teachings of Herbig to determine Mel frequency cepstral coefficients (MFCCs) constituting a feature vector by transforming the power spectrum of an acoustic signal into the Mel-Scale using a Mel filterbank, taking the logarithm of the Mel frequency bands, and performing a discrete cosine transform. Doing so would allow for identifying or recognizing a speaker based on a set of speech recognition codebooks comprising a speaker-independent codebook and at least one speaker-dependent codebook (Herbig; Column 3, line 63 - Column 4, line 7).
Smith in view of Herbig does not specifically disclose: wherein the MFCCs are compared to the text file.
Shen teaches
wherein the MFCCs are compared to the text file (Column 3, lines 8-17, "In some embodiments, detecting the input key phrase data in the input audio data includes separating a portion of the audio data into predetermined segments. The processor extracts MFCCs indicative of human speech features present within each segment, and compares the extracted MFCCs with MFCCs corresponding to the key phrase from a Universal Background Model (“UBM”) stored in the memory. The processor further determines that the portion of the audio signal includes the utterance of the key phrase based on the comparison."; Comparing the MFCCs extracted from input audio data with MFCCs corresponding to a key phrase from a model stored in memory reads on comparing the MFCCs to the text file.).
Shen is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith in view of Herbig to incorporate the teachings of Shen to compare MFCCs extracted from input audio data with MFCCs corresponding to a key phrase from a model stored in memory. Doing so would allow for authenticating a speaker as an enrolled user with a speaker recognition system (Shen; Column 1, lines 51-56).
Claims 16 – 17 are rejected under 35 U.S.C. 103 as being unpatentable over Smith in view of Bang et al. (US Patent No. 10,665,222), hereinafter Bang.
Regarding claim 16, Smith discloses the component as claimed in claim 1, but does not specifically disclose: wherein the comparison is performed by computing a dynamic time warping (DTW) of the audio signal and the text file for alignment of the audio signal with the text file for matching the aligned audio signal and text file by measuring similarity between temporal sequences of the matched aligned audio signal and text file varying in speed.
Bang teaches:
wherein the comparison is performed by computing a dynamic time warping (DTW) of the audio signal and the text file for alignment of the audio signal with the text file for matching the aligned audio signal and text file by measuring similarity between temporal sequences of the matched aligned audio signal and text file varying in speed (Column 4, lines 48-55, "A weighted finite state transducer (WFST) unit or decoder 22 uses the acoustic scores to identify one or more utterance hypotheses and compute their scores. The WFST decoder 22 may use a neural network of arcs and states that also are referred to as WFSTs. The WFST neural network or other networks may be based on hidden Markov models (HMMs), dynamic time warping, and many other network techniques or structures."; Using acoustic scores to identify one or more utterance hypotheses and compute their scores using dynamic time warping reads on computing a dynamic time warping of the audio signal and the text file for alignment of the audio signal with the text file.).
Bang is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith to incorporate the teachings of Bang to use acoustic scores to identify one or more utterance hypotheses and compute their scores using dynamic time warping. Doing so would allow for recording audio with a microphone, processing the acoustic data with a speech processing device, and outputting speech or visual information to the user or other application to perform a further action depending on the audio information (Bang; Column 4, lines 6-11).
Regarding claim 17, Smith discloses the component as claimed in claim 1, but does not specifically disclose: wherein the circuitry is further configured to implement a Hidden Markov Model (HMM) that obtains at least one observation in response to an input of the audio signals, and matching comprises assigning the at least one command to the at least one observation according to patterns extracted from the data.
Bang teaches:
wherein the circuitry is further configured to implement a Hidden Markov Model (HMM) that obtains at least one observation in response to an input of the audio signals, and matching comprises assigning the at least one command to the at least one observation according to patterns extracted from the data (Column 4, lines 48-55, "A weighted finite state transducer (WFST) unit or decoder 22 uses the acoustic scores to identify one or more utterance hypotheses and compute their scores. The WFST decoder 22 may use a neural network of arcs and states that also are referred to as WFSTs. The WFST neural network or other networks may be based on hidden Markov models (HMMs), dynamic time warping, and many other network techniques or structures."; Using acoustic scores to identify one or more utterance hypotheses and compute their scores using a neural network based on hidden Markov models reads on implementing a Hidden Markov Model (HMM) that obtains at least one observation in response to an input of the audio signals, and matching comprises assigning the at least one command to the at least one observation according to patterns extracted from the data.).
Bang is considered to be analogous to the claimed invention because it is in the same field of automatic speech recognition. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Smith to incorporate the teachings of Bang to use acoustic scores to identify one or more utterance hypotheses and compute their scores using a neural network based on hidden Markov models. Doing so would allow for recording audio with a microphone, processing the acoustic data with a speech processing device, and outputting speech or visual information to the user or other application to perform a further action depending on the audio information (Bang; Column 4, lines 6-11).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Han et al. (US Patent No. 12,424,203) teaches an electronic device for maintaining a voice recognition performance and effectively reducing the time or the processing amount spent for voice recognition.
Salarian et al. (US Patent No. 11,763,814) teaches a method for on-device automatic speech recognition (ASR) using a limited command set.
Lyon et al. (US Patent No. 10,672,387) teaches a method for recognizing speech, such as user commands, including receiving audio input data via one or more microphones, generating a feature vector for the audio input data, and obtaining recognized speech from the audio input utilizing the feature vector.
Mutagi et al. (US Patent No. 10,026,401) teaches a method for allowing users to associate functional identifiers with devices via voice commands.
Wu et al. ("A Framework for Off-Line Operation of Smart and Traditional Devices of IoT Services”) teaches a framework that ensures the reliability of Internet of Things (IoT) system services by providing the services offline.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to James Boggs whose telephone number is (571)272-2968. The examiner can normally be reached M-F 8:00 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571)272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAMES BOGGS/Examiner, Art Unit 2657