DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/15/26 has been entered.
Response to Arguments
Applicant’s arguments with respect to the 35 U.S.C. 103 rejections of claims 1-30 have been considered but are not persuasive.
Regarding Applicant’s argument presented on pages 13-14:
“However, even when combined with Weinstein, Ivanov's processing of "information about the conditions under which observation was made" does not amount to the claimed features. For instance, even if Ivanov's system is "trained in an environment," e.g., "acoustic environment type," Ivanov's "voice recognition configuration selector" does not use any "audio profiles" as claimed, such as by "comparing first characteristics of ambient noise of an audio input to one or more stored audio profiles". Weinstein (individually or when combined with Ivanov) also fails to describe such features.”
The examiner contests that Ivanov describes the claimed “audio profiles” and further describes a step of "comparing first characteristics of ambient noise of an audio input to one or more stored audio profiles".
See the operating environment information/condition that is associated with each voice recognition speech model, as described in ¶s [0019]-[0021]: “[0019] … Thus the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends. Otherwise, the training under the first condition training set is repeated in operation block 201.
[0020] The conditions will be selected so as to cover the intended use as much as possible. The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.), "trained with other factor(s)" such as those affecting the voice recognition engine, or combination thereof. In other words, a "condition" may be related to the training device, the training environment or the training signal conditioning including pre-processing applied to the audio signal.
[0021] In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends. The number of conditions and situations covered is limited by resource availability and can be extended as new conditions and needs are identified” (emphasis added).
Further, see ¶ [0022], which describes a step of comparing first characteristics of ambient noise of an audio input to one or more stored audio profiles.
¶ [0022]: “Once trained, the voice recognition system may operate as illustrated in FIG. 4 which illustrates a method of operation in accordance with various embodiments. In operation block 401, a pre-processing front end will collect a speech sample of interest, and operating-environment logic, in accordance with the embodiments, will measure and identify the condition under which the observation is made as shown in operation block 403. Data collected from the operating-environment logic will be combined with the speech sample and passed to the voice recognition system by, for example, an application programming interface (API) 411. In operation block 405, a voice recognition configuration selector will process the information about the conditions under which observation was made and will select the data-base best representing the condition in which the speech sample was obtained. The database identifier (DB ID 413) identifies the selected speech model from among the collection of databases 409. In operation block 407, the voice recognition engine will then use the selected speech model optimal for the current conditions and will process the sample of speech, after which it will return the result. The method of operation then returns to operation block 401”) (emphasis added).
Finally, see ¶s [0024], [0027], and [0028], which describe examples of operating-environment information (audio profiles) obtained by the operating-environment logic:
“[0024] … Examples of operating-environment information obtained by the operating-environment logic may include, but is not limited to, a device ID for device 610, the signal conditioning algorithm used, a noise environment ID, a signal quality indicator, noise level, signal-to-noise ratio, or other information such as impeding (reflective/absorptive) nearby surfaces, etc. This information may be obtained from the microphone signal pre-processing front end 120, the sensors 132, other dedicated measurement logic, or from network information sources.”
“[0027] Additional examples of the type of condition information that the operating-environment logic 130 may attempt to obtain include conditions such as, but not limited to, a) physical/electrical characteristics of the device; b) level, frequency and temporal characteristics of the desired speech source; c) location of the source with respect to the device and its surroundings; d) location and characteristics of interference sources; e) level, frequency and temporal characteristics of surrounding noise; f) reverberation present in the environment; g) physical location of the device (e.g. on table, hand-held, in-pocket etc.); or h) characteristics of signal enhancement algorithms. In other words, the condition may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples.”
“[0028] Additional examples of operating-environment information 133 sent by the operating-environment logic 130 to the voice recognition configuration selector 140 may include, but is not limited to, a) information to identify what device was used in the speech data observation (configuration decision can be based on selecting a database obtained with the device used, or one with similar characteristics); b) information identifying signal conditioning algorithms used, such as dynamic processors, filters, gain line-up, noise suppressor etc. (allowing determination to use a database trained with similar or identical signal conditioning); c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions); d) information identifying other characteristics of the external environment, affecting data observation such as presence of reflective/absorptive surfaces (portable laying on table, or car seat), high degree of reverberation (portable in highly reverberant/live environment, or on highly reflective surface); or e) information characterizing overall quality of signal, for example: low overall (or too high) signal level, frequency loss with specific characteristics etc. In other words, the operating-environment information 133 has information about at least one condition which may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples. The audio environment may be determined in a variety of ways, such as, but not limited to, collecting and aggregating sensor data from the sensors 132, using location information from location information logic 131, extracting audio environment data observed by the microphone signal pre-processing logic 120 or from other components of the device 610” (emphasis added).
Regarding Applicant’s argument presented on page 14:
“However, Weinstein (individually or when combined with Ivanov) is silent on any "known location" as claimed, much less "determining a type of location associated with where the audio input is received based on the first characteristics and the second characteristics," as claimed.”
The examiner contests that Weinstein describes the claimed network connection being associated with a known location and further describes a step of “determining a type of location associated with where the audio input is received based on the first characteristics and the second characteristics.”
See the description of context information 372, which includes location information 374 and an IP address 378, where wireless network data collected at the client device, such as cellular tower identifiers or Wi-Fi signatures, can be used to determine the location information, and where the IP address can also be used to determine the location information. Thus, the cellular network connection, Wi-Fi network connection, and IP address network connection are all associated with known locations.
Col. 14 lns. 47-63: “The example of FIG. 3B illustrates processing context information to generate inputs suitable for a neural network. The system 370 receives context information 372 that may be from a client device and/or network devices. The context information may include, among other things, location information 374, a user identifier 376, and an IP address 378.
The location information 374 may include a latitude and longitude from a GPS device located at the client device, and/or wireless network data collected at the client device such as cellular tower identifiers or Wi-Fi signatures. The location information 374 is provided as an input to a location resolver module 380 that correlates locations to regions having particular languages and/or accents. The IP address 378 can also generally be resolved to a location associated with the client device, although the location identified by the IP address 378 may not necessarily be the same as (or as accurate as) the location indicated by the location information 374” (emphasis added).
Further, see the description of the statistical classifier 174, which receives audio characteristics 128 and context information 126, and which carries out a step of determining a type of location (category) associated with where the audio input is received based on the first characteristics (audio characteristics 128) and the second characteristics (context information 126).
Col. 5 lns. 28-39: “The computing system 120 also obtains audio characteristics 128 from the audio signal 112. These audio characteristics 128 may be independent of the words spoken by the user 102. For example, the audio characteristics 128 may indicate audio features that correspond to one or more of background noise, recording channel properties, the speaker's speaking style, the speaker's gender, the speaker's age, and/or the speaker's accent. While the feature vectors 122 may be indicative of audio characteristics of specific portions of the particular words spoken, the audio characteristics 128 may therefore be indicative of time-independent characteristics of the audio signal.” (emphasis added).
Col. 3 ln. 58 – Col. 4 ln. 8: “In addition, the client device 110 may obtain various context information 114 contemporaneously with (e.g., shortly before, during, or shortly after) the user's utterance. The context information 114 is information about the user 102, the client device 110, and/or the circumstances in which the utterance was made, which was not derived from the audio signal 112 or another audio signal. For example, when the user 102 initiates a speech recognition session, the client device 110 may obtain context information 114 in response to the initiation of the session. The context information 114 may include, for example, a geographic location, an IP address, accent data about the user 102, and/or a search history of the user 102. To obtain a geographic location, for example, the client device 110 may obtain a GPS location of the client device. Alternatively or in addition, the client device 110 may obtain location information using cellular network data and/or Wi-Fi signals. The client device 110 may also obtain (e.g., retrieve from memory) its current IP address.” (emphasis added).
Col. 5 lns. 17-27: “The computing system 120 receives the audio signal 112 and context information 114 and obtains information about acoustic features of the audio signal 112. […] The computing system 120 also processes the received context information to obtain data 126 that is suitable for input into the neural network 140.” (emphasis added).
Col. 6 lns. 36-56: “FIG. 1B illustrates an example of a system 170 for speech recognition using statistical classifiers. The system 170 is similar to the system 100 described with reference to FIG. 1A. However, instead of a computing system 120 that provides feature vectors 122, context information 126, and audio characteristics 128 to a neural network 140, the system 170 includes a computing system 172 that inputs the context information 126 and optionally the audio characteristics 128 to a statistical classifier 174. The computing system 172 then selects a speech recognition model 176 based on the output of the statistical classifier 174. The speech recognition model 176 may be a particular speech recognition model that corresponds to a language and/or an accent of the speaker 102. For example, the output of the statistical classifier 174 may identify a particular speech recognition model from a collection of speech recognition models, each of which corresponds to a different language (e.g., English, Japanese, or Spanish), and/or accent (e.g., American English, British English, or Indian English). The collection of speech recognition models may be stored on the storage device local to the computing system 172, or may be external to the computing system 172.” (emphasis added).
Col. 6 ln. 57-col. 7 ln. 9: “In some implementations, the statistical classifier 174 may be, for example, a supervised learning model such as a multi-way support vector machine (SVM) or a logistic regression classifier that analyzes the inputs and makes a determination about which speech recognition model should be used to transcribe the particular audio signal 112. The statistical classifier 174 may be trained using a set of training examples, where each training example is marked as belonging to a category [maps to claimed type of location] corresponding to a particular speech recognition model. The statistical classifier 174 then builds a model that assigns new examples to each of the categories. In operation, the statistical classifier 174 receives the set of inputs corresponding to the context information 126 and optionally the audio characteristics 128, and classifies the inputs according to these categories. The classifications may be represented by probabilities or likelihoods that the inputs correspond to a particular speech recognition model. For example, the statistical classifier 174 can indicate an 80% probability that the inputs correspond to Brazilian Portuguese, an 18% probability of French, and a 2% probability of Italian.” (emphasis added).
Claim Objections
Claims 15 and 23 are objected to because of the following informalities: Both claims describe receiving an audio input; however, it’s not clear as to whether this is the same audio input that is discussed in claims 9 and 17, or if it’s an additional audio input (as described in amended claim 7).
Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 7-13, 15-21, 23, and 24 are rejected under 35 U.S.C. 103 as being unpatentable over Weinstein et al. (US 9,311,915, herein “Weinstein”) in view of Ivanov et al. (US 2014/0278415, herein “Ivanov”).
RE claims 1, 9, and 17, Weinstein describes a method of voice or speech recognition executed by a processor of a computing device, a computing device, and a non-transitory processor-readable medium having stored thereon processor-executable instructions configured to cause a processor of a computing device to perform operations, comprising:
a microphone (col. 3 lns. 13-18: “FIG. 1A illustrates an example of a system 100 for speech recognition using neural networks. The system 100 includes a client device 110, a computing system 120, and a network 130. In the system 100, a user 102 speaks an utterance into the client device 110, which generates an audio signal 112 encoding the utterance.”);
a memory (col. 19 lns. 35-48: “Embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable-medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter affecting a machine-readable propagated signal, or a combination of one or more of them.”); and
a processor coupled to the microphone and the memory (col. 19 lns. 50-62: “The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus may include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.”), and configured with processor-executable instructions to:
[identify] first characteristics of ambient noise of an audio input (See col. 5 lns. 28-39: “The computing system 120 also obtains audio characteristics 128 from the audio signal 112. These audio characteristics 128 may be independent of the words spoken by the user 102. For example, the audio characteristics 128 may indicate audio features that correspond to one or more of background noise, recording channel properties, the speaker's speaking style, the speaker's gender, the speaker's age, and/or the speaker's accent. While the feature vectors 122 may be indicative of audio characteristics of specific portions of the particular words spoken, the audio characteristics 128 may therefore be indicative of time-independent characteristics of the audio signal.” (emphasis added).
Also see col. 12 lns. 12-23: “FIG. 3A is a diagram 300 that illustrates an example of processing to generate latent variables of factor analysis. The example of FIG. 3A shows techniques for determining an i-vector, which includes these latent variables of factor analysis. I-vectors are time-independent components that represent overall characteristics of an audio signal rather than characteristics at a specific segment of time within an utterance. I-vectors can summarize a variety of characteristics of audio that are independent of the phonetic units spoken, for example, information indicative of the identity and/or accent of the speaker, the language spoken, recording channel properties, and noise characteristics.” (emphasis added));
identify second characteristics of a network connection of the computing device (See col. 3 ln. 58 – Col. 4 ln. 8: “In addition, the client device 110 may obtain various context information 114 contemporaneously with (e.g., shortly before, during, or shortly after) the user's utterance. The context information 114 is information about the user 102, the client device 110, and/or the circumstances in which the utterance was made, which was not derived from the audio signal 112 or another audio signal. For example, when the user 102 initiates a speech recognition session, the client device 110 may obtain context information 114 in response to the initiation of the session. The context information 114 may include, for example, a geographic location, an IP address, accent data about the user 102, and/or a search history of the user 102. To obtain a geographic location, for example, the client device 110 may obtain a GPS location of the client device. Alternatively or in addition, the client device 110 may obtain location information using cellular network data and/or Wi-Fi signals. The client device 110 may also obtain (e.g., retrieve from memory) its current IP address.” (emphasis added).
Also see col. 5 lns. 17-27: “The computing system 120 receives the audio signal 112 and context information 114 and obtains information about acoustic features of the audio signal 112. […] The computing system 120 also processes the received context information to obtain data 126 that is suitable for input into the neural network 140.” (emphasis added));
determine a type of location associated with where the audio input is received based on the first characteristics of the ambient noise and the second characteristics of the network connection (Col. 6 lns. 36-56: “FIG. 1B illustrates an example of a system 170 for speech recognition using statistical classifiers. The system 170 is similar to the system 100 described with reference to FIG. 1A. However, instead of a computing system 120 that provides feature vectors 122, context information 126, and audio characteristics 128 to a neural network 140, the system 170 includes a computing system 172 that inputs the context information 126 and optionally the audio characteristics 128 to a statistical classifier 174. The computing system 172 then selects a speech recognition model 176 based on the output of the statistical classifier 174. The speech recognition model 176 may be a particular speech recognition model that corresponds to a language and/or an accent of the speaker 102. For example, the output of the statistical classifier 174 may identify a particular speech recognition model from a collection of speech recognition models, each of which corresponds to a different language (e.g., English, Japanese, or Spanish), and/or accent (e.g., American English, British English, or Indian English). The collection of speech recognition models may be stored on the storage device local to the computing system 172, or may be external to the computing system 172.” (emphasis added).
Col. 6 ln. 57-col. 7 ln. 9: “In some implementations, the statistical classifier 174 may be, for example, a supervised learning model such as a multi-way support vector machine (SVM) or a logistic regression classifier that analyzes the inputs and makes a determination about which speech recognition model should be used to transcribe the particular audio signal 112. The statistical classifier 174 may be trained using a set of training examples, where each training example is marked as belonging to a category [maps to claimed type of location] corresponding to a particular speech recognition model. The statistical classifier 174 then builds a model that assigns new examples to each of the categories. In operation, the statistical classifier 174 receives the set of inputs corresponding to the context information 126 and optionally the audio characteristics 128, and classifies the inputs according to these categories. The classifications may be represented by probabilities or likelihoods that the inputs correspond to a particular speech recognition model. For example, the statistical classifier 174 can indicate an 80% probability that the inputs correspond to Brazilian Portuguese, an 18% probability of French, and a 2% probability of Italian.” (emphasis added).);
determine a voice recognition model from a plurality of voice recognition models to use for voice or speech recognition based on the type of location (See col. 7 lns. 17-24: “The computing system 172 then selects a speech recognition model based on the output of the statistical classifier 174. For example, the computing system 172 may select the speech recognition model having the highest probability. To continue the example above, the computing system 172 may select a speech recognition model corresponding to Brazilian Portuguese. In some cases, the computing system 172 may apply a minimum threshold probability to the selection.”); and
perform voice or speech recognition on the audio input using the determined voice recognition model (See col. 7 lns. 32-40: “The computing system 172 then generates a transcription from the acoustic feature vectors 122 using the selected speech recognition model. In some implementations, the speech recognition model may be external to the computing system 172, in which case the computing system 172 routes the acoustic feature vectors 122 to the appropriate speech recognition model and then receives a transcription in response. The transcription 160 may then be transmitted back to the client device 110 via the network 130.”).
Weinstein doesn’t describe a system or method including a step to compare first characteristics of ambient noise of an audio input to one or more stored audio profiles, wherein the one or more stored audio profiles are associated with a respective location.
However, Ivanov describes a system and method including a step to compare first characteristics of ambient noise of an audio input to one or more stored audio profiles, wherein the one or more stored audio profiles are associated with a respective location (See ¶ [0022]: “Once trained, the voice recognition system may operate as illustrated in FIG. 4 which illustrates a method of operation in accordance with various embodiments. In operation block 401, a pre-processing front end will collect a speech sample of interest, and operating-environment logic, in accordance with the embodiments, will measure and identify the condition under which the observation is made as shown in operation block 403. Data collected from the operating-environment logic will be combined with the speech sample and passed to the voice recognition system by, for example, an application programming interface (API) 411. In operation block 405, a voice recognition configuration selector will process the information about the conditions under which observation was made and will select the data-base best representing the condition in which the speech sample was obtained. The database identifier (DB ID 413) identifies the selected speech model from among the collection of databases 409. In operation block 407, the voice recognition engine will then use the selected speech model optimal for the current conditions and will process the sample of speech, after which it will return the result. The method of operation then returns to operation block 401”) (emphasis added).
Further, see “[0024] … Examples of operating-environment information obtained by the operating-environment logic may include, but is not limited to, a device ID for device 610, the signal conditioning algorithm used, a noise environment ID, a signal quality indicator, noise level, signal-to-noise ratio, or other information such as impeding (reflective/absorptive) nearby surfaces, etc. This information may be obtained from the microphone signal pre-processing front end 120, the sensors 132, other dedicated measurement logic, or from network information sources.”
“[0027] Additional examples of the type of condition information that the operating-environment logic 130 may attempt to obtain include conditions such as, but not limited to, a) physical/electrical characteristics of the device; b) level, frequency and temporal characteristics of the desired speech source; c) location of the source with respect to the device and its surroundings; d) location and characteristics of interference sources; e) level, frequency and temporal characteristics of surrounding noise; f) reverberation present in the environment; g) physical location of the device (e.g. on table, hand-held, in-pocket etc.); or h) characteristics of signal enhancement algorithms. In other words, the condition may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples.”
“[0028] Additional examples of operating-environment information 133 sent by the operating-environment logic 130 to the voice recognition configuration selector 140 may include, but is not limited to, a) information to identify what device was used in the speech data observation (configuration decision can be based on selecting a database obtained with the device used, or one with similar characteristics); b) information identifying signal conditioning algorithms used, such as dynamic processors, filters, gain line-up, noise suppressor etc. (allowing determination to use a database trained with similar or identical signal conditioning); c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions); d) information identifying other characteristics of the external environment, affecting data observation such as presence of reflective/absorptive surfaces (portable laying on table, or car seat), high degree of reverberation (portable in highly reverberant/live environment, or on highly reflective surface); or e) information characterizing overall quality of signal, for example: low overall (or too high) signal level, frequency loss with specific characteristics etc. In other words, the operating-environment information 133 has information about at least one condition which may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples. The audio environment may be determined in a variety of ways, such as, but not limited to, collecting and aggregating sensor data from the sensors 132, using location information from location information logic 131, extracting audio environment data observed by the microphone signal pre-processing logic 120 or from other components of the device 610” (emphasis added).).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Weinstein a system and method including a step to compare first characteristics of ambient noise of an audio input to one or more stored audio profiles, wherein the one or more stored audio profiles are associated with a respective location, as taught by Ivanov, in order to improve the quality of the speech recognition system by selecting a model that has been optimized for the current audio conditions (see Ivanov at ¶ [0012]).
RE claims 2, 10, and 18, Weinstein describes the method of claim 1, computing device of claim 9, and non-transitory processor-readable medium of claim 17, further comprising a global positioning system receiver (See col. 14 lns. 53-56: “The location information 374 may include a latitude and longitude from a GPS device located at the client device, and/or wireless network data collected at the client device such as cellular tower identifiers or Wi-Fi signatures.”),
wherein the processor is further configured with processor-executable instructions to use global positioning system information to determine the location where the audio input is received (See col. 3 ln. 67-col. 4 ln. 5: “The context information 114 may include, for example, a geographic location, an IP address, accent data about the user 102, and/or a search history of the user 102. To obtain a geographic location, for example, the client device 110 may obtain a GPS location of the client device.”),
and wherein the processor is configured to determine the type of location further based on the determined location (Col. 6 lns. 36-56: “FIG. 1B illustrates an example of a system 170 for speech recognition using statistical classifiers. The system 170 is similar to the system 100 described with reference to FIG. 1A. However, instead of a computing system 120 that provides feature vectors 122, context information 126, and audio characteristics 128 to a neural network 140, the system 170 includes a computing system 172 that inputs the context information 126 and optionally the audio characteristics 128 to a statistical classifier 174. The computing system 172 then selects a speech recognition model 176 based on the output of the statistical classifier 174. The speech recognition model 176 may be a particular speech recognition model that corresponds to a language and/or an accent of the speaker 102. For example, the output of the statistical classifier 174 may identify a particular speech recognition model from a collection of speech recognition models, each of which corresponds to a different language (e.g., English, Japanese, or Spanish), and/or accent (e.g., American English, British English, or Indian English). The collection of speech recognition models may be stored on the storage device local to the computing system 172, or may be external to the computing system 172.” (emphasis added).).
RE claims 3, 11, and 19, Weinstein doesn’t describe but Ivanov describes the method of claim 1, computing device of claim 9, and non-transitory processor-readable medium of claim 17, wherein the processor is further configured with the processor-executable instructions to
use the first characteristics of the ambient noise to determine the location where the audio input is received ([0024]: “Examples of operating-environment information obtained by the operating-environment logic may include, but is not limited to, a device ID for device 610, the signal conditioning algorithm used, a noise environment ID, a signal quality indicator, noise level, signal-to-noise ratio, or other information such as impeding (reflective/absorptive) nearby surfaces, etc. This information may be obtained from the microphone signal pre-processing front end 120, the sensors 132, other dedicated measurement logic, or from network information sources.”
[0028]: “c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions);”
[0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments.”).
See the rationale in the rejection of claims 1, 9, and 17, as it is equally applicable here.
RE claims 4, 12, and 20, Weinstein describes the method of claim 1, computing device of claim 9, and non-transitory processor-readable medium of claim 17, wherein the processor is further configured with processor-executable instructions to
use communication network information to determine the location where the audio input is received (See col. 3 ln. 67-col. 4 ln. 5: “The context information 114 may include, for example, a geographic location, an IP address, accent data about the user 102, and/or a search history of the user 102. To obtain a geographic location, for example, the client device 110 may obtain a GPS location of the client device.”
Also see col. 14 lns. 53-56: “The location information 374 may include a latitude and longitude from a GPS device located at the client device, and/or wireless network data collected at the client device such as cellular tower identifiers or Wi-Fi signatures.”).
RE claims 5, 13, and 21, Weinstein doesn’t describe but Ivanov describes the method of claim 1, computing device of claim 9, and non-transitory processor-readable medium of claim 17, wherein the processor is further configured with processor-executable instructions to determine a voice recognition model to use for voice or speech recognition by:
selecting the voice recognition model from the plurality of voice recognition models stored in the memory ([0022]: “In operation block 405, a voice recognition configuration selector will process the information about the conditions under which observation was made and will select the data-base best representing the condition in which the speech sample was obtained. The database identifier (DB ID 413) identifies the selected speech model from among the collection of databases 409. In operation block 407, the voice recognition engine will then use the selected speech model optimal for the current conditions and will process the sample of speech, after which it will return the result.”),
wherein each of the plurality of voice recognition models is associated with a different scene category each having a designated audio profile ([0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends.”).
See the rationale in the rejection of claims 1, 9, and 17, as it is equally applicable here.
RE claims 7, 15, and 23, Weinstein doesn’t describe but Ivanov describes the method of claim 1, computing device of claim 9, and non-transitory processor-readable medium of claim 17, wherein the processor is further configured with the processor-executable instructions to:
receive, via the microphone, an additional audio input/audio input associated with ambient noise sampling at the location (¶ [0016]: “The present disclosure also provides a device that includes a microphone signal pre-processing front end and operating-environment logic, operatively coupled to the microphone signal pre-processing front end, and operative to identify at least one condition related to pre-processing applied to obtained speech samples by the microphone signal pre-processing front end”.
Also see ¶ [0019]: “Turning to FIG. 2, a flowchart provides an example method of operation for speech model creation for a given processing condition. In one embodiment, a voice recognition system will be trained under a number of different conditions. The voice recognition system achieves optimal performance for observations obtained under the training condition, but not necessarily optimal if the observation came for another condition different than that used in training. Thus the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends. Otherwise, the training under the first condition training set is repeated in operation block 201.”);
associate the location or a location category with the received additional audio input/audio input (¶ [0020]: “The conditions will be selected so as to cover the intended use as much as possible. The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.)”
Also see ¶ [0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments.”); and
transmit the audio input and associated location or location category information to a remote computing device for generating the voice recognition model for the associated location or location category based on the received audio input (¶ [0021]: “FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends.”
Also see ¶ [0023]: “as shown in FIG. 5, voice recognition front end processing may be on a various mobile devices (e.g. smartphone 509, tablet 507, laptop 511, desktop computer 513 and PDA 505), while a networked server 501 is operative to process requests from the multiple front-ends, which be mobile devices, or other networked systems as shown in FIG. 5 (such as other computers, or embedded systems). In this example embodiment, the front-end will send packetized information containing speech and description of the conditions, over a network link 503 of a network 500 (such as the Internet) and will receive the response from the server 501, as illustrated in FIG. 5. Each user may represent a different condition as shown, such that the voice recognition configuration selector on server 501 may select different speech models according to each device's specific conditions including its pre-processing, etc.”
Further, see ¶ [0024]: “the operating environment logic 150 and the voice recognition configuration selector 140 may be located on the device, while the voice recognition logic 150 and voice recognition configuration database 160 are located on a server”).
See the rationale in the rejection of claims 1, 9, and 17, as it is equally applicable here.
RE claims 8, 16, and 24, Weinstein doesn’t describe but Ivanov describes the method of claim 1, computing device of claim 9, and non-transitory processor-readable medium of claim 17, wherein the processor is further configured with the processor-executable instructions to:
compile an audio profile from an audio input associated with ambient noise at the location (¶ [0019]: “Turning to FIG. 2, a flowchart provides an example method of operation for speech model creation for a given processing condition. In one embodiment, a voice recognition system will be trained under a number of different conditions. The voice recognition system achieves optimal performance for observations obtained under the training condition, but not necessarily optimal if the observation came for another condition different than that used in training. Thus the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends.”
Also see ¶ [0021]: “FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends.”
Further, see ¶ [0024] “Examples of operating-environment information obtained by the operating-environment logic may include, but is not limited to, a device ID for device 610, the signal conditioning algorithm used, a noise environment ID, a signal quality indicator, noise level, signal-to-noise ratio, or other information such as impeding (reflective/absorptive) nearby surfaces, etc. This information may be obtained from the microphone signal pre-processing front end 120, the sensors 132, other dedicated measurement logic, or from network information sources.”
¶ [0027]: “Additional examples of the type of condition information that the operating-environment logic 130 may attempt to obtain include conditions such as, but not limited to, a) physical/electrical characteristics of the device; b) level, frequency and temporal characteristics of the desired speech source; c) location of the source with respect to the device and its surroundings; d) location and characteristics of interference sources; e) level, frequency and temporal characteristics of surrounding noise; f) reverberation present in the environment; g) physical location of the device (e.g. on table, hand-held, in-pocket etc.); or h) characteristics of signal enhancement algorithms. In other words, the condition may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples.”
¶ [0028]: “Additional examples of operating-environment information 133 sent by the operating-environment logic 130 to the voice recognition configuration selector 140 may include, but is not limited to, a) information to identify what device was used in the speech data observation (configuration decision can be based on selecting a database obtained with the device used, or one with similar characteristics); b) information identifying signal conditioning algorithms used, such as dynamic processors, filters, gain line-up, noise suppressor etc. (allowing determination to use a database trained with similar or identical signal conditioning); c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions); d) information identifying other characteristics of the external environment, affecting data observation such as presence of reflective/absorptive surfaces (portable laying on table, or car seat), high degree of reverberation (portable in highly reverberant/live environment, or on highly reflective surface)” (emphasis added).);
associate the location or a location category with the compiled audio profile (¶ [0020]: “The conditions will be selected so as to cover the intended use as much as possible. The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.)”
Also see ¶ [0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments.”); and
transmit the audio profile associated with the location or location category to a remote computing device for generating the voice recognition model for the location or location category based on the compiled audio profile (¶ [0021]: “FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends.”
Also see ¶ [0023]: “as shown in FIG. 5, voice recognition front end processing may be on a various mobile devices (e.g. smartphone 509, tablet 507, laptop 511, desktop computer 513 and PDA 505), while a networked server 501 is operative to process requests from the multiple front-ends, which be mobile devices, or other networked systems as shown in FIG. 5 (such as other computers, or embedded systems). In this example embodiment, the front-end will send packetized information containing speech and description of the conditions, over a network link 503 of a network 500 (such as the Internet) and will receive the response from the server 501, as illustrated in FIG. 5. Each user may represent a different condition as shown, such that the voice recognition configuration selector on server 501 may select different speech models according to each device's specific conditions including its pre-processing, etc.”
Further, see ¶ [0024]: “the operating environment logic 150 and the voice recognition configuration selector 140 may be located on the device, while the voice recognition logic 150 and voice recognition configuration database 160 are located on a server”).
See the rationale in the rejection of claims 1, 9, and 17, as it is equally applicable here.
Claim(s) 6, 14, and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Weinstein in view of Ivanov, as applied to claims 1, 9, and 17 above, and further in view of Miyazawa (US 2004/0138882).
RE claims 6, 14, and 22, Weinstein in view of Ivanov doesn’t explicitly describe the method of claim 1, computing device of claim 9, and non-transitory processor-readable medium of claim 17, wherein the processor is further configured with process-executable instructions to perform voice or speech recognition on the audio input using the determined voice recognition model by: using the determined voice recognition model to adjust the audio input for ambient noise; and performing voice and/or speech recognition on the adjusted audio input.
However, Miyazawa describes a system and method including using the determined voice recognition model to adjust the audio input for ambient noise (¶ [0027]: “the speech recognition apparatus of the present invention performs the noise data determination for determining which noise data of the plural types of noise data corresponds to the current noise. The noise removal is performed on the noise-superposed speech data based on the result of determination of the noise data. And then, the speech recognition is performed on the noise-removed speech using the acoustic model corresponding to the noise data.”); and
performing voice and/or speech recognition on the adjusted audio input (¶ [0027]: “the speech recognition apparatus of the present invention performs the noise data determination for determining which noise data of the plural types of noise data corresponds to the current noise. The noise removal is performed on the noise-superposed speech data based on the result of determination of the noise data. And then, the speech recognition is performed on the noise-removed speech using the acoustic model corresponding to the noise data.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Weinstein in view of Ivanov a system and method wherein the processor is further configured with process-executable instructions to perform voice or speech recognition on the audio input using the determined voice recognition model by: using the determined voice recognition model to adjust the audio input for ambient noise; and performing voice and/or speech recognition on the adjusted audio input, as taught by Miyazawa, in order to generate very accurate recognized speech, by first removing noise from the speech signal, and then performing speech recognition on the noise-removed speech.
Claim(s) 25-30 are rejected under 35 U.S.C. 103 as being unpatentable over Ivanov in view of Weinstein, and further in view of Yoshizawa (US 2003/0050783).
RE claim 25, Ivanov describes a method performed by a computing device for generating a speech recognition model, comprising:
receiving, from user equipment remote from the computing device, an audio input and location information associated with a location where the audio input was recorded ([0023]: “as shown in FIG. 5, voice recognition front end processing may be on a various mobile devices (e.g. smartphone 509, tablet 507, laptop 511, desktop computer 513 and PDA 505), while a networked server 501 is operative to process requests from the multiple front-ends, which be mobile devices, or other networked systems as shown in FIG. 5 (such as other computers, or embedded systems). In this example embodiment, the front-end will send packetized information containing speech and description of the conditions, over a network link 503 of a network 500 (such as the Internet) and will receive the response from the server 501, as illustrated in FIG. 5. Each user may represent a different condition as shown, such that the voice recognition configuration selector on server 501 may select different speech models according to each device's specific conditions” Also see [0027] and [0028], which describe location information),
wherein the audio input includes ambient noise associated with a type of location where the audio input is received [0022]:” In operation block 401, a pre-processing front end will collect a speech sample of interest, and operating-environment logic, in accordance with the embodiments, will measure and identify the condition under which the observation is made as shown in operation block 403. Data collected from the operating-environment logic will be combined with the speech sample and passed to the voice recognition system by, for example, an application programming interface (API) 411. In operation block 405, a voice recognition configuration selector will process the information about the conditions under which observation was made and will select the data-base best representing the condition in which the speech sample was obtained. The database identifier (DB ID 413) identifies the selected speech model from among the collection of databases 409. In operation block 407, the voice recognition engine will then use the selected speech model optimal for the current conditions and will process the sample of speech, after which it will return the result.” Also see [0027], which describes characteristics of ambient noise: “d) location and characteristics of interference sources; e) level, frequency and temporal characteristics of surrounding noise; f) reverberation present in the environment;” and [0028], which describes determining various types of locations: “c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions);”);
comparing characteristics of the ambient noise to one or more stored audio profiles, wherein the one or more stored audio profiles are associated with a respective location ([0019]: “Thus the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends. Otherwise, the training under the first condition training set is repeated in operation block 201.” (emphasis added).
[0020]: “The conditions will be selected so as to cover the intended use as much as possible. The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.), "trained with other factor(s)" such as those affecting the voice recognition engine, or combination thereof. In other words, a "condition" may be related to the training device, the training environment or the training signal conditioning including pre-processing applied to the audio signal.” (emphasis added).
[0021] “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends. The number of conditions and situations covered is limited by resource availability and can be extended as new conditions and needs are identified” (emphasis added).
Further, see ¶ [0022], which describes a step of comparing first characteristics of ambient noise of an audio input to one or more stored audio profiles, wherein the one or more stored audio profiles are associated with a respective location.
¶ [0022]: “Once trained, the voice recognition system may operate as illustrated in FIG. 4 which illustrates a method of operation in accordance with various embodiments. In operation block 401, a pre-processing front end will collect a speech sample of interest, and operating-environment logic, in accordance with the embodiments, will measure and identify the condition under which the observation is made as shown in operation block 403. Data collected from the operating-environment logic will be combined with the speech sample and passed to the voice recognition system by, for example, an application programming interface (API) 411. In operation block 405, a voice recognition configuration selector will process the information about the conditions under which observation was made and will select the data-base best representing the condition in which the speech sample was obtained. The database identifier (DB ID 413) identifies the selected speech model from among the collection of databases 409. In operation block 407, the voice recognition engine will then use the selected speech model optimal for the current conditions and will process the sample of speech, after which it will return the result. The method of operation then returns to operation block 401”) (emphasis added).
Finally, see ¶s [0024], [0027], and [0028], which describe examples of operating-environment information (audio profiles) obtained by the operating-environment logic:
“[0024] … Examples of operating-environment information obtained by the operating-environment logic may include, but is not limited to, a device ID for device 610, the signal conditioning algorithm used, a noise environment ID, a signal quality indicator, noise level, signal-to-noise ratio, or other information such as impeding (reflective/absorptive) nearby surfaces, etc. This information may be obtained from the microphone signal pre-processing front end 120, the sensors 132, other dedicated measurement logic, or from network information sources.”
“[0027] Additional examples of the type of condition information that the operating-environment logic 130 may attempt to obtain include conditions such as, but not limited to, a) physical/electrical characteristics of the device; b) level, frequency and temporal characteristics of the desired speech source; c) location of the source with respect to the device and its surroundings; d) location and characteristics of interference sources; e) level, frequency and temporal characteristics of surrounding noise; f) reverberation present in the environment; g) physical location of the device (e.g. on table, hand-held, in-pocket etc.); or h) characteristics of signal enhancement algorithms. In other words, the condition may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples.”
“[0028] Additional examples of operating-environment information 133 sent by the operating-environment logic 130 to the voice recognition configuration selector 140 may include, but is not limited to, a) information to identify what device was used in the speech data observation (configuration decision can be based on selecting a database obtained with the device used, or one with similar characteristics); b) information identifying signal conditioning algorithms used, such as dynamic processors, filters, gain line-up, noise suppressor etc. (allowing determination to use a database trained with similar or identical signal conditioning); c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions); d) information identifying other characteristics of the external environment, affecting data observation such as presence of reflective/absorptive surfaces (portable laying on table, or car seat), high degree of reverberation (portable in highly reverberant/live environment, or on highly reflective surface); or e) information characterizing overall quality of signal, for example: low overall (or too high) signal level, frequency loss with specific characteristics etc. In other words, the operating-environment information 133 has information about at least one condition which may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples. The audio environment may be determined in a variety of ways, such as, but not limited to, collecting and aggregating sensor data from the sensors 132, using location information from location information logic 131, extracting audio environment data observed by the microphone signal pre-processing logic 120 or from other components of the device 610” (emphasis added));
using the characteristics of the ambient noise of the audio input the received audio input to generate a voice recognition model associated with the location for use in voice and/or speech recognition ([0019]: “the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends.”
Also see [0020]: “The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.), "trained with other factor(s)" such as those affecting the voice recognition engine, or combination thereof. In other words, a "condition" may be related to the training device, the training environment or the training signal conditioning including pre-processing applied to the audio signal”
And [0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends.”)
.
Ivanov doesn’t describe a system or method wherein the audio input includes ambient noise associated with a type of location where the audio input is received and wherein the location information is associated with a local network connection of the user equipment.
However, Weinstein describes a system and method wherein the audio input includes ambient noise associated with a type of location where the audio input is received and wherein the location information is associated with a local network connection of the user equipment (see FIG. 6 and col. 18 lns. 43-50: “In step 602, the computing system receives context information associated with an utterance. The utterance is encoded as an audio signal that is also received by the computing system. The context information may include, for example, an IP address of the client device from which the audio signal originated, a geographic location of the client device which the audio signal originated, and/or a search history associated with the speaker of the utterance.” (emphasis added)
Also see col. 18 ln. 59 – col. 19 ln. 1: “Optionally, in step 604, the computing system receives a set of data corresponding to time-independent characteristics of the audio signal that is derived from the received audio signal and/or another audio signal. This set of data may be data indicative of latent variables of multivariate factor analysis. In some implementations, this set of data may be [an] i-vector. The computing system then provides the set of data derived from the audio signal along with the data corresponding to the audio signal in the context information as inputs to the neural network.” (emphasis added)
Further, see col. 19 lns. 13-15: “Then, in step 606, the computing system provides the context information and optionally the time-independent characteristics of the audio signal to a statistical classifier.”
Still further, see col. 19 lns. 21-22: “Finally, in step 608, the computing system selects a speech recognizer based on the output of the statistical classifier.”
Additionally, Weinstein provides further detail regarding the context information at col. 14 lns. 47-67: “The example of FIG. 3B illustrates processing context information to generate inputs suitable for a neural network. The system 370 receives context information 372 that may be from a client device and/or network devices. The context information may include, among other things, location information 374, a user identifier 376, and an IP address 378.
The location information 374 may include a latitude and longitude from a GPS device located at the client device, and/or wireless network data collected at the client device such as cellular tower identifiers or Wi-Fi signatures. The location information 374 is provided as an input to a location resolver module 380 that correlates locations to regions having particular languages and/or accents. The IP address 378 can also generally be resolved to a location associated with the client device, although the location identified by the IP address 378 may not necessarily be the same as (or as accurate as) the location indicated by the location information 374. However, the IP address may indicate a place of origin of a client device, which may more accurately correlate with the user's probabl[e] language and/or accent than the current location of the client device.” (emphasis added)
Finally, Weinstein also provides further detail regarding the i-vector at col. 12 lns. 12-23: “FIG. 3A is a diagram 300 that illustrates an example of processing to generate latent variables of factor analysis. The example of FIG. 3A shows techniques for determining an i-vector, which includes these latent variables of factor analysis. I-vectors are time-independent components that represent overall characteristics of an audio signal rather than characteristics at a specific segment of time within an utterance. I-vectors can summarize a variety of characteristics of audio that are independent of the phonetic units spoken, for example, information indicative of the identity and/or accent of the speaker, the language spoken, recording channel properties, and noise characteristics.”) (emphasis added).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Ivanov a system and method wherein the audio input includes ambient noise associated with a type of location where the audio input is received and wherein the location information is associated with a local network connection of the user equipment, as taught by Weinstein, in order to use the user’s location information to identify relevant languages and/or accents, which improves the voice recognition selection process, as it includes the additional selection criteria of selecting a voice recognition model that is best suited for a particular language and/or accent.
Ivanov in view of Weinstein doesn’t describe a system or method including a step to provide the generated voice recognition model associated with the location to the user equipment. However, Yoshizawa describes a system and method including a step to provide the generated voice recognition model associated with the location to the user equipment (¶ [0014]: “According to one aspect of the present invention, a terminal device includes a transmitting means, a receiving means, a first storage means, and a speech recognition means. The transmitting means transmits a voice produced by a user and environmental noises to a server device. The receiving means receives from the server device an acoustic model adapted to the voice of the user and the environmental noises. The first storage means stores the acoustic model received by the receiving means. The speech recognition means conducts speech recognition using the acoustic model stored in the first storage means.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Ivanov in view of Weinstein a system and method including a step to provide the generated voice recognition model associated with the location to the user equipment, as taught by Yoshizawa, in order to dynamically provide the best voice recognition model to the user equipment based on environmental noise conditions, which enables accurate speech recognition while reducing the memory requirements of the user equipment (Yoshizawa at ¶ [0015]).
RE claims 26 and 29, Ivanov describes the method of claim 25 and computing device of claim 28, wherein the processor is further configured with processor-executable instructions to:
receive the audio input and location information further comprises receiving a plurality of audio inputs, each having location information associated with different locations ([0023]: “as shown in FIG. 5, voice recognition front end processing may be on a various mobile devices (e.g. smartphone 509, tablet 507, laptop 511, desktop computer 513 and PDA 505), while a networked server 501 is operative to process requests from the multiple front-ends, which be mobile devices, or other networked systems as shown in FIG. 5 (such as other computers, or embedded systems). In this example embodiment, the front-end will send packetized information containing speech and description of the conditions, over a network link 503 of a network 500 (such as the Internet) and will receive the response from the server 501, as illustrated in FIG. 5. Each user may represent a different condition as shown, such that the voice recognition configuration selector on server 501 may select different speech models according to each device's specific conditions including its pre-processing, etc.”); and
use the received audio input to generate a voice recognition model associated with the location further comprises using the received plurality of audio inputs to generate voice recognition models, wherein each of the generated voice recognition models is configured to be used at a respective one of the different locations ([0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends. The number of conditions and situations covered is limited by resource availability and can be extended as new conditions and needs are identified.”).
RE claims 27 and 30, Ivanov describes the method of claim 25 and the computing device of claim 28, wherein the processor is further configured with processor-executable instructions to:
determine a location category based on the location information received from the user equipment ([0019]: “the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends.”
[0020]: “The conditions will be selected so as to cover the intended use as much as possible. The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.)”); and
associate the generated voice recognition model with the determined location category ([0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends. The number of conditions and situations covered is limited by resource availability and can be extended as new conditions and needs are identified.”).
RE claim 28, Ivanov describes a computing device, comprising:
a processor (FIG. 6 and ¶ [0026]: “The operating-environment logic 130, the voice recognition configuration selector 140 or microphone signal pre-processing front end may be implemented in various ways such as by software and/or firmware executing on one or more programmable processors”) configured with processor-executable instructions to:
receive, from user equipment remote from the computing device, an audio input and location information associated with a location where the audio input was recorded ([0023]: “as shown in FIG. 5, voice recognition front end processing may be on a various mobile devices (e.g. smartphone 509, tablet 507, laptop 511, desktop computer 513 and PDA 505), while a networked server 501 is operative to process requests from the multiple front-ends, which be mobile devices, or other networked systems as shown in FIG. 5 (such as other computers, or embedded systems). In this example embodiment, the front-end will send packetized information containing speech and description of the conditions, over a network link 503 of a network 500 (such as the Internet) and will receive the response from the server 501, as illustrated in FIG. 5. Each user may represent a different condition as shown, such that the voice recognition configuration selector on server 501 may select different speech models according to each device's specific conditions” Also see [0027] and [0028], which describe location information),
wherein the audio input includes ambient noise associated with a type of location where the audio input is received [0022]:” In operation block 401, a pre-processing front end will collect a speech sample of interest, and operating-environment logic, in accordance with the embodiments, will measure and identify the condition under which the observation is made as shown in operation block 403. Data collected from the operating-environment logic will be combined with the speech sample and passed to the voice recognition system by, for example, an application programming interface (API) 411. In operation block 405, a voice recognition configuration selector will process the information about the conditions under which observation was made and will select the data-base best representing the condition in which the speech sample was obtained. The database identifier (DB ID 413) identifies the selected speech model from among the collection of databases 409. In operation block 407, the voice recognition engine will then use the selected speech model optimal for the current conditions and will process the sample of speech, after which it will return the result.” Also see [0027], which describes characteristics of ambient noise: “d) location and characteristics of interference sources; e) level, frequency and temporal characteristics of surrounding noise; f) reverberation present in the environment;” and [0028], which describes determining various types of locations: “c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions);”); [and]
use characteristics of the ambient noise of the audio input the received audio input to generate a voice recognition model associated with the location for use in voice and/or speech recognition ([0019]: “the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends.”
Also see [0020]: “The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.), "trained with other factor(s)" such as those affecting the voice recognition engine, or combination thereof. In other words, a "condition" may be related to the training device, the training environment or the training signal conditioning including pre-processing applied to the audio signal”
And [0021]: “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends.”);
associate the generated voice recognition model with the type of location and an audio profile ( [0019]: “Thus the method of operation begins and in operation block 201, voice recognition engine is trained with a training set under a first condition. In operation block 203, the voice recognition engine is tested with inputs obtained under the first condition. The inputs may or may not include the data used during training. If the test is successful in decision block 205, then the model for the first condition is stored in operation block 207 and the method of operation ends. Otherwise, the training under the first condition training set is repeated in operation block 201.” (emphasis added).
[0020]: “The conditions will be selected so as to cover the intended use as much as possible. The condition may be identified as, for example, "trained on device X" (i.e. a given device type and model), "trained in environment Y" (i.e. noise type/level, acoustic environment type, etc.), "trained with signal conditioning Z" (specifying any relevant pre-processing such as, for example, gain settings, noise reduction applied, etc.), "trained with other factor(s)" such as those affecting the voice recognition engine, or combination thereof. In other words, a "condition" may be related to the training device, the training environment or the training signal conditioning including pre-processing applied to the audio signal.” (emphasis added).
[0021] “In one example, the voice recognition system can be trained on a given mobile device with signal conditioning algorithms turned off in multiple environments (such as in a car, restaurant, airport, etc.), and with signal conditioning enabled in the same environments. Each time a speech-model data-base ensuring optimal voice recognition performance is obtained and stored. FIG. 3 provides an example of such a method of operation for database creation for a set of processing conditions in various environments. As shown in operation block 301, a model is obtained under a first condition, then under a second condition in operation block 303, and so on, until an Nth condition in operation block 305 at which point the method of operation ends. The number of conditions and situations covered is limited by resource availability and can be extended as new conditions and needs are identified” (emphasis added).
Further, see ¶s [0024], [0027], and [0028], which describe examples of operating-environment information (audio profiles) obtained by the operating-environment logic:
“[0024] … Examples of operating-environment information obtained by the operating-environment logic may include, but is not limited to, a device ID for device 610, the signal conditioning algorithm used, a noise environment ID, a signal quality indicator, noise level, signal-to-noise ratio, or other information such as impeding (reflective/absorptive) nearby surfaces, etc. This information may be obtained from the microphone signal pre-processing front end 120, the sensors 132, other dedicated measurement logic, or from network information sources.”
“[0027] Additional examples of the type of condition information that the operating-environment logic 130 may attempt to obtain include conditions such as, but not limited to, a) physical/electrical characteristics of the device; b) level, frequency and temporal characteristics of the desired speech source; c) location of the source with respect to the device and its surroundings; d) location and characteristics of interference sources; e) level, frequency and temporal characteristics of surrounding noise; f) reverberation present in the environment; g) physical location of the device (e.g. on table, hand-held, in-pocket etc.); or h) characteristics of signal enhancement algorithms. In other words, the condition may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples.”
“[0028] Additional examples of operating-environment information 133 sent by the operating-environment logic 130 to the voice recognition configuration selector 140 may include, but is not limited to, a) information to identify what device was used in the speech data observation (configuration decision can be based on selecting a database obtained with the device used, or one with similar characteristics); b) information identifying signal conditioning algorithms used, such as dynamic processors, filters, gain line-up, noise suppressor etc. (allowing determination to use a database trained with similar or identical signal conditioning); c) information identifying noise environment, in terms of characteristics such as stationary/non-stationary, car, babble, airport, level, signal-to-noise ratio etc. (allowing determination to use database trained under similar conditions); d) information identifying other characteristics of the external environment, affecting data observation such as presence of reflective/absorptive surfaces (portable laying on table, or car seat), high degree of reverberation (portable in highly reverberant/live environment, or on highly reflective surface); or e) information characterizing overall quality of signal, for example: low overall (or too high) signal level, frequency loss with specific characteristics etc. In other words, the operating-environment information 133 has information about at least one condition which may be related to pre-processing applied to obtained speech samples by the microphone signal pre-processing logic 120 or may be related to an audio environment of the obtained speech samples. The audio environment may be determined in a variety of ways, such as, but not limited to, collecting and aggregating sensor data from the sensors 132, using location information from location information logic 131, extracting audio environment data observed by the microphone signal pre-processing logic 120 or from other components of the device 610” (emphasis added))
.
Ivanov doesn’t describe a system or method wherein the audio input includes ambient noise associated with a type of location where the audio input is received and wherein the location information is associated with a local network connection of the user equipment, wherein the network connection is associated with a known location.
However, Weinstein describes a system and method wherein the audio input includes ambient noise associated with a type of location where the audio input is received and wherein the location information is associated with a local network connection of the user equipment, wherein the network connection is associated with a known location (See col. 3 ln. 58 – Col. 4 ln. 8: “In addition, the client device 110 may obtain various context information 114 contemporaneously with (e.g., shortly before, during, or shortly after) the user's utterance. The context information 114 is information about the user 102, the client device 110, and/or the circumstances in which the utterance was made, which was not derived from the audio signal 112 or another audio signal. For example, when the user 102 initiates a speech recognition session, the client device 110 may obtain context information 114 in response to the initiation of the session. The context information 114 may include, for example, a geographic location, an IP address, accent data about the user 102, and/or a search history of the user 102. To obtain a geographic location, for example, the client device 110 may obtain a GPS location of the client device. Alternatively or in addition, the client device 110 may obtain location information using cellular network data and/or Wi-Fi signals. The client device 110 may also obtain (e.g., retrieve from memory) its current IP address.” (emphasis added).
Further, see FIG. 6 and col. 18 lns. 43-50: “In step 602, the computing system receives context information associated with an utterance. The utterance is encoded as an audio signal that is also received by the computing system. The context information may include, for example, an IP address of the client device from which the audio signal originated, a geographic location of the client device which the audio signal originated, and/or a search history associated with the speaker of the utterance.” (emphasis added)
Also see col. 18 ln. 59 – col. 19 ln. 1: “Optionally, in step 604, the computing system receives a set of data corresponding to time-independent characteristics of the audio signal that is derived from the received audio signal and/or another audio signal. This set of data may be data indicative of latent variables of multivariate factor analysis. In some implementations, this set of data may be [an] i-vector. The computing system then provides the set of data derived from the audio signal along with the data corresponding to the audio signal in the context information as inputs to the neural network.” (emphasis added)
Further, see col. 19 lns. 13-15: “Then, in step 606, the computing system provides the context information and optionally the time-independent characteristics of the audio signal to a statistical classifier.”
Still further, see col. 19 lns. 21-22: “Finally, in step 608, the computing system selects a speech recognizer based on the output of the statistical classifier.”
Additionally, Weinstein provides further detail regarding the context information at col. 14 lns. 47-67: “The example of FIG. 3B illustrates processing context information to generate inputs suitable for a neural network. The system 370 receives context information 372 that may be from a client device and/or network devices. The context information may include, among other things, location information 374, a user identifier 376, and an IP address 378.
The location information 374 may include a latitude and longitude from a GPS device located at the client device, and/or wireless network data collected at the client device such as cellular tower identifiers or Wi-Fi signatures. The location information 374 is provided as an input to a location resolver module 380 that correlates locations to regions having particular languages and/or accents. The IP address 378 can also generally be resolved to a location associated with the client device, although the location identified by the IP address 378 may not necessarily be the same as (or as accurate as) the location indicated by the location information 374. However, the IP address may indicate a place of origin of a client device, which may more accurately correlate with the user's probabl[e] language and/or accent than the current location of the client device.” (emphasis added)
Finally, Weinstein also provides further detail regarding the i-vector at col. 12 lns. 12-23: “FIG. 3A is a diagram 300 that illustrates an example of processing to generate latent variables of factor analysis. The example of FIG. 3A shows techniques for determining an i-vector, which includes these latent variables of factor analysis. I-vectors are time-independent components that represent overall characteristics of an audio signal rather than characteristics at a specific segment of time within an utterance. I-vectors can summarize a variety of characteristics of audio that are independent of the phonetic units spoken, for example, information indicative of the identity and/or accent of the speaker, the language spoken, recording channel properties, and noise characteristics.”) (emphasis added).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Ivanov a system and method wherein the audio input includes ambient noise associated with a type of location where the audio input is received and wherein the location information is associated with a local network connection of the user equipment, wherein the network connection is associated with a known location, as taught by Weinstein, in order to use the user’s location information to identify relevant languages and/or accents, which improves the voice recognition selection process, as it includes the additional selection criteria of selecting a voice recognition model that is best suited for a particular language and/or accent.
Ivanov in view of Weinstein doesn’t describe a system or method including a step to provide the generated voice recognition model associated with the location to the user equipment. However, Yoshizawa describes a system and method including a step to provide the generated voice recognition model associated with the location to the user equipment (¶ [0014]: “According to one aspect of the present invention, a terminal device includes a transmitting means, a receiving means, a first storage means, and a speech recognition means. The transmitting means transmits a voice produced by a user and environmental noises to a server device. The receiving means receives from the server device an acoustic model adapted to the voice of the user and the environmental noises. The first storage means stores the acoustic model received by the receiving means. The speech recognition means conducts speech recognition using the acoustic model stored in the first storage means.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include in Ivanov in view of Weinstein a system and method including a step to provide the generated voice recognition model associated with the location to the user equipment, as taught by Yoshizawa, in order to dynamically provide the best voice recognition model to the user equipment based on environmental noise conditions, which enables accurate speech recognition while reducing the memory requirements of the user equipment (Yoshizawa at ¶ [0015]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Kim et al. (US 10,770,065)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daniel C Washburn whose telephone number is (571)272-5551. The examiner can normally be reached Monday-Friday 9:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657