DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Applicant’s Arguments and Amendments
This communication is in response to amendment on 05/04/2026. Applicant filed an amendment on 05/04/2026., amending independent claims 1, 11 and 12.
Regarding interpretation under 35 U.S.C. §112(f) removed based on amendment.
Regarding rejection under 35 U.S.C. §101, applicant amendment does not overcome 101. Microphone is just collecting data which is pre-solution activity. Claim does not recite any additional elements which can overcome 101. Applicant mentioned about improvement (Remark page 11, 4th para.), but examiner does not see claim recites improvements. Please see updated 101 rejection.
Applicant’s assertion:
(“Thus, at least because of the improved accuracy in identifying the utterer and uttered
content, and thereby accurately controlling an appliance via voices by outputting a control
command including the identified utterer and content, to the appliance, as recited in claim 1,
Applicant respectfully submits that claim 1, as a whole, improves another technology or
technical field (e.g., control of an appliance by voice), and accordingly, integrates the judicial
exception (if any) into a practical application, and thus, is eligible under Prong II of Step 2A.”) {Remark page 12, first para}
Examiner Note: Examiner respectfully disagrees with this assertion. Claim recites “outputting the output data to an appliance, wherein the registered utterance content includes a control command to control the appliance.” Outputting step is pe-solution activity and the claims use language of “to control the appliance”. This is an intended use of the output.
Regarding 35 U.S.C. §103, examiner changes ground of rejection necessitated by claim amendment. Applicant's arguments have been fully considered but they are moot because examiner used new prior art for amended claim limitation.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1- 12 are rejected under 35 U.S.C. § 101 because the claims are directed to a judicial exception (an abstract idea) and do not recite additional elements that amount to significantly more.
Summary of the statutory framework and guidance applied
The Office applies the two-step framework for subject matter eligibility consistent with current USPTO guidance.
Step 1: Determine whether the claim recites a statutory category (process, machine, manufacture, or composition of matter).
Step 2A, Prong One: If the claim recites a statutory category, determine whether the claim recites a judicial exception (abstract idea, law of nature, or natural phenomenon).
Step 2A, Prong Two: If a judicial exception is recited, determine whether the claim integrates the exception into a practical application.
Step 2B: If the claim does not integrate the exception into a practical application, determine whether the claim recites additional elements that amount to significantly more than the judicial exception.
Step 1 (statutory category)
Claims 1, 2–10 are directed to a method (process).
Claims 11 are directed to a device (machine).
Claim 12 is directed to a non-transitory computer-readable medium (manufacture). Each of these falls within a statutory category.
Step 2A, Prong One (judicial exception)
The claims recite selection of phoneme sequences from certain speaker, selection of databases associated with specific-speaker and selected database based on calculated differences. These elements constitute a judicial exception in the form of abstract mathematical concepts and data processing.
Step 2A, Prong Two (integration into a practical application)
The claims do not integrate the judicial exception into a practical application. The additional elements recited—generic computing components (processors, memory, computing device),—are described at a high level and perform conventional functions of receiving, inputting, and outputting data. The claims do not describe a specific improvement to the functioning of the computer or another technological improvement. They instead use generic computer implementation to carry out the abstract mathematical operations.
Step 2B (significantly more)
The claims recite no additional elements or combination of elements that amount to significantly more than the abstract idea. The claimed processors, memories, device, are well-understood, routine, and conventional activities and components in the field of natural language processing. The claims do not recite a non-conventional arrangement of components or otherwise specify how the claimed elements effect a technological improvement.
Limitation-by-limitation analysis for independent claims
Claim 1 (method) – Limitation analysis
“acquiring input utterance data being utterance data concerning an utterance of a certain utterer, the input utterance data being taken by a microphone;” — This is data gathering/input identification and is a conventional, pre-solution activity using microphone. Microphone is using for collecting input utterance. Human can gather or collect data.
“performing voice recognition from the input utterance data” — Claim recites identify voice of the utterer which human mind can do.
“selecting, from among a plurality of registered utterance contents set in advance, a registered utterance content closest to a recognized utterance content indicated by a result of the voice recognition as a selected utterance content;” — Human can determine utterance content like phrase, phoneme sequence, accent, tone of the user/utterer. Human activity.
“selecting, from among a plurality of databases respectively associated with the registered utterance contents, a database associated with the selected utterance content, each of the databases storing identifiers of a plurality of registered utterers and feature quantities of utterance data concerning the registered utterance content contents having been uttered respectively by the registered utterers in association with each other; ” — Human can create multiple or plural database/libraries based on different speakers, and select dictionary associated with each speaker. Human can create template, table using pen and paper. For example, if there is three speakers, human can create 3 databases/libraries/dictionaries, each database holds/store information associated with each speaker. Human can select database based on user-specific content.
“calculating a similarity between a feature quantities of the input utterance data and the feature quantities of the registered utterers stored in the selected database; and” — This recites math calculation which human can do using paper and pencil and using human mind can determine how much difference or similar are data.
“identifying, as an identical registered utterer being identical to the certain utterer,_a registered utterer having a highest similarity among the similarities between the feature quantity of the input utterance data and the feature quantities of the registered utterers”; -- Human can determine if utterer content are similar with logged one or how much difference is there or utterance content match with any other utterer. Human can analyze this data.
“generating output data including: an identifier of the identical registered utterer that shows a result of the identification, and the registered utterance content; and--- Human can register several user and give them user id. Also, write their contents using pen and paper.
outputting the output data to an appliance, wherein the registered utterance content includes a control command to control the appliance.” Human can recognize specific speaker using command to control some device. Human can ask another human to do some action for them. Claim recites outputting data is just post solution activity.
Conclusion for claim 1: The claim is directed to abstract mental, human and mathematical activity.
conventional computing; it does not integrate the judicial exception into a practical application and does not include additional elements that amount to significantly more.
Claim 11 (device) – Limitation analysis
The claim recites “an acquisition part”, “first selection part”, “second selection part”, “a similarity calculation part” and “an output part” in claims 11, which is generic hardware components (“[0044] The processor 3 includes, for example, a central processing unit, and has the acquisition part …”) configured to execute instructions that perform the same abstract data transformations and training steps recited in claim 1. Implementing the abstract method on generic hardware does not provide a practical application or significantly more.
Conclusion for claim 11: The device claim is ineligible for the same reasons as claim 1.
Claim 12 (computer-readable medium) – Limitation analysis
The claim recites instructions stored on a non-transitory medium that cause a processor to perform the same abstract steps recited in claim 1. The storage of instructions to perform an abstract method on a computer-readable medium does not render the subject matter eligible because mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible.
Conclusion for claim 12: The claim is ineligible for the same reasons as claim 1.
Dependent claims (claims 2--10)
With respect to claim(s) 2, the claim(s) recite(s) “wherein, in the selecting of the selected utterance content, when the registered utterance contents include a registered utterance content identical to the recognized utterance content, the identical registered utterance content is selected as the selected utterance content.” Human identify phoneme sequence matched with registered utterance content and selected. No additional limitations are present.
With respect to claim(s) 3, the claim(s) recite(s) “wherein,in the selecting of the selected utterance content, when the registered utterance contents include no registered utterance content identical to the recognized utterance content, the closest registered utterance content is selected as the selected utterance content.” Human can identify phoneme sequence or keyword mismatch and closest to the sound element utterance select. No additional limitations are present.
With respect to claim(s) 4, recites “wherein, in the selecting of the selected utterance content, a registered utterance content which includes all sound elements of the recognized utterance content is selected from among the registered utterance contents.” Human recognize all sound element in the utterance like phoneme, sound, sequence, vowel and pronunciation. No additional limitations are present.
With respect to claim(s) 5, recites “wherein, in the selecting of the selected utterance content, a registered utterance content which has configuration data closest to configuration data indicating a configuration of sound elements of the recognized utterance content is selected from among the registered utterance contents.” Human recognize all registered/recognized sound element of the phonetic representation and can analysis those. No additional limitations are present.
With respect to claim(s) 6 and 7, the claim(s) recite(s) “wherein the sound element includes a phoneme” and “wherein the sound element includes a vowel”, which reads on a human mind recognize vowel in the phoneme sequence. No additional limitations are present.
With respect to claim(s) 8, the claim(s) recite(s) “wherein the sound element includes a phoneme sequence in each of n-syllabified phonemic units of an utterance content, "n" being an integer of two or larger.” Human identify how many syllable/s in the phoneme and represent them in the integer values. No additional limitations are present.
With respect to claim(s) 9, the claim(s) recite(s) “wherein the configuration data includes a vector which is defined by allocation of a value corresponding to an occurrence frequency of one or more sound elements of the recognized utterance content or the registered utterance content to a positional arrangement of all sound elements set in advance.” Human can identify phoneme sequence arrangement and location of the vowel/consonant. No additional limitations are present.
With respect to claim(s) 10, the claim(s) recite(s) “wherein the value corresponding to the occurrence frequency is defined by an occurrence frequency proportion of each of the one or more sound elements that occupies a total number of sound elements of the recognized utterance content or the registered utterance content.”, which reads on a human identifying number of times each phoneme occurs in the total phoneme, and define them in math equation. No additional limitations are present.
These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim/s 1-7, 9 11, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Hayakawa et al. US 20180075843 A1 in view of KAJAREKAR, US 20190272831 A1 and in view of SHIN ET AL. US 20200314094 A1 and further in view of Bobbili et al. US 11763809 B1.
Regarding Claim 1 Hayakawa teaches:
1. An utterer identification method for an utterer identification device, comprising: acquiring input utterance data being utterance data concerning an utterance of a certain utterer; the input utterance data being taken by a microphone; Hayakawa teaches (“[0026] The interface unit 11 is an example of a voice input unit and includes an audio interface. The interface unit 11 acquires from, for example, a microphone (not illustrated), a monaural voice signal that is an analog signal and represents a voice that a user uttered. …”) (“[0099] The voice input unit 111 includes, for example, an audio interface and an A/D converter. The voice input unit 111 acquires, for example, a voice signal that is an analog signal from a microphone and digitizes the voice signal by sampling the voice signal at a prescribed sampling rate. The voice input unit 111 outputs the digitized voice signal to the control unit 114.”) (“[0077] By confirming a presented keyword and performing a prescribed input operation by the user, a device connected to the voice recognition device 1 or a device in which the voice recognition device 1 is implemented may execute an operation corresponding to the keyword. Alternatively, the user may utter a voice indicating approval or disapproval. By recognizing the voice, the voice recognition device 1 may determine approval or disapproval. When the voice recognition device 1 determines that the user has uttered a voice indicating approval, the device connected to the voice recognition device 1 or the device in which the voice recognition device 1 is implemented may execute an operation corresponding to the keyword.”) (“[0079] The voice section detection unit 21 detects a voice section from an input voice signal (step S101). With respect to each frame in the voice section, the feature extraction unit 22 calculates a feature vector that includes a plurality of feature amounts representing characteristics of the voice of a user (step S102).”) Hayakawa et al. US 20180075843 A1
performing voice recognition from the input utterance data; Hayakawa teaches (“[0077] By confirming a presented keyword and performing a prescribed input operation by the user, a device connected to the voice recognition device 1 or a device in which the voice recognition device 1 is implemented may execute an operation corresponding to the keyword. …”) by Hayakawa et al. US 20180075843 A1
selecting, from among a plurality of registered utterance contents set in advance, a registered utterance content closest to a recognized utterance content indicated by a result of the voice recognition as a selected utterance content; Hayakawa teaches (“0023] Therefore, the voice recognition device extracts a common phoneme string from voices that are uttered repeatedly by the user, exemplifying a speaker, and compares the extracted phoneme string with the phoneme strings of the respective keywords registered in a keyword dictionary to select the most resembling keyword. The voice recognition device presents the selected keyword to the user. The keyword may be an individual word or a phrase including a plurality of words.”) (“[0027] The processing unit 13 includes, for example, one or a plurality of processors, a memory circuit, and a peripheral circuit. By performing voice recognition processing, the processing unit 13 selects one of the keywords registered in a keyword dictionary on the basis of the voice signal and outputs information representing the selected keyword via the communication interface unit 15. Alternatively, the processing unit 13 may display the selected keyword via a display device (not illustrated). Details of the voice recognition process performed by the processing unit 13 will be described later.”) (“claim 2… wherein selection of the predetermined number of keywords includes selecting the prescribed number of keywords among the plurality of keywords in descending order of the degree of similarity for each keyword.”) (“[0084] As described thus far, when no keyword is recognized among the keywords registered in the keyword dictionary from the voice that the user uttered, the voice recognition device extracts a common phoneme string that appears in common between maximum-likelihood phoneme strings of a plurality of voice sections that have been uttered repeatedly. The voice recognition device calculates degrees of similarity between the common phoneme string and the phoneme strings of the respective keywords registered in the keyword dictionary in accordance with the DP matching, identifies a keyword corresponding to a maximum value among the degrees of similarity, and presents the identified keyword to the user. Thus, even when the user does not correctly utter a keyword registered in the keyword dictionary and utters a different phrase each time, the voice recognition device may identify a keyword that the user intended to make the voice recognition device recognize. Therefore, even when the user does not remember a keyword correctly, the voice recognition device may prevent the user from uttering repeatedly to try to utter the keyword.”) (“[0079] The voice section detection unit 21 detects a voice section from an input voice signal (step S101). With respect to each frame in the voice section, the feature extraction unit 22 calculates a feature vector that includes a plurality of feature amounts representing characteristics of the voice of a user (step S102).”) Hayakawa et al. US 20180075843 A1
Hayakawa does not explicitly teach plurality of databases and selecting database.
KAJAREKAR teaches:
calculating a similarities between a feature quantity of the input utterance data and the featurequantities of the registered utterers stored in the selected database; and [0287] KAJAREKAR teaches (“[0289] … Upon receiving the spoken request, the electronic device compares a voice print of the spoken request to the first set of existing reference voice prints and determines that the spoken request is spoken by the registered user. …”) by KAJAREKAR, US 20190272831 A1
KAJAREKAR is considered to be analogous to the claimed invention because it relates to intelligent automated assistants and, more specifically, to techniques for training a speaker recognition model for intelligent automated assistants.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Hayakawa, to further incorporate the teachings of KAJAREKAR in order to include voice print of the registered user.
One could have been motivated to do so because system can generate a more accurate and reliable speaker profile. (“[0007] Determining whether the plurality of conditions are satisfied and updating the speaker profile based on the voice print in accordance with a determination that the plurality of conditions are satisfied can enable the electronic device to generate a more accurate and reliable speaker profile. In particular, satisfying the plurality of conditions can serve to authenticate that the received user utterance is likely spoken by a registered user of the electronic device. Thus, updating the speaker profile based on a voice print generated from the authenticated user utterance can result in a speaker profile that better represents the voice characteristics of the registered user. This can enhance operability of the electronic device by reducing the rate of error (e.g., false positives or false negatives) associated with voice invocation of the digital assistant using the speaker profile, which in turn improves user experience.”) by KAJAREKAR, US 20190272831 A1
The combination does not explicitly teach plurality of databases and selecting database and databases storing identifiers of a plurality of registered utterers.
Shin teaches:
selecting, from among a plurality of databases respectively associated with the registered utterance contents, a database associated with the selected utterance content, FIGS. 7B and 7C, paragraph [0125]- [0130]. Shin teaches (“[0125] FIG. 7B illustrates an example in which the biometric information DB server 300 is divided into a basic information DB server 300-1 for storing and managing encrypted basic information and a detailed information DB server 300-2 for storing and managing encrypted detailed information, unlike in FIG. 7A. There is only a difference in that the server 200-1 may request the encrypted detailed information to the basic information DB server 300-1, and request the encrypted detailed information to the detailed information DB server 300-2, and the remaining contents are the same as described above with reference to FIG. 7A, and thus a detailed description thereof will be omitted.”) (“[0130] In FIGS. 7B and 7C, since the basic information and detailed information are separately stored and managed, it is preferable that the basic information and the detailed information are encrypted in a separate data format and stored in each of the servers 300-1, 300-2, 200-2, and 200-3, but the embodiment is not limited thereto.”
Shin teaches:
a registered utterer having a highest similarity among the similarities between the feature quantity of the input utterance data and the feature quantities of the registered utterers; and Shin teaches thereafter, the processor 220 may identify or recognize the user corresponding to the detailed information having the greatest detailed information similarity among the plurality of users as a user corresponding to the basic information and the detailed information received from the terminal device 100 (“[0082] Since the basic information and the detailed information received from the terminal device 100 are separately encrypted and received in the storage 210, the processor 220 may decrypt and compare the stored basic information and the received basic information, compare the received basic information with the received basic information, and compare the detailed information corresponding to the basic information having a predetermined value of similarity with the received basic information, among the stored basic information, with the received detailed information, to perform authentication or identification.” (“[0084] First, according to an embodiment, the identification or recognition refers to recognition of the user corresponding to the biometric information presented by the terminal device 100 by the server 200, and the server 200 compares the biometric information of all registered users with the received biometric information to identify and recognize a user corresponding to the received biometric information. As described above, in the identification process, the comparison of all registered biometric information with the received biometric information is because the process of finding the user corresponding to the biometric information having the highest similarity with the biometric information received from the user registered in the server 200 is identification or recognition process.”) (“[0087] Specifically, the processor 220 may compare the received detailed information with at least one detailed information corresponding to the basic information having a basic information similarity equal to or greater than a predetermined value to calculate a detailed information similarity. Thereafter, the processor 220 may identify or recognize the user corresponding to the detailed information having the greatest detailed information similarity among the plurality of users as a user corresponding to the basic information and the detailed information received from the terminal device 100.”) (“[0111] … For example, in order to obtain biological characteristics of a user such as a fingerprint, an iris, a retina, a vein, a hand shape, a DNA, etc., as biometric information, the subject body needs to be brought into close proximity to the sensor, and to obtain the biological information of the user such as a signature/handwriting, a voice, a keyboard input, a gait, or the like, a specific character input or a specific voice utterance of the user is required. …”) by Shin et al. US 20200314094 A1
Shin is considered to be analogous to the claimed invention because it relates to a server, a method for controlling a server, and a terminal device. More particularly, the disclosure relates to a server for performing authentication or identification using biometric information, a method for controlling a server, and a terminal device.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Hayakawa, and KAJAREKAR to further incorporate the teachings of Shin.
One could have been motivated to do so because system will be faster to authenticate user which leads to improve performance. (“[0148] … As the number of reference points are relatively small compared to minutiae points and the performance could be dramatically improved. As a result of the performance evaluation, the time for a single user authentication takes less than one second. This result is at least 60 times faster than two-way setting which exploits Garbled circuit for outsourced minutiae-based fingerprint authentication.”) by Shin et al. US 20200314094 A1
Bobbili teaches:
each of the databases storing identifiers of a plurality of registered utterers and [[a]] feature quantities of utterance data concerning [[a]] the registered utterance content contents having been uttered respectively by [[a]] the registered utterers in association with each other; Bobbili teaches (“(68) The profile storage 270 may include one or more device profiles. Each device profile may be associated with a different device identifier. Each device profile may include various device identifying information. Each device profile may also include one or more user identifiers, representing one or more users associated with the device. For example, a household device's profile may include the user identifiers of users of the household.” Col.16, Lines 20-27 ) (“(80) The device 110 and/or the system 120 may associate a unique identifier with each natural language user input. The device 110 may include the unique identifier when sending the audio data 211 to the system 120, and the response data from the system 120 may include the unique identifier to identify which natural language user input the response data corresponds.” Col. 19, lines 27-34) (“(165) As indicated, a user profile may indicate which skills a corresponding user has enabled (e.g., authorized to execute using data associated with the user). Such indications may be stored in the profile storage 270. When the shortlister component 1050 receives the ASR output data 1110, the shortlister component 1050 may determine whether profile data associated with the user and/or device 110 that originated the command includes an indication of enabled skills.” Col. 37, lines 60-67) (“(189) The result data 1130 may include various portions. For example, the result data 1130 may include content (e.g., audio data, text data, and/or video data) to be output to a user. The result data 1130 may also include a unique identifier used by the system(s) 120 and/or the skill system(s) 125 to locate the data to be output to a user. The result data 1130 may also include an instruction. For example, if the user input corresponds to “turn on the light,” the result data 1130 may include an instruction causing the system to turn on a light associated with a profile of the device (110a/110b) and/or user.” Col. 42, lines 59-67) by Bobbili et al. US 11763809 B1
Bobbili teaches:
identifying, as an identical registered utterer being identical to the certain utterer, Bobbili teaches (“(68) The profile storage 270 may include one or more device profiles. Each device profile may be associated with a different device identifier. Each device profile may include various device identifying information. Each device profile may also include one or more user identifiers, representing one or more users associated with the device. For example, a household device's profile may include the user identifiers of users of the household.” Col.16, Lines 20-27 ) by Bobbili et al. US 11763809 B1
generating output data including: an identifier of the identical registered utterer that shows a result of the identification, and the registered utterance content; and Bobbili teaches the result data 1130 may also include a unique identifier used by the system(s) 120 and/or the skill system(s) 125 to locate the data to be output to a user. The result data 1130 may also include an instruction. For example, if the user input corresponds to “turn on the light,” the result data 1130 may include an instruction causing the system to turn on a light associated with a profile of the device (110a/110b) and/or user. (“(68) The profile storage 270 may include one or more device profiles. Each device profile may be associated with a different device identifier. Each device profile may include various device identifying information. Each device profile may also include one or more user identifiers, representing one or more users associated with the device. For example, a household device's profile may include the user identifiers of users of the household.” Col.16, Lines 20-27 ) (“(80) The device 110 and/or the system 120 may associate a unique identifier with each natural language user input. The device 110 may include the unique identifier when sending the audio data 211 to the system 120, and the response data from the system 120 may include the unique identifier to identify which natural language user input the response data corresponds.” Col. 19, lines 27-34) (“(165) As indicated, a user profile may indicate which skills a corresponding user has enabled (e.g., authorized to execute using data associated with the user). Such indications may be stored in the profile storage 270. When the shortlister component 1050 receives the ASR output data 1110, the shortlister component 1050 may determine whether profile data associated with the user and/or device 110 that originated the command includes an indication of enabled skills.” Col. 37, lines 60-67) (“(189) The result data 1130 may include various portions. For example, the result data 1130 may include content (e.g., audio data, text data, and/or video data) to be output to a user. The result data 1130 may also include a unique identifier used by the system(s) 120 and/or the skill system(s) 125 to locate the data to be output to a user. The result data 1130 may also include an instruction. For example, if the user input corresponds to “turn on the light,” the result data 1130 may include an instruction causing the system to turn on a light associated with a profile of the device (110a/110b) and/or user.” Col. 42, lines 59-67) by Bobbili et al. US 11763809 B1
Bobbili teaches:
outputting the output data to an appliance, wherein the registered utterance content includes a control command to control the appliance. Bobbili teaches Each device profile may also include one or more user identifiers, representing one or more users associated with the device. For example, a household device's profile may include the user identifiers of users of the household. And if the user input corresponds to “turn on the light,” the result data 1130 may include an instruction causing the system to turn on a light associated with a profile of the device (110a/110b) and/or user. (“(68) The profile storage 270 may include one or more device profiles. Each device profile may be associated with a different device identifier. Each device profile may include various device identifying information. Each device profile may also include one or more user identifiers, representing one or more users associated with the device. For example, a household device's profile may include the user identifiers of users of the household.” Col.16, Lines 20-27 ) (“(80) The device 110 and/or the system 120 may associate a unique identifier with each natural language user input. The device 110 may include the unique identifier when sending the audio data 211 to the system 120, and the response data from the system 120 may include the unique identifier to identify which natural language user input the response data corresponds.” Col. 19, lines 27-34) (“(165) As indicated, a user profile may indicate which skills a corresponding user has enabled (e.g., authorized to execute using data associated with the user). Such indications may be stored in the profile storage 270. When the shortlister component 1050 receives the ASR output data 1110, the shortlister component 1050 may determine whether profile data associated with the user and/or device 110 that originated the command includes an indication of enabled skills.” Col. 37, lines 60-67) (“(189) The result data 1130 may include various portions. For example, the result data 1130 may include content (e.g., audio data, text data, and/or video data) to be output to a user. The result data 1130 may also include a unique identifier used by the system(s) 120 and/or the skill system(s) 125 to locate the data to be output to a user. The result data 1130 may also include an instruction. For example, if the user input corresponds to “turn on the light,” the result data 1130 may include an instruction causing the system to turn on a light associated with a profile of the device (110a/110b) and/or user.” Col. 42, lines 59-67) by Bobbili et al. US 11763809 B1
Bobbili is considered to be analogous to the claimed invention because it relates to vehicle sharing service platforms, and more particularly, to systems and methods for providing vehicle service to users and billing the users in the vehicle sharing service system.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Hayakawa, KAJAREKAR, and SHIN, to further incorporate the teachings of Bobbili in order to include command controlling appliances.
One could have been motivated to do so because system have accuracy of user recognition operations. (“(63) … The user-recognition component 295 also determines an overall confidence regarding the accuracy of user recognition operations.”) by Bobbili et al. US 11763809 B1
Claim 11 is a device claim with a limitation similar to the limitation of method Claim 1 and is rejected under similar rationale. Additionally,
Regarding Claim 11 Hayakawa further teaches:
11. An utterer identification device, comprising: a circuit configured to: Hayakawa teaches (“[0027] The processing unit 13 includes, for example, one or a plurality of processors, a memory circuit, and a peripheral circuit. By performing voice recognition processing, the processing unit 13 selects one of the keywords registered in a keyword dictionary on the basis of the voice signal and outputs information representing the selected keyword via the communication interface unit 15. …”) (“[0029] The communication interface unit 15 includes a communication interface circuit for connecting the voice recognition device 1 to another device, for example, a navigation system….”) by Hayakawa US 20180075843 A1
Claim 12 is a non-transitory computer readable medium claim with a limitation similar to the limitation of method Claim 1 and is rejected under similar rationale.
Regarding Claim 12 Hayakawa teaches:
12. A non-transitory computer readable recording medium storing an utterer identification program that causes a computer to serve as an utterer identification device, the utterer identification program comprising: causing the computer to execute: Hayakawa teaches (“A non-transitory computer-readable recording …”) by Hayakawa US 20180075843 A1
Regarding Claim 2 the combination teaches the method claim 1 as identified above.
Hayakawa further teaches:2. The utterer identification method according to claim 1, wherein, in the selecting of the selected utterance content, when the registered utterance contents include a registered utterance content identical to the recognized utterance content, the identical registered utterance content is selected as the selected utterance content. Hayakawa teaches selects one of the keywords registered in a keyword dictionary on the basis of the voice signal. (“[0027] The processing unit 13 includes, for example, one or a plurality of processors, a memory circuit, and a peripheral circuit. By performing voice recognition processing, the processing unit 13 selects one of the keywords registered in a keyword dictionary on the basis of the voice signal and outputs information representing the selected keyword via the communication interface unit 15. Alternatively, the processing unit 13 may display the selected keyword via a display device (not illustrated). ...”) (“[0050] By comparing the maximum-likelihood phoneme string of the voice section with the phoneme strings representing the utterances of the keywords registered in the keyword dictionary, the determination unit 24 determines whether or not the user uttered any keyword in the voice section.”) by Hayakawa et al. US 20180075843 A1
Regarding Claim 3 the combination teaches the method claim 1 as identified above.
Hayakawa further teaches:
The combination teaches the method claim 1 as identified above.
3. The utterer identification method according to claim 1, wherein, in the selecting of the selected utterance content, when the registered utterance contents include no registered utterance content identical to the recognized utterance content, the closest registered utterance content is selected as the selected utterance content.
Hayakawa teaches (“[0084] As described thus far, when no keyword is recognized among the keywords registered in the keyword dictionary from the voice that the user uttered, the voice recognition device extracts a common phoneme string that appears in common between maximum-likelihood phoneme strings of a plurality of voice sections that have been uttered repeatedly. The voice recognition device calculates degrees of similarity between the common phoneme string and the phoneme strings of the respective keywords registered in the keyword dictionary in accordance with the DP matching, identifies a keyword corresponding to a maximum value among the degrees of similarity, and presents the identified keyword to the user. Thus, even when the user does not correctly utter a keyword registered in the keyword dictionary and utters a different phrase each time, the voice recognition device may identify a keyword that the user intended to make the voice recognition device recognize. Therefore, even when the user does not remember a keyword correctly, the voice recognition device may prevent the user from uttering repeatedly to try to utter the keyword.”) Hayakawa et al. US 20180075843 A1
Regarding Claim 4 the combination teaches the method claim 1 as identified above.
Hayakawa further teaches:
4. The utterer identification method according to claim 1, wherein, in the selecting of the selected utterance content, a registered utterance content which includes all sound elements of the recognized utterance content is selected from among the registered utterance contents. Hayakawa teaches (“[0027] The processing unit 13 includes, for example, one or a plurality of processors, a memory circuit, and a peripheral circuit. By performing voice recognition processing, the processing unit 13 selects one of the keywords registered in a keyword dictionary on the basis of the voice signal and outputs information representing the selected keyword via the communication interface unit 15. …”) (“[0050] By comparing the maximum-likelihood phoneme string of the voice section with the phoneme strings representing the utterances of the keywords registered in the keyword dictionary, the determination unit 24 determines whether or not the user uttered any keyword in the voice section.”) (“[0051] FIG. 3 is a diagram illustrating an example of a keyword dictionary. In keyword dictionary 300, with respect to each keyword, a character string representing a written form of the keyword and a phoneme string representing a pronunciation of the keyword are registered. For example, for a keyword “Jitaku e kaeru (Japanese pronunciation, meaning “Return my home” in English)”, a phoneme string “jitakuekaeru” of the keyword is registered.”) (“[0084] As described thus far, when no keyword is recognized among the keywords registered in the keyword dictionary from the voice that the user uttered, the voice recognition device extracts a common phoneme string that appears in common between maximum-likelihood phoneme strings of a plurality of voice sections that have been uttered repeatedly. The voice recognition device calculates degrees of similarity between the common phoneme string and the phoneme strings of the respective keywords registered in the keyword dictionary in accordance with the DP matching, identifies a keyword corresponding to a maximum value among the degrees of similarity, and presents the identified keyword to the user. …”) by Hayakawa et al. US 20180075843 A1
Regarding Claim 5 the combination teaches the method claim 1 as identified above.
Hayakawa further teaches:
5. The utterer identification method according to claim 1, wherein, in the selecting of the selected utterance content, a registered utterance content which has configuration data closest to configuration data indicating a configuration of sound elements of the recognized utterance content is selected from among the registered utterance contents. Hayakawa teaches (“[0045] Specifically, with respect to each frame in the voice section, by inputting the feature vector of the frame to the GMM, the maximum-likelihood phoneme string search unit 23 calculates output probabilities of the respective HMM states corresponding to the respective phonemes for the frame. In addition, before inputting a feature vector into the GMM, the maximum-likelihood phoneme string search unit 23 may, for the feature vector calculated from each frame, perform normalization, referred to as cepstral mean normalization (CMN), in which, with respect to each dimension of the feature vector, a mean value is estimated and the estimated mean value is subtracted from a value at the dimension.”) (“[0080] On the basis of the feature vectors of the respective frames, the maximum-likelihood phoneme string search unit 23 searches for a maximum-likelihood phoneme string corresponding to a voice uttered in the voice section (step S103). On the basis of the maximum-likelihood phoneme string and a keyword dictionary, the determination unit 24 determines whether or not any keyword registered in the keyword dictionary is detected in the voice section (step S104). When any keyword is detected (Yes in step S104), the processing unit 13 outputs information representing the keyword and finishes the voice recognition process.”) by Hayakawa et al. US 20180075843 A1
Regarding Claim 9 the combination teaches the method claim 5 as identified above.
Hayakawa further teaches:
9. The utterer identification method according to claim 5, wherein the configuration data includes a vector which is defined by allocation of a value corresponding to an occurrence frequency of one or more sound elements of the recognized utterance content or the registered utterance content to a positional arrangement of all sound elements set in advance. Hayakawa teaches (“[0071] Alternatively, the matching unit 26 may calculate the degree of similarity P based on a degree of coincidence between the phoneme string of the keyword of interest and the common phoneme string in accordance with the equation (2). In this case, C is the number of coincident phonemes between the common phoneme string and the phoneme string of the keyword of interest and D is the number of phonemes that are included in the phoneme string of the keyword of interest but not included in the common phoneme string. In addition, S is the number of phonemes that are included in the phoneme string of the keyword of interest and are different from the phonemes at corresponding positions in the common phoneme string.”) (“[0043] … The maximum-likelihood phoneme string is a phoneme string in which respective phonemes included in a voice are arranged in a sequence of utterances thereof and that are estimated to be most probable.”) (“In the equation, C is the number of coincident phonemes between the maximum-likelihood phoneme string and the phoneme string of the keyword of interest and D is the number of phonemes that are included in the phoneme string of the keyword of interest but not included in the maximum-likelihood phoneme string. In addition, S is the number of phonemes that are included in the phoneme string of the keyword of interest and are different from phonemes at corresponding positions in the maximum-likelihood phoneme string.”) (“[0056] When two or more maximum-likelihood phoneme strings have been saved in the storage unit 14, i.e., the user has uttered keywords repeatedly while no keyword has been recognized, the common phoneme string extraction unit 25 extracts a string in which common phonemes to the maximum-likelihood phoneme strings are arranged in a sequence of utterances (hereinafter, simply referred to as common phoneme string).”) (“[0058] After a phoneme(s) representing silence and/or a phoneme(s) that appear(s) in only either one of the maximum-likelihood phoneme strings has/have been deleted from the respective maximum-likelihood phoneme strings, the common phoneme string extraction unit 25 extracts coincident phonemes between the two maximum-likelihood phoneme strings in order from the heads of the two maximum-likelihood phoneme strings. The common phoneme string extraction unit 25 sets a string in which the extracted phonemes are arranged from the head as a common phoneme string.”) (“[0093] In the variation, the common phoneme string extraction unit 25 may extract a common phoneme string by extracting phonemes each of which is common to a majority of maximum-likelihood phoneme strings among three or more maximum-likelihood phoneme strings and arranging the extracted phonemes in a sequence of utterances. …”) by Hayakawa et al. US 20180075843 A1
Regarding Claim 6 the combination teaches the method claim 4 as identified above.
Hayakawa further teaches:
6. The utterer identification method according to claim 4, wherein the sound element includes a phoneme. Hayakawa teaches (“[0023] Therefore, the voice recognition device extracts a common phoneme string from voices that are uttered repeatedly by the user, exemplifying a speaker, and compares the extracted phoneme string with the phoneme strings of the respective keywords registered in a keyword dictionary to select the most resembling keyword. The voice recognition device presents the selected keyword to the user. The keyword may be an individual word or a phrase including a plurality of words.”) by Hayakawa et al. US 20180075843 A1
Regarding Claim 7 the combination teaches the method claim 4 as identified above.
Hayakawa further teaches:
7. The utterer identification method according to claim 4, wherein the sound element includes a vowel. the maximum-likelihood phoneme strings 401 and 402, each of the phonemes “sp”, “silB”, and “silE” is a phoneme representing silence. Hayakawa teaches (“[0059] FIG. 4 is a diagram illustrating an example of maximum-likelihood phoneme strings and a common phoneme string. Illustrated as FIG. 4, it is assumed that, in the first utterance, a user uttered, “Etto jitaku, ja nakatta, ie ni kaeru (Japanese pronunciation, meaning “Uh, my home, no, return to a house” in English)”. For the utterance, a maximum-likelihood phoneme string 401 is calculated. In the second utterance, it is assumed that, the user uttered, “Chigau ka. Jitaku, jibun no sunde iru tokoro, ni kaeru (Japanese pronunciation, meaning “No, that's wrong. My home, the place where I live, return there” in English)”. For the utterance, a maximum-likelihood phoneme string 402 is calculated. In the maximum-likelihood phoneme strings 401 and 402, each of the phonemes “sp”, “silB”, and “silE” is a phoneme representing silence.”) (“[0060] … a common phoneme string (“oitakuertknikaeuq”) 420 to be obtained.”) Notes: each of the phonemes “sp”, “silB”, and “silE” and “oitakuertknikaeuq” this common phoneme has vowel in it. by Hayakawa et al. US 20180075843 A1
Claim/s 8 are rejected under 35 U.S.C. 103 as being unpatentable over Hayakawa, KAJAREKAR, SHIN, and Bobbili in view of Ogawa et al. US 7657430 B2.
Regarding Claim 8 the combination teaches the method claim 4 as identified above.
The combination does not explicitly teach the sound element includes a phoneme sequence in each of n-syllabified phonemic units of an utterance content, "n" being an integer of two or larger.
Ogawa teaches teaches:
8. The utterer identification method according to claim 4, wherein the sound element includes a phoneme sequence in each of n-syllabified phonemic units of an utterance content, "n" being an integer of two or larger. FIG. 1-2,Ogawa teaches (“(37) For example, three sounds "AKA", "AO", and "MIDORI" are input to the word extracting unit 2. The word extracting unit 2 classifies these three sounds to three corresponding clusters, an "AKA" cluster 21, an "AO" cluster 22, and a "MIDORI" cluster 23, respectively. Concurrently, the word extracting unit 2 assigns representative syllable sequences ("A/KA", "A/O", and "MI/DO/RI" in the case shown in FIG. 6) and IDs ("1", "2", and "3" in the case shown in FIG. 6) to the clusters.”) (“(47) The phonetic typewriter 45 further performs speech recognition of the input sound on a syllable basis using the feature parameters supplied from the feature extraction module 43 while referencing the acoustic model database 51, and then outputs the syllable sequence obtained by the speech recognition to both matching module 44 and network generating module 47. For example, from the utterance "WATASHINONAMAEWAOGAWADESU", a syllable sequence "WA/TA/SHI/NO/NA/MA/E/WA/O/GA/WA/DE/SU" is obtained. Any commercially available phonetic typewriter can be used as the phonetic typewriter 45.”) (“(52) The acoustic model database 51 stores an acoustic model representing acoustic features of individual phonemes and syllables of a language for the utterance to be recognized. For example, a Hidden Markov Model (HMM) may be used as an acoustic model. The dictionary database 52 stores a word dictionary describing information about pronunciations and a model describing chains of the phonemes and syllables for the words or phrases to be recognized.”) (“(70) Referring back to FIG. 10, at step S55, the phonetic typewriter 45 recognizes the feature parameters extracted by the feature extraction module 43 in the process of step S53 on a phoneme basis independently from the process of step S54, and outputs the acquired syllable sequence to the matching module 44. For example, when an utterance "WATASHINONAMAEWAOGAWADESU", where "OGAWA" is an unknown word, is input to the phonetic typewriter 45, the phonetic typewriter 45 outputs a syllable sequence "WA/TA/SHI/NO/NA/MA/E/WA/O/GA/WA/DE/SU". At step S55, a syllable sequence may be acquired using the processing result at step S54.”) (“7. A computer-readable recording medium storing a program, the program processing an input utterance and registering an unknown word contained in the input utterance into a dictionary database on the basis of the processing result, the program including the steps of: (a) recognizing the input utterance; (b) determining whether the recognition result of the input utterance obtained by step (a) contains an unknown word on the basis of an acoustic model representing acoustic features of individual phonemes and syllables of a language; (c) determining whether the recognition result determined at step (b) to contain an unknown word is rejected or not for acquisition and registering into the dictionary database; …”) by Ogawa et al. US 7657430 B2
Ogawa is considered to be analogous to the claimed invention because it relates to voice quality conversion devices.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Hayakawa, KAJAREKAR, SHIN, and Bobbili to further incorporate the teachings of Ogawa in order to include Vowel feature in the system.
One could have been motivated to do so because system recognition accuracy is improved.(“(128) … As can be seen from the experimental result in FIG. 18, the recognition accuracy was 48.5%, which is improved compared to that of 40.2% by the <OOV> pronunciation acquiring method by use of a sub-word sequence shown in FIG. 3. …” col. 15, lines 14-18) by Ogawa et al. US 7657430 B2
Claim/s 10 are rejected under 35 U.S.C. 103 as being unpatentable over Hayakawa, KAJAREKAR, SHIN, and Bobbili in view of Beach et al. US 10395640 B1.
Regarding Claim 10 the combination teaches the method claim 9 as identified above.
The combination does not explicitly teach an occurrence frequency proportion of each of the one or more sound elements that occupies a total number of sound elements.
Beach teaches:
10. The utterer identification method according to claim 9, wherein the value corresponding to the occurrence frequency is defined by an occurrence frequency proportion of each of the one or more sound elements that occupies a total number of sound elements of the recognized utterance content or the registered utterance content. Beach teaches (“ For example, it is possible to generate a total count of how many times each phoneme occurs in the sample set (training text and audio pairs), how often the phoneme was recognized correctly, the phoneme's overall accuracy (as well as the accuracy in total—the average of all the phoneme's average accuracy), and the “incorrect phonemes” the “correct phoneme” was confused with. Additionally, aggregate statistics can be generated, such as the average accuracy of all phonemes. Additionally, the comparison may be used to identify the phonemes with the highest and lowest accuracies. …”) (“When this comparison is performed for a number of text and audio samples, it is possible to generate the statistics described above for each phoneme. For example, it is possible to generate a total count of how many times each phoneme occurs in the sample set (training text and audio pairs), how often the phoneme was recognized correctly, the phoneme's overall accuracy (as well as the accuracy in total—the average of all the phoneme's average accuracy) …” col. 9, lines 27-34) (“… (26) FIGS. 8A and 8B show exemplary charts generated using one possible measurement consistent with the technology of the present application. The charts are presented as simple graphs with the phoneme accuracy as the Y-axis and the 40 possible phonemes associated with the English language as the X-axis. The phoneme accuracy is presented as a simple percentage, which is generated by the number of correct phoneme identifications divided by the total number of phoneme presentations times. …” col. 9, lines 45-55) Beach et al. US 10395640 B1
Beach is considered to be analogous to the claimed invention because it relates to relates generally to speech recognition systems.
Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Hayakawa, KAJAREKAR, SHIN, and Bobbili to further incorporate the teachings of Beach in order to include total number of sound element.
One could have been motivated to do so because system recognize phoneme accuracy. (“(6) … The audio phoneme sequence and the text phoneme sequence are compared to determine a phoneme average accuracy.” col. 2, lines 42-43) Beach et al. US 10395640 B1
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FOUZIA HYE SOLAIMAN whose telephone number is (571)270-5656. The examiner can normally be reached M-F (8-5)AM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/F.H.S./Examiner, Art Unit 2653
/Paras D Shah/Supervisory Patent Examiner, Art Unit 2653
07/22/2026