DETAILED ACTION
This Office Action is in response to the correspondence filed by the applicant on 5/26/2026.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s argument, pages 9-12, with respect to the rejection of claims under 103 have been fully considered and are moot upon a further consideration and a new ground(s) of rejection made under AIA 35 U.S.C. 103 as being unpatentable over HARIF (US 6,820,056 B1), and in further view of LARGEY (US 2013/0226589 A1). Please see the rejection below for more details.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 7, 8, 14, 15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over HARIF (US 6,820,056 B1), and in further view of LARGEY (US 2013/0226589 A1).
REGARDING CLAIM 1, HARIF discloses a method for using a non-speech sound to control a feature associated with a hearable device, the method comprising:
detecting a first pattern a combination of different types of non-speech sounds by a user of the hearable device (HARIF Fig. 2 – “Sound Commands 52”; Col 6:18-6:43 – “If the decision from step 81 is Yes, then a further determination is made in decision step 82 as to whether a non-verbal sound has been recognized.”) that are sequentially produced (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”), wherein each non-speech sound is created by one or more of breath, nose, tongue, lips, and throat of the user with the intention to control the feature (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”);
identifying the first pattern of non-speech sounds as a control gesture corresponding to a [single] control command for a particular adjustment of the feature associated with the hearable device (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”), by applying one or more sound factors (HARIF Col 4:24-54 – “Thus, a comparison 55 is made of an input of non-verbal sound to the stored non-verbal sound commands 52 and recognized non-verbal sounds are input via display adapter 36 to display 38 for verification, as will hereinafter be described.”);
based, at least in part, on identifying the control gesture, adjusting the feature according to the [single] control command (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”); and
outputting to the user, a feedback indicator to describe the adjusting of the feature (HARIF Col 4:55-5:14 – “When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62. The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down.”; Col 6:18-43 – “If the determination from step 82 is Yes, a non-verbal sound is recognized, then, step 84, that sound is compared with the stored command sounds. If there is No compare, then there is displayed to user: “Do Not Recognize”, step 90. If Yes, there is a compare, then, step 86, the command is displayed for confirmation. In the confirmation decision step 87, if there is No confirmation, then, again, there is displayed to user: “Do Not Recognize”, step 90.”).
HARIF does not explicitly teach the [square-bracketed] limitation. In other words, HARIF teaches a combination of different types of non-speech sounds (e.g., a sequence of “Mouth-Tongue Clack” and “”Knock” as shown in the table in Col 5:40-50) to move a cursor to an upper left direction by moving the cursor left then up. However, HARIF does not explicitly teach the sequence of commands adjusts feature according to a [single] control command, i.e., moving the cursor to a upper-left direction as one single command, not left then up.
LARGEY discloses the [square-bracketed] limitation. LARGEY discloses a method/system for controlling a device with non-verbal commands comprising:
detecting a first pattern a combination of different types of non-speech sounds by a user of the hearable device that are sequentially produced (LARGEY Par 42 – “For example, a single vocalized click may be translated to “no”, and a double click (e.g. two clicks occurring within a predetermined period) may be translated as “yes”. Of course other combinations of clicks, or other nonphonetic commands, may be translated to other synthesized phonetic commands.”), wherein each non-speech sound is created by one or more of breath, nose, tongue, lips, and throat of the user with the intention to control the feature (LARGEY Par 21 – “As used herein the term “nonphonetic command” is used to encompass non-linguistic sounds produced by human vocalization.”; Par 22 – “As used herein, “language” explicitly excludes languages that rely on click consonants, such as languages colloquially referred to as “click-tongue”, of which Xhosa is one example.”);
identifying the first pattern of non-speech sounds as a control gesture corresponding to a [single] control command (LARGEY Par 42 – “For example, a single vocalized click may be translated to “no”, and a double click (e.g. two clicks occurring within a predetermined period) may be translated as “yes”. Of course other combinations of clicks, or other nonphonetic commands, may be translated to other synthesized phonetic commands.”) for a particular adjustment of the feature associated with the hearable device (LARGEY Par 53 – “In other embodiments the command discriminator 430 includes the command synthesizer 470 and provides a phonetic command to the GPS receiver 630 in response to the nonphonetic command. Similarly, other embodiments of the functional block 620, e.g. recorder or smart phone, may be configured to receive from the command discriminator 430 an electronic signal indicating the occurrence of a nonphonetic command, or may receive a synthesized voice command, and then operate to perform its core functionality, respectively for example recording and calling.”), by applying one or more sound factors (LARGEY Par 56 – “In the step 530, the command discriminator 430 attempts to match the spectrum determined in the step 520 to one of a number of model spectra, or mathematical descriptions of model spectra. The model spectra or their mathematical descriptions may be stored, e.g. in the memory 435. The matching may include, e.g. a determination of various metrics describing quality of fit, and a match probability.”);
based, at least in part, on identifying the control gesture, adjusting the feature according to the [single] control command (LARGEY Par 42 – “For example, a single vocalized click may be translated to “no”, and a double click (e.g. two clicks occurring within a predetermined period) may be translated as “yes”. Of course other combinations of clicks, or other nonphonetic commands, may be translated to other synthesized phonetic commands.”; Par 53 – “Similarly, other embodiments of the functional block 620, e.g. recorder or smart phone, may be configured to receive from the command discriminator 430 an electronic signal indicating the occurrence of a nonphonetic command, or may receive a synthesized voice command, and then operate to perform its core functionality, respectively for example recording and calling.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF to include identifying a sequence of non-speech sounds as a control gesture corresponding to a single control command, as taught by LARGEY.
One of ordinary skill would have been motivated to include identifying a sequence of non-speech sounds as a control gesture corresponding to a single control command, in order to allow a user to control multiple functions of an electronic device seamlessly.
REGARDING CLAIM 7, HARIF in view of LARGEY discloses the method of claim 1, further comprising:
outputting an inquiry for user control (HARIF Col 4:55-5:15 – “The window has speak 64 and clear 65 buttons, as well as an on/off button 63 to end the speech recognition session. When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62.”);
detecting the first pattern of the [non-speech] sounds (HARIF Col 4:55-5:15 – “The window has speak 64 and clear 65 buttons, as well as an on/off button 63 to end the speech recognition session. When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62.”); and
determining the first pattern of [non-speech] sounds is responsive to the inquiry (HARIF Col 4:55-5:15 – “When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62. The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”; LARGEY Par 56 – “In this manner, the system 700 may provide the caller the ability to use the nonphonetic commands to communicate with the VRS 730 when the caller is in a noisy environment. In some embodiments the functionality of the command discriminator 720 may be tightly integrated with the VRS 730, such that the command discriminator 720 directly communicates the received nonphonetic command to the VRS 730 without the need for synthesizing the phonetic command. In some embodiments the nonphonetic command may be communicated to the command discriminator 720 using out-of-band signaling, thereby bypassing the voice band.”).
HARIF does not explicitly teach the [square-bracketed] limitation. In other words, HARIF teaches prompting a user for confirmation for non-speech commands, but does not explicitly teach receiving a user’s confirmation through a non-speech response.
LARGEY discloses the [square-bracketed] limitation. LARGEY discloses a method/system for controlling a device with non-verbal commands comprising:
outputting an inquiry for user control (HARIF Col 4:55-5:15 – “The window has speak 64 and clear 65 buttons, as well as an on/off button 63 to end the speech recognition session. When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62.”; LARGEY Par 55 – “Another embodiment is described by FIG. 7, which illustrates a system 700, for example an embodiment of a VRS as might be used by a bank or other service provider that prompts a caller to provide voice responses to navigate a tree of features available to the caller.”);
detecting the first pattern of the [non-speech] sounds (HARIF Col 4:55-5:15 – “The window has speak 64 and clear 65 buttons, as well as an on/off button 63 to end the speech recognition session. When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62.”; LARGEY Par 56 – “If the command discriminator 720 instead determines the occurrence of a nonphonetic command as described previously, the command discriminator 720 may control a synthesizer 760 to synthesize the corresponding phonetic command, and control the MUX 750 to pass the synthesized phonetic command to the VRS 730.”); and
determining the first pattern of [non-speech] sounds is responsive to the inquiry (HARIF Col 4:55-5:15 – “When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62. The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”; LARGEY Par 56 – “In this manner, the system 700 may provide the caller the ability to use the nonphonetic commands to communicate with the VRS 730 when the caller is in a noisy environment. In some embodiments the functionality of the command discriminator 720 may be tightly integrated with the VRS 730, such that the command discriminator 720 directly communicates the received nonphonetic command to the VRS 730 without the need for synthesizing the phonetic command. In some embodiments the nonphonetic command may be communicated to the command discriminator 720 using out-of-band signaling, thereby bypassing the voice band.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF to include a non-speech response from a user, as taught by LARGEY.
One of ordinary skill would have been motivated to include a non-speech response from a user, in order to provide the user the ability to use the non-speech commands to communicate with the VRS when the user is in a noisy environment.
REGARDING CLAIM 8, HARIF discloses a sound gesture control system to adjust a feature associated with a hearable device, the sound gesture control system comprising: at least one sensor to detect at least one non-speech sound of a user using the hearable device (HARIF Fig. 1; Col 3: 57-62 – “The speech and/or non-verbal sound input is made through input device 27, which is diagrammatically depicted as a microphone, which accesses the system through an appropriate interface adapter 22.”); a hearable device of a user (HARIF Figs. 1-3) comprising: one or more processors (Fig. 1 – “CPU 10; Operating System 41”); and logic encoded in one or more non-transitory media for execution by the one or more processors and when executed, operable to perform operations (Col 6:44-59 – “One of the implementations of the present invention is as an application program 40 made up of programming steps or instructions resident in RAM 14, FIG. 1, during computer operations. Until required by the computer system, the program instructions may be stored in another readable medium, e.g. in disk drive 20 or in a removable memory, such as an optical disk for use in a CD ROM computer input or in a floppy disk for use in a floppy disk drive computer input..”) comprising:
detecting a first pattern a combination of different types of non-speech sounds by a user of the hearable device (HARIF Fig. 2 – “Sound Commands 52”; Col 6:18-6:43 – “If the decision from step 81 is Yes, then a further determination is made in decision step 82 as to whether a non-verbal sound has been recognized.”) that are sequentially produced (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”), wherein each non-speech sound is created by one or more of breath, nose, tongue, lips, and throat of the user with the intention to control the feature (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”);
identifying the first pattern of non-speech sounds as a control gesture corresponding to a [single] control command for a particular adjustment of the feature associated with the hearable device (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”), by applying one or more sound factors (HARIF Col 4:24-54 – “Thus, a comparison 55 is made of an input of non-verbal sound to the stored non-verbal sound commands 52 and recognized non-verbal sounds are input via display adapter 36 to display 38 for verification, as will hereinafter be described.”);
based, at least in part, on identifying the control gesture, adjusting the feature according to the [single] control command (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”), wherein the feature is selected from the group of: setting, mode, audio content player, audio beam focus, calling interaction, and smart assistant operation (HARIF Col 5:16-20 – “Of course, the commands may relate to functions other than the above-described cursor and cursor movements. Illustrative commands are: “Show Background”, “Underline”, Hide Menu”, “Delete Last Word”, “Next Paragraph” or “Close Session”.”); and
outputting to the user, a feedback indicator to describe the adjusting of the feature (HARIF Col 4:55-5:14 – “When a non-verbal sound is input, e.g. hand clap=cursor, FIG. 4, the cursor command is displayed, 66, along with a dialog line 67 requesting “Yes or No” confirmation. Then, if the user confirms the cursor command, the cursor 68 appears in an initial position in the text string 62. The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down.”; Col 6:18-43 – “If the determination from step 82 is Yes, a non-verbal sound is recognized, then, step 84, that sound is compared with the stored command sounds. If there is No compare, then there is displayed to user: “Do Not Recognize”, step 90. If Yes, there is a compare, then, step 86, the command is displayed for confirmation. In the confirmation decision step 87, if there is No confirmation, then, again, there is displayed to user: “Do Not Recognize”, step 90.”).
HARIF does not explicitly teach the [square-bracketed] limitation. In other words, HARIF teaches a combination of different types of non-speech sounds (e.g., a sequence of “Mouth-Tongue Clack” and “”Knock” as shown in the table in Col 5:40-50) to move a cursor to an upper left direction by moving the cursor left then up. However, HARIF does not explicitly teach the sequence of commands adjusts feature according to a [single] control command, i.e., moving the cursor to a upper-left direction as one single command, not left then up.
LARGEY discloses the [square-bracketed] limitation. LARGEY discloses a method/system for controlling a device with non-verbal commands comprising:
detecting a first pattern a combination of different types of non-speech sounds by a user of the hearable device that are sequentially produced (LARGEY Par 42 – “For example, a single vocalized click may be translated to “no”, and a double click (e.g. two clicks occurring within a predetermined period) may be translated as “yes”. Of course other combinations of clicks, or other nonphonetic commands, may be translated to other synthesized phonetic commands.”), wherein each non-speech sound is created by one or more of breath, nose, tongue, lips, and throat of the user with the intention to control the feature (LARGEY Par 21 – “As used herein the term “nonphonetic command” is used to encompass non-linguistic sounds produced by human vocalization.”; Par 22 – “As used herein, “language” explicitly excludes languages that rely on click consonants, such as languages colloquially referred to as “click-tongue”, of which Xhosa is one example.”);
identifying the first pattern of non-speech sounds as a control gesture corresponding to a [single] control command (LARGEY Par 42 – “For example, a single vocalized click may be translated to “no”, and a double click (e.g. two clicks occurring within a predetermined period) may be translated as “yes”. Of course other combinations of clicks, or other nonphonetic commands, may be translated to other synthesized phonetic commands.”) for a particular adjustment of the feature associated with the hearable device (LARGEY Par 53 – “In other embodiments the command discriminator 430 includes the command synthesizer 470 and provides a phonetic command to the GPS receiver 630 in response to the nonphonetic command. Similarly, other embodiments of the functional block 620, e.g. recorder or smart phone, may be configured to receive from the command discriminator 430 an electronic signal indicating the occurrence of a nonphonetic command, or may receive a synthesized voice command, and then operate to perform its core functionality, respectively for example recording and calling.”), by applying one or more sound factors (LARGEY Par 56 – “In the step 530, the command discriminator 430 attempts to match the spectrum determined in the step 520 to one of a number of model spectra, or mathematical descriptions of model spectra. The model spectra or their mathematical descriptions may be stored, e.g. in the memory 435. The matching may include, e.g. a determination of various metrics describing quality of fit, and a match probability.”);
based, at least in part, on identifying the control gesture, adjusting the feature according to the [single] control command (LARGEY Par 42 – “For example, a single vocalized click may be translated to “no”, and a double click (e.g. two clicks occurring within a predetermined period) may be translated as “yes”. Of course other combinations of clicks, or other nonphonetic commands, may be translated to other synthesized phonetic commands.”; Par 53 – “Similarly, other embodiments of the functional block 620, e.g. recorder or smart phone, may be configured to receive from the command discriminator 430 an electronic signal indicating the occurrence of a nonphonetic command, or may receive a synthesized voice command, and then operate to perform its core functionality, respectively for example recording and calling.”), wherein the feature is selected from the group of: setting, mode, audio content player, audio beam focus, calling interaction, and smart assistant operation (LARGEY Par 56 – “In this manner, the system 700 may provide the caller the ability to use the nonphonetic commands to communicate with the VRS 730 when the caller is in a noisy environment.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF to include identifying a sequence of non-speech sounds as a control gesture corresponding to a single control command, as taught by LARGEY.
One of ordinary skill would have been motivated to include identifying a sequence of non-speech sounds as a control gesture corresponding to a single control command, in order to allow a user to control multiple functions of an electronic device seamlessly.
CLAIM 14 is similar to claim 7; thus, it is rejected under the same rationale.
REGARDING CLAIM 15, ZHANG in view of HARIF discloses a non-transitory computer-readable storage medium carrying program instructions thereon for using sound gesture to control a feature associated with a hearable device, the instructions when executed by one or more processors cause the one or more processors to perform operations comprising: steps of claim 8; thus, it is rejected under the same rationale.
CLAIM 20 is similar to claim 7; thus, it is rejected under the same rationale.
Claims 4 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over HARIF (US 6,820,056 B1) in view of LARGEY (US 2013/0226589 A1), and in further view of ZHANG (US 2025/0358559 A1).
REGARDING CLAIM 4, HARIF in view of LARGEY discloses the method of claim 1.
HARIF in view of LARGEY does not explicitly teach a tactile feedback.
ZHANG discloses a method/system for controlling a device with non-speech sounds further comprising:
producing a tactile feedback by moving one or more hearable components proximal to a user ear, wherein the tactile feedback is associated with outputting of the feedback indicator (ZHANG Par 106 – “The motor 191 may generate a vibration prompt. The motor 191 may be configured to provide an incoming call vibration prompt and a touch vibration feedback. For example, touch operations performed on different disclosures (for example, photographing and audio playing) may correspond to different vibration feedback effect. The motor 191 may also correspond to different vibration feedback effect for touch operations performed on different areas of the display 194. Different disclosure scenarios (for example, a time reminder, information receiving, an alarm clock, and a game) may also correspond to different vibration feedback effect. Touch vibration feedback effect may be further customized.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF in view of LARGEY to include a tactile feedback, as taught by ZHANG.
One of ordinary skill would have been motivated to include a tactile feedback, in order to notify a user with better accessibility and stronger confidence.
CLAIM 11 is similar to claim 4; thus, it is rejected under the same rationale.
Claims 2-3, 6, 9-10, 13, 16-17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over HARIF (US 6,820,056 B1), and in further view of LARGEY (US 2013/0226589 A1), and in further view of KEITH (US 2024/0022565 A1).
REGARDING CLAIM 2, HARIF in view of LARGEY discloses the method of claim 1.
LARGEY further discloses a method/system for controlling a device with non-verbal commands further comprising: receiving output from an [artificial intelligence (AI)] model (LARGEY Par 46 – “The model spectra or their mathematical descriptions may be stored, e.g. in the memory 435. The matching may include, e.g. a determination of various metrics describing quality of fit, and a match probability.”) trained, at least in part, on datasets related to non-gesture sound patterns that can be correlated to control gestures (LARGEY Par 40 – “Training may be accomplished via a training mode. The training mode may, for example, prompt the user with a desired synthesized phonetic command, after which the user may produce one or more nonphonetic commands that the device 400 will henceforth translate to the desired synthesized command. Those skilled in the pertinent art are familiar with various training methods.”), to predict that the detected first pattern of non-speech sounds intended by the user to control the feature is the control gesture rather than a non-gesture sound (LARGEY Par 36 – “The command discriminator 430 may determine the occurrence of an audio command when a detected spectral signature matches one of several model signatures stored in the memory 435. The audio command may be spectrally compact, providing a high degree of confidence that the audio command is present in the received audio stream.”; Par 53 – “Similarly, other embodiments of the functional block 620, e.g. recorder or smart phone, may be configured to receive from the command discriminator 430 an electronic signal indicating the occurrence of a nonphonetic command, or may receive a synthesized voice command, and then operate to perform its core functionality, respectively for example recording and calling.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF to include training a model with gesture and non-gesture sounds, as taught by LARGEY.
One of ordinary skill would have been motivated to include training a model with gesture and non-gesture sounds, in order to allow a user to train a model tailored for the user.
HARIF in view of LARGEY does not explicitly an [artificial intelligence (AI)] model.
KEITH discloses a method/system for AI model trained on non-gestures sounds and gesture sounds further comprising:
receiving output from an [artificial intelligence (AI)] model trained (KEITH Par 587 – “Using the method described, the technology is able to collect the sleeping and breathing conditions, and using advanced machine learning/AI, is able to record various kinds of apneic events and other normal or abnormal sleeping activities.”; Par 597 – “In the step 5904, results of the analysis are utilized in performing a function. The function is able to be an alarm during or after sleep, providing the data to the user, a doctor or other professional, and/or any other function. For example, if a child is being monitored for sleep apnea, a signal is able to be sent to another device (e.g., in the parents' room) to alert the parents that the child is having an apneic episode.”; Par 587 – “Using the method described, the technology is able to collect the sleeping and breathing conditions, and using advanced machine learning/AI, is able to record various kinds of apneic events and other normal or abnormal sleeping activities.”), at least in part, on datasets related to non-gesture sound patterns that can be correlated to control gestures (KEITH Par 586 – “The sleep apnea method/device described herein is able to be located near the sleeping patient and uses sound, motion and several other sensors to monitor the patient for a number of factors including the position of the patient during sleep, the breath patterns, the movements of the patient such as toss/turn movements and leg flapping.”; Par 515 – “The device is able to continuously gather data including various input and behavioral output which are then able to be analyzed via machine learning to determine if any correlations or patterns are determined. …The device is able to detect when someone is walking slower than usual, breathing differently, has more accidents, performs poorly on tests, and/or any other behavioral performance changes.”; Par 570 – “Generate personal behavioral baseline models. The specific human is monitored to generate a baseline behavioral model. This baseline would be considered normal behaviors. Activities or human conditions beyond a threshold of normalcy would be compared to a common model to identify conditions for potential undesired outcomes.”; Par 576 – “In the step 5802, a personal behavioral baseline model for each user is generated. The specific user/human is monitored to generate a baseline behavioral model. The baseline is considered “normal” behaviors (e.g., the user is not having a psychological event).”; Par 587 – “Using the method described, the technology is able to collect the sleeping and breathing conditions, and using advanced machine learning/AI, is able to record various kinds of apneic events and other normal or abnormal sleeping activities.”), to predict that the detected first pattern of non-speech sounds intended by the user to control the feature is the control gesture rather than a non-gesture sound (KEITH Par 587 – “Using the method described, the technology is able to collect the sleeping and breathing conditions, and using advanced machine learning/AI, is able to record various kinds of apneic events and other normal or abnormal sleeping activities.”; Par 591 – “For example, machine learning is implemented by analyzing many datasets of sleep apnea to learn what sounds, movements, patterns occur during sleep apnea. The currently monitored (e.g., real-time) information is then compared with that stored information to determine if an apneic event is currently occurring. For example, if the historical data indicates that a sign of sleep apnea is no breathing for a period of time above a threshold followed by a gasping (or similar) sound, then when a user is sleeping, and no breathing sound is detected for seconds (or another threshold) followed by a loud gasping/inhalation sound, it is able to be considered an apneic episode.”; Par 588 – “Apneic episodes (the cessation of breathing) are able to be identified by the specific sound patterns (or lack thereof). Typically, the cessation of breath for a period of time, then a gasp or bodily jerk, then the normal continuation of breathing is one example. There are several types of sleep apnea, and each can be identified using this technique. The identification of the breath pattern is able to be identified by machine learning models and AI technologies. The system is able to further identify bodily movements including: restless sleep movements, tossing/turning, leg tossing, and others.”; Par 597 – “In the step 5904, results of the analysis are utilized in performing a function. The function is able to be an alarm during or after sleep, providing the data to the user, a doctor or other professional, and/or any other function. For example, if a child is being monitored for sleep apnea, a signal is able to be sent to another device (e.g., in the parents' room) to alert the parents that the child is having an apneic episode.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF in view of LARGEY to include training an AI model with gesture and non-gesture sounds, as taught by KEITH.
One of ordinary skill would have been motivated to include training an AI model with gesture and non-gesture sounds, in order to more accurately detect a gesture sound.
REGARDING CLAIM 3, HARIF in view of LARGEY discloses the method of claim 1.
HARIF in view of LARGEY does not explicitly teach a distinct pattern of breathing different from regular breathing patterns of the user.
KEITH further disclose the method/system, wherein the control gesture includes a distinct pattern of breathing that is different from regular breathing patterns of the user (KEITH Par 515 – “The device is able to continuously gather data including various input and behavioral output which are then able to be analyzed via machine learning to determine if any correlations or patterns are determined. …The device is able to detect when someone is walking slower than usual, breathing differently, has more accidents, performs poorly on tests, and/or any other behavioral performance changes.”; Par 591 – “For example, machine learning is implemented by analyzing many datasets of sleep apnea to learn what sounds, movements, patterns occur during sleep apnea. The currently monitored (e.g., real-time) information is then compared with that stored information to determine if an apneic event is currently occurring. For example, if the historical data indicates that a sign of sleep apnea is no breathing for a period of time above a threshold followed by a gasping (or similar) sound, then when a user is sleeping, and no breathing sound is detected for seconds (or another threshold) followed by a loud gasping/inhalation sound, it is able to be considered an apneic episode.”), wherein the distinct pattern includes at least one variation in a particular rate of inhale and/or exhale and includes a predefined hold time after exhale and/or after inhale (KEITH Par 432 –“ Breathing patterns are able to be detected at other times as well (e.g., when the user is not talking). Breath(ing) patterns are able to be detected by measuring a duration between each breath (e.g., start to start or end to end), the volume of each breath (e.g., in decibels), and detecting the duration and volume over a period of time (e.g., 5 s, 30 s) to determine a pattern similar to a heart beat. Detecting a breath is able to be performed by sound matching (e.g., machine learning learns what a breath sound is. Moreover, each aspect of a breath is able to be detected. For example, a breath in makes a different sound than a breath out, and each is able to be detected. Similarly, there are pauses between each breath, where the amount of time of the pause is able to be slightly different for each user.”; Par 588 – “Typically, the cessation of breath for a period of time, then a gasp or bodily jerk, then the normal continuation of breathing is one example.”; Par 591 – “For example, if the historical data indicates that a sign of sleep apnea is no breathing for a period of time above a threshold followed by a gasping (or similar) sound, then when a user is sleeping, and no breathing sound is detected for seconds (or another threshold) followed by a loud gasping/inhalation sound, it is able to be considered an apneic episode.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF in view of LARGEY to include a distinct breathing pattern.
One of ordinary skill would have been motivated to include a distinct breathing pattern, in order to accurately detect a sleep disorder of a user and to inform the user.
REGARDING CLAIM 6, HARIF in view of LARGEY discloses the method of claim 1, further comprising:
receiving a second pattern of non-speech sounds (HARIF Fig. 2 – “Sound Commands 52”; Col 6:18-6:43 – “If the decision from step 81 is Yes, then a further determination is made in decision step 82 as to whether a non-verbal sound has been recognized.”);
[gathering context information associated with the second pattern of non-speech sounds];
applying one or more non-gesture sound factors (HARIF Col 4:24-54 – “Thus, a comparison 55 is made of an input of non-verbal sound to the stored non-verbal sound commands 52 and recognized non-verbal sounds are input via display adapter 36 to display 38 for verification, as will hereinafter be described.”) to identify the second pattern of non-speech sounds as a non-gesture sound (HARIF Col 4:55-5:15 – “The cursor 68 may then be moved by commands, e.g. hand clap moves cursor to the right, tongue/mouth clack moves cursor to the left, knocking on desk moves cursor up and metallic tapping moves cursor down. The sound recognition and command execution may be set up so that a sequence of the command sounds (claps, clacks, knocks and taps) accelerate the cursor in the selected direction.”); and
rejecting the second pattern of non-speech sounds for control of the feature (HARIF Col 6:18-43 – “If the determination from step 82 is Yes, a non-verbal sound is recognized, then, step 84, that sound is compared with the stored command sounds. If there is No compare, then there is displayed to user: “Do Not Recognize”, step 90.”).
HARIF in view of LARGEY does not explicitly teach the [square-bracketed] limitation.
KEITH discloses the [square-bracketed] limitation. KEITH discloses a method/system for AI model trained on non-gestures sounds and gesture sounds further comprising:
receiving a second pattern of non-speech sounds (KEITH Par 178 – “For example, if a user has been continuously using his device as he normally does, his gait matches the stored information, and his resulting trust score is 100 (out of 100) and there have been no anomalies with the user's device (e.g., the risk score is 0 out of 100), then there may be no need for further authentication / verification of the user.”; Par 441 – “FIG. 44 illustrates a diagram of performing breath pattern analytics according to some embodiments. The user is able to hold a mobile device 4400 (e.g., a smart phone) and talk as the user typically would. The microphone, camera, and/or sensors of the mobile device 4400 are able to detect and capture the user's breath information. In some embodiments, the mobile device 4400 processes the breath information using the processor and memory of the device. Processing is able to include sound/signal processing such as using filters, masks and machine learning to determine specific breath information among other sound information. The processed information is able to be compared with stored breath information to determine if the currently acquired information is a match of previously stored information.”);
[gathering context information associated with the second pattern of non-speech sounds] (KEITH Par 178 – “In some embodiments, MFA includes behavioral analytics, where the device continuously analyzes the user's behavior as described herein to determine a trust score for the user. The device (or system) determines a risk score for the user based on environmental factors such as where the device currently is, previous logins/locations, and more, and the risk score affects the user's confidence score. In some embodiments, the scan of a dynamic optical mark is only implemented if the user's trust score (or confidence score) is below a threshold. For example, if a user has been continuously using his device as he normally does, his gait matches the stored information, and his resulting trust score is 100 (out of 100) and there have been no anomalies with the user's device (e.g., the risk score is 0 out of 100), then there may be no need for further authentication/verification of the user.”; Par 445 – “Analytics performed in this class can quickly and accurately identify the specific user. Examples of these analytics include: live face recognition—since the user is probably staring at the personal or stationary access device, the face will likely be available to the built-in device camera; voice pattern and quality analytics; breath pattern and quality analysis; external factors including location patterns, user height, environmental and weather; and micro-motion analytics.”);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF in view of LARGEY to include gathering context information, as taught by KEITH.
One of ordinary skill would have been motivated to include gathering context information, in order to more accurately identify a specific user.
CLAIM 9 is similar to claim 2; thus, it is rejected under the same rationale.
CLAIM 10 is similar to claim 3; thus, it is rejected under the same rationale.
CLAIM 13 is similar to claim 6; thus, it is rejected under the same rationale.
CLAIM 16 is similar to claim 2; thus, it is rejected under the same rationale.
CLAIM 17 is similar to claim 3; thus, it is rejected under the same rationale.
CLAIM 19 is similar to claim 6; thus, it is rejected under the same rationale.
Claims 5, 12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over HARIF (US 6,820,056 B1), and in further view of LARGEY (US 2013/0226589 A1), and in further view of SCHUSTER (US 20140350926 A1).
REGARDING CLAIM 5, HARIF in view of LARGEY discloses the method of claim 1.
HARIF further disclose the method/system, wherein the feature includes [audio beam focusing] movement of an object and wherein the feedback indicator includes a notification of a section of [a sound field] the object that the [audio beam focusing] movement of an object is directed (HARIF Col 5:34-49 – “For example, if the user desires to control cursor movements, he may be presented with a default menu: hand clap move to right; mouth-tongue crack move to left … ”; Col 4:24-54 – “Some examples are vocal: long and short whistles, coughs or hacks, teeth clicks, mouth-tongue clacks and hisses; or manual-physical: knocking on a desk, tapping on a computer case with a metallic object, clapping hands or rubbing sounds. These sounds may be discerned by the above-described voice recognition apparatus based upon digitized sound patterns. Since the sounds are more distinct from each other and from speech words than the standard distinctions between speech words and verbal commands, such sounds are easily recognizable and distinguished by the recognition apparatus and programs. Thus, a comparison 55 is made of an input of non-verbal sound to the stored non-verbal sound commands 52 and recognized non-verbal sounds are input via display adapter 36 to display 38 for verification, as will hereinafter be described.”).
HARIF in view of LARGEY does not explicitly teach the [square-bracketed] limitations, and teaches the underlined features instead. In other words, HARIF teaches displaying a recognized non-verbal sound. Thus, when a “mouth-tongue crack” sound is recognized, the command “move to left” will be displayed as a feedback for verification.
SCHUSTER discloses the [square-bracketed] limitations. SCHUSTER teaches a method/system for receiving control commands, wherein the feature includes [audio beam focusing] (SCHUSTER Par 41 – “For example, the apparatus 100 operator may command the beamformer 120 to change the direction of the beamform using commands such as “focus left”, “focus right”, “focus forward” (or “focus ahead”), etc. In response to these or similar voice commands, the beamformer controller 140 will accordingly adjust one or more of the filters 121, 123 or 125 to fulfill the command. In some embodiments, the beamformer controller 140 may access system memory 170 to obtain predetermined filter coefficient settings related to beamforms corresponding to given commands. For example, a set of predetermined filter coefficients may be stored in system memory 170 for beamforms focused in various directions (“left”, “right”, “up”, “down”, “straight ahead”, etc.) that may be accessed by the beamformer controller 140 in response to corresponding commands.”).
Since HARIF in view of LARGEY already teaches generating a feedback associated a non-speech command (e.g., “move to the left”), the combination of HARIF in view of LARGEY, and SCHUSTER teaches a non-speech command to control a direction of the audio beam focus (“focus left”) and generating a feedback associated with the command (“focus left”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of HARIF in view of LARGEY to include audio beam focusing, as taught by SCHUSTER.
One of ordinary skill would have been motivated to include audio beam focusing, in order to allow a user to efficiently control a variety types of electronic devices with desired functionalities including audio beam focusing.
CLAIM 12 is similar to claim 5; thus, it is rejected under the same rationale.
CLAIM 18 is similar to claim 5; thus, it is rejected under the same rationale.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN C KIM whose telephone number is (571)272-3327. The examiner can normally be reached Monday to Friday 8:00 AM thru 4:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONATHAN C KIM/Primary Examiner, Art Unit 2655