DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claims 11 and 12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-10 and 13-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by
Hunt (US 2020/0344544 A1).
As per claims 1, 16 and 17, Hunt teaches a method/non-transitory computer readable medium to implement said method/device, comprising:
transmitting, during a first time period, an ultrasound transmit signal that propagates within at least a portion of an ear canal of a person (0045, 0042);
receiving, during the first time period, an ultrasound receive signal, the ultrasound receive signal representing a version of the ultrasound transmit signal with one or more characteristics modified based on the propagation within the ear canal and based on the person speaking a phrase during at least a portion of the first time period (0045-0046, 0078, 0085, 0081-0092); and
recognizing the spoken phrase based on the ultrasound receive signal (0008, 0054, 0059).
As per claim 2, Hunt teaches the method of claim 1, further comprising at least one of the following: generating a control signal that controls an operation of a device based on the spoken phrase; or generating text based on the spoken phrase (0060, 0052).
As per claim 3, Hunt teaches the method of claim 2, wherein: the device comprises a hearable; the transmitting of the ultrasound transmit signal comprises transmitting the ultrasound transmit signal using the hearable (0056); and the receiving of the ultrasound receive signal comprises receiving the ultrasound receive signal using the hearable (0056, 0057).
As per claim 4, Hunt teaches the method of claim 2, wherein: the device comprises a computing device that is coupled to a hearable (0056, 0057); the transmitting of the ultrasound transmit signal comprises transmitting the ultrasound transmit signal using the hearable (0056); and the receiving of the ultrasound receive signal comprises receiving the ultrasound receive signal using the hearable (0056, 0057).
As per claim 5, Hunt teaches the method of claim 1, further comprising: receiving an audio signal comprising the spoken phrase (0007-0008); and wherein the recognizing of the spoken phrase comprises recognizing the spoken phrase based on the ultrasound receive signal and the audio signal (007-0008).
As per claim 6, Hunt teaches the method of claim 5, wherein the recognizing of the spoken phrase further comprises: generating a first spectrogram of a signal derived from the ultrasound receive signal; generating a second spectrogram of the audio signal (0076-0078, 0080-0082, 0095); and generating a feature vector using a machine-learned model by providing the machine-learned model the first spectrogram and the second spectrogram (0076-0078, 0080-0082, 0095); and recognizing the spoken phrase based on the feature vector (0095).
As per claim 7, Hunt teaches the method of claim 6, further comprising: generating a stacked spectrogram comprising a combination of the first spectrogram and the second spectrogram, wherein the generating of the feature vector comprises generating the feature vector using the machine-learned model by providing the machine-learned model the stacked spectrogram as an input (0076-0078, 0080-0082, 0095).
As per claim 8, Hunt teaches the method of claim 7, wherein the machine-learned model comprises: a convolutional neural network (0051) ; or a single-channel transformer having a convolutional layer (0051, 0074, 0095).
As per claim 9, Hunt teaches the method of claim 6, wherein: the generating of the feature vector comprises generating the feature vector using the machine-learned model by providing the first spectrogram and the second spectrogram as separate inputs to the machine-learned model (0076-0078, 0080-0082, 0095); and the machine-learned model comprises: a multiple-input convolutional neural network (0068, 0081, 0098); or a multi-channel transformer comprising separate convolutional layers (0068, 0081, 0098).
As per claim 10, Hunt teaches the method of claim 1, wherein the spoken phrase is silently spoken by the person during at least the portion of the first time period (0047).
As per claim 13, Hunt teaches the method of claim 1, wherein the recognizing of the spoken phrase comprises recognizing the spoken phrase using the ultrasound receive signal and without using one or more of the following: an audio signal that includes the spoken phrase and is captured using passive audio sensing; or another signal obtained from another sensor that is different from an ultrasound sensor (0045-0046, 0078, 0085, 0081-0092).
As per claim 14, Hunt teaches the method of claim 1, further comprising: rendering audible content during the first time period, the rendering causing an audible signal to propagate within at least a portion of the ear canal of the person (0045-0046, 0078, 0085, 0085, 0081-0092).
As per claim 15, Hunt teaches the method of claim 14, wherein: the rendering of the audible content comprises transmitting, during the first time period, audible signal that propagates within at least a portion of the ear canal of the person (0045-0046, 0078, 0085, 0081-92); the ultrasound receive signal comprises an internal noise component caused by interference generated by the rendering of the audible signal (0071); the method further comprises generating a denoised signal by filtering the internal noise component within the ultrasound receive signal based on a version of the audible signal (0071); and the recognizing of the spoken phrase comprises recognizing the spoken phrase based on the denoised signal (0051).
As per claim 18, Hunt teaches the device of claim 17, further comprising: a speaker; and an active-noise-cancellation circuit comprising a feedback microphone, wherein the at least one transducer comprises the speaker and the feedback microphone (0006-0007, 0113).
As per claim 19, Hunt teaches the device of claim 17, wherein: the at least one transducer comprises a speaker and a microphone (0007); the speaker is configured to be positioned proximate to a first ear of a person (0007, 0056, 0050); and the microphone is configured to be positioned proximate to a second ear of the person ( Fig.2, 0007, 0056, 0050).
As per claim 20, Hunt teaches the device of claim 17, wherein the device comprises: at least one earbud (Fig.2)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See attached form PTO-892.
Fan et al., (US 2024/0312478 A1) teach techniques and apparatuses that perform voice activity detection using active acoustic sensing. By transmitting and receiving acoustic signals, a hearable can recognize changes in an acoustic circuit to perform voice activity detection. With active acoustic sensing, the hearable can detect a vocalization made by a user in a noisy and/or loud environment. As such, the hearable can support a voice user interface (VUI) by providing an indication of when the user is speaking. The hearable can also support multi-factor voice authentication to enhance security and provide robust protection from voice attacks. In addition to being relatively unobtrusive, some hearables can be configured to support voice activity detection using active acoustic sensing without the need for additional hardware.
Gong et al., (US 11,334,157 B1) teach a wearable device and system for detecting data generated by a user's contact with a surface and using that data to determine that a user contacted the surface and, in some examples, the location of the contact. The location of the contact may then be used by another device, such as an artificial reality headset, to select a user interface element, activate a function or perform a task in an artificial reality environment, or execute another type of function typically performed in response to an input to a user interface. In one example, the wearable device includes an ultrasound transmitter (Tx) and an ultrasound receiver (Rx). In some embodiments, the wearable device may also include a processing element capable of processing the signals detected or received by the receiver to determine that a contact occurred.
Petrank (US 10,659,867 B2) teaches an audio processing device plays a source audio signal with an electroacoustic transducer of a user earpiece, and records an aural signal that is sensed by same said electroacoustic transducer. The audio processing device determines values of one or more features of the aural signal that indicate a characteristic of a space in which the user earpiece is located. The audio processing device compares the determined values of the one or more features of the aural signal with pre-defined values of the one or more features. Based on a result of the comparing, the audio processing device determines whether the user earpiece is located at a user's ear.
Usher et al., (US 2020/0152185 A1) teach a method and device for voice operated control with learning. The method can include measuring a first sound received from a first microphone, measuring a second sound received from a second microphone, detecting a spoken voice based on an analysis of measurements taken at the first and second microphone, learning from the analysis when the user is speaking and a speaking level in noisy environments, training a decision unit from the learning to be robust to a detection of the spoken voice in the noisy environments, mixing the first sound and the second sound to produce a mixed signal, and controlling the production of the mixed signal based on the learning of one or more aspects of the spoken voice and ambient sounds in the noisy environments.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIJAY B CHAWAN whose telephone number is (571)272-7601. The examiner can normally be reached 7-5 Monday thru Thursday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/VIJAY B CHAWAN/Primary Examiner, Art Unit 2658