Detailed Action
This communication is in response to the Application filed on 1/21/2025.
Claims 1-18 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 8/30/2026 3/12/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent claim 1 recites,
1. A system for generating audio feedback to silent speech, the system comprising: a speaker; and processing circuitry, which is configured to:
generate speech output including the articulated words of a test subject from sensed movements of skin of a face of the test subject in response to words articulated silently by the test subject and without contacting the skin; [This relates to a series of information processing/mental like steps that a human can perform using speech]
convert the speech output into an audio output; [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.]
convey the audio output to the speaker as audio feedback while reducing latency; and [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.]
play the audio feedback with reduced latency to the test subject on the speaker. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.]
Regarding independent Claim 10, Claim 10 is a Method claim with limitations similar to that of claim 1 and is rejected under the same rationale.
The Dependent Claims do not include additional limitations that could incorporate the abstract idea into a practical application or cause the Claim as a whole to amount to significantly more than the underlying abstract idea.
This judicial exception is not integrated into a practical application. In particular, claim 1 recites the additional element of “speaker” and “processor”. For example, in paragraph 20 of the as filed specification, there is the description
earbud (e.g., a second electronic device) may perform computations based on the sensed facial skin micromovements. Distributing the operations thus may reduce the size, cost, and/or heat generated by the first and/or second earbuds. In some instances, one or more processing and/or memory operations may be distributed to additional (e.g., non-wearable) computing resources.
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a processor and speaker is noted as a general computer. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Further, the additional limitation in the claims noted above are directed towards insignificant solution activity. The claims are not patent eligible.
Dependent claim 2 recites,
“2. The system to claim 1, wherein the processing circuitry is configured to reduce latency by achieving 25mS latency. [reducing latency is a mental-like process a human can perform] Processing circuitry is noted as an additional limitation.
Dependent claim 3 recites,
“3. The system according to claim 1, wherein the speaker is comprised in a wearable device. A speaker and device are noted as additional limitations.
Dependent claim 4 recites,
“4. The system according to claim 1, wherein the speaker and the processing circuitry are comprised in a wearable device. A speaker, processing circuitry and device are noted as additional limitations.
Dependent claim 5 recites,
“5. The system to claimn1, wherein the processing circuitry is configured to generate the audio output with the reduced latency by:
running a feature extraction (FE) algorithm to generate FE output; [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.]
inputting the FE output to a neural network (NN) algorithm; [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.]
running the NN algorithm to generate speech data output; and [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.]
inputting the speech data output to an inter-integrated sound (I2S) interface to generate the audio output with the reduced latency.
[This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] processing circuitry is noted as an additional limitation.
Dependent claim 6 recites,
“6. The system according to claim 1, wherein the processing circuitry is comprised in a wearable device. Processing circuitry is noted as an additional limitation.
Dependent claim 7 recites,
“7. The system according to claim 6, wherein the wearable device comprises an earpiece. Wearable device and earpiece are noted as additional limitations.
Dependent claim 8 recites,
“8. The system to claim 1, wherein the processing circuitry Is configured to convey the audio output to the speaker as audio feedback while reducing latency by using a split architecture system for detecting facial skin micromovements for speech detection, comprising:
a first device configured to perform a first subset of functionalities associated with detecting facial skin micromovements for determining non-vocalized speech, the first device comprising the circuitry configured to provide audio feedback While reducing latency; and [audio feedback relates to a series of information processing/mental like steps and mathematical operations that a human can perform.]
at least one second device paired with the first device and configured to perform a second subset of functionalities associated with detecting facial skin for determining non-vocalized speech, in coordination with the first device; and [detecting facial skin relates to a series of information processing/mental like steps that a human can perform using observation.]
convey the audio output to the speaker as audio feedback While reducing latency. [This relates to a series of information processing/mental like steps that a human can perform using speech.] A speaker, processing circuitry and devices are noted as additional limitations.
Dependent claim 9 recites,
“9. The system according to claim 8, wherein the first device comprises an earbud. Device and earbuds are noted as additional limitations.
As to dependent Claim 11, Claim 11 is a method claim with limitations similar to that of claim 2 and is rejected under the same rationale.
As to dependent Claim 12, Claim 12 is a method claim with limitations similar to that of claim 3 and is rejected under the same rationale.
As to dependent Claim 13, Claim 13 is a method claim with limitations similar to that of claim 4 and is rejected under the same rationale.
As to dependent Claim 14, Claim 14 is a method claim with limitations similar to that of claim 5 and is rejected under the same rationale.
As to dependent Claim 15, Claim 15 recites
“15. The method according to claim 10, wherein generating the speech output, converting the speech output into the audio output, and conveying the audio output to a speaker is performed in a wearable device. [this relates to a series of information processing/mental like steps that a human can perform speaking.] A device and speaker are noted as additional limitations.
As to dependent Claim 16, Claim 16 is a method claim with limitations similar to that of claim 7 and is rejected under the same rationale.
As to dependent Claim 17, Claim 17 is a method claim with limitations similar to that of claim 8 and is rejected under the same rationale.
As to dependent Claim 18, Claim 18 is a method claim with limitations similar to that of claim 9 and is rejected under the same rationale.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 5, 8, 10, 11, 14 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over MURAI VON BÜNAU (U.S. Patent Number US 20230154450 A1), in view of ZALEVSKY (U.S. Patent Number US 20170209047 A1).
Regarding Claim 1, MURAI VON BÜNAU teaches 1. A system for generating audio feedback to silent speech, the system comprising: a speaker; and processing circuitry, which is configured to: (see MURAI VON BÜNAU [0030] “…output this waveform as an acoustic wave via a loudspeaker.”) generate speech output including the articulated words of a test subject (see MURAI VON BÜNAU [0060] “…fed into the trained machine learning algorithm (32) to be converted to an acoustic speech waveform in real time, i.e., one or more live audio signals are generated by the machine-learning algorithm…”)(see MURAI VON BÜNAU [0103] “The main elements of the corresponding voice prosthesis are shown in FIG. 7. The patient (53) is confined to the patient bed (54), typically an ICU bed. A power supply (55), radar transmission and receiving electronics (56), signal processing electronics (57), a computing device (58), and audio amplifier (59) are contained in a bedside unit (60). A portable touchscreen device (61) such as a tablet serves as the user interface through which patient and care staff can interact with the system. convert the speech output into an audio output; convey the audio output to the speaker as audio feedback while reducing latency; and (see MURAI VON BÜNAU [0060] “…to be converted to an acoustic speech waveform in real time, i.e., one or more live audio signals are generated by the machine-learning algorithm. The acoustic speech waveform (35) is output using the device loudspeaker.”) (see MURAI VON BÜNAU [0031] “Many of the techniques described herein can employ a process labeled “voice grafting” hereinafter. Voice grafting can be understood in terms of the source-filter model of speech production as follows: For a patient who has partially or completely lost the ability to phonate, but retained at least a partial ability to articulate, the techniques described herein computationally “graft” the patient's time varying filter function, i.e. articulation, onto a source function, i.e. phonation, which is based on the speech output of one or more healthy speakers, in order to synthesize natural sounding speech in real time. Real time can correspond to a processing delay of less than 0.5 seconds, optionally of less than 50 ms.”) play the audio feedback with reduced latency to the test subject on the speaker. (see MURAI VON BÜNAU [0031] “Many of the techniques described herein can employ a process labeled “voice grafting” hereinafter. Voice grafting can be understood in terms of the source-filter model of speech production as follows: For a patient who has partially or completely lost the ability to phonate, but retained at least a partial ability to articulate, the techniques described herein computationally “graft” the patient's time varying filter function, i.e. articulation, onto a source function, i.e. phonation, which is based on the speech output of one or more healthy speakers, in order to synthesize natural sounding speech in real time. Real time can correspond to a processing delay of less than 0.5 seconds, optionally of less than 50 ms.”)
MURAI VON BÜNAU does not specifically teach from sensed movements of skin of a face of the test subject in response to words articulated silently by the test subject and without contacting the skin; However, ZELEVSKY does teach this limitation (see ZELEVSKY [0078] “The source of coherent light 202 emits a light beam 104 to illuminate the object 102 during a certain time period (continuously or by multiple timely separated sessions). The object constitutes a body region of a subject (e.g. individual) whose movement is affected by a change in the body condition, typically a flow of a fluid of interest (i.e. a fluid having a property that is to be measured). The object's diffusive surface responds to coherent illumination by a speckle pattern which propagates toward the imaging optics 112 and is captured by the PDA 111 during said certain time period, to generate output measured data.”) (See ZELEVSKY [0079] “As shown more specifically in FIGS. 2A and 2B, the imaging unit is configured for focusing coherent light on a plane 108 which is displaced from a plane of an object 102 to be monitored. In other words, the back focal plane of the lens 112 is displaced from the object plane thus producing a defocused image of the object. A coherent light beam 104 (e.g., a laser beam) illuminates an object 102, and a secondary speckle pattern is formed as the reflection/scattering of the coherent light beam 104 from the object 102. The secondary speckle pattern is generated because of the diffusive surface of the object 102. The speckle pattern propagates toward the in-focus plane 108, where it takes a form 106. The speckle pattern propagates in a direction along the optical axis of the system, is collected by the imaging lens 112 and is collected by the PDA 111.”) (See ZELEVSKY [0109] “…If the illuminated tissue is part of his head, then the vibrations are proportional to the voice produced by the speaker. Only this signal will be input to the amplification device …”)
MURAI VON BÜNAU and ZELEVSKY are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of MURAI VON BÜNAU to incorporate from sensed movements of skin of a face of the test subject in response to words articulated silently by the test subject and without contacting the skin of ZELEVSKY. This allows for detection of the object's motion as recognized by ZELEVSKY [0080].
Regarding independent Claim 10, Claim 10 is a method claim with limitations similar to that of claim 1 and is rejected under the same rationale.
As to Claim 2, MURAI VON BÜNAU in view of ZELEVSKY teaches: The system to claim 1,
Furthermore, MURAI VON BÜNAU teaches 2. wherein the processing circuitry is configured to reduce latency by achieving 25ms latency. (see MURAI VON BÜNAU [0031] Many of the techniques described herein can employ a process labeled “voice grafting” hereinafter. Voice grafting can be understood in terms of the source-filter model of speech production as follows: For a patient who has partially or completely lost the ability to phonate, but retained at least a partial ability to articulate, the techniques described herein computationally “graft” the patient's time varying filter function, i.e. articulation, onto a source function, i.e. phonation, which is based on the speech output of one or more healthy speakers, in order to synthesize natural sounding speech in real time. Real time can correspond to a processing delay of less than 0.5 seconds, optionally of less than 50 ms. (examiner interprets 25mS latency as “less than 50mS”)”)
As to Claim 5 MURAI VON BÜNAU in view of ZELEVSKY teaches 5. The system to claim 1,
Furthermore, MURAI VON BÜNAU teaches wherein the processing circuitry is configured to generate the audio output with the reduced latency by: running a feature extraction (FE) algorithm to generate FE output; (see MURAI VON BÜNAUa [0086] “…feature vector can be extracted from each frame. The task of the machine learning algorithm is then to transform the time series of feature vectors (cf. blocks 72)”) inputting the FE output to a neural network (NN) algorithm; (see MURAI VON BÜNAU [0091] “Next, various examples for implementing the machine learning algorithm are described. The recognition task can be accomplished using the types of statistical algorithms commonly used in state-of-the-art speech recognition, such as Hidden Markov Models (HMMs), Gaussian Mixture Models (GMMs), or Deep Neural Networks (DNNs). In each of these cases, a statistical model is created the predict the probabilities that a certain time series of feature vectors corresponds to a certain representation, e.g. a certain phoneme, syllable or word. The probabilities are “learned” during the training process by using the feature vectors corresponding to the vocal tract signals in the training data set and the representations of the corresponding sample text as a statistical sample.”) running the NN algorithm to generate speech data output; and (see MURAI VON BÜNAU[0057] Step 2: Training the algorithm (FIG. 4d), i.e., training phase. The second step in the overall process is training a machine learning algorithm, to transform vocal tract signals into the acoustic waveform of acoustic speech output. The audio training data (27) and the vocal tract training data (31) are used to train the machine learning algorithm (32) such as a deep neural net to transform vocal tract data into acoustic speech output. [0058] Thus, as a general rule, a method includes training a machine learning algorithm based on, firstly, one or more reference audio signals. The reference audio signals include a speech output of a reference text. The machine learning algorithm is, secondly, trained based on one or more vocal-tract signals. The one or more vocal-tract signals are associated with an articulation of the reference text by a patient. [0059] The training of the machine learning algorithm according to step 1 and step 2 is distinct from inference using the trained machine learning algorithm. Inference is described in connection with step 3 below. inputting the speech data output to an inter-integrated sound (I2S) interface to generate the audio output with the reduced latency. (see MURAI VON BÜNAU [0104] Two or more antennas (36) are used to collect reflected and transmitted radar signals that encode the time-varying vocal tract shape. The antennas are placed in proximity to the patient's vocal tract, e.g. under the right and left jaw bone. To keep their position stable relative to the vocal tract they can be attached directly to the patient's skin as patch antennas. Each antenna can send and receive modulated electromagnetic signals in a frequency band between 1 kHz and 12 GHz, optionally 1 GHz and 12 GHz, so that (complex) reflection and transmission can be measured. Possible modulations of the signal are: frequency sweep-, stepped frequency sweep-, pulse-, frequency comb-, frequency-, phase-, or amplitude modulation. In addition, a video camera (48) captures a video stream of the patient's face, containing information about the patient's lip and facial movements. The video camera is mounted in front of the patient's face on a cantilever (62) attached to the patient bed. The same cantilever can support the loudspeaker (63) for acoustic speech output.”) (see MURAI VON BÜNAU [0071] “…A multiple input, multiple output (MIMO) configuration can be employed as well.”)
As to Claim 8, MURAI VON BÜNAU and ZELEVSKY teach 8. The system to claim 1,
Furthermore, MURAI VON BÜNAU teaches wherein the processing circuitry Is configured to convey the audio output to the speaker as audio feedback while reducing latency by using a split architecture system (see MURAI VON BÜNAU [0031] Many of the techniques described herein can employ a process labeled “voice grafting” hereinafter. Voice grafting can be understood in terms of the source-filter model of speech production as follows: For a patient who has partially or completely lost the ability to phonate, but retained at least a partial ability to articulate, the techniques described herein computationally “graft” the patient's time varying filter function, i.e. articulation, onto a source function, i.e. phonation, which is based on the speech output of one or more healthy speakers, in order to synthesize natural sounding speech in real time. Real time can correspond to a processing delay of less than 0.5 seconds, optionally of less than 50 ms.”) (see MURAI VON BÜNAU [0078] An example of an electromagnetic auxiliary sensing modality is a video camera recording the motion of the lips and facial features during speech (FIG. 5d). The video camera (48) captures light reflected (47) scattered off the patient's face under ambient or emitted illumination (46). As can be seen from some deaf people's ability to “lip read”, lip and facial movements are a rich source of information about what is being said and how. Also, a camera can be realized in a compact, light-weight, unobtrusive setup. Since in most of the types of impairments considered here, lip and facial movements remain unimpaired, a video camera is thus a preferred implementation for an auxiliary sensing modality. Also, multiple video cameras or depth-sensing cameras, such as cameras using time-of-flight technology, can be used in order to reconstruct three-dimensional facial geometry. [0079] Another example of an electromagnetic auxiliary sensing modality is surface electromyography (EMG) (FIG. 5e). Surface EMG can measure the action potentials of the musculature involved in speech production, providing complementary information to the vocal tract configuration, e.g. by encoding intended loudness. A combination of surface EMG sensors for the extrinsic laryngeal musculature (49) and the neck and facial musculature (50) can be used. Surface EMG can be particularly useful in cases where the extrinsic laryngeal musculature is present and active.”) determining non-vocalized speech, the first device comprising the circuitry configured to provide audio feedback While reducing latency; and (see MURAI VON BÜNAU [0031] “Many of the techniques described herein can employ a process labeled “voice grafting” hereinafter. Voice grafting can be understood in terms of the source-filter model of speech production as follows: For a patient who has partially or completely lost the ability to phonate, but retained at least a partial ability to articulate, the techniques described herein computationally “graft” the patient's time varying filter function, i.e. articulation, onto a source function, i.e. phonation, which is based on the speech output of one or more healthy speakers, in order to synthesize natural sounding speech in real time. Real time can correspond to a processing delay of less than 0.5 seconds, optionally of less than 50 ms.”) convey the audio output to the speaker as audio feedback while reducing latency. (see MURAI VON BÜNAU [0060] “…fed into the trained machine learning algorithm (32) to be converted to an acoustic speech waveform in real time, i.e., one or more live audio signals are generated by the machine-learning algorithm. The acoustic speech waveform (35) is output using the device loudspeaker.”)
Furthermore, ZELEVSKY teaches for detecting facial skin micromovements for speech detection, comprising: a first device configured to perform a first subset of functionalities associated with detecting facial skin micromovements for (see ZELEVSKY [0078] “The source of coherent light 202 emits a light beam 104 to illuminate the object 102 during a certain time period (continuously or by multiple timely separated sessions). The object constitutes a body region of a subject (e.g. individual) whose movement is affected by a change in the body condition, typically a flow of a fluid of interest (i.e. a fluid having a property that is to be measured). The object's diffusive surface responds to coherent illumination by a speckle pattern which propagates toward the imaging optics 112 and is captured by the PDA 111 during said certain time period, to generate output measured data.”) at least one second device paired with the first device and configured to perform a second subset of functionalities associated with detecting facial skin for determining non-vocalized speech, in coordination with the first device; and (See ZELEVSKY [0079] “As shown more specifically in FIGS. 2A and 2B, the imaging unit is configured for focusing coherent light on a plane 108 which is displaced from a plane of an object 102 to be monitored. In other words, the back focal plane of the lens 112 is displaced from the object plane thus producing a defocused image of the object. A coherent light beam 104 (e.g., a laser beam) illuminates an object 102, and a secondary speckle pattern is formed as the reflection/scattering of the coherent light beam 104 from the object 102. The secondary speckle pattern is generated because of the diffusive surface of the object 102. The speckle pattern propagates toward the in-focus plane 108, where it takes a form 106. The speckle pattern propagates in a direction along the optical axis of the system, is collected by the imaging lens 112 and is collected by the PDA 111.”) (See ZELEVSKY [0109] “…If the illuminated tissue is part of his head, then the vibrations are proportional to the voice produced by the speaker. Only this signal will be input to the amplification device …”)
MURAI VON BÜNAU in view of ZELEVSKY are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of MURAI VON BÜNAU to incorporate for detecting facial skin micromovements for speech detection, comprising: a first device configured to perform a first subset of functionalities associated with detecting facial skin micromovements for…at least one second device paired with the first device and configured to perform a second subset of functionalities associated with detecting facial skin for determining non-vocalized speech, in coordination with the first device; and of ZELEVSKY. This allows for detection of the object's motion as recognized by ZELEVSKY [0080].
As to dependent Claim 11, Claim 11 is a method claim with limitations similar to that of claim 2 and is rejected under the same rationale.
As to dependent Claim 14, Claim 14 is a method claim with limitations similar to that of claim 5 and is rejected under the same rationale.
As to dependent Claim 17, Claim 17 is a method claim with limitations similar to that of claim 8 and is rejected under the same rationale.
Claims 3, 4, 6, 12, 13 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over MURAI VON BÜNAU (U.S. Patent Number US 20230154450 A1), in view of ZELEVSKY (U.S. Patent Number US 20170209047 A1) and further in view of Rameau (US Patent Number US 20220208194 A1).
As to Claim 3, MURAI VON BÜNAU in view of ZELEVSKY teaches: 3. The system according to claim 1,
MURAI VON BÜNAU in view of ZELEVSKY do not specifically teach wherein the speaker is comprised in a wearable device. However, Rameau does teach this limitation (see Rameau [0046] Referring to FIG. 1, in various embodiments, a system 100 may include a computing device 110 (or multiple computing devices, co-located or remote to each other) and a personal device 160 (which may be, for example, a wearable device for sensing sEMG signals). In potential embodiments, the personal device 160 may be integrated with the computing device 110 or components thereof. The computing device 110 (or multiple computing devices) may receive and analyze signals acquired via personal device 160. In certain implementations, computing system 110 may be used to control personal device 160. The computing device 110 may include one or more processors and one or more volatile and non-volatile memories for storing computing code and data that are captured, acquired, recorded, and/or generated. The computing device 110 may include a controller 115 that may be configured to exchange control signals with personal device 160 and/or control the analysis of data and interaction with users (e.g., so as to provide text or synthesized speech). The computing device 110 may also include a predictive model training module 120 for training predictive models, and a predictive model application module 130 for applying trained models.”)
MURAI VON BÜNAU in view of ZELEVSKY and Rameau are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of combination of MURAI VON BÜNAU and ZELEVSKY to incorporate wherein the speaker is comprised in a wearable device of Rameau. This allows for provision of synthesized speech as recognized by Rameau [0046].
As to claim 4, MURAI VON BÜNAU in view of ZELEVSKY teaches 4. The system according to claim 1,
MURAI VON BÜNAU in view of ZELEVSKY do not specifically teach wherein the speaker and the processing circuitry are comprised in a wearable device. However, Rameau teach this limitation (See Rameau [0046] “Referring to FIG. 1, in various embodiments, a system 100 may include a computing device 110 (or multiple computing devices, co-located or remote to each other) and a personal device 160 (which may be, for example, a wearable device for sensing sEMG signals). In potential embodiments, the personal device 160 may be integrated with the computing device 110 or components thereof. The computing device 110 (or multiple computing devices) may receive and analyze signals acquired via personal device 160. In certain implementations, computing system 110 may be used to control personal device 160. The computing device 110 may include one or more processors and one or more volatile and non-volatile memories for storing computing code and data that are captured, acquired, recorded, and/or generated. The computing device 110 may include a controller 115 that may be configured to exchange control signals with personal device 160 and/or control the analysis of data and interaction with users (e.g., so as to provide text or synthesized speech). The computing device 110 may also include a predictive model training module 120 for training predictive models, and a predictive model application module 130 for applying trained models.”)
MURAI VON BÜNAU in view of ZELEVSKY and Rameau are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of combination of MURAI VON BÜNAU and ZELEVSKY to incorporate the speaker and the processing circuitry are comprised in a wearable device of Rameau. This allows for provision of synthesized speech as recognized by Rameau [0046].
As to Claim 6, MURAI VON BÜNAU in view of ZELEVSKY teaches: 6. The system according to claim 1,
MURAI VON BÜNAU in view of ZELEVSKY do not specifically teach wherein the processing circuitry is comprised in a wearable device. However, Rameau does teach this limitation (see Rameau [0046] Referring to FIG. 1, in various embodiments, a system 100 may include a computing device 110 (or multiple computing devices, co-located or remote to each other) and a personal device 160 (which may be, for example, a wearable device for sensing sEMG signals). In potential embodiments, the personal device 160 may be integrated with the computing device 110 or components thereof. The computing device 110 (or multiple computing devices) may receive and analyze signals acquired via personal device 160. In certain implementations, computing system 110 may be used to control personal device 160. The computing device 110 may include one or more processors and one or more volatile and non-volatile memories for storing computing code and data that are captured, acquired, recorded, and/or generated. The computing device 110 may include a controller 115 that may be configured to exchange control signals with personal device 160 and/or control the analysis of data and interaction with users (e.g., so as to provide text or synthesized speech). The computing device 110 may also include a predictive model training module 120 for training predictive models, and a predictive model application module 130 for applying trained models.”)
MURAI VON BÜNAU in view of ZELEVSKY and Rameau are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of combination of MURAI VON BÜNAU and ZELEVSKY to incorporate the processing circuitry is comprised in a wearable device of Rameau. This allows for provision of synthesized speech as recognized by Rameau [0046].
As to dependent Claim 12, Claim 12 is a method claim with limitations similar to that of claim 3 and is rejected under the same rationale.
As to dependent Claim 13, Claim 13 is a method claim with limitations similar to that of claim 4 and is rejected under the same rationale.
As to Claim 15, MURAI VON BÜNAU in view of ZELEVSKY teaches: 15. The method according to claim 10,
MURAI VON BÜNAU in view of ZELEVSKY do not specifically teach wherein generating the speech output, converting the speech output into the audio output, and conveying the audio output to a speaker is performed in a wearable device. However, Rameau does teach this limitation (see Rameau [0046] “Referring to FIG. 1, in various embodiments, a system 100 may include a computing device 110 (or multiple computing devices, co-located or remote to each other) and a personal device 160 (which may be, for example, a wearable device for sensing sEMG signals). In potential embodiments, the personal device 160 may be integrated with the computing device 110 or components thereof. The computing device 110 (or multiple computing devices) may receive and analyze signals acquired via personal device 160. In certain implementations, computing system 110 may be used to control personal device 160. The computing device 110 may include one or more processors and one or more volatile and non-volatile memories for storing computing code and data that are captured, acquired, recorded, and/or generated. The computing device 110 may include a controller 115 that may be configured to exchange control signals with personal device 160 and/or control the analysis of data and interaction with users (e.g., so as to provide text or synthesized speech). The computing device 110 may also include a predictive model training module 120 for training predictive models, and a predictive model application module 130 for applying trained models.”) (see Rameau [0056] “…At 660, the dataset may be fed to a trained predictive model to generate predicted words or phrases corresponding to the sEMG signals recorded during utterances by the subject. At 665, the output of the model (comprising, e.g., predictions) may be presented as, for example, text and/or synthesized speech, and/or may be stored (e.g., in database 155). Process 600 may return to step 650 so as to continue recording sEMG signals to recognize subsequent speech, or may end at 695.”) (see Rameau [0019] “…and/or as audible synthesized speech from an audio source (e.g., a speaker of the computing device or other device).”)
MURAI VON BÜNAU in view of ZELEVSKY and Rameau are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of combination of MURAI VON BÜNAU and ZELEVSKY to incorporate generating the speech output, converting the speech output into the audio output, and conveying the audio output to a speaker is performed in a wearable device of Rameau. This allows for provision of synthesized speech as recognized by Rameau [0046].
Claims 7, 9, 16 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over MURAI VON BÜNAU (U.S. Patent Number US 20230154450 A1), in view of ZELEVSKY (U.S. Patent Number US 20170209047 A1) and further in view of Rafii (U.S. Patent Number US 10997970 B1)
As to Claim 7 MURAI VON BÜNAU in view of ZELEVSKY teaches 7. The system according to claim 6,
MURAI VON BÜNAU in view of ZELEVSKY do not specifically teach wherein the wearable device comprises an earpiece. However, Rafii does teach this limitation (see Rafii (15:15-21) “(58) …Such model and signal processing transformation preferably can be embedded in a hearing aid device that Mary can wear, or can be built into (or embedded in) a speaker system, a smart speaker, a smart headset, earbuds, a mobile device, a virtual assistant,”) (17:1-15) (66) Returning to the Mary-Paul sessions, although the pitch and tone of the trained voice need not be exactly the same as Paul's voice, preferably the pitch and tone of the trained voice is selected from a small population of available voices closest to Paul. For instance, if Paul voice has a low pitch (as in a male voice), the trained voice should be selected from a low pitch male voice, and so forth. Alternatively, in another embodiment of the present invention, a machine learning model may be used that performs a generic transformation on the content (or transcription) of the voice, and then add the pitch and timbre (or acoustic content) to the reconstruction stage of the output voice similar to the function of block 70-2 in FIG. 1B.”)
MURAI VON BÜNAU in view of ZELEVSKY and Rafii are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of combination of MURAI VON BÜNAU and ZELEVSKY to incorporate the wearable device comprises an earpiece of Rafii. This allows users to hear speech spoken more clearly as recognized by Rafii (15:15-21).
As to Claim 9, MURAI VON BÜNAU and ZELEVSKY teach 9. The system according to claim 8,
MURAI VON BÜNAU ZELEVSKY Li do not specifically teach wherein the first device comprises an earbud. However, Rafii does teach this limitation (see Rafii “(15:15-21) “(58) …Such model and signal processing transformation preferably can be embedded in a hearing aid device that Mary can wear, or can be built into (or embedded in) a speaker system, a smart speaker, a smart headset, earbuds, a mobile device, a virtual assistant,”) (17:1-15) “(66) Returning to the Mary-Paul sessions, although the pitch and tone of the trained voice need not be exactly the same as Paul's voice, preferably the pitch and tone of the trained voice is selected from a small population of available voices closest to Paul. For instance, if Paul voice has a low pitch (as in a male voice), the trained voice should be selected from a low pitch male voice, and so forth. Alternatively, in another embodiment of the present invention, a machine learning model may be used that performs a generic transformation on the content (or transcription) of the voice, and then add the pitch and timbre (or acoustic content) to the reconstruction stage of the output voice similar to the function of block 70-2 in FIG. 1B.”)
MURAI VON BÜNAU in view of ZELEVSKY and Rafii are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of combination of MURAI VON BÜNAU and ZELEVSKY to incorporate the first device comprises an earbud of Rafii. This allows This allows users to hear speech spoken more clearly as recognized by Rafii (15:15-21).
As to dependent Claim 16, Claim 16 is a method claim with limitations similar to that of claim 7 and is rejected under the same rationale.
As to dependent Claim 18, Claim 18 is a method claim with limitations similar to that of claim 9 and is rejected under the same rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Begum (US Publication No. US-20160093284-A1) METHOD AND APPARATUS TO SYNTHESIZE VOICE BASED ON FACIAL STRUCTURES
WEISSBERG (US Publication No. US-20170103748-A1) SYSTEM AND METHOD FOR EXTRACTING AND USING PROSODY FEATURES
Trisnadi (US Publication No. US-20160209929-A1) METHOD AND SYSTEM FOR THREE-DIMENSIONAL MOTION-TRACKING
SYRDAL (US Publication No. US-20210201911-A1) SYSTEM AND METHOD FOR DYNAMIC FACIAL FEATURES FOR SPEAKER RECOGNITION
Moghadamfalahi (US Publication No. US-20190295566-A1) METHODS, SYSTEMS AND APPARATUSES FOR INNER VOICE RECOVERY FROM NEURAL ACTIVATION RELATING TO SUB-VOCALIZATION
Vemury (US Publication No. US-20210279492-A1) SKIN REFLECTANCE IMAGE CORRECTION IN BIOMETRIC IMAGE CAPTURE
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KRISTEN MICHELLE MASTERS whose telephone number is (703)756-1274. The examiner can normally be reached M-F 8:30 AM - 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KRISTEN MICHELLE MASTERS/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659