Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAIL ACTION
Priority
This application claims priority to U.S provisional Patent Application No. 63558821, filed on 2/28/2024 and is hereby incorporated by references.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lovchinsky et al (US 12598434 B1) in view of Coppo et al (US 20250006175 A1).
Regarding claim 1, Lovchinsky discloses a system [e.g. FIG. 3-4] for training a speech enhancement neural network personalized for a target speaker's voice [e.g. FIG. 1 and 3-4; a hearing aid having a de-noising model trainable with data collected by the hearing aid during operation includes a microphone], the system comprising:
an ear-worn device [e.g. FIG. 1; 110] configured to run the speech enhancement neural network [e.g. FIG. 3; column 1 lines 31-40; de-noises the acoustic signal using a neural network]:
a processing device in communication with an ear-worn device [e.g. FIG. 1 and 3; 320]; and
one or more servers in communication with the processing device [e.g. 2 and 5; a separate electronic device, such as a mobile phone or tablet or a cloud-based server], in communication with the ear-worn device; wherein:
the one or more servers are configured to:
receive one or more recordings of the target speaker's voice [e.g. FIG. 1 and 3-5; collecting training data by microphone for speaker recognition; training data including clean speech and speech mixed with noise];
generate, using a voice neural network and the one or more recordings of the target speaker's voice, one or more samples of the target speaker's voice [e.g. FIG. 3-5; voice samples may be recorded using a smartphone with the speaker being at different distances to obtain samples of different quality. Voice signatures may be used by the machine learning model implementing]; and
generate, using the one or more samples of the target speaker's voice, one or more personalized parameters [e.g. FIG. 4-5; the training may modify weights applied at different levels of a neural network. The training results in the external device or server obtaining an updated model] for the speech enhancement neural network; and the processing device is configured to:
transmit the one or more personalized parameters for the speech enhancement neural network to the ear-worn device [e.g. FIG. 5; 510; the updated model with updated weights parameters is provided to the ear-worn device 110].
It is noted that Lovchinsky differs to the present invention in that Lovchinsky fails to explicitly disclose a concept of a target speaker’s cloned voice.
However, Coppo teaches the well-known concept of receiving one or more recordings of the target speaker's voice [e.g. FIG. 1; [0043]; collecting speaker’s audio signal]; generate, using a voice cloning neural network and the one or more recordings of the target speaker's voice [e.g. FIG. 1; speech (voice) cloning system, a neural network-based speech synthesizing system] one or more samples of the target speaker's voice cloned voice [e.g. FIG. 1-2; cloned voice]; and generate, using the one or more samples of the target speaker's cloned voice [cloned speech samples , one or more personalized parameters for the speech enhancement neural network [e.g. FIG. 2; [0006]; the trained configuration (e.g., weights of a neural network-based implementation of one or more of the components/sub-system of the synthesizing system) based on the short speech sample provided for the target speaker and annotation information matched to the speech sample, to provide a more optimal performance with respect to the specific target speaker whose speech is to be cloned].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 2, Lovchinsky and Coppo further disclose the processing device is a first processing device [e.g. Lovchinsky: FIG. 1 and 3; 110], and wherein the one or more servers are configured, when receiving the one or more recordings of the target speaker's voice [e.g. Lovchinsky: FIG. 4-5; the training may modify weights applied at different levels of a neural network. The training results in the external device or server obtaining an updated model], to receive the one or more recordings of the target speaker's voice from the first processing device or a second processing device [e.g. Lovchinsky: FIG. 1 and 3; 110 or 125; Coppo: FIG. 1-2; processor-based devices to generate synthesized speech].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 3, Lovchinsky and Coppo further disclose the first processing device or the second processing device is further configured to provide an instruction to capture or select the one or more recordings of the target speaker's voice [e.g. Lovchinsky: FIG. 1 and 3-4; capturing audio signal from microphone; identify target speaker; Coppo: FIG. 1-2; obtaining a speech sample for a target speaker].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 4, Lovchinsky and Coppo further disclose the first processing device or the second processing device is configured, when providing the instruction to capture or select the one or more recordings of the target speaker's voice [e.g. Lovchinsky: FIG. 1 and 3-4; capturing audio signal from microphone; identify target speaker; Coppo: FIG. 1-2; obtaining a speech sample for a target speaker], to provide the instruction as text displayed on a display screen [e.g. Lovchinsky: FIG. 1; column 8 lines 1-25; text notification via a smartphone; Coppo: FIG. 1-2; a pre-determined text passage is presented to the user via the display].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 5, Lovchinsky and Coppo further disclose providing the instruction to capture or select the one or more recordings of the target speaker's voice, to provide an instruction to capture a recording of a certain minimum length of time [e.g. Coppo: FIG. 1-2; [0002]; target speaker’s voice data for 2-3 minutes].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 6, Lovchinsky and Coppo further disclose the minimum length of time is equal to 10 seconds, equal to 2 minutes [e.g. Coppo: FIG. 1-2; [0002]; target speaker’s voice data for 2-3 minutes], or between 10 seconds and 2 minutes.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 7, Lovchinsky and Coppo further disclose to receive an email or text with a link, and to provide the instruction in response to a user selection of the link [e.g. Coppo: FIG. 1-2; [0044 and 0046]; using a message request sent via a link to re-acquire a new source sample from the target speaker].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 8, Lovchinsky and Coppo further disclose to provide an instruction to upload an already-saved video or audio file that includes the target speaker's voice [e.g. Lovchinsky: FIG. 1 and 3-4; voice samples may be recorded using a smartphone with the speaker being at different distances to obtain samples of different quality; Coppo: FIG. 1-2; uploading audio data].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 9, Lovchinsky and Coppo further disclose the one or more servers are configured when receiving the one or more recordings of the target speaker's voice, to receive only one recording of the target speaker's voice [e.g. Lovchinsky: FIG. 1 and 3; one microphone recoding speaker’s voice; Coppo: FIG. 1-2; speaker’s voice recording].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 10, Lovchinsky and Coppo further disclose the ear-worn device comprises a hearing aid, a cochlear implant, or an earphone [e.g. Lovchinsky: FIG. 1-2; ear-worn devices for hearing aids].
Regarding claim 11, Lovchinsky and Coppo further disclose the one or more servers are configured to generate the one or more personalized parameters for the speech enhancement neural network by training the speech enhancement neural network [e.g. Lovchinsky: FIG. 4-5; the training may modify weights applied at different levels of a neural network. The training results in the external device or server obtaining an updated model],
Regarding claim 12, Lovchinsky and Coppo further disclose the one or more servers are configured, when training the speech enhancement neural network, to retrain the speech enhancement neural network [e.g. Lovchinsky: FIG. 4-5; received audio signals to be used in re-training the machine learning model].
Regarding claim 13, Lovchinsky and Coppo further disclose the one or more personalized parameters comprise neural network weights of the speech enhancement neural network [e.g. Lovchinsky: FIG. 4-5; the training may modify weights applied at different levels of a neural network. The training results in the external device or server obtaining an updated model].
Regarding claim 14, Lovchinsky and Coppo further disclose the one or more personalized parameters comprise an embedding of the target speaker’s voice [e.g. Lovchinsky: FIG. 1 and 4-5; determining an embedding of the audio signal; Coppo: FIG. 1-2; define an embedding space representative of at least some voice characteristics for speakers in the large corpus].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 15, Lovchinsky and Coppo further disclose the one or more servers are configured to generate the embedding of the target speaker’s voice by running a speaker embedding neural network on the one or more samples of the target speaker's cloned voice [e.g. Lovchinsky: FIG. 1 and 4-5; determining an embedding of the audio signal; Coppo: FIG. 1-2; define an embedding space representative of at least some voice characteristics for speakers in the large corpus].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 16, Lovchinsky and Coppo further disclose the one or more samples of the target speaker's cloned voice have a higher quality than the one or more recordings of the target speaker's voice [e.g. Lovchinsky: FIG. 1 and 4-5; de-noise audio signal; Coppo: FIG. 1-2; applying filtering and speech enhancement operations on the speech sample to enhance quality of the speech sample].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 17, Lovchinsky and Coppo further disclose processing device is a first processing device, and wherein the first processing device or a second processing device is configured to capture the one or more recordings of the target speaker's voice using a microphone or microphones of the first processing device or the second processing device [e.g. Lovchinsky: FIG. 1-3; microphone(s); Coppo: FIG. 1-2; microphone].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Regarding claim 18, Lovchinsky and Coppo further disclose the system is further configured to capture the one or more recordings of the target speaker's voice automatically [e.g. Lovchinsky: FIG. 1-3; the controller may determine which data to store; Coppo: FIG. 1-2].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]].
Claim(s) 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lovchinsky et al (US 12598434 B1) in view of Coppo et al (US 20250006175 A1) and Singer et al (US 11871190 B2).
Regarding claim 19, Lovchinsky and Coppo further disclose the one or more recordings are received [e.g. Lovchinsky: FIG. 1 and 3-4; capturing audio signal from microphone; identify target speaker; Coppo: FIG. 1-2; obtaining a speech sample for a target speaker], but Lovchinsky and Coppo fail to disclose the detail of the one or more recordings.
However, Singer teaches the well-known concept of the one or more recordings from the microphones being received synchronously [e.g. FIG. 1-2 and 4; recordings from synchronous microphones].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo and the well-known concept of receiving audio recordings synchronously technique taught by Singer as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]] and performing multi-microphone source separation without resampling [See Singer; column 1 lines 58-67].
Regarding claim 20, Lovchinsky, Coppo and Singer further disclose the one or more recordings are received asynchronously [e.g. Lovchinsky: FIG. 3-5; voice samples may be recorded using a smartphone with the speaker being at different distances to obtain samples of different quality; Coppo: FIG. 1-2; Singer: FIG. 1-2 and 4; recordings from asynchronous microphones].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify the hearing aids system disclosed by Lovchinsky to exploit the well-known concept of a voice cloning technique taught by Coppo and the well-known concept of receiving audio recordings synchronously technique taught by Singer as above, in order to provide improved voice cloning that better imitates the prosody of the target speaker [See Coppo; [0008]] and performing multi-microphone source separation without resampling [See Singer; column 1 lines 58-67].
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
McKinney et al (US 20220279290 A1).
KURIHARA (US 20230239617 A1).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZHUBING REN whose telephone number is (571)272-2788. The examiner can normally be reached Monday-Friday 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZHUBING REN/ Primary Examiner, Art Unit 2658