Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
The terminal disclaimer filed on 06/10/2026 disclaiming the terminal portion of any patent granted on this application which would extend beyond the expiration date of U.S. Patent No. 12086558 has been reviewed and is accepted. The terminal disclaimer has been recorded.
Response to Arguments
Applicant's arguments with respect to claims 1, 8, and 15 have been considered but are moot in view of the new ground(s) of rejection. Applicant’s arguments are directed to the amended subject matter; new prior art citations and explanation from Mahyar are provided in light of the amendments. For instance, col 1 line 45 to col 2 line 27 an overview of extracting and breaking down voice characteristics then recombining or concatenating per se, in col 13 lines 28-38 utterance portions such as words or phrases or utterances themselves are combined/concatenated and particularly in an accent or region as well for non-limiting example using the concept e.g. utterance types under BRI, such as high/low frequencies, silent inputs, durations, frequency range, envelope, etc. col 11 line 62 to col 12 line 30 and including accents per se col 12 line 62 to col 13 lines 20
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US 10930263 B1 Mahyar; Hooman (hereinafter Mahyar).
Re claim 1, Mahyar teaches
1. A computer-implemented method for automated voice casting, the computer-implemented method comprising: (fig. 1 voice dubbing or casting)
retrieving, by one or more processors, a primary voice sample that comprises a plurality of primary utterances in a primary language, wherein the primary voice sample corresponds to a primary speaker; (matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)
determining, by the one or more processors, via a neural network, a primary embedding associated with the primary voice sample; (using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
retrieving, by the one or more processors, a plurality of candidate voice samples each comprising a plurality of candidate utterances from a candidate speaker in a target language different from the primary language, each candidate utterance of the plurality of candidate utterances associated with an utterance type; (utterance types under BRI, such as high/low frequencies, silent inputs, durations, frequency range, envelope, etc. col 11 line 62 to col 12 line 30 and including accents per se col 12 line 62 to col 13 lines 20 … and as in only certain models are initially activated e.g. for a language match, then selected models are chosen based on further criteria Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
selecting, by the one or more processors, a set of candidate utterances from the plurality of candidate utterances, wherein the set of candidate utterances meets at least one predetermined candidate utterance type criterion, wherein each of the at least one predetermined candidate utterance type criterion includes a minimum frequency and/or maximum frequency; based on the selecting, generating, by the one or more processors, a candidate voice sample for each candidate voice sample comprising the set of candidate utterances; (only certain models are initially activated e.g. for a language match, then selected models are chosen based on further criteria Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
selecting, by the one or more processors, a set of candidate utterances from the plurality of candidate utterances, wherein the set of candidate utterances meets at least one predetermined candidate utterance type criterion, wherein each of the at least one predetermined candidate utterance type criterion includes a minimum frequency and/or maximum frequency; (preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
based on the selecting, generating, by the one or more processors, a candidate voice sample for each candidate voice sample comprising the set of candidate utterances by deconstructing and recombining the primary voice sample according to the utterance type of each of the set of candidate utterances; (col 1 line 45 to col 2 line 27 an overview of extracting and breaking down voice characteristics then recombining or concatenating per se, and further in col 13 lines 28-38 utterance portions such as words or phrases or utterances themselves are combined/concatenated and particularly in an accent or region as well for non-limiting example using the concept e.g. utterance types under BRI, such as high/low frequencies, silent inputs, durations, frequency range, envelope, etc. col 11 line 62 to col 12 line 30 and including accents per se col 12 line 62 to col 13 lines 20 … and as in only certain models are initially activated e.g. for a language match, then selected models are chosen based on further criteria Col 11 line 14 – col 12 line 44… selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
determining, by the one or more processors, via the neural network, a candidate embedding for each candidate voice sample; (the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
generating, by the one or more processors, for each candidate voice sample, a similarity score for the primary voice sample by comparing the primary embedding and the candidate embedding; and (scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
identifying, by the one or more processors, a specific candidate speaker as providing a vocal match to the primary speaker based on the similarity score. (selecting the closest sounding match based on language and min/max frequency via vectors/embedding similarity illustrated in fig. 1, i.e. criteria scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
Re claim 8, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope
For instance, see fig. 3 hardware components.
Re claim 15, this claim has been rejected for teaching a broader, or narrower claim based on general inclusion of hardware alone (e.g. processor, memory, instructions), representation of claim 1 omitting/including hardware for instance, otherwise amounting to a virtually identical scope
For instance, see fig. 3 memory on a physical device.
Re claims 2, 9, and 16,
2. The computer-implemented method of claim 1, the computer-implemented method further comprising: determining, by the one or more processors, that at least one predetermined utterance is not present in the plurality of candidate utterances of the candidate voice sample. (some candidates will be excluded initially…selecting the closest sounding match based on language and min/max frequency via vectors/embedding similarity illustrated in fig. 1, i.e. criteria scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
Re claims 3, 10, and 17, Mahyar teaches
The computer-implemented method of claim 2, the computer-implemented method further comprising: accessing, by the one or more processors, a plurality of voice samples of the corresponding candidate stored in a database; and (from a database of samples in models…selecting the closest sounding match based on language and min/max frequency via vectors/embedding similarity illustrated in fig. 1, i.e. criteria scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
determining, by the one or more processors, the at least one predetermined utterance is present in at least one of the plurality of voice samples of the corresponding candidate. (selecting the closest sounding match based on language and min/max frequency via vectors/embedding similarity illustrated in fig. 1, i.e. criteria scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
Re claims 4, 11, and 18, Mahyar teaches
4. The computer-implemented method of claim 2, the computer-implemented method further comprising: generating, by the one or more processors, a notification indicating that the at least one predetermined utterance is not present in the candidate voice sample; and (the results on the customer interface is the notification whether by sound or display, bad matches excluded col 13 line 58 to col 14 line 8, note: spec of present invention silent on “notification” details…selecting the closest sounding match based on language and min/max frequency via vectors/embedding similarity illustrated in fig. 1, i.e. criteria scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
outputting, by the one or more processors, the notification to a user interface of a computing device. (the results on the customer interface is the notification whether by sound or display, bad matches excluded col 13 line 58 to col 14 line 8, note: spec of present invention silent on “notification” details…selecting the closest sounding match based on language and min/max frequency via vectors/embedding similarity illustrated in fig. 1, i.e. criteria scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
Re claims 5, 12, and 19, Mahyar teaches
5. The computer-implemented method of claim 1, the identifying further comprising: identifying, by the one or more processors, a second candidate speaker as providing a vocal match to the primary speaker based on the similarity score of the primary embedding and a second candidate embedding; (criteria scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
generating, by the one or more processors, an utterance similarity score by comparing the primary embedding and the respective candidate embedding for each utterance type of the utterance type criterion; and (language and/or frequency among other criteria, scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
prioritizing, by the one or more processors, the second candidate speaker based on the utterance similarity score. (best selected, language and/or frequency among other criteria, scoring matches using the models where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
Re claims 6, 13, and 20, Mahyar teaches
6. The computer-implemented method of claim 1, wherein determining, via the neural network, the primary embedding associated with the primary voice sample includes: inputting, by the one or more processors, the primary voice sample into an encoder, wherein the encoder includes a trained deep learning neural network; and (using the models + DNN where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
generating, by the one or more processors, via the encoder, the primary embedding, wherein the primary embedding includes a multi-dimensional embedding. (both initial and primary are vectors/embeddings… using the models + DNN where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
Re claims 7 and 14, Mahyar teaches
7. The computer-implemented method of claim 1, wherein determining, via the neural network, the candidate embedding associated with the candidate voice sample includes: inputting, by the one or more processors, the candidate voice sample into an encoder, wherein the encoder includes a trained deep learning neural network; and (using the models + DNN where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
generating, by the one or more processors, via the encoder, the candidate embedding, wherein the candidate embedding includes a multi-dimensional embedding. (both initial and primary are vectors/embeddings… using the models + DNN where the transformation from initial encoding, e.g. col 7 lines 16-29 & col 8 lines 31-47, into a vector/embedding is maintained through selecting the best match from preferred models selected based on further criteria e.g. max/min frequencies, following that only certain models are initially activated e.g. for a language match, then selected models are chosen Col 11 line 14 – col 12 line 44… matching the target or primary speaker to a set of candidates in various languages fig. 1 col 2 lines 5-24 and col 11 line 62 to col 12 line 30)… and using 2-dimensional vectors or embeddings for the original target speaker col 10 lines 20-31 with fig. 1 as well as the candidate or closest matches via DNN col 8 lines 22-30)
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20230076258 A1 GABRYJELSKI; Henry et al.
Voice dubbing
US 20200169591 A1 INGEL B A et al.
Revoicing
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL COLUCCI whose telephone number is (571)270-1847. The examiner can normally be reached on M-F 9 AM - 7 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL COLUCCI/Primary Examiner, Art Unit 2655 (571)-270-1847
Examiner FAX: (571)-270-2847
Michael.Colucci@uspto.gov