Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-8 and 10-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Ingel et al. (U.S. Patent Application Pub. No. 2020/0213680, hereinafter “Ingel”).
In regard to claim 1, Ingel discloses an electronic device (Fig. 2, 160) comprising:
a display (touch screen 246, paragraph [0121]);
a speaker (speaker 228, paragraph [0120]);
memory storing one or more instructions (memory 250, paragraph [0125]); and
one or more processors operatively coupled to the display, the speaker, and the memory, and configured to execute the one or more instructions (processors 204, paragraph [0118]),
wherein the one or more instructions, when executed by the one or more processors, cause the electronic device to:
control the display to display video data including subtitle data in a target language (video comprising captions is displayed, paragraph [0196]);
obtain context information of an utterer from the video data and audio data corresponding to the video data (see Fig. 6, a media analysis unit 620 analyzes an audio stream 610 and a video stream 615 to determine context information, paragraph [0192]);
obtain audio track data corresponding to the target language based on the subtitle data and the context information of the utterer (transcript data determined from the captions is translated to a target language, and a revoiced audio stream is generated in the target language with properties determined from the context information, paragraph [0193]); and
control the speaker to output the obtained audio track data (speaker 228 outputs the dubbed version of the audio stream in the target language, paragraph [0120]).
In regard to claim 2, Ingel discloses the one or more instructions, when executed by the one or more processors, cause the electronic device to input the subtitle data and the context information of the utterer into a trained Artificial Intelligence (AI) model to obtain the audio track data, and wherein the trained Al model is a Text-to-Speech (TTS) Al model trained to receive text data as an input and convert the text data to speaker adaptive audio data based on the context information of the utterer (voice generation unit 355 utilizes artificial neural network algorithms to generate the audio data in the target language based on the context information, paragraphs [0193] and [0198]).
In regard to claim 3, Ingel discloses the trained AI model is configured to:
obtain a characteristic parameter of the utterer based on the context information of the utterer (voice properties of each individual are determined, paragraph [0192]), and
output the audio track data in which the subtitle data is converted to the speaker adaptive audio data based on the characteristic parameter of the utterer (the voice properties of each individual are used to generate the target language audio, paragraph [0193]).
In regard to claim 4, Ingel discloses the one or more instructions, when executed by the one or more processors, cause the electronic device to identify the characteristic parameter of the utterer based on the context information of the utterer and convert the subtitle data to the speaker adaptive audio data based on the characteristic parameter of the utterer to obtain the audio track data (from the context information, voice properties including intonation, etc. are determined, paragraphs [0151] and [0194]).
In regard to claim 5, Ingel discloses the characteristic parameter of the utterer comprises at least one of a voice type, a voice intonation, a voice pitch, a voice speech speed, or a voice volume, and
wherein the context information of the utterer comprises at least one of gender information, age information, emotion information, character information, or speech volume information of the utterer (context information includes gender, age, volume, etc., paragraphs [0150], [0192], and [0194]; characteristic parameters include intonation, speech pitch, volume, etc., paragraphs [0151] and [0194]).
In regard to claim 6, Ingel discloses the one or more instructions, when executed by the one or more processors, cause the electronic device to:
obtain, based on at least one of the video data or the audio data, timing data related to a speech start of the utterer, identification data of the utterer, and emotion data of the utterer (beginning times of speech segments and emotional state repeated for each of a plurality of speakers, paragraphs [0149] and [0151]), and
identify the characteristic parameter of the utterer based on the timing data, the identification data, and the emotion data (the information is used to generate a unique voice profile for each individual speaker, paragraph [0192]).
In regard to claim 7, Ingel discloses the one or more instructions, when executed by the one or more processors, cause the electronic device to:
obtain the subtitle data streaming from a specific data channel (transcript data received from an associated database, paragraph [0138]), or
obtain the subtitle data through text recognition with respect to frames included in the video data (optical character recognition (OCR) of caption data, paragraph [0196]).
In regard to claim 8, Ingel discloses the one or more instructions, when executed by the one or more processors, cause the electronic device to:
separate the audio data corresponding to the video data into background audio data and speech audio data (source separation algorithms to separate audio background from speech audio, paragraph [0149]),
obtain the context information of the utterer from the speech audio data (see Fig. 6, a media analysis unit 620 analyzes an audio stream 610, paragraph [0192]), and
control the speaker to mix the audio track data and the background audio data obtained based on the subtitle data and the context information of the utterer (the target language audio and background sound-emanating object are mixed together at an appropriate ratio, paragraphs [0372-0374]), and
output the mixed audio track data and background audio data (speaker 228 outputs the dubbed version of the audio stream in the target language, paragraph [0120]).
In regard to claim 10, Ingel discloses a method performed by an electronic device for obtaining an audio track, the method comprising:
displaying, via a display of the electronic device, video data including subtitle data in a target language (video comprising captions is displayed, paragraph [0196]);
obtaining context information of an utterer from the video data and audio data corresponding to the video data (see Fig. 6, a media analysis unit 620 analyzes an audio stream 610 and a video stream 615 to determine context information, paragraph [0192]);
obtaining audio track data corresponding to the target language based on the subtitle data and the context information of the utterer (transcript data determined from the captions is translated to a target language, and a revoiced audio stream is generated in the target language with properties determined from the context information, paragraph [0193]); and
outputting, via a speaker of the electronic device, the obtained audio track data (speaker 228 outputs the dubbed version of the audio stream in the target language, paragraph [0120]).
In regard to claim 11, Ingel discloses the obtaining the audio track comprises inputting the subtitle data and the context information of the utterer into a trained Artificial Intelligence (AI) model to obtain the audio track data, and wherein the trained Al model is a Text-to-Speech (TTS) Al model trained to receive text data as an input and convert the text data to speaker adaptive audio data based on the context information of the utterer (voice generation unit 355 utilizes artificial neural network algorithms to generate the audio data in the target language based on the context information, paragraphs [0193] and [0198]).
In regard to claim 12, Ingel discloses the trained AI model is configured to:
obtain a characteristic parameter of the utterer based on the context information of the utterer (voice properties of each individual are determined, paragraph [0192]), and
output the audio track data in which the subtitle data is converted to the speaker adaptive audio data based on the characteristic parameter of the utterer (the voice properties of each individual are used to generate the target language audio, paragraph [0193]).
In regard to claim 13, Ingel discloses the obtaining the audio track data comprises identifying the characteristic parameter of the utterer based on the context information of the utterer and converting the subtitle data to the speaker adaptive audio data based on the characteristic parameter of the utterer to obtain the audio track data (from the context information, voice properties including intonation, etc. are determined, paragraphs [0151] and [0194]).
In regard to claim 14, Ingel discloses the characteristic parameter of the utterer comprises at least one of a voice type, a voice intonation, a voice pitch, a voice speech speed, or a voice volume, and
wherein the context information of the utterer comprises at least one of gender information, age information, emotion information, character information, or speech volume information of the utterer (context information includes gender, age, volume, etc., paragraphs [0150], [0192], and [0194]; characteristic parameters include intonation, speech pitch, volume, etc., paragraphs [0151] and [0194]).
In regard to claim 15, Ingel discloses a non-transitory computer readable medium having instructions stored therein (paragraph [0686]), which when executed by a processor of an electronic device, cause the electronic device to:
display, via a display of the electronic device, video data including subtitle data in a target language (video comprising captions is displayed, paragraph [0196]);
obtain context information of an utterer from the video data and audio data corresponding to the video data (see Fig. 6, a media analysis unit 620 analyzes an audio stream 610 and a video stream 615 to determine context information, paragraph [0192]);
obtain audio track data corresponding to the target language based on the subtitle data and the context information of the utterer (transcript data determined from the captions is translated to a target language, and a revoiced audio stream is generated in the target language with properties determined from the context information, paragraph [0193]); and
output, via a speaker of the electronic device, the obtained audio track data (speaker 228 outputs the dubbed version of the audio stream in the target language, paragraph [0120]).
In regard to claim 16, Ingel discloses the instructions further cause the electronic device to input the subtitle data and the context information of the utterer into a trained Artificial Intelligence (AI) model to obtain the audio track data, and wherein the trained Al model is a Text-to-Speech (TTS) Al model trained to receive text data as an input and convert the text data to speaker adaptive audio data based on the context information of the utterer (voice generation unit 355 utilizes artificial neural network algorithms to generate the audio data in the target language based on the context information, paragraphs [0193] and [0198]).
In regard to claim 17, Ingel discloses the trained AI model is configured to:
obtain a characteristic parameter of the utterer based on the context information of the utterer (voice properties of each individual are determined, paragraph [0192]), and
output the audio track data in which the subtitle data is converted to the speaker adaptive audio data based on the characteristic parameter of the utterer (the voice properties of each individual are used to generate the target language audio, paragraph [0193]).
In regard to claim 18, Ingel discloses the instructions further cause the electronic device to identify the characteristic parameter of the utterer based on the context information of the utterer and convert the subtitle data to the speaker adaptive audio data based on the characteristic parameter of the utterer to obtain the audio track data (from the context information, voice properties including intonation, etc. are determined, paragraphs [0151] and [0194]).
In regard to claim 19, Ingel discloses the characteristic parameter of the utterer comprises at least one of a voice type, a voice intonation, a voice pitch, a voice speech speed, or a voice volume, and
wherein the context information of the utterer comprises at least one of gender information, age information, emotion information, character information, or speech volume information of the utterer (context information includes gender, age, volume, etc., paragraphs [0150], [0192], and [0194]; characteristic parameters include intonation, speech pitch, volume, etc., paragraphs [0151] and [0194]).
In regard to claim 20, Ingel discloses the instructions further cause the electronic device to:
obtain, based on at least one of the video data or the audio data, timing data related to a speech start of the utterer, identification data of the utterer, and emotion data of the utterer (beginning times of speech segments and emotional state repeated for each of a plurality of speakers, paragraphs [0149] and [0151]), and
identify the characteristic parameter of the utterer based on the timing data, the identification data, and the emotion data (the information is used to generate a unique voice profile for each individual speaker, paragraph [0192]).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ingel, in view of Cohen et al. (U.S. Patent No. 11,417,343, hereinafter “Cohen”).
In regard to claim 9, Ingel discloses the one or more instructions, when executed by the one or more processors, cause the electronic device to:
obtain first timing data related to a speech start of the first utterer and second timing data related to a speech start of the second utterer (speaker diarization algorithms and/or speaker recognition detects segments of particular speakers, paragraph [0149]),
perform Text-to-Speech (TTS) conversion on first subtitle data corresponding to the first utterer based on the first timing data, first identification data of the first utterer, and first emotion data of the first utterer, to obtain first audio track data corresponding to the first utterer (voice generation of transcript corresponding to source language audio, paragraph [0193]; the source language audio determined based on beginning times of speech segments and emotional state repeated for each of a plurality of speakers, paragraphs [0149] and [0151]);
and perform TTS conversion on second subtitle data corresponding to the second utterer based on the second timing data, second identification data of the second utterer, and second emotion data of the second utterer, to obtain second audio track data corresponding to the second utterer (voice generation of transcript corresponding to source language audio, paragraph [0193]; the source language audio determined based on beginning times of speech segments and emotional state repeated for each of a plurality of speakers, paragraphs [0149] and [0151]).
Ingel does not expressly disclose the timing data is determined based on the video data including the first utterer and the second utterer.
Cohen discloses a method for identifying speakers comprising determining timing data based on the video data including the first utterer and the second utterer (Fig. 10, voice segments are determined for identified speakers based on video and/or image related speaker identification parameters, column 2, line 67 to column 3, line 3; column 17, line 60 to column 18 line 20; and column 20, lines 32-39).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to determine timing data based on video data, because analyzing the video would provide additional cues to properly identify the utterer, as taught by Cohen (column 3, line 51 to column 4, line 14).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Levine et al., de Juan et al., Buckley et al., Wang et al., Chicote et al., Mahyar, Ishikawa et al., Hu, DuBose, Rossano et al., and Adami et al. disclose additional methods for generating target language audio from subtitle information and/or identifying utterers in video data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIAN LOUIS ALBERTALLI whose telephone number is (571)272-7616. The examiner can normally be reached M-F 8AM-3PM, 4PM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
BLA 9/3/26
/BRIAN L ALBERTALLI/ Primary Examiner, Art Unit 2656