DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
Applicant is reminded of the proper content of an abstract of the disclosure.
A patent abstract is a concise statement of the technical disclosure of the patent and should include that which is new in the art to which the invention pertains. The abstract should not refer to purported merits or speculative applications of the invention and should not compare the invention with the prior art.
If the patent is of a basic nature, the entire technical disclosure may be new in the art, and the abstract should be directed to the entire disclosure. If the patent is in the nature of an improvement in an old apparatus, process, product, or composition, the abstract should include the technical disclosure of the improvement. The abstract should also mention by way of example any preferred modifications or alternatives.
Where applicable, the abstract should include the following: (1) if a machine or apparatus, its organization and operation; (2) if an article, its method of making; (3) if a chemical compound, its identity and use; (4) if a mixture, its ingredients; (5) if a process, the steps.
Extensive mechanical and design details of an apparatus should not be included in the abstract. The abstract should be in narrative form and generally limited to a single paragraph within the range of 50 to 150 words in length.
Examiner suggests that the current abstract be modified to provide a concise statement of the technical disclosure of the patent and should include that which is new in the art to which the invention pertains. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details. The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, "The disclosure concerns," "The disclosure defined by this invention," "The disclosure describes," etc. In addition, the form and legal phraseology often used in patent claims, such as "means" and "said," should be avoided.
See MPEP § 608.01(b) for guidelines for the preparation of patent abstracts.
Election/Restrictions
Applicant presented persuasive argument on 5/1/2026 that there may not be serious burden in searching. Therefore, the restriction requirement for claims 1-20 has been withdrawn and all claims 1-20 has been examined.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 6, and 8 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Chumkesornkulkit et al (WO 2025133682 A1), hereinafter Chumkesornkulkit.
Regarding claim 1, Chumkesornkulkit teaches a method comprising: receiving a request for a target language associated with a video asset, wherein: the video asset comprises at least one first dialog in a native language; and the target language is different from the native language;
"One of the objects of the invention is to provide a system and method for processing a video clip so that the original audio track therein (which is in the first language) can be replaced by another audio track in the desired language (which is different from the first language)." - Pg 1, Lines 45-47
determining at least one phoneme associated with at least one second dialog in the target language; converting, based on the at least one phoneme, the at least one second dialog in a script;
"Upon receipt, the speaker encoder (702) may be instantiated and trained to extract a fixed dimensional speaker embedding from the audio portion of the video clip or more particularly the original speeches included within the audio portion. A text processor (704) may also be provided for receiving the translated text from the language translation module (127) and converting the received text into phoneme or grapheme embedding sequences. " - Pg 18, Lines 5-9
NOTE: Chumkesornkulkit discloses taking the original audio (first language) and using a text processor that takes the translated text second/target language) to convert the translated text into phoneme embedding sequences. The language translation module 127 would need to provide the text processor a text transcript of the second dialog in order for the text processor to convert the received text into phoneme sequences.
generating, based on the script and a model database, facial data associated with an actor in the at least one second dialog;
“A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions, the speaker’s mel- spectrogram, the original speeches, the translated speeches, etc.) received from the respective components in the system (100). The data may also be sorted and stored in the database (150), in some embodiments.” – Pg 8, Lines 14-17
NOTE: Chumkesornkulkit discloses a database used to store information regarding a speaker’s face. This is used to accurately overlay the created visemes to the speaker’s face region to create the lip-syncing modification so that the video can be modified to replace the lip movement and speech. This modification would also require the script of the second language to generate the appropriate visemes.
and integrating the facial data and the at least one second dialog into the video asset.
"At Step S5812, the rendering engine (130) may feed the visemes obtained from Step S5810 to any suitable speech2video model (such as superS2V model or particularly DeepFaceLab pre-built face extraction model) which may modify the speaker’s face region based on the received visemes, so that the speaker’s mouth region movements match the movements of the translated speech in the second language. Thereafter, the desired video clip may be obtained and displayed to the user of the system (100), as at Step S5814." - Pg 15, Lines 20-24
Regarding claim 6, Chumkesornkulkit teaches the method of claim 1. Chumkesornkulkit further teaches wherein the integrating the facial data and the at least one second dialog further comprises: replacing the at least one first dialog with the at least one second dialog; and superimposing the facial data onto at least one frame of the video asset.
“using a machine learning model (such as the DeepFaceLab model) to replace or overlay the face of a speaker speaking in a first language with the face of the same speaker speaking in a second language in a video frame” – Pg 4, Lines 54-55
NOTE: Chumkesornkulkit discloses using a machine learning model to overlay the actor speaking the first language with mouth movements and facial data of the same actor speaking in a second language. This functionally corresponds to superimposing the facial data onto a frame of the video.
Regarding claim 8, Chumkesornkulkit teaches wherein the facial data comprises at least one of: mouth movements; geometric features; texture information; or temporal information.
“The rendering engine (130) may create visemes that are synchronized to the translated speeches and then modify the face or face region of the speaker based on the created visemes, so that the speaker’s mouth region movements match the movements of the translated speeches in the second language” – Pg 7, Lines 58-59 and Pg 8, Line 1
NOTE: Chumkesornkulkit discloses modifying the face region of the speaker in the video so that the mouth region movements matches the translated speeches in the second languages. This would require that the speaker’s facial data comprises mouth movements must be obtained. Pg 16, Lines 17-19 also disclose locating facial features such as the eyes, nose, and mouth within the face region in the video frame.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chumkesornkulkit and Foster (US 20250372124 A1), hereinafter Foster.
.
Regarding claim 2, Chumkesornkulkit teaches the method of claim 1. Chumkesornkulkit does not teach receiving a request for an updated context associated with the video asset, wherein: the video asset comprises at least one third dialog in an original context; and the updated context is different from the original context; determining at least one second phoneme associated with at least one fourth dialog in the updated context; converting, based on the at least one second phoneme, the at least one fourth dialog in a second script; generating, based on the second script and the model database, second facial data associated with an actor in the at least one fourth dialog; and integrating the second facial data and the at least one fourth dialog into the video asset. However, Foster teaches receiving a request for an updated context associated with the video asset, wherein: the video asset comprises at least one third dialog in an original context; and the updated context is different from the original context;
For example, if a character in a video game or movie says a profane word that the end-user has flagged for obfuscation, the deepfake generator model can be executed to replace the audio of the character saying the profane word with audio in the same character's voice but saying something non-profane such as “shoot” or “darn” instead.” – Par 73, Lines 6-11
NOTE: Foster discloses an apparatus that may alter the original context of a video based on a user inputted list of content they want to be modified. This can be understood as a request from the user for an updated context. The dialogue refers to the audio regarding the video. The third dialog can be understood as audio of a speech the actor says in the original context. Foster further discloses that one example of a type of modification of dialogue in an original context is to replace curse words/profanity.
determining at least one second phoneme associated with at least one fourth dialog in the updated context; converting, based on the at least one second phoneme, the at least one fourth dialog in a second script;
"Upon receipt, the speaker encoder (702) may be instantiated and trained to extract a fixed dimensional speaker embedding from the audio portion of the video clip or more particularly the original speeches included within the audio portion. A text processor (704) may also be provided for receiving the translated text from the language translation module (127) and converting the received text into phoneme or grapheme embedding sequences. " – Chumkesornkulkit Pg 18, Lines 5-9
NOTE: Foster teaches a obtaining a transcript to identify which content needs to be modified, see Foster par 27. With the original transcript of the audio-video content, one of ordinary skill could make a new script that has the updated context which can be understood as the second script. After the combination, the script reflecting the updated context as taught by Foster can be received by the text processor to convert the received text into phoneme sequences as taught by Chumkesornkulkit. This would then allow Chumkesornkulkit to convert the at least one second dialogue (updated context) in a script based on at least one phenome. The steps of obtaining the second phoneme associated with a fourth dialog would be the same for obtaining the first phoneme associated with the second dialog as explained in the rejection of claim 1, but would also consider the updated context reflected in the fourth dialog.
generating, based on the second script and the model database, second facial data associated with an actor in the at least one fourth dialog;
“A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions, the speaker’s mel- spectrogram, the original speeches, the translated speeches, etc.) received from the respective components in the system (100). The data may also be sorted and stored in the database (150), in some embodiments.” – Pg 8, Lines 14-17
NOTE: Chumkesornkulkit discloses a database used to store information regarding a speaker’s face. This is used to accurately overlay the created visemes to the speaker’s face region to create the lip-syncing modification so that the video can be modified to replace the lip movement and speech. This modification would also require the script with the updated context as taught by Foster to generate the appropriate visemes which can be understood as the second script. After the modification, the replacement of an original context with an updated context as taught by Foster can then be used to create visemes to modify lip movements as taught by Chumkesornkulkit for the updated context instead of the different language since the method of doing so would be the same. The second facial data is generated using the same steps for generating the first facial data as explained in claim 1 but would also consider the updated context reflected in the fourth dialog.
and integrating the second facial data and the at least one fourth dialog into the video asset.
"At Step S5812, the rendering engine (130) may feed the visemes obtained from Step S5810 to any suitable speech2video model (such as superS2V model or particularly DeepFaceLab pre-built face extraction model) which may modify the speaker’s face region based on the received visemes, so that the speaker’s mouth region movements match the movements of the translated speech in the second language. Thereafter, the desired video clip may be obtained and displayed to the user of the system (100), as at Step S5814." - Pg 15, Lines 20-24
NOTE: Chumkesornkulkit teaches integrating facial data and dialogue to modify a speaker’s face so that the actor appears to speak a different language compared to the original language. However, the technique to replace dialogue of an original context with an updated context dialogue would be the same. Integrating the second facial data and the fourth dialogue to the video asset would follow the same steps as integrating the first facial data and the second dialogue to the video asset as explained in the rejection of claim 1.
It would have been obvious for one of ordinary skill in the art before the effective filing date of the present invention to modify Chumkesornkulkit by incorporating the teachings of Foster to receive a request for updated context associated with a video, convert the updated text dialogue into a script based on a determined phoneme, generate facial data associated with the script and model database, and integrate the facial data and dialog into the video asset. One would be motivated to make this combination to seamlessly alter the context of a video that may be unfavorable to the user’s preference. The user would then be able to view a video that replaces the original context with a preferred context without obvious censorship blocking or noises that could distract from the viewing experience.
Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chumkesornkulkit and Bhat et al (US 11582519 B1), hereinafter Bhat.
Regarding claim 3, Chumkesornkulkit teaches the method of claim 1. Chumkesornkulkit does not teach receiving a request for a target actor associated with the video asset, wherein: the video asset comprises at least one fifth dialog associated with an original actor; and the target actor is different from the original actor; determining at least one third phoneme associated with at least one sixth dialog associated with the target actor; converting, based on the at least one third phoneme, the at least one sixth dialog in a third script; generating, based on the third script and the model database, third facial data associated with the target actor in the at least one sixth dialog; and integrating the third facial data and the at least one sixth dialog into the video asset. However, Bhat teaches receiving a request for a target actor associated with the video asset, wherein: the video asset comprises at least one fifth dialog associated with an original actor; and the target actor is different from the original actor;
"For example, the video synthesis system may use facial recognition and actor identification to identify the lead actor within the frame. The system may also maintain a mapping (e.g., a pre-determined mapping) that maps the source person to a replacement target person. Accordingly, the video synthesis system may determine that the regional actor in India corresponds to the appropriate replacement for the source person shown in the video based on the determined identity" - Col 3, Lines 58-65
NOTE: Bhat teaches an audiovisual input depicting a video or movie of an original actor in which the original actor is replaced by a regional replacement actor. The fifth dialog can be understood as audio of a speech that is said by the original actor. One example is taking an English-speaking movie and replacing the English speaking actor with a regional actor in India. The audio may also be dubbed and mouth movements adjusted so that the second language reflects the replacement regional actor, see col 4, Lines 44-67.
determining at least one third phoneme associated with at least one sixth dialog associated with the target actor; converting, based on the at least one third phoneme, the at least one sixth dialog in a third script;
"Upon receipt, the speaker encoder (702) may be instantiated and trained to extract a fixed dimensional speaker embedding from the audio portion of the video clip or more particularly the original speeches included within the audio portion. A text processor (704) may also be provided for receiving the translated text from the language translation module (127) and converting the received text into phoneme or grapheme embedding sequences." – Chumkesornkulkit Pg 18, Lines 5-9
NOTE: Chumkesornkulkit discloses taking the original audio (first language) and using a text processor that takes the translated text second/target language) to convert the translated text into phoneme embedding sequences. The language translation module 127 would need to provide the text processor a transcript of the second dialog in order for the text processor to convert the received text into phoneme sequences. After the combination, Bhat’s method for replacing an original actor with a replacement actor can modify Chumkesornkulkit’s system for determining a phoneme associated with the sixth dialog and converting it into a third script. Bhat also discloses dubbing the original audio with a different language, see Bhat Col 4, Lines 44-67, which would naturally imply that the second dialog (second dialog) is associated with the target actor. This combination would then teach determining a phoneme with at least one second dialogue associated with the target actor; converting, based on the at least one phoneme, the at least one second dialog in a script. The steps of obtaining the third phoneme associated with a sixth dialog would be the same for obtaining the first phoneme associated with the second dialog as explained in the rejection of claim 1 but would also consider the target actor reflected in the sixth dialog.
generating, based on the third script and the model database, third facial data associated with the target actor in the at least one sixth dialog;
A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions” – Pg 8, Lines 14-25
NOTE: Chumkesornkulkit discloses a database used to store information regarding a speaker’s face. This is used to accurately overlay the created visemes to the speaker’s face region to create the lip-syncing modification so that the video can be modified to replace the lip movement and speech. This modification would also require the script of the second language obtained with the methods as explained in claim 1 to generate the appropriate visemes which can be understood as the third script. After the combination, Bhat’s method for replacing an original actor in a video with a replacement actor can modify Chumkesornkulkit’s system for generating the facial data based on the script and model database as taught by Chumkesornkulkit. This combination would then generate facial data associated with the target actor rather than the original actor. The third facial data is generated using the same steps for generating the first facial data as explained in claim 1 but would also consider the target actor reflected in the sixth dialog.
and integrating the third facial data and the at least one sixth dialog into the video asset.
"At Step S5812, the rendering engine (130) may feed the visemes obtained from Step S5810 to any suitable speech2video model (such as superS2V model or particularly DeepFaceLab pre-built face extraction model) which may modify the speaker’s face region based on the received visemes, so that the speaker’s mouth region movements match the movements of the translated speech in the second language. Thereafter, the desired video clip may be obtained and displayed to the user of the system (100), as at Step S5814." - Pg 15, Lines 20-24
NOTE: Chumkesornkulkit teaches integrating facial data and dialogue to modify a speaker’s face so that the actor appears to speak a different language compared to the original language. After the combination, Bhat’s method of replacing an original actor with a new actor can modify Chumkesornkulkit’s integration of facial data and the second dialogue into the video asset. This combination would then allow facial data associated with the new target actor to be integrated into the video asset instead of the original actor. Integrating the third facial data and the sixth dialogue to the video asset would follow the same steps as integrating the first facial data and the second dialogue to the video asset as explained in the rejection of claim 1.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Bhat by incorporating the teachings of Chumkesornkulkit to determine a phoneme associated with a second dialog associated with the target actor, convert the second dialogue to a script, generate facial data associated with the target actor, and integrate the facial data into the video asset. One would be motivated to make this combination to provide a more regionally accurate performer that matches the new replaced language compared to the original actor. This creates a more authentic viewing experience since the replacement actor will match the appearance of the new target language.
Claim(s) 4 and 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chumkesornkulkit and Li (US 20210248801 A1), hereinafter Li.
Regarding claim 4, Chumkesornkulkit teaches the method of claim 1. Chumkesornkulkit does not teach wherein the generating the facial data further comprises: mapping the determined at least one phoneme to at least one viseme, wherein the determined at least one phoneme is associated with the script. However, Li teaches wherein the generating the facial data further comprises: mapping the determined at least one phoneme to at least one viseme, wherein the determined at least one phoneme is associated with the script.
“extracting phonemes from an input audio signal or textual transcript, and then mapping the extracted phonemes to corresponding visemes” – Par 16 Line 5-7
NOTE: After the combination, Li’s technique of mapping a phoneme to at least one viseme can modify Chumkesornkulkit’s system so that the phoneme sequences generated from a transcript input as taught in the rejection of claim 1 can then be mapped to an appropriate viseme to replace the speaker/actor’s lip movements in the video.
and translating, based on a facial image model, the at least one viseme to at least one mouth movement associated to the actor.
“The rendering engine (130) may create visemes that are synchronized to the translated speeches and then modify the face or face region of the speaker based on the created visemes, so that the speaker’s mouth region movements match the movements of the translated speeches in the second language.” – Chumkesornkulkit Pg 7, Lines 58-59, and Pg 8, Line 1
NOTE: Chumkesornkulkit discloses using the generated visemes to modify the face region of the speaker/actor in the video so that the mouth movements reflect the new second language. This functionally corresponds to translating the viseme to at least one mouth movement
It would have been obvious to one of ordinary skill in the art before the effective filing dates of the present invention to modify Chumkesornkulkit by incorporating the teachings of Li to generate facial data by mapping the phoneme to a viseme and translating the viseme to a mouth movement. One would be motivated to make this combination because it is well known in the art that phonemes is the vocal sound of the language and visemes are the visually counterpart of the sound. Naturally, the phoneme and viseme would be mapped together to accurately match the mouth movement and voice of the speaker in the video.
Regarding claim 5, Chumkesornkulkit teaches the method of claim 1. Chumkesornkulkit does not teach wherein the generating the facial data further comprises: mapping the determined at least one phoneme to at least one viseme, wherein the determined at least one phoneme is associated with the script; generating at least one mouth movement, wherein the at least one mouth movement associated with the at least one viseme is not defined in a facial image model; and storing the at least one mouth movement in the facial image model. However, Li teaches wherein the generating the facial data further comprises: mapping the determined at least one phoneme to at least one viseme, wherein the determined at least one phoneme is associated with the script;
“extracting phonemes from an input audio signal or textual transcript, and then mapping the extracted phonemes to corresponding visemes” – Par 16 Line 5-7
NOTE: After the combination, Li’s technique of mapping a phoneme to at least one viseme can modify Chumkesornkulkit’s system so that the phoneme sequences generated from a transcript input as taught in the rejection of claim 1 can then be mapped to an appropriate viseme to replace the speaker/actor’s lip movements in the video.
generating at least one mouth movement, wherein the at least one mouth movement associated with the at least one viseme is not defined in a facial image model; and storing the at least one mouth movement in the facial image model.
“A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions, the speaker’s mel- spectrogram, the original speeches, the translated speeches, etc.) received from the respective components in the system (100). The data may also be sorted and stored in the database (150), in some embodiments.” – Pg 8, Lines 14-17
NOTE: When generating a new mouth movement associated to a viseme that has not been defined yet, one of ordinary skill would be required to store the new mouth movement associated with the viseme depicting a different language in order to fully replace the original mouth movements in the video with mouth movements of the second language. After the combination, the extraction of phonemes to map to a viseme as taught by Li can modify Chumkesornkulkit so that the visemes can be stored in the database that holds the facial image model information.
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Chumkesornkulkit by incorporating the teachings of Li to store mouth movements that are not defined in the facial image model. One would be motivated to make this combination so that the facial image model can have a comprehensive set of facial data needed to fully portray the mouth movements of a second language
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chumkesornkulkit and Cohen-Or et al (US 20260119854 A1), hereinafter Cohen-Or.
Regarding claim 7, Chumkesornkulkit teaches the method of claim 1. Chumkesornkulkit does not teach wherein the script lists the at least one phoneme and at least one corresponding timestamp. However, Cohen-Or teaches wherein the script lists the at least one phoneme and at least one corresponding timestamp.
“For a particular audio track, phonemes and visemes can be time-coded as they appear on screen or on audio, and this process can be automatically conducted or manually conducted.” – Par 71, Lines 4-7
NOTE: Cohen-or discloses an audio track wherein the phonemes and visemes are time-coded. Time-coding functionally corresponds to a timestamp as both record an exact moment in a video. After the combination, the system of Chumkesornkulkit can take the audio input, convert it into a transcript and list the timecodes the phonemes as taught by Cohen-Or
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Chumkesornkulkit by incorporating the teachings of Cohen-or to have the script lists the at least one phoneme and at least one corresponding timestamp. One would be motivated to make this combination to obtain accurate lip-syncing when inserting the new mouth movements and sounds in the second language.
Claim(s) 9-13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Foster and Chumkesornkulkit.
Regarding claim 9, Foster teaches receiving a request for an updated context associated with a video asset, wherein: the video asset comprises at least one first dialog in an original context; and the updated context is different from the original context;
“For example, if a character in a video game or movie says a profane word that the end-user has flagged for obfuscation, the deepfake generator model can be executed to replace the audio of the character saying the profane word with audio in the same character's voice but saying something non-profane such as “shoot” or “darn” instead.” – Par 73, Lines 6-11
NOTE: Foster discloses an apparatus that may alter the original context of a video based on a user inputted list of content they want to be modified. This can be understood as a request from the user for an updated context. Foster further discloses that one example of a type of modification of dialogue in an original context is to replace curse words/profanity.
Foster does not teach determining at least one phoneme associated with at least one second dialog in the updated context; converting, based on the at least one phoneme, the at least one second dialog in a script; generating, based on the script and a model database, facial data associated with an actor in the at least one second dialog; and integrating the facial data and the at least one second dialog into the video asset. However, Chumkesornkulkit teaches determining at least one phoneme associated with at least one second dialog in the updated context; converting, based on the at least one phoneme, the at least one second dialog in a script;
"Upon receipt, the speaker encoder (702) may be instantiated and trained to extract a fixed dimensional speaker embedding from the audio portion of the video clip or more particularly the original speeches included within the audio portion. A text processor (704) may also be provided for receiving the translated text from the language translation module (127) and converting the received text into phoneme or grapheme embedding sequences. " – Chumkesornkulkit Pg 18, Lines 5-9
NOTE: Foster teaches a obtaining a transcript to identify which content needs to be modified, see Foster par 27. With the original transcript of the audio-video content, one of ordinary skill could make a new script that has the updated context. After the combination, the script reflecting the updated context as taught by Foster can be received by the text processor to convert the received text into phoneme sequences as taught by Chumkesornkulkit. This would then allow Foster’s apparatus to convert the at least one second dialogue (updated context) in a script based on at least one phenome.
generating, based on the script and a model database, facial data associated with an actor in the at least one second dialog;
A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions” – Pg 8, Lines 14-25
NOTE: Chumkesornkulkit discloses a database used to store information regarding a speaker’s face. This is used to accurately overlay the created visemes to the speaker’s face region to create the lip-syncing modification so that the video can be modified to replace the lip movement and speech. This modification would also require the script of the second language to generate the appropriate visemes.
and integrating the facial data and the at least one second dialog into the video asset.
"At Step S5812, the rendering engine (130) may feed the visemes obtained from Step S5810 to any suitable speech2video model (such as superS2V model or particularly DeepFaceLab pre-built face extraction model) which may modify the speaker’s face region based on the received visemes, so that the speaker’s mouth region movements match the movements of the translated speech in the second language. Thereafter, the desired video clip may be obtained and displayed to the user of the system (100), as at Step S5814." - Pg 15, Lines 20-24
NOTE: Chumkesornkulkit teaches integrating facial data and dialogue to modify a speaker’s face so that the actor appears to speak a different language compared to the original language. After the combination, the replacing of a first language to a second language by integrating facial data into the video asset as taught by Chumkesornkulkit can be substituted with the replacing of original context with updated context as taught by Foster and then the facial data can be integrated into the video asset to reflect the updated context.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Foster by incorporating the teachings of Chumkesornkulkit to convert the second dialogue into a script based on a determined phoneme, generate facial data based on a script and model database, and integrate the facial data into the video asset. One would be motivated to make this combination to seamlessly alter the context of a video that may be unfavorable to the user’s preference. The user would then be able to view a video that replaces the original context with a preferred context without obvious censorship blocking or noises that could distract from the viewing experience.
Regarding claim 10, Foster in view of Chumkesornkulkit teaches the method of claim 9. Foster further teaches identifying the original context in the at least one first dialog, wherein the original context indicates at least one of: profanity; violence; cultural expression; or adult activity.
“For example, if a character in a video game or movie says a profane word that the end-user has flagged for obfuscation, the deepfake generator model can be executed to replace the audio of the character saying the profane word with audio in the same character's voice but saying something non-profane such as “shoot” or “darn” instead.” – Par 73, Lines 6-11
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Foster by incorporating the teachings of Chumkesornkulkit to identify profanity, violence, cultural expression, or adult activity in the first dialogue’s original context. One would be motivated to make this combination in order to find all of the unwanted context from the original video so that the system can replace it with context that is more appropriate to the user’s preferences.
Regarding claim 11, Foster in view of Chumkesornkulkit teaches the method of claim 9. Foster further teaches wherein the identifying the at least one first dialog in the original context further comprises: converting the at least one first dialog into text; and identifying, based on the converted text, the original context.
“Providing even more detail consistent with present principles, offline content other than the AV content itself (like transcripts) may also be used to identify content that should be altered, as may the audio and/or video of the AV content itself” – Par 27, Lines 1-5
NOTE: Foster discloses using transcripts of the original context in order to determine what should be altered. This would imply that the audio input must be converted into some textual transcript.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Foster by incorporating the teachings of Chumkesornkulkit to convert the first dialogue into text and identify based on the converted text, the original context. One would be motivated to make this change because converting the audio input into text will make the search for unwanted context within the original context easier to find. Then the original context can be replaced with the updated context according to the user’s preferences.
Regarding claim 12, Foster in view of Chumkesornkulkit teaches the method of claim 9. Foster does not teach identifying, based on an image model, the actor in the at least one first dialog; and generating, based on a speech model associated with the actor, the at least one second dialog. Chumkesornkulkit teaches identifying, based on an image model, the actor in the at least one first dialog; and generating, based on a speech model associated with the actor, the at least one second dialog.
“A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions, the speaker’s mel- spectrogram, the original speeches, the translated speeches, etc.) received from the respective components in the system (100). The data may also be sorted and stored in the database (150), in some embodiments.” – Pg 8, Lines 14-17
NOTE: Chumkesornkulkit discloses a database that stores the face image and speech associated with the speaker. This can be understood as the corresponding image model and speech model that is accessed through the actor model database. The second dialog can then be generated can then be generated using the steps as explain in the rejection of claim 9. Par 36, Lines 12-13 and Par 39, Lines 8-10 of Applicant’s specifications show that the image model and speech model can both be accessed from the actor model database which is taught by Chumkesornkulkit. After the combination, Chumkesornkulkit’s system for using a database corresponding to an actor can store the face and speech model information can modify Foster’s system for updating context of video. This would allow the face image model of the actor and the original (first dialog) and updated context (second dialog) to be used rather than the translated speech as taught by Chumkesornkulkit.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Foster by incorporating the teachings of Chumkesornkulkit to identify the actor in the first dialogue based on the image model and generate the second dialogue with the speech model. One would be motivated to make this combination to improve how the mouth movements of the face associated with the actor will look in the video and to update the context of the video while keeping the same voice of the actor.
Regarding claim 13, Foster in view of Chumkesornkulkit teaches the method of claim 9. Foster does not teach wherein the integrating the facial data and the at least one second dialog further comprises: replacing the at least one first dialog with the at least one second dialog; and superimposing the facial data onto at least one frame of the video asset.
“using a machine learning model (such as the DeepFaceLab model) to replace or overlay the face of a speaker speaking in a first language with the face of the same speaker speaking in a second language in a video frame” – Pg 4, Lines 54-55
NOTE: Chumkesornkulkit discloses using a machine learning model to overlay the actor speaking the first language with mouth movements and facial data of the same actor speaking in a second language. This functionally corresponds to superimposing the facial data onto a frame of the video.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Foster by incorporating the teachings of Chumkesornkulkit to integrate the facial date and second dialog by replacing the first dialogue with the second dialogue and superimposing the facial data to at least one frame of the video asset. One would be motivated to make this combination because it will improve visual consistency when a viewer looks at the actor’s mouth movements and sees that it matches the new updated context (second dialog). This would lead to a predictable result of an enhanced viewing experience of the video and accurate lip-syncing.
Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Foster and Chumkesornkulkit and Lev-Ami et al (US 11550879 B2), hereinafter Lev-Ami.
Regarding claim 14, Foster in view of Chumkesornkulkit teaches the method of claim 9. Foster does not teach verifying, with a digital right server, a license agreement for manipulating the facial data associated with the actor. However, Lev-Ami teaches verifying, with a digital right server, a license agreement for manipulating the facial data associated with the actor.
“some embodiments may further be used for monitoring, tracking, validating and/or verifying copyright (or other legal right or Intellectual Property (IP) right) in a content item (e.g., an original content item, as well as derived versions or modified versions or edited versions” – Col 1, Lines 46-51
NOTE: Lev-Ami teaches systems and methods for validating copyright, legal right, and IP rights of a content item and modified versions of the content. After the combination, the system for verifying a license agreement for content items can modify Foster system for updating content and replacing mouth movements of the actor in the video to reflect the update. This combination would allow verifying licensing rights associated to the facial data associated to the actor and blocking access to manipulating the facial data if not the license is not granted.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Foster by incorporating the teachings of Lev-Ami to verify, with a digital right server, a license agreement for manipulating the facial data associated with the actor. One would be motivated to make this combination because it is a commonly known requirement to establish permission before any modification of an actor for commercial use. This would also prevent any legal issues related to infringing on licensing rules.
Claim(s) 15-16, 18 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhat and Chumkesornkulkit.
Regarding claim 15, Bhat teaches a method comprising: receiving a request for a target actor associated with a video asset, wherein: the video asset comprises at least one first dialog associated with an original actor; and the target actor is different from the original actor;
"For example, the video synthesis system may use facial recognition and actor identification to identify the lead actor within the frame. The system may also maintain a mapping (e.g., a pre-determined mapping) that maps the source person to a replacement target person. Accordingly, the video synthesis system may determine that the regional actor in India corresponds to the appropriate replacement for the source person shown in the video based on the determined identity" - Col 3, Lines 58-65
NOTE: Bhat teaches an audiovisual input depicting a video or movie of an original actor in which the original actor is replaced by a regional replacement actor. One example is taking an English-speaking movie and replacing the English speaking actor with a regional actor in India. The audio may also be dubbed and mouth movements adjusted so that the second language reflects the replacement regional actor, see col 4, Lines 44-67.
Bhat does not teach determining at least one phoneme associated with at least one second dialog associated with the target actor; converting, based on the at least one phoneme, the at least one second dialog in a script; generating, based on the script and a model database, facial data associated with the target actor in the at least one second dialog; and integrating the facial data and the at least one second dialog into the video asset. However, Chumkesornkulkit teaches determining at least one phoneme associated with at least one second dialog associated with the
"Upon receipt, the speaker encoder (702) may be instantiated and trained to extract a fixed dimensional speaker embedding from the audio portion of the video clip or more particularly the original speeches included within the audio portion. A text processor (704) may also be provided for receiving the translated text from the language translation module (127) and converting the received text into phoneme or grapheme embedding sequences. " – Chumkesornkulkit Pg 18, Lines 5-9
NOTE: Chumkesornkulkit discloses taking the original audio (first language) and using a text processor that takes the translated text second/target language) to convert the translated text into phoneme embedding sequences. The language translation module 127 would need to provide the text processor a transcript of the second dialog in order for the text processor to convert the received text into phoneme sequences. After the combination, the conversion of the second dialogue into a script as taught by Chumkesornkulkit can modify Bhat’s method for replacing an original actor with a replacement actor. Bhat also discloses dubbing the original audio with a different language, see Bhat Col 4, Lines 44-67, which would naturally imply that the second dialog (second dialog) is associated with the target actor. This combination would then teach determining a phoneme with at least one second dialogue associated with the target actor; converting, based on the at least one phoneme, the at least one second dialog in a script;
generating, based on the script and a model database, facial data associated with the target actor in the at least one second dialog;
A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions” – Pg 8, Lines 14-25
NOTE: Chumkesornkulkit discloses a database used to store information regarding a speaker’s face. This is used to accurately overlay the created visemes to the speaker’s face region to create the lip-syncing modification so that the video can be modified to replace the lip movement and speech. This modification would also require the script of the second language to generate the appropriate visemes. After the combination, the facial data generated based on the script and model database as taught by Chumkesornkulkit can modify Bhat’s method for replacing an actor in a video. This combination would then generate facial data associated with the target actor rather than the original actor.
and integrating the facial data and the at least one second dialog into the video asset.
"At Step S5812, the rendering engine (130) may feed the visemes obtained from Step S5810 to any suitable speech2video model (such as superS2V model or particularly DeepFaceLab pre-built face extraction model) which may modify the speaker’s face region based on the received visemes, so that the speaker’s mouth region movements match the movements of the translated speech in the second language. Thereafter, the desired video clip may be obtained and displayed to the user of the system (100), as at Step S5814." - Pg 15, Lines 20-24
NOTE: Chumkesornkulkit teaches integrating facial data and dialogue to modify a speaker’s face so that the actor appears to speak a different language compared to the original language. After the combination, the integration of facial data and the second dialog as taught by Chumkesornkulkit can modify Bhat’s method of replacing an original actor with a new actor so that facial data associated with the new target actor is integrated into the video asset instead of the original actor.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Bhat by incorporating the teachings of Chumkesornkulkit to determine a phoneme associated with a second dialog associated with the target actor, convert the second dialogue to a script, generate facial data associated with the target actor, and integrate the facial data into the video asset. One would be motivated to make this combination to provide a more regionally accurate performer that matches the new replaced language compared to the original actor. This creates a more authentic viewing experience since the replacement actor will match the appearance of the new target language.
Regarding claim 16, , Bhat in view of Chumkesornkulkit teaches the method of claim 15. Bhat further teaches receiving at least one parameter associated with the target actor; determining, based on the at least one parameter, a facial model associated with the replacement actor; and generating, based on the facial model, at least one actor image associated with the target actor, wherein the at least one actor image comprises the facial data.
“The video synthesis system may also determine both face and body parameters for the regional replacement actor, for example, based on a corpus of images and/or video of the regional replacement actor. The video synthesis system may then determine a 3D model of the replacement actor within the particular frame.” – Col 2 Lines 57-62
NOTE: Bhat discloses obtaining face and/or body parameters associated with the replacement actor. A 3D model of the replacement actor is then used to replace the original actor in the video. The 3D model would naturally include the facial model associated with the replacement actor and the facial model would also include the facial data in order to perform the lip-syncing/dubbing disclosed by Bhat.
Regarding claim 18, Bhat in view of Chumkesornkulkit teaches the method of claim 15. Bhat does not teach identifying the original actor associated with the video asset, wherein the identifying the original actor further comprises: generating at least one region proposal associated with the original actor; extracting, based on the at least one region proposal, at least one feature; and classifying, based on the at least one feature, the original actor. However, Chumkesornkulkit teaches identifying the original actor associated with the video asset, wherein the identifying the original actor further comprises: generating at least one region proposal associated with the original actor;
“To locate the speaker’s face or face region in the image in each video frame, the face detection module (111) may be configured to perform operations shown in Figure 9, for instance and without limitation: (a) To perform “person segmentation” on each video frame by deploying a machine learning model such as BodyPix model, in which the image in each video frame may be segmented into pixels that are part of the speaker and those that are not, in order to crop the speaker from the image; (b) To track and crop the speaker’s face or face region from each video frame, while or after performing “person segmentation”. The speaker’s face or face region may be tracked and cropped from the image in each frame by using a machine learning model such as MediaPipe Holistic model; - Pg 4, Lines 37-46
NOTE: Chumkesornkulkit teaches locating the facial region of the speaker in the video frame using face detection algorithms. The algorithm comprises segmentation the video frame so identify pixels that make up the actor and pixels that make up the background. The segmentation technique functionally corresponds to generating a region proposal associated with the original actor in the video.
extracting, based on the at least one region proposal, at least one feature;
“locate facial features in the cropped face or face region. The facial features may encompass left eye, right eye, nose tip, mouth, left eye tragion, right eye tragion, etc.” – Pg 4, lines 48-50
and classifying, based on the at least one feature, the original actor.
“The face identification module (113) may be configured to be in data communication with the face detection module (111) for receiving the located face or face region and then determining the speaker’s identity based on the face or face region as received.” – Pg 4, Lines 58-60
NOTE: Chumkesornkulkit teaches a face identification module that receives the located face of the speaker in the video frame and determines the speaker’s identity. This functionally corresponds to classifying the original actor based on the received facial features from locating the face region. After the combination, the face identification module as taught by Chumkesornkulkit can modify Bhat’s system for replacing the first actor with a replacement actor. This combination would allow locating the region where the original actor is positioned in the video frame, identify a feature, and use the feature to classify the original actor before performing the replacement of the original actor with the replacement actor.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Bhat by incorporating the teachings of Chumkesornkulkit to identify the original actor in the vide asset by generating one region proposal associated with the region actor, extracting a feature of the original actor, ad classifying the original actor with the extracted feature. One would be motivated to make this combination to accurately identify the who and where the actor is in the video asset so that the replacement actor can take their place.
Regarding claim 20, Bhat in view of Chumkesornkulkit teaches the method of claim 15. Bhat does not teach wherein the integrating the facial data and the at least one second dialog further comprises: replacing the at least one first dialog with the at least one second dialog; and superimposing the facial data onto at least one frame of the video asset. However Chumkesornkulkit teaches wherein the integrating the facial data and the at least one second dialog further comprises: replacing the at least one first dialog with the at least one second dialog; and superimposing the facial data onto at least one frame of the video asset.
“using a machine learning model (such as the DeepFaceLab model) to replace or overlay the face of a speaker speaking in a first language with the face of the same speaker speaking in a second language in a video frame” – Pg 4, Lines 54-55
NOTE: Chumkesornkulkit discloses using a machine learning model to overlay the actor speaking the first language with mouth movements and facial data of the same actor speaking in a second language. This functionally corresponds to superimposing the facial data onto a frame of the video.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Foster by incorporating the teachings of Chumkesornkulkit to integrate the facial date and second dialog by replacing the first dialogue with the second dialogue and superimposing the facial data to at least one frame of the video asset. One would be motivated to make this combination because it will improve visual consistency when a viewer looks at the actor’s mouth movements and sees that it matches the new updated context (second dialog). This would lead to a predictable result of an enhanced viewing experience of the video and accurate lip-syncing.
Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhat, Chumkesornkulkit and Ramesh et al (US 11749311 B1), hereinafter Ramesh.
Regarding claim 17, Bhat in view of Chumkesornkulkit teaches the method of claim 16. Bhat does not teach wherein the receiving at least one parameter associated with the target actor further comprises: receiving, based on a region of a viewer of the video asset, the at least one parameter, wherein the region of the viewer is different from a region associated with the video asset. However, Ramesh teaches wherein the receiving at least one parameter associated with the target actor further comprises: receiving, based on a region of a viewer of the video asset, the at least one parameter, wherein the region of the viewer is different from a region associated with the video asset.
“The one or more user attributes can include a demographic attribute (e.g., an age, gender, geographic location, etc.). Additionally or alternatively, the one or more user attributes can include a viewing history specifying videos previously viewed by the viewer. With this approach, the computing system can select the replacement actor based on feedback or records indicating that other viewers having user attributes similar to the viewer selected that same replacement actor.”– Col 13, Lines 16-24
NOTE: Ramesh discloses obtaining user attributes of a viewer such as demographic, age, gender, geography, etc. These user attributes functionally correspond to the at least one parameter. The user attributes are used for selecting a replacement actor for the video which could include selecting a replacement actor of the same region as the viewer. After the combination, the obtaining of user attributes to determine a replacement actor for a video as taught by Ramesh can modify Bhat’s method for replacing a target actor to then modify the facial model’s mouth movements. This combination would then allow a receiving a user attribute that can determine the appropriate replacement actor for the video.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Bhat by incorporating the teachings of Ramesh to receive a parameter based on the region of the viewer wherein the region of the viewer is different from the region associated with the video asset. One would be motivated to make this combination in order to improve video quality and regional consistency. By acquiring at least one viewer parameter such as geography, an appropriate replacement actor of the similar parameters can then replace the original actor which results in a personalized viewing experience.
Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhat, Chumkesornkulkit and Puniyani et al US 11995118 B2, hereinafter Puniyani.
Regarding claim 19, Bhat in view of Chumkesornkulkit teaches the method of claim 15. Bhat does not teach updating a manifest file with an identifier of the target actor. However, Puniyani teaches
updating a manifest file with an identifier of the target actor
“In some embodiments, control circuitry 404 may replace the textual value of the associated first data field with the textual value of the text segment. For example, the ordered pair {Lead Actress, “ ”} may be updated to includes values {Lead Actress, “Kate Winslet”}” – Col 25, Lines 41-45
NOTE: Puniyani discloses a metadata item associated to media content such as a video/movie is in the form of a data structure that stores a label field and the corresponding text value, see Col 25, Lines 31-40. This can be understood as the manifest file. Puniyani further discloses an example of updating the contents of the metadata item by adding a new actor name “Kate Winslet” into the “Lead Actress” data field. The name “Kate Winslet” that was used to update the metadata functionally corresponds to an identifier of a target actor. One of ordinary skill could simply create another target actor identifier as taught by Puniyani for the replacement actor generated by the methods of Bhat after the combination is made.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to modify Bhat by incorporating the teachings of Puniyani to update the manifest file with an identifier of the target actor. One would be motivated to make this combination to maintain consistency of the video asset’s metadata after replacing the original actor with the new replacement actor.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID V. NGUYEN whose telephone number is (571)272-6111. The examiner can normally be reached M-F 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Y Poon can be reached at 571-270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DAVID VAN NGUYEN/
Examiner, Art Unit 2617 /KING Y POON/Supervisory Patent Examiner, Art Unit 2617