DETAILED ACTION
This Office Action is in response to the correspondence filed by the applicant on 1/3/2025.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5, 8, 9, and 14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claims 5 and 14 recite, “wherein training of the discriminator module comprises: …” There is insufficient antecedent basis for the bolded limitation in the claim.
Claims 8 and 9 recite, “wherein the mark-ups are …” There is insufficient antecedent basis for the bolded limitation in the claim.
Allowable Subject Matter
Claims 4-5 and 13-14 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and if rewritten to overcome the 112 (b) rejection above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 6, 10, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG (US 2023/0042654 A1), and in further view of WAIBEL (US 2025/0315631 A1).
REGARDING CLAIM 1, ZHANG discloses a computer-implemented method for generating a video for rendering speech, the method comprising:
receiving an original video, wherein the original video comprises a human face rendering a speech (ZHANG Fig. 6A S401; Par 55 – “In step S401, a source audio and a target video may be acquired. The target video includes the target object.”);
processing the original video to extract a mouth area of the human face (ZHANG Fig. 6A S603; Par 102 – “In step S603, mouth parameter extraction and facial parameter extraction may be performed sequentially on the target object in a current video frame of the target video to correspondingly obtain a target mouth key point parameter and a target facial parameter.”; Par 103 – “The target mouth key point parameter and the target facial parameter are parameters of the target object. When the target video includes multiple video frames, the target mouth key point parameter and the target facial parameter of the target object may be extracted from each video frame.”);
generating a speech audio from an input text using a text-to-speech (TTS) system (ZHANG Par 147 – “If the input is the source text 801, a corresponding source audio is generated through a text-to-speech module 803, and then a corresponding facial parameter is obtained from the source audio through a speech-to-facial parameter network 804.”);
generating a plurality of frames of the mouth area corresponding to a plurality of audio segments in the speech audio (ZHANG Par 116 – “In step S6055, mouth region rendering may be performed on the replaced facial parameter at each moment and the replaced mouth key point parameter at each moment through a first rendering network in the image rendering model, to obtain a mouth region texture image at each moment.”; Fig. 8 – “Speech-to-facial parameter network 804 (2D mouth key points) 804 -> Facial model 806 -> First-stage rendering network -> r1”), [wherein the mouth area in each frame is lip-synched with a corresponding audio segment of the plurality of audio segments in the speech audio] (ZHANG Par 98 – “In step S602, convolution processing and full connection processing may be performed sequentially on the audio feature vector to obtain the expression parameter and the mouth key point parameter of the source audio at the corresponding moment.”; Par 146 –“ In an embodiment of this disclosure, 2D/3D parameters may be learned by using a piece of text or audio, and thus a realistic lip-sync speech video of a specific target character is synthesized. During the implementation, an inputted text is first converted into a corresponding audio by using the TTS technology, 2D/3D facial parameters are then learned from an audio feature by using a convolutional neural network, 2D/3D facial parameters are also extracted from a video of a target character, the parameters of the target character are replaced with the learned parameters to reconstruct a new 2D/3D facial model, and the reconstructed facial model (e.g., a reconstructed image) is inputted into a rendering network to generate a video frame, thereby generating a lip-sync speech video of the target character.”; Par 155 – “FIG. 12 is a block diagram of an image rendering model, such as image rendering model 807, according to an embodiment of this disclosure. As shown in FIGS. 12, 2D mouth key points, a UV map, and a background image are given. The image rendering model is configured to synthesize a final lip-sync speech video frame. During the implementation, 20 reconstructed mouth key points may be first connected to obtain a polygon as a mouth contour KR, and then a UV map UR is mapped from 3D facial parameters based on a specific algorithm. The resolutions of KR and UR are both 256×256. KR and UR are concatenated to be used as an input of the image rendering model. The image rendering model is divided into two stages. The first stage (e.g., the first rendering network) synthesizes mouth region texture r1, and r1 a target video background frame bg (e.g., a background image) are concatenated as an input of the second rendering network.”) using a pre-trained generative neural network (ZHANG Fig. 12 – “First rendering network”; Par 36 – “Advantages of using the two-stage rendering network may include the following: (1) Training two rendering networks separately can reduce the difficulty of training and ensure the accuracy of the mouth texture generated by the first rendering network;”);
overlaying the plurality of frames of the mouth area over the human face in the original video (ZHANG Par 118 – “In step S6056, the mouth region texture image and the background image may be concatenated through a second rendering network in the image rendering model, to obtain a synthetic image at each moment.”; Fig 12 – “r1 + bg”; Par 155 – “The first stage (e.g., the first rendering network) synthesizes mouth region texture r 1, and r1 a target video background frame bg (e.g., a background image) are concatenated as an input of the second rendering network.”);
performing facial reconstruction of the human face in each frame of the original video using a trained facial reconstruction model (ZHANG Fig. 12 – “Second rendering network”; Par 119 – “In some embodiments, the second rendering network includes at least one third convolution layer, at least one second downsampling layer, and at least one second upsampling layer. … to restore resolutions of the mouth region texture image and the background image, and to obtain the synthetic image at a current moment.”) in response to the overlaying, wherein the facial reconstruction corrects distorted and blurry areas around the mouth area to achieve a natural, sharp, and expressive facial representation (ZHANG Par 36 – “(2) When the second rendering network is trained, the mouth region is penalized again to correct the mouth and optimize the details of teeth and wrinkles. In addition, when the rendering network is trained, a video frame similarity loss is further used to ensure that there is little difference between the outputted previous and next frames, avoiding the problems that the video is shaking, not smooth, or unnatural.”); and
adding the speech audio to the original video in response to performing the facial reconstruction for generating the video for rendering speech (ZHANG Par 120 – “In step S6057, the synthetic video including the target object and the source audio may be determined according to the synthetic image at each moment.”; Par 155 – “FIG. 12 is a block diagram of an image rendering model, such as image rendering model 807, according to an embodiment of this disclosure. As shown in FIGS. 12, 2D mouth key points, a UV map, and a background image are given. The image rendering model is configured to synthesize a final lip-sync speech video frame.”; Pars 66-67 – “In step S405, a synthetic video may be generated through the reconstructed image. The synthetic video includes the target object, and an action of the target object corresponds to the source audio. In an example, a synthetic video is generated based on the reconstructed image, the synthetic video including the target object, and the action of the target object being synchronized with the source audio.”; Par 143 – “For example, for streaming applications, a target object may be a virtual streamer. Through the action driving method of a target object provided in the embodiments of this disclosure, according to a text or audio inputted by a streamer, a sync-speech virtual streaming video is automatically generated. The virtual streamer can broadcast a game live to attract attention, enhance interaction through chat programs, and obtain high clicks through cover dance, thereby improving the efficiency of streaming.”).
ZHANG does not explicitly teach the [square-bracketed] limitations. ZHANG teaches generating the plurality frames of mouth area based on the source audio 802 (Fig. 8). ZHANG further teaches generating the lip-synced videos by the image rendering model (807) using the learned 2D/3D parameters. Thus, the generated mouth images after the first-stage rendering network seem to be synced with the audio. Although ZHANG suggests the [square-bracketed] limitations, the Examiner provides WAIBEL for the clarity of the rejection.
WAIBEL explicitly teaches the [square-bracketed] limitations. WAIBEL discloses a method/system for lip-synchronization comprising:
receiving an original video, wherein the original video comprises a human face rendering a speech (WAIBEL Fig. 1 Input video 12; Par 31 – “This is achieved by pipelining multiple models. With reference to the multimodal system 10 shown in FIG. 1 , first, from the original input video 12, … ”);
generating a speech audio from an input text using a text-to-speech (TTS) system (WAIBEL Par 54 – “Subsequently, the TTS module 18 is given this translated text and the resulting Mel spectrogram is turned into a waveform file by the HiFi-GAN vocoder. The final audio is now created by the voice conversion module 20, which gets the waveform of German speech that the vocoder produced as input and uses the original English audio of the input video as target speaker.”);
generating a plurality of frames of the mouth area corresponding to a plurality of audio segments in the speech audio (WAIBEL Fig. 1 – “Lip Generation 24”; Fig. 2; Par 48 – “The task is to generate the masked area of xm with respect to the audio sequence. Besides, reference image xr is useful to inject identity information to the G. Otherwise, it would be challenging for the generator to preserve the identity. Audio and image features can be concatenated along the depth to feed the face decoder 36.”; Par 54 – “The lip generation module 24 is given the detected faces in every frame of the input video as well as the speech produced by the voice conversion module and generates new video frames of the speaker's face with the lips of the speaker adapted to the given German audio.”), [wherein the mouth area in each frame is lip-synched with a corresponding audio segment of the plurality of audio segments in the speech audio] (WAIBEL Par 52 – “The lip generation module 24 preferably synchronizes as closely as possible the lip movements of the speaker in the video frames generated by the lip generation module 24 (and ultimately in the output video 26) to adapted speech from the voice conversion module 20.”; Par 96 – “a lip generation module trained, through machine learning, to generate, based on the face of the first speaker in the input video from the face detection module and from the adapted speech from the voice conversion module, new video frames of face and lips of the output speaker that are synchronized to the adapted speech from the voice conversion module;”) using a pre-trained generative neural network (WAIBEL Fig. 2; Par 19 – “FIG. 2 is a diagram of a lip generation module of the video generation system of FIG. 1 according to various embodiments of the present invention.”; Par 48 – “The task is to generate the masked area of xm with respect to the audio sequence. Besides, reference image xr is useful to inject identity information to the G. Otherwise, it would be challenging for the generator to preserve the identity. Audio and image features can be concatenated along the depth to feed the face decoder 36.”);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of ZHANG to include generating lip-synched mouth area images, as taught by WAIBEL.
One of ordinary skill would have been motivated to include generating lip-synched mouth area images, in order to generate the final video images more effectively so that rendering the final video would result in a less computational cost.
REGARDING CLAIM 6, ZHANG in view of WAIBEL discloses the method of claim 1, wherein training of the facial reconstruction model (ZHANG Par 122 – “FIG. 7 is a schematic flowchart of a method for training an image rendering model according to an embodiment of this disclosure. As shown in FIG. 7 , the method may include the following steps.”) comprises:
receiving a face image dataset capturing the user with clear visibility of lips during speech (ZHANG Par 123 – “In step S701, an image rendering model may be called based on a reconstructed image sample and a target image sample.”; Par 126 –“The target image sample includes a target object sample, and a finally generated synthetic image sample also includes the target object sample.”; Par 156 – “For the rendering network, a predicted value F (e.g., a synthetic image F) and a real value R (e.g., a real image R) are respectively concatenated to an input I (e.g., an input image I) of the rendering network and then enter a discriminator 1301, to obtain two losses LD_fake and LD_real of the real value and the predicted value.”);
modifying an original facial image using one or more image manipulating operations and artifacts to obtain one or more modified facial images (ZHANG Par 124 – “In some embodiments, the reconstructed image sample may be obtained through the following steps: performing facial parameter conversion on an audio parameter of an audio sample at a current moment to obtain an audio parameter sample; performing parameter extraction on the target image sample to obtain a target parameter sample; and combining the audio parameter sample and the target parameter sample to obtain a combined parameter sample and performing image reconstruction on a target object in the target image sample according to the combined parameter sample, to obtain the reconstructed image sample.”);
utilizing a reconstruction neural network to generate the original image with at least one modified facial image as input (ZHANG Par 118 – “In step S6056, the mouth region texture image and the background image may be concatenated through a second rendering network in the image rendering model, to obtain a synthetic image at each moment.”; Fig 12 – “r1 + bg”; Par 155 – “The first stage (e.g., the first rendering network) synthesizes mouth region texture r 1, and r1 a target video background frame bg (e.g., a background image) are concatenated as an input of the second rendering network.”; Fig. 12 – “Second rendering network”; Par 119 – “In some embodiments, the second rendering network includes at least one third convolution layer, at least one second downsampling layer, and at least one second upsampling layer. … to restore resolutions of the mouth region texture image and the background image, and to obtain the synthetic image at a current moment.”);
computing a reconstruction loss (ZHANG Par 136 – “In some embodiments, step S704 may be implemented in the following manner: acquiring a real synthetic image corresponding to the reconstructed image sample and the target image sample; and concatenating the synthetic image sample and the real synthetic image, inputting the concatenated image into the preset loss model, and calculating a previous-next frame similarity loss for the synthetic image sample and the real synthetic image through the preset loss model, to obtain the loss result.”); and
updating weights for the reconstruction neural network based on the reconstruction loss (ZHANG Par 138 – “In step S705, parameters in the first rendering network and the second rendering network may be modified according to the loss result to obtain the image rendering model after training.”).
REGARDING CLAIM 10, ZHANG in view of WAIBEL discloses a computer system comprising a processor and one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by the processor to cause the processor (ZHANG Fig. 3 Par 45 – “processor .. a memory …”) to perform operations comprising: performing the steps of claim 1; thus, it is rejected under the same rationale.
Claim 15 is similar to claim 6; thus, it is rejected under the same rationale.
Claims 2-3 and 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG (US 2023/0042654 A1) in view of WAIBEL (US 2025/0315631 A1), and in further view of FU (US 2026/0094469 A1).
REGARDING CLAIM 2, ZHANG in view of WAIBEL discloses the method of claim 1.
[aligning each frame of the original video based on detected mouth position of the human face];
recognizing the mouth area of the human face at each frame of the original video and representing the mouth area as a mask (ZHANG Par 102 – “In step S603, mouth parameter extraction and facial parameter extraction may be performed sequentially on the target object in a current video frame of the target video to correspondingly obtain a target mouth key point parameter and a target facial parameter.”; Par 154 – “Facial parameter extraction module 805 is configured to extract 2D mouth key point positions and 3D facial parameters of a target character from video frames of the target character.”); and
cropping an image of the mouth area according to the mask (ZHANG Par 115 – “The replaced facial parameter, the replaced mouth key point parameter, and the background image corresponding to the target video at each moment are inputted into the image rendering model. The reconstructed image includes the replaced facial parameter and the replaced mouth key point parameter.”).
ZHANG does not explicitly teach the [square-bracketed] limitations.
FU discloses the [square-bracketed] limitations. FU discloses a method/system for facial image processing, wherein the processing of the original video comprises:
[aligning each frame of the original video based on detected mouth position of the human face] (FU Par 84 – “Based on the above disclosure of the second transformed image, for some scenarios, after the face image is obtained, the face image is first processed to detect facial key points, to obtain the facial key points of the face image, such that the facial key points include key points of respective parts in the face, e.g., key points of the target part; then, the key points of the target part are aligned with the key points of the preset second image space (i.e., average mouth) to obtain the second transformation matrix, thereby minimizing the errors between the key points obtained from mapping the key points of the target part to the second image space according to the second transformation matrix and the key points of the second image space; accordingly, the second transformation matrix can indicate alignment between the pixel points in the image space of the face image and the pixel points in the second image space. Next, the second transformation matrix is considered as the affine transformation matrix, and the face image is transformed into the second image space by affine transformation, to obtain the second transformed image, such that the second transformed image can display the target part based on the cropping constraints described by the second image space (e.g., facial dimension, display position of mouth and the like). Accordingly, the second transformed image can represent the area of the target part cropped from the face image and the information other than the target part information in the second transformed image is far little than the information other than the target part information in the face image. This can effectively avoid the interference caused by other information.”);
recognizing the mouth area of the human face at each frame of the original video and representing the mouth area as a mask (FU Par 113 – “Based on the contents related to the above steps 11-13, for some scenarios, e.g., application scenarios with low computation load, after the first transformed image and the second transformed image are obtained, the two images may first be spliced to obtain a splicing result, e.g., the splicing result shown in FIG. 2 , such that the splicing result can describe the facial information carried by the first transformed image and the target part information carried by the second transformed image with as few pixel points as possible; then, the splicing result is input to the machine learning model with area prediction function, to obtain a first mask (e.g., the facial mask shown in FIG. 2 ) output by the model for representing the position of the visible area of the face in the first transformed image and a second mask (e.g., the mouth mask shown in FIG. 2 ) for indicating the position of the visible area of the target part in the second transformed image; afterwards, the first mask is transformed back to the image space of the face image according to the inverse transformation matrix corresponding to the first transformation matrix, to obtain an identification result for the visible area of the face; and the second mask is transformed back to the image space of the face image according to the inverse transformation matrix corresponding to the second transformation matrix, to obtain an identification result for the visible area of the target part.”); and
cropping an image of the mouth area according to the mask (FU Par 113 – “Based on the contents related to the above steps 11-13, for some scenarios, e.g., application scenarios with low computation load, after the first transformed image and the second transformed image are obtained, the two images may first be spliced to obtain a splicing result, e.g., the splicing result shown in FIG. 2 , such that the splicing result can describe the facial information carried by the first transformed image and the target part information carried by the second transformed image with as few pixel points as possible; then, the splicing result is input to the machine learning model with area prediction function, to obtain a first mask (e.g., the facial mask shown in FIG. 2 ) output by the model for representing the position of the visible area of the face in the first transformed image and a second mask (e.g., the mouth mask shown in FIG. 2 ) for indicating the position of the visible area of the target part in the second transformed image; afterwards, the first mask is transformed back to the image space of the face image according to the inverse transformation matrix corresponding to the first transformation matrix, to obtain an identification result for the visible area of the face; and the second mask is transformed back to the image space of the face image according to the inverse transformation matrix corresponding to the second transformation matrix, to obtain an identification result for the visible area of the target part.”; Par 74 – “In this way, the target part described by the second transformed image can satisfy the cropping constraints of the target part depicted by the second image space as much as possible, and the second transformed image can better represent the result obtained from the target part cropping on the face image, e.g., the cropped image of the mouth shown in FIG. 2 .”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of ZHANG in view of WAIBEL to include aligning the frames, as taught by FU.
One of ordinary skill would have been motivated to include aligning the frames, in order to effectively reduce the interference caused by the background image (Par 76).
REGARDING CLAIM 3, ZHANG in view of WAIBEL and FU discloses the method of claim 2, wherein during the processing of the original video, a closed mouth is generated in the mouth area using a trained Al model (ZHANG Par 155 – “The resolutions of KR and UR are both 256×256. KR and UR are concatenated to be used as an input of the image rendering model. The image rendering model is divided into two stages. The first stage (e.g., the first rendering network) synthesizes mouth region texture r1, and r1 a target video background frame bg (e.g., a background image) are concatenated as an input of the second rendering network. The second stage (e.g., the second rendering network) and the background image are combined to synthesize a final output r 2.”; Par 150 – “The text analysis module 901 is configured to parse an inputted text (e.g., a source text) to determine pronunciation, tone, intonation, and the like of each word, and map the text to a linguistic feature. The linguistic feature includes, but is not limited to: pinyin, pause, punctuation, tone, and the like. The acoustic model module 902 is configured to map the linguistic feature to an acoustic parameter. The acoustic parameter is parametric representation of the source text on a time domain. The vocoder module 903 is configured to convert the acoustic parameter to an audio waveform. The audio waveform is parametric representation of the source text on a frequency domain.”; In other words, since the source mouth key points are generated based on the source text, a closed mouth (for pronunciation “m”) is generated during the processing of the original video to generate the targe video synced with the source text.).
Claim 11 is similar to claim 2; thus, it is rejected under the same rationale.
Claim 12 is similar to claim 3; thus, it is rejected under the same rationale.
Claims 7 is rejected under 35 U.S.C. 103 as being unpatentable over ZHANG (US 2023/0042654 A1) in view of WAIBEL (US 2025/0315631 A1), and in further view of FILEV (US 2008/0269958 A1).
REGARDING CLAIM 7, ZHANG in view of WAIBEL discloses the method of claim 1.
WAIBEL further teaches mark-ups (WAIBEL Par 44 – "The possibility to add information input is added by supporting Speech Synthesis Markup Language (see Paul Taylor and Amy Isard, 1997, “Ssml: A speech synthesis markup language,” Speech communication, 21 (1-2): 123-133, which is incorporated herein by reference in its entirety) (SSML) tags regarding the controllable aspects of prosody in the input text. Using SSML, emphasis tags can be added to words in the translation that correspond to words in the original transcript that were emphasized by the speaker. The system will then, in various embodiments, adapt the prosodic control values for the phonemes of that word to create an emphasis in the output.”), but does not explicitly teach all the limitations of claim 7.
FILEVE discloses a method/system for generating speech from textual data, wherein the input text is enriched with mark-ups denoting a plurality of signs, wherein the plurality of signs comprises emotion (FILEV Par 64 – “Emotively tagged text includes marked-up phrases that indicate emotional content associated with certain words of the phrase. The avatar appearance is dynamically altered to express emotion, indicate speech is taking place and/or convey information, etc. The avatar expression is controlled by manipulating specific points on the surface of the avatar. In a computer generated avatar, a mathematical representation of a 3D surface of a physical avatar is made, typically using polygonal modeling techniques/algorithms.”; Par 142 – “This vector representation is used by the emotional speech synthesizer 164 of FIG. 13 to mark-up portions of the text to be spoken by the avatar with indicators of emotion.”), word emphasis, sound volume (FILEV Par 142 – “These emotional markers are later interpreted by a text to speech engine, discussed below, in order to dynamically alter the prosody, tone, inflection, etc., with which the marked words are spoken. Such dynamic alteration of word pronunciation may convey emotional content in the speech.”; Par 144 – “Other speech markers are defined in the Speech Synthesis Markup Language specification from the World Wide Web Consortium (W3C). Prosodic and emphasis elements that indicate avatar emotion are of the form "Have a <emphasis> nice </emphasis> day!" which would put the stress on the word nice.”) and special articulation indicating specific lip movements and shape of mouth (Par 68 – “The lip movements of the avatar may be animated using a set of predefined lip positions that are correlated to each allophone of speech. A number corresponding to a viseme may be used to index each position which is either morphed or concatenated to the rest of the avatar's face. There are standard viseme sets such as the Disney visemes and several others that are in common use. The text-to-speech engine produces a stream of visemes that are time synchronized to the speech that is produced. The visemes are streamed to the rendering engine to affect lip movement.”), and facial expressions, wherein the mark-ups are employed by the facial reconstruction model (Par 63 – “The avatar gestures and text-to-speech control inform the avatar controller 92 as to how to render movement and facial expressions of the avatar as the avatar speaks. For example, the avatar gestures may control hand movements, gaze direction, etc., of the avatar. The text-to-speech control may control when to begin, end, suspend, abort or resume any text-to-speech operations. The avatar emotion and emotively tagged text, as discussed in detail below, inform the avatar controller 92 as to how to render movement and facial expressions of the avatar as the avatar expresses emotion.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of ZHANG in view of WAIBEL to include mark-ups for indicating emotions, as taught by FILEV.
One of ordinary skill would have been motivated to include mark-ups for indicating emotions, in order to generate more natural synthetic audio-visual data with parameters.
Claims 8 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG (US 2023/0042654 A1) in view of WAIBEL (US 2025/0315631 A1), and in further view of SUBRAMANIAN (US 2007/0208569 A1).
REGARDING CLAIM 8, ZHANG in view of WAIBEL discloses the method of claim 1.
WAIBEL further teaches mark-ups (WAIBEL Par 44 – "The possibility to add information input is added by supporting Speech Synthesis Markup Language (see Paul Taylor and Amy Isard, 1997, “Ssml: A speech synthesis markup language,” Speech communication, 21 (1-2): 123-133, which is incorporated herein by reference in its entirety) (SSML) tags regarding the controllable aspects of prosody in the input text. Using SSML, emphasis tags can be added to words in the translation that correspond to words in the original transcript that were emphasized by the speaker. The system will then, in various embodiments, adapt the prosodic control values for the phonemes of that word to create an emphasis in the output.”), but does not explicitly teach all the limitations of claim 9.
SUBRAMANIAN discloses a method/system for voice/text generation and communication, wherein the mark-ups are specific to a speaker (SUBRAMANIAN Par 63 – "The text mining solution improves accuracy and speed by using text mining databases particular for languages and over voice analysis alone. In cases where text mining emotion-text/phrase dictionary 220 is used for analysis of speech from a particular person, the dictionary can be further trained either manually or automatically to provide higher weights to the user's most frequently used phrases and learned emotional content of those phrases. That information can be saved in the user's profile.”; Par 64 – “As discussed above, emotion markup component 210 derives the emotion from a voice communication stream using two separate emotion analyses, voice pattern analysis (voice analyzer 232) and text pattern analysis (text/phrase analyzer 236).”, wherein the mark-ups are generated by a trained mark-up model, the mark-up model being trained to recognize and replicate a speaking style of a speaker (SUBRAMANIAN Par 47 – “The personality attributes for a speaker are learned emotional content of words and phrases that personal the speaker. These attributes are also used for modifying the dictionary definitions for words and speech patterns that the speaker uses to convey emotion to an audience, but often the personality attributes are learned emotional content of words and phrases that may be inconsistent or even contradictory to their generally accepted emotion content.”; Par 41 – "The textual data with emotion markup produced by emotion markup component 210 can be archived in a database for future searching or training, or transmitted to other devices that include emotion translation component 250 for reproducing the speech that preserves the emotion of the original communication.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of ZHANG in view of WAIBEL to include user-specific markups, as taught by SUBRAMANIAN.
One of ordinary skill would have been motivated to include user-specific markups, in order to provide more customized synthetic audio-visual data to a user according to the user’s specific needs.
REGARDING CLAIM 9, ZHANG in view of WAIBEL discloses the method of claim 1,
WAIBEL further teaches mark-ups (WAIBEL Par 44 – "The possibility to add information input is added by supporting Speech Synthesis Markup Language (see Paul Taylor and Amy Isard, 1997, “Ssml: A speech synthesis markup language,” Speech communication, 21 (1-2): 123-133, which is incorporated herein by reference in its entirety) (SSML) tags regarding the controllable aspects of prosody in the input text. Using SSML, emphasis tags can be added to words in the translation that correspond to words in the original transcript that were emphasized by the speaker. The system will then, in various embodiments, adapt the prosodic control values for the phonemes of that word to create an emphasis in the output.”), but does not explicitly teach all the limitations of claim 9.
SUBRAMANIAN discloses a method/system for voice/text generation and communication, wherein the mark-ups are generic across speakers (SUBRAMANIAN Par 43 – “In accordance with one embodiment of the present invention, emotion is deduced from a communication by text pattern analysis and voice analysis. Emotion-voice pattern dictionary 222 contains emotion to voice pattern definitions for deducing emotion from voice patterns in a communication, while emotion-text/phrase dictionary 220 contains emotion to text pattern definitions for deducing emotion from text patterns in a communication. The dictionary definitions can be generic and abstracted across speakers, or specific to a particular speaker, audience and circumstance of a communication.”), wherein the mark-ups are generated by a trained mark-up model based on the input text (SUBRAMANIAN Par 44 – “A generic, or default, will provide acceptable mainstream results for deducing emotion in a communication. The dictionary definitions can be optimized for a particular speaker, audience and circumstance of a communication and achieve highly accurate emotion recognition results in the context of the optimization, but the mainstream results suffer dramatically. The generic dictionaries can be optimized by training, either manually or automatically, to provide higher weights to the most frequently used text patterns (words and phrases) and voice patterns, and to provide learned emotional content to text and voice patterns.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the method/system of ZHANG in view of WAIBEL to include generic markups, as taught by SUBRAMANIAN.
One of ordinary skill would have been motivated to include generic markups, in order to provide more universal synthetic audio-visual data to a user.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN C KIM whose telephone number is (571)272-3327. The examiner can normally be reached Monday to Friday 8:00 AM thru 4:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew C Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONATHAN C KIM/Primary Examiner, Art Unit 2655