Prosecution Insights
Last updated: October 01, 2026
Application No. 19/006,056

DEVICE AND METHOD FOR GENERATING AVATAR LIP-SYNC ANIMATION BASED ON MULTIMODAL BIOSIGNALS

Non-Final OA §103§112
Filed
Dec 30, 2024
Priority
Jun 25, 2024 — RE 10-2024-0082346
Examiner
LI, JAI WEI TOMMY
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Korea University Research and Business Foundation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
32 currently pending
Career history
33
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claim(s) 2 objected to because of the following informalities: “the coordinate values” was not disclosed in the current/dependent claim, and should be properly introduced prior. Appropriate correction is required. Specification The disclosure is objected to because of the following informalities: Figure 1 exposes each label to include the "s" whereas, the specifications do not. Appropriate correction is required. Claim Rejections - 35 USC § 112 Claim 8 recites the limitation "the user" in "brain waves emitted by the user". There is insufficient antecedent basis for this limitation in the claim, as claim 8 does not previously recite a user. Claim 8 recites the limitation "the mouth shape and facial movement " in "predicting the mouth shape and facial movement after the multimodal data is collected". There is insufficient antecedent basis for this limitation in the claim, as claim 8 does not previously recite a mouth shape and facial movement. Claim 8 recites the limitation "the lip-sync reconstruction step" in "the lip-sync reconstruction step to the avatar generated in the avatar generation step". There is insufficient antecedent basis for this limitation in the claim, as claim 8 does not previously recite a lip-sync reconstruction step. Claim 9 recites the limitation "the embedding convergence vector" in “by inputting the embedding convergence vector into the pre-prepared lip-sync reconstruction model”. There is insufficient antecedent basis for this limitation in the claim, as claim 9 does not previously recite an embedding convergence vector. Claim 9 recites the limitation " the lip-sync reconstruction step" in “wherein the lip-sync reconstruction step predicts the mouth shape and facial movement”. There is insufficient antecedent basis for this limitation in the claim, as claim 9 does not previously recite a lip-sync reconstruction step. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, and 7-9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (U.S. Pub. No. 20210248801) in view of Lee et al. (U.S. Pub. No. 20200142481) and Chong et al. (U.S. Pub. No. 20210005003). Regarding claim 1, Lee discloses a multimodal biosignal-based avatar lip-sync animation generation device comprising (para 32, “Referring now to FIG. 1, a block diagram of an example computing system for audio-driven talking head animation is shown. Generally, environment 100 is suitable for audio-driven animation, and, among other things, facilitates automatic generation of an animation from an audio signal of speech and an image or some other structure representing a head to animate. Environment 100 includes client device 105 and server 120.”; also, para 34, “Server 120 includes audio-driven talking head animator 125. Generally, audio-driven talking head animator 125 automatically generates an animation of a talking head from an audio signal of speech and an image or some other structure representing a head to animate.”): a multimodal data collection circuit configured to collect multimodal data including image data (para 35, “At a high level, a user operating application 107 on client device 105 may identify an audio signal of speech and an image or some other structure representing a head to animate, and application 107 may transmit the input speech and input image (or corresponding paths at which the input speech and input image may be accessed) to audio-driven talking head animator 125 via network 110.”; also, para 38, “In this example, the inputs into audio-driven talking head animator 200 are audio 205 and an image of a head to animate 220.”; also, para 62, “With reference to FIG. 7, computing device 700 includes bus 710 that directly or indirectly couples the following devices: memory 712, one or more processors 714, one or more presentation components 716, input/output (I/O) ports 718, input/output components 720, and illustrative power supply 722.”); feature extraction circuit configured to extract feature vectors including facial feature of the user from the preprocessed multimodal data (para 38, “facial landmark extractor 225 may extract a set of template facial landmarks 230 from the image of the head 220, and audio feature extractor 210 may extract an audio feature vector 215 from a window of audio 205”; also, para 43, “Generally, landmark extractor 225 may use any known technique to detect and/or extract a set of facial landmarks from an input image, as will be understood by those of ordinary skill in the art.”); lip-sync reconstruction circuit configured to predict a mouth shape and facial movement after the multimodal data is collected, by inputting the extracted feature vectors into a pre-prepared lip-sync reconstruction model (para 44, “Generally, landmark predictor 240 accepts as inputs the set of template facial landmarks 230 and audio feature vector 215, and generates a set of predicted 3D facial landmarks 250 corresponding to the window of audio 205 from which audio feature vector 215 was extracted.”; also, para 19, “the present audio-driven animation techniques can synchronize multiple components of a talking head, including lips, nose, eyes, ears, and head pose”; also, para 48, “the predicted 3D facial landmarks may be compared with ground truth 3D facial landmarks extracted from a corresponding video frame from a training video, and single-style landmark predictor 300 may be updated accordingly (e.g., by minimizing L2 distance between predicted and ground-truth facial landmarks)”); and a lip-sync animation implementation circuit configured to implement an avatar lip- sync animation by applying the mouth shape and facial movement predicted by the lip-sync reconstruction circuit to the avatar generated by the avatar generation circuit (para 38, “Image warper 260 uses the predicted 3D facial landmarks 250 to warp the image of the head 220 to generate an animation frame. The process may be repeated for successive windows of audio 205 to generate successive animation frames, which are compiled by animation compiler 270 into an animation video, using audio 205 as the audio track.”; also, para 54, “In some embodiments, image warper 260 may apply Delaunay triangulation on the predicted facial landmarks to derive a set of corresponding triangles, which can be used to subdivide the image into corresponding triangle regions. Each landmark point can be inserted into a corresponding triangle region, and image warper 260 can warp the image to align each landmark point with its corresponding triangle region.”; also, para 55, “As such, animation compiler 270 may compile animation 280 using the sequence of warped images generated by image warper 260 and audio 205.”). Li does not disclose a biosignal data which includes brain waves emitted from a user, including speaking, a preprocessing circuit configured to preprocess the multimodal data, a biosignal feature, and an avatar generation circuit configured to generate an avatar that represents an appearance of the user. However, in a similar field of endeavor, Lee discloses a biosignal data which includes brain waves emitted from a user, including speaking (para 44, “the sensor module 150 detects the brainwave related to the imagined speech in a state worn by the user”; also, para 37, “measure a brainwave of imagined speech related to the conversational intention of a user”), a preprocessing circuit configured to preprocess the multimodal data (para 51, “performs filtering at a frequency of 0.5 Hz to 50 Hz band through a preprocessor (not shown) for the brainwave data measured by the sensor module 150 to minimize the influence of ambient noise and power DC noise”), a biosignal feature (para 52, “extracts features for common relationship between the imagined speech brainwave, the audible act brainwave, the speech act brainwave, and the soundwave data by using a signal feature extraction algorithm including Wavelet Transform, Common Spatial Pattern (CSP), Principal Component Analysis (PCA), or Independent Component Analysis (ICA)”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention of an avatar lip-sync animation generation device comprising a multimodal data collection circuit that collects image data, a feature extraction circuit that extracts a facial feature of the user, a lip-sync reconstruction circuit that predicts a mouth shape and facial movement by inputting the extracted feature vectors into a pre-prepared neural-network lip-sync reconstruction model, and a lip-sync animation implementation circuit that implements the animation by warping to the predicted facial-landmark coordinates, with the features of Lee's invention of collecting biosignal data that includes brain waves emitted from a user during imagined speech, preprocessing the multimodal data by filtering the brainwave data, and extracting a biosignal feature from the preprocessed data. The combination would have been obvious because Lee teaches that the conversational intention of a user can be decoded from the brain waves measured while the user imagines speaking, so supplying that biosignal-derived feature as an input to Li's animation pipeline would drive the lip-sync from the user's imagined speech rather than from recorded audio, yielding the predictable result of an avatar lip-sync animation generated from multimodal biosignals. Chong discloses an avatar generation circuit configured to generate an avatar that represents an appearance of the user (para 45, “An avatar indicates a 3D animation model that resembles a specific person within a two-dimensional (2D) image.”; also, para 49, “In response thereto, the server 130 extracts a face from the input 2D image and extracts a landmark from the extracted face. Thereafter, the server 130 generates a first mesh model using the landmark.”; also, para 95, “According to the method of generating a 3D avatar from a 2D image according to the present disclosure, a user can generate and show an avatar that very resembles himself or herself, that has a highlighted beauty and character characteristics, and that is fun.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention as modified by Lee with the features of Chong's invention of an avatar generation circuit that generates an avatar representing the appearance of the user. The combination would have been obvious because Chong teaches that a three-dimensional avatar built from the facial landmarks extracted from a two-dimensional image lets a user generate and show an avatar that very resembles himself or herself, while Li already extracts those same facial landmarks from the input image, so applying Li's predicted mouth shape and facial movement to Chong's generated avatar rather than to the original photograph would yield the predictable result of an avatar lip-sync animation that represents the user's appearance. Regarding claim 2, Li as modified by Lee and Chong discloses the device according to claim 1, wherein Li further discloses lip-sync animation implementation circuit is configured to implement an avatar lip-sync animation by applying the mouth shape and facial movement predicted by the lip-sync reconstruction circuit to the avatar generated by the avatar generation circuit (Li: para 38, “Image warper 260 uses the predicted 3D facial landmarks 250 to warp the image of the head 220 to generate an animation frame. The process may be repeated for successive windows of audio 205 to generate successive animation frames, which are compiled by animation compiler 270 into an animation video, using audio 205 as the audio track.”; also, para 54, “In some embodiments, image warper 260 may apply Delaunay triangulation on the predicted facial landmarks to derive a set of corresponding triangles, which can be used to subdivide the image into corresponding triangle regions. Each landmark point can be inserted into a corresponding triangle region, and image warper 260 can warp the image to align each landmark point with its corresponding triangle region.”; also, para 55, “As such, animation compiler 270 may compile animation 280 using the sequence of warped images generated by image warper 260 and audio 205.”), based on the coordinate values of the facial landmark (Li: para 47, “if 68 facial landmark coordinates in 3D space are desired for each window of audio clip 305”; also, para 4, “the input image of the head can be warped to fit the predicted 3D facial landmarks”). Li does not disclose wherein the avatar generation circuit is configured to generate an avatar in a two-dimensional or three-dimensional form from the image data of the user using computer vision technology, and maps the facial feature extracted by the feature extraction circuit to the generated avatar to thereby specify a facial landmark. However, in a similar field of endeavor, Chong discloses wherein the avatar generation circuit is configured to generate an avatar in a two-dimensional or three-dimensional form from the image data of the user using computer vision technology (para 46, “The user terminal 110 may generate a 3D avatar from a 2D image captured by the camera or a 2D image obtained from another apparatus.”; also, para 54, “The first mesh generation module 212 generates a first mesh model 208 by applying a mesh topology indicative of a 3D geometrical characteristic of the face using the extracted landmark.”; also, para 70, “Face attribute information indicative of such detailed characteristics of faces may be inferred through a deep learning algorithm (model) related to image analysis.”), and maps the facial feature extracted by the feature extraction circuit to the generated avatar to thereby specify a facial landmark (para 64, “The extracted landmark information may be represented as coordinates. A mesh model to which the landmark has been applied is generated by matching the coordinates of the landmark with the defined vertexes of the mesh topology.”; also, para 93, “As described above, a blend shape defined in relation to a size and motion adjustment with respect to the mean face mesh model is also naturally applied to a 3D avatar”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention of a lip-sync animation implementation circuit that applies the predicted mouth shape and facial movement based on the coordinate values of the facial landmark with the features of Chong's invention of an avatar generation circuit that generates a two-dimensional or three-dimensional avatar from the image data of the user using computer vision and maps the extracted facial feature to the generated avatar to thereby specify a facial landmark. The combination would have been obvious because Chong teaches matching the coordinates of the extracted landmark with the defined vertexes of the mesh topology, so that the generated avatar carries the very landmark coordinates that Li's animation already drives, yielding the predictable result of a lip-sync animation applied by landmark coordinate value to an avatar that resembles the user. Regarding claim 3, Li as modified by Lee and Chong discloses the device according to claim 1, further comprising a feature convergence circuit configured to converge the feature vectors extracted from the feature extraction circuit and converting them into an embedding convergence vector (Li: para 40, “the N columns can be flattened by concatenating the values of the columns (e.g., magnitude and/or phase values) into a single dimensional vector (e.g., audio feature vector 215)”; also, para 47, “speech content encoder 330 may encode an extracted audio feature vector 320 and a set of template 3D facial landmarks 325 into a speech content embedding, and single-style landmark decoder 340 may decode the speech content embedding into a set of predicted 3D facial landmarks 350”), wherein the lip-sync reconstruction circuit is configured to predict mouth shape and facial movement by inputting the embedding convergence vector into the pre-prepared lip-sync reconstruction model (Li: para 47, “speech content encoder 330 may encode an extracted audio feature vector 320 and a set of template 3D facial landmarks 325 into a speech content embedding, and single-style landmark decoder 340 may decode the speech content embedding into a set of predicted 3D facial landmarks 350”). Regarding claim 7, Li as modified by Lee and Chong discloses the device according to claim 1, wherein Li further discloses the lip-sync reconstruction model comprises of any one of: a first prediction circuit configured to predict the mouth shape and facial movement from the extracted feature vectors (Li: para 44, “Generally, landmark predictor 240 accepts as inputs the set of template facial landmarks 230 and audio feature vector 215, and generates a set of predicted 3D facial landmarks 250 corresponding to the window of audio 205 from which audio feature vector 215 was extracted.”; also, para 44, “Generally, landmark predictor 240 may be designed to account for a single speaking style (e.g., single-style landmark predictor 130 of FIG. 1) or multiple speaking styles (e.g., multi-style landmark predictor 140 of FIG. 1).”; also, para 19, “the present audio-driven animation techniques can synchronize multiple components of a talking head, including lips, nose, eyes, ears, and head pose”), or mouth shape and facial movement based on the classified intentions (Li: para 47, “speech content encoder 330 may encode an extracted audio feature vector 320 and a set of template 3D facial landmarks 325 into a speech content embedding, and single-style landmark decoder 340 may decode the speech content embedding into a set of predicted 3D facial landmarks 350”; also, para 19, “the present audio-driven animation techniques can synchronize multiple components of a talking head, including lips, nose, eyes, ears, and head pose"). Li does not disclose a second prediction circuit configured to identify and classify intentions of the user from the extracted feature vectors. However, in a similar field of endeavor, Lee discloses a second prediction circuit configured to identify and classify intentions of the user from the extracted feature vectors (para 54, “a decoding model 200 includes a classification model 210 of classifying the words intended by the user through features for the brainwave or the soundwave generated when performing the imagined speech, the speech act, or the audible act”; also, para 37, “automatically generate a candidate sentence by combining classified words after classifying words intended by the user from the measured brainwave”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention of a first prediction circuit that predicts the mouth shape and facial movement from the extracted feature vectors, and of a speech content encoder that encodes the extracted feature vector into a speech content embedding and a landmark decoder that decodes that embedding into the predicted facial landmarks, with the features of Lee's invention of a second prediction circuit that identifies and classifies the intentions of the user from the extracted feature vectors. The combination would have been obvious because Lee teaches a classification model that classifies the words the user intends from the features of the brainwave, and Li teaches that the mouth shape is predicted by decoding an encoded speech content embedding rather than the raw signal, so encoding the words classified by Lee as the speech content that Li's landmark decoder converts into predicted facial landmarks would yield the predictable result of a mouth shape and facial movement predicted from the intention the user was classified as having. Regarding claim 8, Li discloses a multimodal biosignal-based avatar lip-sync animation generation method comprising (para 58, “Turning initially to FIG. 5, FIG. 5 illustrates a method 500 for audio-driven animation, in accordance with embodiments described herein.”; also, para 57, “With reference now to FIGS. 5-6, flow diagrams are provided illustrating some example methods for audio-driven animation. Each block of the methods 500 and 600 and any other methods described herein comprise a computing process performed using any combination of hardware, firmware, and/or software.”): collecting multimodal data including image data (para 35, “At a high level, a user operating application 107 on client device 105 may identify an audio signal of speech and an image or some other structure representing a head to animate, and application 107 may transmit the input speech and input image (or corresponding paths at which the input speech and input image may be accessed) to audio-driven talking head animator 125 via network 110.”; also, para 38, “In this example, the inputs into audio-driven talking head animator 200 are audio 205 and an image of a head to animate 220.”); extracting feature vectors including a facial feature of the user from the preprocessed multimodal data (para 38, “facial landmark extractor 225 may extract a set of template facial landmarks 230 from the image of the head 220, and audio feature extractor 210 may extract an audio feature vector 215 from a window of audio 205”; also, para 43, “Generally, landmark extractor 225 may use any known technique to detect and/or extract a set of facial landmarks from an input image, as will be understood by those of ordinary skill in the art.”); mouth shape and facial movement after the multimodal data is collected by inputting the extracted feature vectors into a pre-prepared lip-sync reconstruction model (para 44, “Generally, landmark predictor 240 accepts as inputs the set of template facial landmarks 230 and audio feature vector 215, and generates a set of predicted 3D facial landmarks 250 corresponding to the window of audio 205 from which audio feature vector 215 was extracted.”; also, para 19, “the present audio-driven animation techniques can synchronize multiple components of a talking head, including lips, nose, eyes, ears, and head pose”; also, para 48, “the predicted 3D facial landmarks may be compared with ground truth 3D facial landmarks extracted from a corresponding video frame from a training video, and single-style landmark predictor 300 may be updated accordingly (e.g., by minimizing L2 distance between predicted and ground-truth facial landmarks)”); and implementing an avatar lip-sync animation by applying the mouth shape and facial movement predicted in the lip-sync reconstruction step to the avatar generated in the avatar generation step (para 38, “Image warper 260 uses the predicted 3D facial landmarks 250 to warp the image of the head 220 to generate an animation frame. The process may be repeated for successive windows of audio 205 to generate successive animation frames, which are compiled by animation compiler 270 into an animation video, using audio 205 as the audio track.”; also, para 55, “As such, animation compiler 270 may compile animation 280 using the sequence of warped images generated by image warper 260 and audio 205.”). Li does not disclose a biosignal data which includes brain waves emitted by the user, including speaking, preprocessing the multimodal data, a biosignal feature, generating an avatar that represents an appearance of the user based on facial features among the extracted feature vectors. However, in a similar field of endeavor, Lee discloses a biosignal data which includes brain waves emitted by the user, including speaking (para 44, “the sensor module 150 detects the brainwave related to the imagined speech in a state worn by the user”; also, para 37, “measure a brainwave of imagined speech related to the conversational intention of a user”), preprocessing the multimodal data (para 51, “performs filtering at a frequency of 0.5 Hz to 50 Hz band through a preprocessor (not shown) for the brainwave data measured by the sensor module 150 to minimize the influence of ambient noise and power DC noise”), a biosignal feature (para 52, “extracts features for common relationship between the imagined speech brainwave, the audible act brainwave, the speech act brainwave, and the soundwave data by using a signal feature extraction algorithm including Wavelet Transform, Common Spatial Pattern (CSP), Principal Component Analysis (PCA), or Independent Component Analysis (ICA)”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention of an avatar lip-sync animation generation method that collects image data, extracts a facial feature of the user, predicts a mouth shape and facial movement by inputting the extracted feature vectors into a pre-prepared neural-network lip-sync reconstruction model, and implements the animation by warping to the predicted facial-landmark coordinates, with the features of Lee's invention of collecting biosignal data that includes brain waves emitted by the user during imagined speech, preprocessing the multimodal data by filtering the brainwave data, and extracting a biosignal feature from the preprocessed data. The combination would have been obvious because Lee teaches that the conversational intention of a user can be decoded from the brain waves measured while the user imagines speaking, so supplying that biosignal-derived feature as an input to Li's animation pipeline would drive the lip-sync from the user's imagined speech rather than from recorded audio, yielding the predictable result of an avatar lip-sync animation generated from multimodal biosignals. Chong discloses generating an avatar that represents an appearance of the user based on facial features among the extracted feature vectors (para 49, “In response thereto, the server 130 extracts a face from the input 2D image and extracts a landmark from the extracted face. Thereafter, the server 130 generates a first mesh model using the landmark.”; also, para 45, “An avatar indicates a 3D animation model that resembles a specific person within a two-dimensional (2D) image.”; also, para 95, “According to the method of generating a 3D avatar from a 2D image according to the present disclosure, a user can generate and show an avatar that very resembles himself or herself, that has a highlighted beauty and character characteristics, and that is fun.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention as modified by Lee with the features of Chong's invention of generating an avatar that represents the appearance of the user from the facial landmarks extracted from the image. The combination would have been obvious because Chong teaches that a three-dimensional avatar built from the facial landmarks extracted from a two-dimensional image lets a user generate and show an avatar that very resembles himself or herself, while Li already extracts those same facial landmarks from the input image, so applying Li's predicted mouth shape and facial movement to Chong's generated avatar rather than to the original photograph would yield the predictable result of an avatar lip-sync animation that represents the user's appearance. Regarding claim 9, Li as modified by Lee and Chong discloses the method according to claim 8, further comprising converging the extracted feature vectors (Li: para 40, “the N columns can be flattened by concatenating the values of the columns (e.g., magnitude and/or phase values) into a single dimensional vector (e.g., audio feature vector 215)”), wherein the lip-sync reconstruction step predicts the mouth shape and facial movement by inputting the embedding convergence vector into the pre-prepared lip-sync reconstruction model (Li: para 47, “speech content encoder 330 may encode an extracted audio feature vector 320 and a set of template 3D facial landmarks 325 into a speech content embedding, and single-style landmark decoder 340 may decode the speech content embedding into a set of predicted 3D facial landmarks 350”). Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (U.S. Pub. No. 20210248801) as modified by Lee et al. (U.S. Pub. No. 20200142481) and Chong et al. (U.S. Pub. No. 20210005003), further in view of Chi (U.S. Pub. No. 20150358096). Regarding claim 4, Li as modified by Lee and Chong discloses the device according to claim 1, wherein the multimodal data collection circuit includes: presented sentence transfer display circuit configured to transfer a presented sentence to the user; a biosignal collection circuit configured to collect biosignal data by measuring biosignals including brain waves of a user; an image collection circuit configured to collect image data by photographing a facial image of the user; and a data storage circuit configured to store the biosignal data of the user in response to the transferred presented sentence and the image data, together with a trigger value being recorded over time. However, in a similar field of endeavor, Chi discloses a presented sentence transfer display circuit configured to transfer a presented sentence to the user (para 41, “The second database 142 includes a plurality of example sentences for generating the candidate sentences”; also, para 47, “The output module 180 provides the candidate sentence or conversation sentence to the user in a form of voice or text through a speaker or a display under the control of the processor 130”); a biosignal collection circuit configured to collect biosignal data by measuring biosignals including brain waves of a user (para 44, “the sensor module 150 detects the brainwave related to the imagined speech in a state worn by the user”); a data storage circuit configured to store the biosignal data of the user in response to the transferred presented sentence (para 36, “In addition, the memory 120 performs a function of temporarily or permanently storing data processed by the processor 130.”; also, para 50, “The brain-computer interface system 100 measures the imagined speech brainwave, speech act brainwave, and audible act brainwave for the example word while the sound of the example word is output, measures the soundwave generated when the user speaks directly, and then converts the measured brainwaves and soundwaves into digital signals and stored the converted result in the memory (S120).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention of a multimodal data collection circuit configured to collect multimodal data including image data with the features of Lee's invention of a presented sentence transfer display circuit that transfers a presented sentence to the user, a biosignal collection circuit that collects brain waves of the user, and a memory that stores the measured brainwaves for the presented example word. The combination would have been obvious because Lee teaches presenting an example sentence to the user in order to elicit an imagined-speech brain wave and storing the resulting brainwave for that word, so incorporating that presentation, collection, and storage into Li's data collection circuit would yield the predictable result of a collection circuit that elicits and retains the user's imagined-speech biosignal. Chong discloses an image collection circuit configured to collect image data by photographing a facial image of the user (para 46, “The user terminal 110 may include a camera. Alternatively, the user terminal 110 may have stored a 2D image including a face.”; also, para 45, “An avatar indicates a 3D animation model that resembles a specific person within a two-dimensional (2D) image.”; also, para 49, “In response thereto, the server 130 extracts a face from the input 2D image and extracts a landmark from the extracted face. Thereafter, the server 130 generates a first mesh model using the landmark.”) the image data (para 46, “Alternatively, the user terminal 110 may have stored a 2D image including a face.”), together with a trigger value being recorded over time (para 15, “the trigger generator 200 is the unit responsible for delivering stimuli (e.g., flashing light) along with trigger signals to mark the time location of such events; and the receiving host 300 would be a PC or laptop that records both EEG data with trigger makers”; also, para 18, “The digital protocol allowed transmission of multi-bit trigger and synchronization codes, increasing the utility of the technique.”; also, para 19, “The low latency of the link 2 allows for the microprocessor 104 to combine the signal source 103 and the stimulus 201 with a high degree of temporal precision-comparable to traditional wireline methods. The aggregated, time synchronized data is transmitted from the data acquisition unit 100 to the receiving host 300 across the wireless link 1 using transceivers 101 and 301.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention as modified by Lee with the features of Chong's invention of an image collection circuit having a camera that captures the face of the user and a storage that holds the resulting two-dimensional image including the face. The combination would have been obvious because Chong teaches capturing the user's face with a camera and storing the resulting image so that the avatar generated from that image resembles the user, so collecting and storing the facial image alongside the biosignal already being collected would yield the predictable result of a collection circuit that retains both modalities of the multimodal data. Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (U.S. Pub. No. 20210248801) as modified by Lee et al. (U.S. Pub. No. 20200142481), Chong et al. (U.S. Pub. No. 20210005003), and Chi (U.S. Pub. No. 20150358096), further in view of Liu (U.S. Pub. No. 20240104183). Regarding claim 5, Li as modified by Lee, Chong, and Chi discloses the device according to claim 4, biosignal collection circuit further includes an electromyography in the measured biosignal of the user; and wherein the lip-sync reconstruction circuit is configured to predict the mouth shape and facial movement by inferring articulatory organ movement trajectories based on the electromyography. However, in a similar field of endeavor, Liu discloses wherein the biosignal collection circuit further includes an electromyography in the measured biosignal of the user (para 133, “In nonlimiting examples, the sensor 197 can be a surface electrode configured to be placed in physical contact with at least a portion of a user's body and configured to capture EMG electrical signals, EOG electrical signals, or a combination thereof.”; also, para 155, “FIG. 6 depicts example biosignals collected from a side of masseter around ears 600 of user. Part (a) of FIG. 6 depicts example EMG signals of facial activities.”); and wherein the lip-sync reconstruction circuit is configured to predict the mouth shape and facial movement by inferring articulatory organ movement trajectories based on the electromyography (para 166, “During the training of the biosignal network, the provided apparatus and methods take the aligned 2D facial landmarks from the vision network as ground truth and train a 1D CNN network to regress the facial landmarks directly from four channels of time-series biosignals (i.e., two EMG and two EOG streams).”; also, para 164, “To reduce computational complexity, the provided apparatus and methods can keep 53 landmarks that cover major facial components such as eyes, eyebrows, nose, and mouth.”; also, para 159, “Then the fine-tuned biosignal network can continuously reconstruct 2D facial landmarks from the biosignal stream, without any visual input. To ensure a fluent 3D avatar animation, the provided apparatus and methods can then apply Landmark Smoothing via Kalman Filter to stabilize the facial landmark movement across successive frames. Next, the provided apparatus and methods can generate 3D facial animation from the stabilized landmarks using a FLAME (Faces Learned with an Articulated Model and Expressions) model. The generated sequence of fitted head models can then be used for rendering a 3D facial animation that recovers the user's facial movements.”; also, para 173, “In the second stage, the provided apparatus and methods optimize the model parameters (e.g., pose, shape, and expression) by optimizing the L2 distance while regularizing the shape coefficients, pose coefficients (including neck, jaw, and eyeballs), and expression coefficients by penalizing their L2 norms.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention as modified by Lee, Chong, and Chi with the features of Liu's invention of a surface electrode that captures electromyography signals among the measured biosignals of the user and a biosignal network that regresses the facial landmarks, including the landmarks of the mouth, directly from the electromyography stream and stabilizes the facial landmark movement across successive frames to drive a three-dimensional facial animation. The combination would have been obvious because Liu teaches that the facial landmarks can be continuously reconstructed from the biosignal stream without any visual input and that the resulting landmark movement, together with the inferred jaw pose, generates a three-dimensional facial animation recovering the user's facial movements, so adding that electromyographic channel to the biosignals already being collected would supply the lip-sync reconstruction model with a direct measurement of the articulator muscles, yielding the predictable result of a mouth shape and facial movement predicted from the muscle activity of the user. Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (U.S. Pub. No. 20210248801) as modified by Lee et al. (U.S. Pub. No. 20200142481) and Chong et al. (U.S. Pub. No. 20210005003), further in view of Hori et al. (U.S. Pub. No. 20180189572). Regarding claim 6, Li as modified by Lee and Chong discloses the device according to claim 3, wherein the feature convergence circuit is configured to: weight, based on a predetermined standard, to the feature vectors extracted by the feature extraction circuit; and converge the feature vectors to which the weight has been applied and converts them into an embedding convergence vector. However, in a similar field of endeavor, Hori discloses apply a weight, based on a predetermined standard, to the feature vectors extracted by the feature extraction circuit (para 40, “Given multimodal video data including K modalities such that K≥2 and some of the modalities may be the same, Modal-1 data are converted to a fixed-dimensional content vector using the feature extractor 211, the attention estimator 212 and the weighted-sum processor 213 for the data, where the feature extractor 211 extracts multiple feature vectors from the data, the attention estimator 212 estimates each weight for each extracted feature vector, and the weighted-sum processor 213 outputs (generates) the content vector computed as a weighted sum of the extracted feature vectors with the estimated weights.”); and converge the feature vectors to which the weight has been applied and converts them into an embedding convergence vector (para 42, “The K transformed N-dimensional vectors are summed into a single N-dimensional content vector in the simple multimodal method of FIG. 2A, whereas the vectors are converted to a single N-dimensional content vector using the modal attention estimator 255 and the weighted-sum processor 245 in the multimodal attention method of FIG. 2B, wherein the modal attention estimator 255 estimates each weight for each transformed N-dimensional vector, and the weighted-sum processor 245 outputs (generates) the N-dimensional content vector computed as a weighted sum of the K transformed N-dimensional vectors with the estimated weights.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Li's invention as modified by Lee and Chong with the features of Hori's invention of an attention estimator that estimates a weight for each extracted feature vector and a weighted-sum processor that converges the weighted feature vectors into a single content vector. The combination would have been obvious because Hori teaches that the content vector is computed as a weighted sum of the extracted feature vectors using the weights estimated for each vector, and applying that weighting to the biosignal feature vector and the facial feature vector already being converged in the combination would yield the predictable result of an embedding convergence vector in which each modality contributes in proportion to its estimated weight. Conclusion The following prior art made of record and not relied upon is considered pertinent to applicant’s disclosure: Kapur et al. (U.S. Pub. No. 20190074012), Jorgensen et al. (U.S. Doc. No. 7574357), Pradeep et al. (U.S. Doc. No. 8655428). Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jai Li whose telephone number is (571)272-1170. The examiner can normally be reached Mon-Thu between 06:00-16:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571)272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAI W LI/ Junior Examiner, Art Unit 2613 /DAVID T WELCH/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Dec 30, 2024
Application Filed
Jul 20, 2026
Non-Final Rejection mailed — §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month