Prosecution Insights
Last updated: October 02, 2026
Application No. 19/019,061

ARTIFICIAL INTELLIGENCE DEVICE FOR EMOTION RECOGNITION USING JOINT AND UNIMODAL REPRESENTATION LEARNING WITH LATE FUSION AND METHOD THEREOF

Non-Final OA §103§112
Filed
Jan 13, 2025
Priority
Jan 12, 2024 — provisional 63/620,184
Examiner
SULTANA, NADIRA
Art Unit
Tech Center
Assignee
LG Electronics Inc.
OA Round
1 (Non-Final)
73%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
80 granted / 110 resolved
+12.7% vs TC avg
Strong +36% interview lift
Without
With
+35.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
22 currently pending
Career history
135
Total Applications
across all art units

Statute-Specific Performance

§101
26.1%
-13.9% vs TC avg
§103
58.8%
+18.8% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
3.3%
-36.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 110 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 2. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 3, 4, 14, 15 are rejected under 35 U.S.C. 112, second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention. Claim 3 recites in line 3, “ a second weight multiplied to the joint visual embedding vector” , in line 6 recites, “ a second weight multiplied to the joint audio embedding vector” . There is insufficient antecedent basis for this limitations in the claim. Also it’s not clear if the joint visual embedding vector and the joint audio embedding vector are multiplied by same second weight or not. Claim 3 also recites in line 5,” generating the second partially fused visual embedding vector based on a sum of a third…..joint audio embedding vector”. In claim 2, “a second partially fused audio embedding vector” was defined, no “second partially fused visual embedding vector” was defined. It’s not clear if there is a typo. If it’s not typo, there is insufficient antecedent basis for this limitations in the claim. For examination purpose, ” generating the second partially fused visual embedding vector” is interpreted as ” generating the second partially fused audio embedding vector”. Depended claim 4 is rejected based on it’s dependency on rejected claim 3. Claim 14 recites in line 3, “ a second weight multiplied to the joint visual embedding vector” , in line 6 recites, “ a second weight multiplied to the joint audio embedding vector” . There is insufficient antecedent basis for this limitations in the claim. Also it’s not clear if the joint visual embedding vector and the joint audio embedding vector are multiplied by same second weight or not. Claim 14 also recites in line 5,” generating the second partially fused visual embedding vector based on a sum of a third…..joint audio embedding vector”. In claim 13, “a second partially fused audio embedding vector” was defined, no “second partially fused visual embedding vector” was defined. It’s not clear if there is a typo. If it’s not typo, there is insufficient antecedent basis for this limitations in the claim. For examination purpose, ” generating the second partially fused visual embedding vector” is interpreted as ” generating the second partially fused audio embedding vector”. Depended claim 15 is rejected based on it’s dependency on rejected claim 14. Applicant is requested to amend the claims to clarify the claimed invention. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7, 10, and 12-18 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. ( Transformer-Based Multimodal Emotional Perception for Dynamic Facial Expression Recognition in the Wild, IEEE, 2023), hereinafter referenced as Zhang, in view of Wasnik et al. (US 20230154172 A1), hereinafter referenced as Wasnik. Regarding Claim 1, Zhang teaches a method for controlling an artificial intelligence (Al) device, the method comprising: receiving, [by a processor in the Al device,] a video segment including a plurality of frames and an audio signal corresponding to the video segment ( Zhang: Section IV. A. Para.[1]-[3], receiving dataset consists of audio and vision modalities of video clips and includes a diverse range of sources, such as movies, TV shows, news, reality shows); processing the video segment, by a visual encoder, to generate a visual embedding ( Zhang: Section III. B. 2) Vision Encoder: Fig. 2, Vision Encoder takes a video clip as input, generate visual embedding ( fv)); processing the audio signal, by an audio encoder, to generate an audio embedding ( Zhang: Section III. B. 1) Audio Encoder: Fig. 2, Audio Encoder takes an audio waveform a of duration t seconds as input, generate audio embedding(fa)); processing the visual embedding, by a unimodal visual transformer encoder, to generate an independent visual embedding vector based on self-attention ( Zhang: Section III. C.1) Bimodal T-MIF: Figs. 2, 3, Eq. 2 illustrates processing of visual feature, independently fed into the vision transformer encoder β, by self-attention, vector f’v = Self-Att(fv)); processing the audio embedding, by a unimodal audio transformer encoder, to generate an independent audio embedding vector based on self-attention ( Zhang: Section III. C.1) Bimodal T-MIF: Figs. 2, 3, Eq. 2 illustrates processing of audio feature , independently fed into the audio transformer encoder α, by self-attention, vector f’a = Self-Att(fa)); processing the visual embedding, by a joint visual transformer encoder, to generate a joint visual embedding vector using cross-attention based on information from the audio embedding ( Zhang: Section III. C.1) Bimodal T-MIF: Figs. 2, 3, Eq. 3 illustrates cross-modal interaction using cross-attention. Visual feature vector f’v, audio feature vector f’a, independently fed into separate transformer blocks and cross-modal interaction is enabled by using cross-attention, to generate joint visual embedding f”v = Co-Att(f’v, f’a )); processing the audio embedding, by a joint audio transformer encoder, to generate a joint audio embedding vector using cross-attention based on information from the video embedding ( Zhang: Section III. C.1) Bimodal T-MIF: Figs. 2, 3, Eq. 3 illustrates cross-modal interaction using cross-attention. Audio feature vector f’a ,visual feature vector f’v, independently fed into separate transformer blocks and cross-modal interaction is enabled by using cross-attention, to generate joint audio embedding f”a = Co-Att(f’a, f’v )); processing the independent visual embedding vector, the independent audio embedding vector, the joint visual embedding vector and the joint audio embedding vector, by a fusion module, to generate a global fused embedding vector (Zhang: Section III. C.1) Bimodal T-MIF: Fig. 3, Eq. 4 illustrates further cross-modal interaction using feed forward, f”’a = FFN(f”a + f’a ), f”’v = FFN(f”v + f’v). Section III. A. Framework Overview: Fig. 2, the modality-specific features are input into a multimodal fusion module to generate cross-modal representations); and generating an emotion prediction using a classifier module that analyzes the global fused embedding vector, and outputting the emotion prediction (Zhang: Section III. A. Framework Overview: Fig. 2, the modality-specific features are input into a multimodal fusion module to generate cross-modal representations, which are further fed into a classifier to produce the final outputs. Section III. C.1) Bimodal T-MIF: Fig. 3, Eqs.2- 4, by repeatedly applying Eqs.(2) - (4), M times, the T-MIF approach generate rich emotion-relevant representations by mutually complementing audio and visual features). Zhang while teaching the method of claim 1, fails to explicitly teach the claimed, receiving, by a processor in the Al device, a video segment including a plurality of frames and an audio signal corresponding to the video segment. However, Wasnik does teach the claimed, receiving, by a processor in the Al device, a video segment including a plurality of frames and an audio signal corresponding to the video segment ( Wasnik: Para.[0020]-[0022],Fig. 1 illustrates a network environment 100, which includes a system 102, for emotion recognition in multimedia videos using multi-modal fusion based-deep neural network diagram. The system 102 includes circuitry 104, memory 106. The circuitry 104 may be implemented based on a number of processor technologies. The memory include a multimodal fusion network 108. The circuitry 104 may be configured to receive and input the multimodal input 124 ( acoustics with video clips)) ; Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wasnik’s teaching of a system and method for emotion recognition in multimedia videos using multi-modal fusion-based deep neural network, into the system and method of transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild, taught by Zhang, because, this would improve the quality and robustness to detect the emotional state of subject by using different modalities (Wasnik, Para.[0018]). Claim 12 is an artificial intelligence (AI) device claim, comprising: a memory configured to store video and audio information; and a controller configured to ( Wasnik: Para.[0021], [0022], Fig. 1, memory 106 may store input data ( video and audio information ) for the multimodal fusion network 108. The circuitry 104 may be implemented based on a number of processor technologies), perform the steps in method claim 1 above and as such, claim 12 is similar in scope and content to claim 1 and therefore, claim 12 is rejected under similar rationale as presented against claim 1 above. Regarding Claim 2, Zhang in view of Wasnik teach the method of claim 1. Zhang further teaches, wherein the processing the independent visual embedding vector, the independent audio embedding vector, the joint visual embedding vector and the joint audio embedding vector, by the fusion module includes: inputting the independent visual embedding vector and the joint visual embedding vector to a first partial fusion module to generate a first partially fused visual embedding vector (Zhang: Section III. C.1) Bimodal T-MIF: Fig. 3, Eq. 4 illustrates further cross-modal interaction using feed forward ( first partial fusion module), f”’v = FFN(f”v + f’v) (first partially fused visual embedding vector)); inputting the independent audio embedding vector and the joint audio embedding vector to a second partial fusion module to generate a second partially fused audio embedding vector(Zhang: Section III. C.1) Bimodal T-MIF: Fig. 3, Eq. 4 illustrates further cross-modal interaction using feed forward( second partial fusion module), f”’a = FFN(f”a + f’a ) (second partially fused audio embedding vector)); and processing the first partially fused visual embedding vector and the second partially fused audio embedding vector, by a global fusion module, to generate the global fused embedding vector ( Zhang: Section III. A. Framework Overview: Fig. 2, the modality-specific features (first partially fused visual embedding vector and the second partially fused audio embedding vector) are input into a multimodal fusion module to generate cross-modal representations ( global fused embedding vector)). Claim 13 is an artificial intelligence (AI) device claim performing the steps in method claim 2 above and as such, claim 13 is similar in scope and content to claim 2 and therefore, claim 13 is rejected under similar rationale as presented against claim 2 above. Regarding Claim 3, Zhang in view of Wasnik teach the method of claim 2. Zhang further teaches, further comprising: generating the first partially fused visual embedding vector based on a sum of a first weight multiplied to the independent visual embedding vector and a second weight multiplied to the joint visual embedding vector ( Zhang: Section III, C. 1): a residual structure is incorporated to enhance the emotion representation in Eq. 4, f”’v = FFN(f”v + f’v), which is first partially fused visual embedding vector, based on the sum of f”v (joint visual embedding vector) and f’v (independent visual embedding vector)); and generating the second partially fused visual embedding vector based on a sum of a third weight multiplied to the independent audio embedding vector and a second weight multiplied to the joint audio embedding vector ( Zhang: Section III, C. 1): a residual structure is incorporated to enhance the emotion representation in Eq. 4, f”’a = FFN(f”a + f’a ), which is first partially fused audio embedding vector, based on the sum of f”a (joint audio embedding vector) and f’a (independent audio embedding vector)); Claim 14 is an artificial intelligence (AI) device claim performing the steps in method claim 3 above and as such, claim 14 is similar in scope and content to claim 3 and therefore, claim 14 is rejected under similar rationale as presented against claim 3 above. Regarding Claim 4, Zhang in view of Wasnik teach the method of claim 3. Wasnik further teaches, further comprising: generating the global fused embedding vector based on a sum of a fifth weight multiplied to the first partially fused visual embedding vector and a sixth weight multiplied to the second partially fused visual embedding vector ( Wasnik: Para.[0080]-[0083], Figs. 3, 4, the system 402 includes the multimodal fusion network 302, which may include the one or more feature extractors 304. The system 402 may generate the third embedding of the input embeddings as the output of the acoustic-visual feature extractor 304C or the visual feature extractor 304D, which is a weighted sum based on features associated with each of the detected one or more faces and the corresponding normalized areas. The weighted sum may be mathematically represented using below equation which is given as follows: FIV = F1 W1 + F2W2, where, FIV represents the third embedding of the input embeddings; F 1 represents the features associated with the detected first face 406, F 2 represents the features associated with the detected second face 408). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wasnik’s teaching of a system and method for emotion recognition in multimedia videos using multi-modal fusion-based deep neural network, into the system and method of transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild, taught by Zhang, because, this would improve the quality and robustness to detect the emotional state of subject by using different modalities (Wasnik, Para.[0018]). Claim 15 is an artificial intelligence (AI) device claim performing the steps in method claim 4 above and as such, claim 15 is similar in scope and content to claim 4 and therefore, claim 15 is rejected under similar rationale as presented against claim 4 above. Regarding Claim 5, Zhang in view of Wasnik teach the method of claim 1. Zhang further teaches, wherein the processing the visual embedding, by the joint visual transformer encoder, to generate the joint visual embedding vector using cross-attention includes using a query and a key received from the joint audio transformer encoder based on the audio embedding and a value internally generated by the joint visual transformer encoder based on the visual embedding ( Zhang: Section III. C.1) Bimodal T-MIF: Fig. 3, Eq. 3 illustrates cross-modal interaction using cross-attention, f”v = Co-Att(f’v, f’a ), to generate joint visual embedding. Here, the vector f’v serves as the query, while f’a serves as both the key and the value. The fundamental idea behind this operation is that f’a is supervised to organize and compress relevant information within the modality, and then share it with f’v to achieve semantic alignment from the audio to the vision modality, with f’v acting as the limiting condition). Claim 16 is an artificial intelligence (AI) device claim performing the steps in method claim 5 above and as such, claim 16 is similar in scope and content to claim 5 and therefore, claim 16 is rejected under similar rationale as presented against claim 5 above. Regarding Claim 6, Zhang in view of Wasnik teach the method of claim 1. Zhang further teaches, wherein the processing the audio embedding, by the joint audio transformer encoder, to generate the joint audio embedding vector using cross-attention includes using a query and a key received from the joint visual transformer encoder based on the visual embedding and a value internally generated by the joint audio transformer encoder based on the audio embedding ( Zhang: Section III. C.1) Bimodal T-MIF: Fig. 3, Eq. 3 illustrates cross-modal interaction using cross-attention, f”a = Co-Att(f’a, f’v ), to generate joint audio embedding. Here, the vector f’a serves as the query, while f’v serves as both the key and the value. The fundamental idea behind this operation is that f’v is supervised to organize and compress relevant information within the modality, and then share it with f’a to achieve semantic alignment from the vision to the audio modality, with f’a acting as the limiting condition). Claim 17 is an artificial intelligence (AI) device claim performing the steps in method claim 6 above and as such, claim 17 is similar in scope and content to claim 6 and therefore, claim 17 is rejected under similar rationale as presented against claim 6 above. Regarding Claim 7, Zhang in view of Wasnik teach the method of claim 1. Zhang further teaches, wherein the processing the video segment, by the visual encoder, to generate the visual embedding includes: resizing the plurality of frames within the video segment to a size of NxN pixels to generate resized frames where N is a positive number ( Zhang: Section IV, B. 2) Training setting: for vision encoder, video clips are adjusted to 112X 112); dividing the resized frames into MxM non-overlapping patches where M is less than N ( Zhang: Section IV, B. 2) Training setting: cropping the frames to 4X4); applying a first convolutional neural network (CNN) to project the MxM non-overlapping patches into fixed-sized visual vectors, the visual embedding being one of the fixed-sized visual vectors ( Zhang: Section III, B. 2) Vision encoder: Fig. 2, CNNs are used to process the visual vectors). Claim 18 is an artificial intelligence (AI) device claim performing the steps in method claim 7 above and as such, claim 18 is similar in scope and content to claim 7 and therefore, claim 18 is rejected under similar rationale as presented against claim 7 above. Regarding Claim 10, Zhang in view of Wasnik teach the method of claim 1. Zhang further teaches, wherein one or more of the unimodal visual transformer encoder, the unimodal audio transformer encoder, the joint visual transformer encoder and the joint audio transformer encoder includes a transformer tower including a plurality of transformer blocks ( Zhang: Section III, C. 1): Fig. 3 illustrates a bimodal transformer-based multimodal information fusion (T-MIF) module, which shows multiple transformer blocks) . Claims 8, 9, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. ( Transformer-Based Multimodal Emotional Perception for Dynamic Facial Expression Recognition in the Wild, IEEE, 2023), hereinafter referenced as Zhang, in view of Wasnik et al. (US 20230154172 A1), hereinafter referenced as Wasnik, further in view of Le et al. (Multi-Label Multimodal Emotion Recognition With Transformer-Based Fusion and Emotion-Level Representation Learning, IEEE Access, 2023), hereinafter referenced as Le. Regarding Claim 8, Zhang in view of Wasnik teach the method of claim 1. Zhang in view of Wasnik fail to explicitly teach the claimed, wherein the processing the audio signal, by the audio encoder, to generate the audio embedding includes: converting the audio signal to mel-spectrograms; and applying a second convolutional neural network (CNN) to extract fixed-sized audio based vectors, the audio embedding being one of the fixed-sized audio based vectors. However, Le does tech the claimed, wherein the processing the audio signal, by the audio encoder, to generate the audio embedding includes: converting the audio signal to mel-spectrograms ( Le: Section III. A, Fig. 1, audio signals are converted to Mel spectrogram); and applying a second convolutional neural network (CNN) to extract fixed-sized audio based vectors, the audio embedding being one of the fixed-sized audio based vectors ( Le: Section III. A, Fig. 1, CNN networks are used to process the audio spectrograms. Feature vectors of size d captured from audio spectrogram chunks). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Le’s teaching of multi-label multimodal emotion recognition with transformer-based fusion and emotion-level representation learning, into the system and method, taught by Zhang in view of Wasnik, because, by using multi-head attention in the transformers to concurrently fuse the multimodal features of video data at multiple layers, would improve the multimodal emotion recognition process. (Le, section V. Para.[1]). Claim 19 is an artificial intelligence (AI) device claim performing the steps in method claim 8 above and as such, claim 19 is similar in scope and content to claim 8 and therefore, claim 19 is rejected under similar rationale as presented against claim 8 above. Regarding Claim 9, Zhang in view of Wasnik teach the method of claim 1. Zhang in view of Wasnik fail to explicitly teach the claimed, further teaches, wherein each of the unimodal visual transformer encoder, the unimodal audio transformer encoder, the joint visual transformer encoder and the joint audio transformer encoder includes only a single transformer block. However, Le does tech the claimed, further teaches, wherein each of the unimodal visual transformer encoder, the unimodal audio transformer encoder, the joint visual transformer encoder and the joint audio transformer encoder includes only a single transformer block ( Le: Section III, Para.[2], Fig. 1 illustrates the overall architecture of multi-label multimodal emotion recognition taking a video clip as an input and outputting the revealed emotions in the video. The model consists of three modules: the feature extraction module, the multimodal fusion module, and the emotion-level embedding module. The transformer based multimodal fusion module can be seen from fig. 1 as a single transformer block). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Le’s teaching of multi-label multimodal emotion recognition with transformer-based fusion and emotion-level representation learning, into the system and method, taught by Zhang in view of Wasnik, because, by using multi-head attention in the transformers to concurrently fuse the multimodal features of video data at multiple layers, would improve the multimodal emotion recognition process. (Le, section V. Para.[1]). Claim 20 is an artificial intelligence (AI) device claim performing the steps in method claim 9 above and as such, claim 20 is similar in scope and content to claim 9 and therefore, claim 20 is rejected under similar rationale as presented against claim 9 above. Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. ( Transformer-Based Multimodal Emotional Perception for Dynamic Facial Expression Recognition in the Wild, IEEE, 2023), hereinafter referenced as Zhang, in view of Wasnik et al. (US 20230154172 A1), hereinafter referenced as Wasnik, further in view of Dontcheva et al. ( US 12,367,238 B2), hereinafter referenced as Dontcheva. Regarding Claim 11, Zhang in view of Wasnik teach the method of claim 1. Wasnik further teaches, wherein the generating the emotion prediction using the classifier module includes: mapping the global fused embedding to a set of probabilities corresponding to a plurality of emotions ( Wasnik: Para.[0090],[0091], Figs. 1, 3, The circuitry 104 may be configured to apply the one or more multi-head attention layers on the set of emotion-relevant features to determine an inter-feature mapping within the set of emotion-relevant features and the circuitry 104 may be configured to generate the fused-feature representation of the set of emotion-relevant features. With the mapping, each respective modality of the plurality of modalities may be mapped to a text vector space. The circuitry 104 may be configured to concatenate the set of emotion-relevant features into a latent representation of the set of emotion-relevant features, based on the inter-feature mapping ); Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Wasnik’s teaching of a system and method for emotion recognition in multimedia videos using multi-modal fusion-based deep neural network, into the system and method of transformer-based multimodal emotional perception for dynamic facial expression recognition in the wild, taught by Zhang, because, this would improve the quality and robustness to detect the emotional state of subject by using different modalities (Wasnik, Para.[0018]). Zhang in view of Wasnik fail to explicitly teach the claimed, and selecting an emotion among the plurality of emotions having a highest probability as the emotion prediction. However, Dontcheva does tech the claimed, and selecting an emotion among the plurality of emotions having a highest probability as the emotion prediction ( Dontcheva: Column 24, lines 38-43, each of the candidate images is assessed for a particular emotion, such as happiness (e.g., by applying each image to an emotion classifier to generate a measure of detected emotion, such as a happiness score), and the image with the highest predicted emotion score is selected as the representative image for that speaker). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Dontcheva’s teaching of a visual and text search interface used to navigate a video transcript, into the system and method, taught by Zhang in view of Wasnik, because, by using audio-visual speaker diarization technique, the user’s experience can be improved. (Dontcheva, Column 2, lines 6-46). Conclusion Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant's disclosure. Zhu et al. (US 12,367,240 B1) teaches systems, methods, and computer-readable media for systems and methods multimodal indexing of video using machine learning. An example method may include deceiving, by a video encoder of an audio-video transformer neural network comprising one or more computer processors coupled to memory, a first frame and a second frame associated with a first segment of a video. The example method may also include receiving, by an audio encoder of the audio-video transformer neural network, an audio spectrogram comprising first audio data associated with the first segment of the video. generating, by the video encoder, a first video embedding. The example method may also include generating, by the audio encoder, a first audio embedding. The example method may also include determining a fusion of the first video embedding and the first audio embedding using a multimodal bottleneck token. The example method may also include determining an output including the first video embedding and the first audio embedding. The example method may also include determining a classification of the first portion of the video based on the output. Yang et al. (US 20240380949 A1) teaches a system and a method that include a processor executing a caption generation program to receive an input video, sample video frames from the input video, extract video frames from the input video, extract video embeddings and audio embeddings from the video frames, including local video tokens and local audio tokens, respectively, input the local video tokens and the local audio tokens into at least a transformer layer of a cross-modal encoder to generate multi-modal embeddings, and generate video captions based on the multi-modal embeddings using a caption decoder. Jeong et al. (US 20220207262 A1) teaches a mouth shape synthesis device and a method using an artificial neural network. To this end, an original video encoder that encodes original video data which is a target of a mouth shape synthesis as a video including a face of a synthesis target; an audio encoder that encodes audio data that is a basis for the mouth shape synthesis and outputs an audio embedding vector; and a synthesized video decoder that uses the original video embedding vector and the audio embedding vector as input data, and outputs synthesized video data in which a mouth shape corresponding to the audio data is synthesized on the synthesis target face may be provided. Qin et al. (TMFER: Multimodal Fusion Emotion Recognition Algorithm Based on Transformer, IEEE, 2023 ) teaches a multimodal fusion emotion recognition algorithm based on Transformer (TMFER), which fuses three modalities of text, speech and image information for emotion recognition. For the different characteristics of each modal information, Bert model pre training processing, MFCC feature extraction and CNN feature extractor extraction methods are used to extract features for each modality respectively, to explore deeper features. To address the problem of unreasonable combination of multi-modal features, the Transformer Encode multi-headed attention mechanism is used to build a feature fusion module to extract and combine potential feature information in different modalities in parallel. The fused data are fed into the algorithm classification module for sentiment recognition classification, and a joint supervised loss function based on large margin learning is customized to solve the problem of unbalanced classification and feature confounding in the baseline model. Finally, based on the IEMOCAP and MELD multimodal datasets, the TMFER algorithm is experimentally compared with current algorithms in the field that are more effective in emotion recognition classification. The experimental results show that the TMFER algorithm outperforms other algorithms in all evaluation metrics. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NADIRA SULTANA whose telephone number is (571)272-4048. The examiner can normally be reached M-F,7:30 am-5:00pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached on (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NADIRA SULTANA/Examiner, Art Unit 2653
Read full office action

Prosecution Timeline

Jan 13, 2025
Application Filed
Aug 25, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737534
SYSTEMS AND METHODS FOR LANGUAGE MODEL-BASED TEXT EDITING
2y 5m to grant Granted Sep 15, 2026
Patent 12731599
ADAPTIVE NOISE ESTIMATION
3y 6m to grant Granted Sep 08, 2026
Patent 12726544
COORDINATING A CONVERSATIONAL AGENT WITH A LARGE LANGUAGE MODEL FOR CONVERSATION REPAIR
2y 7m to grant Granted Sep 01, 2026
Patent 12681967
SYSTEM AND METHOD FOR OPTIMIZING QUERY RESOLUTION ON DIGITAL CHANNELS IN A CONTACT CENTER
2y 3m to grant Granted Jul 14, 2026
Patent 12676157
THREE-DIMENSIONAL AUDIO SIGNAL CODING METHOD AND APPARATUS, AND ENCODER
2y 7m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
73%
Grant Probability
99%
With Interview (+35.7%)
2y 11m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 110 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month