Prosecution Insights
Last updated: October 04, 2026
Application No. 19/093,146

EMOTION CLASSIFICATION METHOD AND APPARATUS

Non-Final OA §101§102
Filed
Mar 27, 2025
Priority
Dec 06, 2024 — CN 202411793531.0
Examiner
KIM, ETHAN DANIEL
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Mashang Consumer Finance Co. Ltd.
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
88 granted / 114 resolved
+15.2% vs TC avg
Strong +25% interview lift
Without
With
+25.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
12 currently pending
Career history
131
Total Applications
across all art units

Statute-Specific Performance

§101
10.0%
-30.0% vs TC avg
§103
49.6%
+9.6% vs TC avg
§102
35.2%
-4.8% vs TC avg
§112
1.4%
-38.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 114 resolved cases

Office Action

§101 §102
Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 2. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 3. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The independent claim 1 recites “A method for training an emotion classification model, the method comprising: obtaining a training object and an emotion classification model; adding an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model; obtaining a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtaining a transformed text feature vector of the sample of the training object based on the transformation layer; performing classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model; obtaining an adjustment based on an emotion classification result of the sample of the training object; and updating the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment”. The limitations “obtaining a training object and an emotion classification model; adding an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model; obtaining a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtaining a transformed text feature vector of the sample of the training object based on the transformation layer; performing classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model; obtaining an adjustment based on an emotion classification result of the sample of the training object; and updating the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper. This judicial exception is not integrated into a practical application. Claim 1 recites “A method for training an emotion classification model, the method comprising:…”; this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regard to claim 1 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 1 does not recite any additional limitations. The claim as drafted, is not patent eligible. The independent claim 9 recites “An apparatus for training an emotion classification model, the apparatus comprising: processing circuitry configured to obtain a training object and an emotion classification model; add an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model; obtain a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtain a transformed text feature vector of the sample of the training object based on the transformation layer; perform classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model; obtain an adjustment based on an emotion classification result of the sample of the training object; and update the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment”. The limitations “processing circuitry configured to obtain a training object and an emotion classification model; add an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model; obtain a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtain a transformed text feature vector of the sample of the training object based on the transformation layer; perform classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model; obtain an adjustment based on an emotion classification result of the sample of the training object; and update the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper. This judicial exception is not integrated into a practical application. Claim 9 “An apparatus for training an emotion classification model, the apparatus comprising:…”; this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regard to claim 9 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 9 does not recite any additional limitations. The claim as drafted, is not patent eligible. The independent claim 17 recites “A non-transitory computer-readable storage medium, storing instructions which when executed by a processor cause the processor to perform: obtaining a training object and an emotion classification model; adding an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model; obtaining a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtaining a transformed text feature vector of the sample of the training object based on the transformation layer; performing classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model; obtaining an adjustment based on an emotion classification result of the sample of the training object; and updating the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment”. The limitations “obtaining a training object and an emotion classification model; adding an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model; obtaining a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtaining a transformed text feature vector of the sample of the training object based on the transformation layer; performing classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model; obtaining an adjustment based on an emotion classification result of the sample of the training object; and updating the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment” as drafted, covers a mental process, as this could be done by mentally or by hand with pen and paper. This judicial exception is not integrated into a practical application. Claim 17 recites “A non-transitory computer-readable storage medium, storing instructions which when executed by a processor cause the processor to perform:…”; this limitation directs towards using a computer for the method, and does not impose any meaningful limits on practicing the abstract idea. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The addition of the generic computer components recited above with regard to claim 17 do not amount to more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 17 does not recite any additional limitations. The claim as drafted, is not patent eligible. Claim Rejections - 35 USC § 102 4. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. 5. Claims 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Wu (U.S. Publication No. 20200320116). Regarding claim 1, Wu discloses a method for training an emotion classification model, the method comprising: obtaining a training object and an emotion classification model ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); adding an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model ([0047] - The messages may be in various multimedia forms, such as, text, voice, image, video, etc. [0059] - The module set 270 may comprise an emotion information extraction module 274. The emotion information extraction module 274 may be configured for extracting emotion information in a multimedia document through emotion analysis. The emotion information extraction module 274 may be implemented based on, such as, a recurrent neural network (RNN) together with a SoftMax layer. The emotion information may be represented, e.g., in a form of vector, and may be further used for deriving an emotion category); obtaining a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtaining a transformed text feature vector of the sample of the training object based on the transformation layer ([0083] - Data in a CKG may be in forms of text, image, voice and video. In some implementations, image, voice and video documents may be transformed into textual form and stored in the CKG to be used for CKG mining. For example, as for image documents, images may be transformed into texts by using image-to-text transforming model, such as a recurrent neural network-convolutional neural network (RNN-CNN) model. As for voice documents, voices may be transformed into texts by using voice recognition models. As for video documents which comprise voice parts and image parts, the voice parts may be transformed into texts by using the voice recognition models, and the image parts may be transformed into texts by using 2D/3D convolutional networks to project video information into dense space vectors and further using the image-to-text models. Videos may be processed through attention-based encoding-decoding models to capture text descriptions for the videos [0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); performing classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); obtaining an adjustment based on an emotion classification result of the sample of the training object ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0103] - The SVM classifier 880 may make a secondary judgment to the candidate training data obtained based on the emotion lexicon 850. Through the operation of the SVM classifier 880, those sentences having a relatively high confidence probability in the candidate training data may be finally appended to a training dataset 890. The training dataset 890 may be used for training the emotion classifying models); and updating the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0103] - The SVM classifier 880 may make a secondary judgment to the candidate training data obtained based on the emotion lexicon 850. Through the operation of the SVM classifier 880, those sentences having a relatively high confidence probability in the candidate training data may be finally appended to a training dataset 890. The training dataset 890 may be used for training the emotion classifying models [0150] - Layer 2 is a bi-directional RNN layer for performing recurrent operations among words of sentences corresponding to the same topic. The purpose of Layer 2 is to convert a whole sentence set into a vector. A vector h.sub.t+1 in Layer 2 may be computed by linearly combining h.sub.t and x.sub.t and attaching an element-wise non-linear transformation function such as RNN(.). Although RNN(.) is adopted here, it should be appreciated that the element-wise non-linear transformation function may also adopt, e.g., tanh or sigmoid, or a LSTM/GRU computing block). Regarding claim 2, Wu discloses the method, wherein the transformation layer comprises a first attention mechanism and a second attention mechanism ([0113] - The softmax layer may be configured for different emotion classifying strategies. In a first strategy, emotion categories may be defined based on 32 emotions in the wheel of emotions 700, including 8 basic emotions with “middle” strength, 8 weak emotions, 8 strong emotions and 8 combined emotions. The softmax layer may be a full connection layer which outputs an emotion vector corresponding to the 32 emotion categories. In a second strategy, emotion categories may be defined based on a combination of emotion and strength. For example, according to the wheel of emotions 700, 8 basic emotions and 8 combined emotions may be defined, wherein each of the 8 basic emotions is further defined with a strength level, while the 8 combined emotions are not defined with any strength level. In this case, the softmax layer may be a full connection layer which outputs an emotion vector corresponding to the 8 basic emotions, strength levels of the 8 basic emotions, and the 8 combined emotions. For both the first and second strategies, the emotion vector output by the softmax layer may be construed as emotion information of the input text sentence [0157] - exemplary neural network language model (NNLM) 1110 according to an embodiment. The NNLM 1110 may be used for generating a summary of a text document. The NNLM 1110 adopts encoder-decoder structure, which includes an attention-based encoder); the transformed audio feature vector is obtained based on the first attention mechanism ([0083] - Data in a CKG may be in forms of text, image, voice and video. In some implementations, image, voice and video documents may be transformed into textual form and stored in the CKG to be used for CKG mining. For example, as for image documents, images may be transformed into texts by using image-to-text transforming model, such as a recurrent neural network-convolutional neural network (RNN-CNN) model. As for voice documents, voices may be transformed into texts by using voice recognition models. As for video documents which comprise voice parts and image parts, the voice parts may be transformed into texts by using the voice recognition models, and the image parts may be transformed into texts by using 2D/3D convolutional networks to project video information into dense space vectors and further using the image-to-text models. Videos may be processed through attention-based encoding-decoding models to capture text descriptions for the videos [0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0157] - exemplary neural network language model (NNLM) 1110 according to an embodiment. The NNLM 1110 may be used for generating a summary of a text document. The NNLM 1110 adopts encoder-decoder structure, which includes an attention-based encoder); and the transformed text feature vector is obtained based on the second attention mechanism ([0083] - Data in a CKG may be in forms of text, image, voice and video. In some implementations, image, voice and video documents may be transformed into textual form and stored in the CKG to be used for CKG mining. For example, as for image documents, images may be transformed into texts by using image-to-text transforming model, such as a recurrent neural network-convolutional neural network (RNN-CNN) model. As for voice documents, voices may be transformed into texts by using voice recognition models. As for video documents which comprise voice parts and image parts, the voice parts may be transformed into texts by using the voice recognition models, and the image parts may be transformed into texts by using 2D/3D convolutional networks to project video information into dense space vectors and further using the image-to-text models. Videos may be processed through attention-based encoding-decoding models to capture text descriptions for the videos [0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0157] - exemplary neural network language model (NNLM) 1110 according to an embodiment. The NNLM 1110 may be used for generating a summary of a text document. The NNLM 1110 adopts encoder-decoder structure, which includes an attention-based encoder). Regarding claim 3, Wu discloses the method, wherein the obtaining the transformed audio feature vector further comprises: concatenating the audio vector and the audio feature vector to obtain a first concatenated vector ([0151] - Considering that recurrent operations are performed in two directions, i.e., left-to-right and right-to-left, h.sub.T may be formed by a concatenation of vectors in the two directions, for example, h.sub.T=[h.sub.left-to-right, h.sub.right-to-left].sup.T); determining a first query vector, a first key vector, and a first value vector of the first attention mechanism based on the first concatenated vector ([0152] - Layer 3 is another bi-directional RNN layer for performing recurrent operations among m topics in the session. The purpose of Layer 3 is to obtain a dense vector representation of the whole session. The bi-directional RNN Layer 3 takes h.sub.T from Layer 2 as inputs. The output of Layer 3 may be denoted as h.sub.m.sup.2, where m is the number of sentence sets in the input session); and determining the transformed audio feature vector based on the first query vector, the first key vector, and the first value vector ([0150] - Layer 2 is a bi-directional RNN layer for performing recurrent operations among words of sentences corresponding to the same topic. The purpose of Layer 2 is to convert a whole sentence set into a vector. A vector h.sub.t+1 in Layer 2 may be computed by linearly combining h.sub.t and x.sub.t and attaching an element-wise non-linear transformation function such as RNN(.). Although RNN(.) is adopted here, it should be appreciated that the element-wise non-linear transformation function may also adopt, e.g., tanh or sigmoid, or a LSTM/GRU computing block). Regarding claim 4, Wu discloses the method, wherein the determining the first query vector, the first key vector, and the first value vector of the first attention mechanism further comprises: multiplying the audio feature vector by a first weight matrix as the first query vector ([0148] - exemplary structure 1000 for classifying lifecycles of topics in a conversation session according to an embodiment. The structure 1000 comprises a multiple-layer neural network, e.g., a four-layer neural network, wherein rectangles may represent vectors and arrows may represent functions, such as matrix-vector multiplication [0158] - word embedding matrix, each word with D dimensions, U ∈ R .sup.(CD)*H, V ∈ R.sup.V*H, W ∈ R .sup.V*H are weight matrices, and h is a hidden layer of size H. The black-box function unit enc is a contextual encoder which returns a vector of size H representing input and current output context. [0160] - The major part in the attention-based encoder may be to learn a soft alignment P between the input x and the output y. The soft alignment may be used for weighting the smoothed version of the input x when constructing a representation. For example, if the current context aligns well with the position i, then the words x.sub.i−Q, . . . , x.sub.i+Q are highly weighted by the encoder); multiplying the first concatenated vector by a second weight matrix as the first key vector ([0148] - exemplary structure 1000 for classifying lifecycles of topics in a conversation session according to an embodiment. The structure 1000 comprises a multiple-layer neural network, e.g., a four-layer neural network, wherein rectangles may represent vectors and arrows may represent functions, such as matrix-vector multiplication [0158] - word embedding matrix, each word with D dimensions, U ∈ R .sup.(CD)*H, V ∈ R.sup.V*H, W ∈ R .sup.V*H are weight matrices, and h is a hidden layer of size H. The black-box function unit enc is a contextual encoder which returns a vector of size H representing input and current output context. [0160] - The major part in the attention-based encoder may be to learn a soft alignment P between the input x and the output y. The soft alignment may be used for weighting the smoothed version of the input x when constructing a representation. For example, if the current context aligns well with the position i, then the words x.sub.i−Q, . . . , x.sub.i+Q are highly weighted by the encoder); and multiplying the first concatenated vector by a third weight matrix as the first value vector ([0148] - exemplary structure 1000 for classifying lifecycles of topics in a conversation session according to an embodiment. The structure 1000 comprises a multiple-layer neural network, e.g., a four-layer neural network, wherein rectangles may represent vectors and arrows may represent functions, such as matrix-vector multiplication [0158] - word embedding matrix, each word with D dimensions, U ∈ R .sup.(CD)*H, V ∈ R.sup.V*H, W ∈ R .sup.V*H are weight matrices, and h is a hidden layer of size H. The black-box function unit enc is a contextual encoder which returns a vector of size H representing input and current output context. [0160] - The major part in the attention-based encoder may be to learn a soft alignment P between the input x and the output y. The soft alignment may be used for weighting the smoothed version of the input x when constructing a representation. For example, if the current context aligns well with the position i, then the words x.sub.i−Q, . . . , x.sub.i+Q are highly weighted by the encoder). Regarding claim 5, Wu discloses the method, wherein the determining the transformed audio feature vector further comprises: obtaining a transposed first key vector based on the first key vector; multiplying the transposed first key vector by the first query vector to obtain a first result; dividing the first result by a square root of a dimension of the first key vector to obtain a second result normalizing the second result to obtain a normalized second result; and multiplying the normalized second result by the first value vector as the transformed audio feature vector. Regarding claim 6, Wu discloses the method, wherein the transformation layer further comprises a pre-trained transformer model ([0083] - transformed into textual form and stored in the CKG to be used for CKG mining. For example, as for image documents, images may be transformed into texts by using image-to-text transforming model, such as a recurrent neural network-convolutional neural network (RNN-CNN) model. As for voice documents, voices may be transformed into texts by using voice recognition models. As for video documents which comprise voice parts and image parts, the voice parts may be transformed into texts by using the voice recognition models, and the image parts may be transformed into texts by using 2D/3D convolutional networks to project video information into dense space vectors and further using the image-to-text models. Videos may be processed through attention-based encoding-decoding models to capture text descriptions for the videos); and the updating further comprises: inputting the emotion classification result and a pre-added label into a loss function, to obtain the adjustment object ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0103] - The SVM classifier 880 may make a secondary judgment to the candidate training data obtained based on the emotion lexicon 850. Through the operation of the SVM classifier 880, those sentences having a relatively high confidence probability in the candidate training data may be finally appended to a training dataset 890. The training dataset 890 may be used for training the emotion classifying models [0153] - Process in this layer may comprise firstly computing y which is a linear function of h.sub.m.sup.2, and then using a softmax function to project y into a probability space and ensure P=[p.sub.1, p.sub.2, . . . , p.sub.|P|].sup.T follows a definition of probability. For error back-propagation, cross-entropy loss which corresponds to a minus log function of P for each topic word may be applied); and keeping parameters of the pre-trained transformer model unchanged, and updating a parameter of the linear layer, the audio vector, and the text vector based on the adjustment ([0150] - Layer 2 is a bi-directional RNN layer for performing recurrent operations among words of sentences corresponding to the same topic. The purpose of Layer 2 is to convert a whole sentence set into a vector. A vector h.sub.t+1 in Layer 2 may be computed by linearly combining h.sub.t and x.sub.t and attaching an element-wise non-linear transformation function such as RNN(.). Although RNN(.) is adopted here, it should be appreciated that the element-wise non-linear transformation function may also adopt, e.g., tanh or sigmoid, or a LSTM/GRU computing block [0153] - Layer 4 is an output layer. Layer 4 may be configured for determining probabilities of possible states of each topic in a pre-given state list, e.g., a lifecycle state list for one topic. A topic word may correspond to a list of probability p.sub.i, where i ranges from 1 to the number of states |P|. Process in this layer may comprise firstly computing y which is a linear function of h.sub.m.sup.2, and then using a softmax function to project y into a probability space and ensure P=[p.sub.1, p.sub.2, . . . , p.sub.|P|].sup.T follows a definition of probability. For error back-propagation, cross-entropy loss which corresponds to a minus log function of P for each topic word may be applied). Regarding claim 7, Wu disclsoes the method, wherein the updating the parameter further comprises: performing backpropagation based on the adjustment, and determining a gradient of the parameter of the linear layer, a gradient of the audio vector, and a gradient of the text vector ([0153] - Process in this layer may comprise firstly computing y which is a linear function of h.sub.m.sup.2, and then using a softmax function to project y into a probability space and ensure P=[p.sub.1, p.sub.2, . . . , p.sub.|P|].sup.T follows a definition of probability. For error back-propagation, cross-entropy loss which corresponds to a minus log function of P for each topic word may be applied [0154] - The above discussed structure 1000 is easy to be implemented. However, gradient will vanish as T grows bigger and bigger. For example, gradients in (0, 1) from h.sub.T back to h.sub.1 will gradually close to zero, making Stochastic Gradient Descent (SGD)-style updating of parameters infeasible. Thus, in some implementations, to alleviate this problem occurred when using simple non-linear functions, e.g., tanh or sigmoid, other types of functions for expressing h.sub.t+1 by h.sub.t and x.sub.t may be adopted, such as, Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), etc.); and updating the parameter of the linear layer, the audio vector, and the text vector based on the gradients respectively ([0150] - Layer 2 is a bi-directional RNN layer for performing recurrent operations among words of sentences corresponding to the same topic. The purpose of Layer 2 is to convert a whole sentence set into a vector. A vector h.sub.t+1 in Layer 2 may be computed by linearly combining h.sub.t and x.sub.t and attaching an element-wise non-linear transformation function such as RNN(.). Although RNN(.) is adopted here, it should be appreciated that the element-wise non-linear transformation function may also adopt, e.g., tanh or sigmoid, or a LSTM/GRU computing block [0153] - Layer 4 is an output layer. Layer 4 may be configured for determining probabilities of possible states of each topic in a pre-given state list, e.g., a lifecycle state list for one topic. A topic word may correspond to a list of probability p.sub.i, where i ranges from 1 to the number of states |P|. Process in this layer may comprise firstly computing y which is a linear function of h.sub.m.sup.2, and then using a softmax function to project y into a probability space and ensure P=[p.sub.1, p.sub.2, . . . , p.sub.|P|].sup.T follows a definition of probability. For error back-propagation, cross-entropy loss which corresponds to a minus log function of P for each topic word may be applied). Regarding claim 8, Wu discloses the method, wherein the training object includes audio information and text information; the text information is generated from speech of the audio information ([0050] - For a voice message, the voice processing unit 254 may perform a voice-to-text conversion on the voice message to obtain text sentences, the text processing unit 252 may perform text understanding on the obtained text sentences, and the core processing module 220 may further determine a text response. If it is determined to provide a response in voice, the voice processing unit 254 may perform a text-to-voice conversion on the text response to generate a corresponding voice response); and the training object is a conversation between a service provider and a user ([0030] - historical conversation records of an individual or group chat with a chatbot or other people, etc., they may want to check important or interested information in the text documents). Regarding claim 9, Wu discloses an apparatus for training an emotion classification model, the apparatus comprising: processing circuitry configured to obtain a training object and an emotion classification model ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); add an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model ([0047] - The messages may be in various multimedia forms, such as, text, voice, image, video, etc. [0059] - The module set 270 may comprise an emotion information extraction module 274. The emotion information extraction module 274 may be configured for extracting emotion information in a multimedia document through emotion analysis. The emotion information extraction module 274 may be implemented based on, such as, a recurrent neural network (RNN) together with a SoftMax layer. The emotion information may be represented, e.g., in a form of vector, and may be further used for deriving an emotion category); obtain a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtain a transformed text feature vector of the sample of the training object based on the transformation layer ([0083] - Data in a CKG may be in forms of text, image, voice and video. In some implementations, image, voice and video documents may be transformed into textual form and stored in the CKG to be used for CKG mining. For example, as for image documents, images may be transformed into texts by using image-to-text transforming model, such as a recurrent neural network-convolutional neural network (RNN-CNN) model. As for voice documents, voices may be transformed into texts by using voice recognition models. As for video documents which comprise voice parts and image parts, the voice parts may be transformed into texts by using the voice recognition models, and the image parts may be transformed into texts by using 2D/3D convolutional networks to project video information into dense space vectors and further using the image-to-text models. Videos may be processed through attention-based encoding-decoding models to capture text descriptions for the videos [0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); perform classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); obtain an adjustment based on an emotion classification result of the sample of the training object ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0103] - The SVM classifier 880 may make a secondary judgment to the candidate training data obtained based on the emotion lexicon 850. Through the operation of the SVM classifier 880, those sentences having a relatively high confidence probability in the candidate training data may be finally appended to a training dataset 890. The training dataset 890 may be used for training the emotion classifying models); and update the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0103] - The SVM classifier 880 may make a secondary judgment to the candidate training data obtained based on the emotion lexicon 850. Through the operation of the SVM classifier 880, those sentences having a relatively high confidence probability in the candidate training data may be finally appended to a training dataset 890. The training dataset 890 may be used for training the emotion classifying models [0150] - Layer 2 is a bi-directional RNN layer for performing recurrent operations among words of sentences corresponding to the same topic. The purpose of Layer 2 is to convert a whole sentence set into a vector. A vector h.sub.t+1 in Layer 2 may be computed by linearly combining h.sub.t and x.sub.t and attaching an element-wise non-linear transformation function such as RNN(.). Although RNN(.) is adopted here, it should be appreciated that the element-wise non-linear transformation function may also adopt, e.g., tanh or sigmoid, or a LSTM/GRU computing block). Dependent claims 10-12 and 14-16 are analogous in scope to claims 2-4 and 6-8, and are rejected according to the same reasoning. Regarding claim 17, Wu discloses a non-transitory computer-readable storage medium, storing instructions which when executed by a processor cause the processor to perform: obtaining a training object and an emotion classification model ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); adding an audio vector of the training object and a text vector of the training object to a transformation layer of the emotion classification model ([0047] - The messages may be in various multimedia forms, such as, text, voice, image, video, etc. [0059] - The module set 270 may comprise an emotion information extraction module 274. The emotion information extraction module 274 may be configured for extracting emotion information in a multimedia document through emotion analysis. The emotion information extraction module 274 may be implemented based on, such as, a recurrent neural network (RNN) together with a SoftMax layer. The emotion information may be represented, e.g., in a form of vector, and may be further used for deriving an emotion category); obtaining a transformed audio feature vector of a sample of the training object based on the transformation layer, and obtaining a transformed text feature vector of the sample of the training object based on the transformation layer ([0083] - Data in a CKG may be in forms of text, image, voice and video. In some implementations, image, voice and video documents may be transformed into textual form and stored in the CKG to be used for CKG mining. For example, as for image documents, images may be transformed into texts by using image-to-text transforming model, such as a recurrent neural network-convolutional neural network (RNN-CNN) model. As for voice documents, voices may be transformed into texts by using voice recognition models. As for video documents which comprise voice parts and image parts, the voice parts may be transformed into texts by using the voice recognition models, and the image parts may be transformed into texts by using 2D/3D convolutional networks to project video information into dense space vectors and further using the image-to-text models. Videos may be processed through attention-based encoding-decoding models to capture text descriptions for the videos [0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); performing classification based on the transformed audio feature vector, the transformed text feature vector, and a linear layer of the emotion classification model ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700); obtaining an adjustment based on an emotion classification result of the sample of the training object ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0103] - The SVM classifier 880 may make a secondary judgment to the candidate training data obtained based on the emotion lexicon 850. Through the operation of the SVM classifier 880, those sentences having a relatively high confidence probability in the candidate training data may be finally appended to a training dataset 890. The training dataset 890 may be used for training the emotion classifying models); and updating the audio vector, the text vector, and the linear layer of the emotion classification result based on the adjustment ([0102] - a support vector machine (SVM) classifier 880 may be used for filtering out interference sentences from the candidate training data or correcting improper emotion annotations of some candidate training data. The SVM classifier 880 may use trigram characters as features. A set of seed training data may be obtained for training the SVM classifier 880. For example, the seed training data may comprise 1,000 manually-annotated instances for each emotion 8. In one case, a sentence in an instance may be annotated by one of the 8 basic emotions or one of the 8 combined emotions, and if one basic emotion is annotated, a strength level shall be further annotated. In another case, a sentence in an instance may be annotated directly by one of the 32 emotions in the wheel of emotions 700 [0103] - The SVM classifier 880 may make a secondary judgment to the candidate training data obtained based on the emotion lexicon 850. Through the operation of the SVM classifier 880, those sentences having a relatively high confidence probability in the candidate training data may be finally appended to a training dataset 890. The training dataset 890 may be used for training the emotion classifying models [0150] - Layer 2 is a bi-directional RNN layer for performing recurrent operations among words of sentences corresponding to the same topic. The purpose of Layer 2 is to convert a whole sentence set into a vector. A vector h.sub.t+1 in Layer 2 may be computed by linearly combining h.sub.t and x.sub.t and attaching an element-wise non-linear transformation function such as RNN(.). Although RNN(.) is adopted here, it should be appreciated that the element-wise non-linear transformation function may also adopt, e.g., tanh or sigmoid, or a LSTM/GRU computing block). Dependent claims 18-20 are analogous in scope to claims 2-4, and are rejected according to the same reasoning. Allowable Subject Matter 6. Claims 5 and 13 are rejected under 35 U.S.C. 101, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: The prior art could not overcome or render obvious the limitations of “obtaining a transposed first key vector based on the first key vector; multiplying the transposed first key vector by the first query vector to obtain a first result; dividing the first result by a square root of a dimension of the first key vector to obtain a second result; normalizing the second result to obtain a normalized second result; and multiplying the normalized second result by the first value vector as the transformed audio feature vector” as claimed. Conclusion 7. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Mittal (U.S. Patent No. 11861940) teaches human emotion recognition in images or video. Sohn (U.S. Publication No. 20250285641) teaches an apparatus for estimating emotion using multimodal model and method of training the same. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ETHAN DANIEL KIM whose telephone number is (571) 272-1405. The examiner can normally be reached on Monday - Friday 9:00 - 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached on (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ETHAN DANIEL KIM/ Examiner, Art Unit 2658 /RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Mar 27, 2025
Application Filed
Aug 26, 2026
Non-Final Rejection mailed — §101, §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731573
ENHANCED SPOKEN DIALOGUE MODIFICATION
3y 3m to grant Granted Sep 08, 2026
Patent 12706101
ROTATION OF SOUND COMPONENTS FOR ORIENTATION-DEPENDENT CODING SCHEMES
3y 2m to grant Granted Aug 11, 2026
Patent 12688220
MODEL PERFORMANCE THROUGH TEXT-TO-TEXT TRANSFORMATION VIA DISTANT SUPERVISION FROM TARGET AND AUXILIARY TASKS
5y 2m to grant Granted Jul 21, 2026
Patent 12682894
SYSTEM AND METHOD FOR SIMULTANEOUSLY IDENTIFYING INTENT AND SLOTS IN VOICE ASSISTANT COMMANDS
3y 11m to grant Granted Jul 14, 2026
Patent 12664993
INTEGRATION OF HIGH FREQUENCY RECONSTRUCTION TECHNIQUES WITH REDUCED POST-PROCESSING DELAY
1y 6m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+25.4%)
2y 10m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 114 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month