DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-18, 24-25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xu U.S. PAP 2022/0028371 A1 in view of Liu “Topic-Aware Contrastive learning for abstractive Dialogue Summarization” (Applicant admitted prior art).
Regarding claim 1 Xu teaches a method for generating a summary of a dialogue ( a method for processing speech dialogue may be implemented on a computing device having one or more processors and one or more storage devices, see par. [0004]), comprising:
determine topic information of the dialogue and generate a model for generating the summary of the dialogue (after the pre-training of the speech dialogue coding model is completed, the speech dialogue coding model may be adjusted based on a downstream task model (e.g., a classification model, a summary extraction model, a translation model) and sample data with annotations (which correspond to a corresponding downstream task, for example, “category,” “summary,” “translation”), which may improve the processing effect of a downstream task, see par. [0112]);
and inputting a target dialogue into the model for generating the summary of the dialogue to obtain the summary of the target dialogue (obtaining target speech dialogue data. The method may include obtaining a text vector representation sequence, a phonetic symbol vector representation sequence, and a role vector representation sequence by performing a vector transformation on the target speech dialogue data based on a text embedding model, a phonetic symbol embedding model, and a role embedding model, respectively. The method may include determining a representation vector corresponding to the target speech dialogue data by inputting the text vector representation sequence, the phonetic symbol vector representation sequence, and the role vector representation sequence into a trained speech dialogue coding model. The method may include determining a summary of the target speech dialogue data by inputting the representation vector into a classification model, see par. [0004]).
However Xu does not teach modeling semantic coherence of dialogue sentences by a contrastive learning mode to determine topic information of the dialogue and generate a model for generating the summary of the dialogue.
In the same field of endeavor Liu teaches To capture the various topic information of a conversation and outline salient facts for the captured topics, this work proposes two topic-aware contrastive learning objectives, namely coherence detection and sub-summary generation objectives, which are expected to implicitly model the topic change and handle information scattering challenges for the dialogue summarization task, see abstract.
It would have been obvious to one of ordinary skill in the art to combine the Xu invention with the teachings of Liu for the benefit of handling information scattering challenges for the dialogue summarization task, see abstract.
Regarding claim 2 Liu teaches the method according to claim 1, wherein the modeling the semantic coherence of dialogue sentences by the contrastive learning mode to determine the topic information of the dialogue and generate the model for generating the summary of the dialogue, comprises:
constructing a model for detecting the semantic coherence of the dialogue sentences, to model a switching relationship between different topics in the dialogue by learning the semantic coherence of dialogue sentences to obtain topic segmentation information of the dialogue (obtain the topical information of a dialogue by modeling the coherence change among utterances. The assumption behind this is that utterances within the same topic are more coherent than those spanning across different topics, see section 2.2);
constructing a model for generating a sub-summary, to generate a sub-summary corresponding to each of topics of the dialogue (we introduce the contrastive sub-summary generation objective, see section 2.2);
Xu teaches constructing a model for generating a full-text summary, to generate a summary of full-text of the dialogue (determine a trained model 125, see par. [0051]).
Regarding claim 3 Xu teaches the method according to claim 2, wherein the modeling the semantic coherence of dialogue sentences by the contrastive learning mode to determine the topic information of the dialogue and generate the model for generating the summary of the dialogue, further comprises:
performing model training on the model for detecting the semantic coherence of the dialogue sentences, the model for generating the sub-summary and the model for generating the full-text summary by using an alternating parameter updating mode (the speech dialogue coding model may be determined according to a training process. The training process may include obtaining sample speech dialogue data. The training process may include obtaining a text vector representation sequence, a phonetic symbol vector representation sequence, and a role vector representation sequence by performing a vector transformation on the sample speech dialogue data based on a text embedding model, a phonetic symbol embedding model, and a role embedding model, respectively, see par. [0008]).
Regarding claim 4 Liu teaches the method according to claim 3, wherein the performing model training on the model for detecting the semantic coherence of the dialogue sentences, the model for generating the sub-summary and the model for generating the full-text summary by using an alternating parameter updating mode, comprises:
sequentially updating parameters by using an objective function of the model for detecting the semantic coherence of the dialogue sentences, an objective function of the model for generating the sub-summary and an objective function of the model for generating the full-text summary, to train the model for detecting the semantic coherence of the dialogue sentences, the model for generating the sub-summary and the model for generating the full-text summary in a training process (The summary of a long dialogue always consists of multiple sentences each of which is regarded as a sub-summary. Considering the fact that one dialogue may contain more than one topics, we assume that each sub-summary is related to one topic. Hence, we introduce the contrastive sub-summary generation objective. The sub-summary objective can be used to update the parameters in the encoder and decoder, see section 2.2);
and taking the model for detecting the semantic coherence of the dialogue sentences and the model for generating the sub-summary as auxiliary tasks to improve the quality of generating the summary by the model for generating the full-text summary in the training process (our model significantly improves the quality of generated summaries when the dialogue summary comprises of more than one sub-summaries, see section 3.5).
Regarding claim 5 Liu teaches the method according to claim 2, wherein the modeling the semantic coherence of dialogue sentences by the contrastive learning mode to determine the topic information of the dialogue and generate the model for generating the summary of the dialogue, further comprises:
pre-processing the dialogue (pre-processing procedure, see section A.3) ;
constructing corresponding model training data according to requirements of a model for understanding the dialogue, the model for generating the sub-summary and the model for generating the full-text summary (design two contrastive objectives as auxiliary task, i.e., coherence detection and sub-summary generation objectives, working together with the primary summarization task during training, see section 5);
Xu teaches and constructing the model for understanding the dialogue, to semantically encode the dialogue, and training the model for understanding the dialogue (by merging text information, phonetic symbol information, and role information of target speech dialogue data, the accuracy of semantic understanding of the target speech dialogue data can be improved, see par. [0045]).
Regarding claim 6 Liu teaches the method according to claim 5, wherein the pre-processing the dialogue, comprises: adding speaker information to each of speech contents of different speakers in the dialogue, and splicing the speech contents of the different speakers together (employ hierarchical models to capture features from different turns of different speakers, see section 1); and tokenizing the spliced dialogue by using a tokenizer of a pre-training model, and reserving a first predetermined number of words as a model input (we add interlocutors information before concatenating utterances, and then truncate the dialogues to keep only first 1024 tokens as input., see section A.2).
Regarding claim 7 Liu teaches the method according to claim 5, wherein the constructing the corresponding model training data according to the requirements of the model for understanding the dialogue, the model for generating the sub-summary and the model for generating the full-text summary, comprises:
taking a window constructed according to consecutive dialogue sentences in the dialogue as positive model training data, and taking data obtained by scrambling and splicing again dialogue sentences in window content as negative model training data, for the model for understanding the dialogue (We introduce a window comprising a subsequence of k (k < |D|) utterances of a dialogue D, as a snippet, denoted as SD k . For instance, (uj,uj+1,...,uj+k) is an example snippet for dialogue D where j ∈ [1,|D| − k] is an integer utterance index. Such a snippet is regarded as a positive example, while the corresponding negative snippet f SD k is constructed by shuffling the order of sentences inside SD k . Given a pair of positive and negative examples, denoted as PDco = (SD k , f SD k ), the contextual representations of each snippet can be obtained through the last layer of the Transformer encoder, denoted as ESD k , Ef SD k , individually, see section 2.2);
generating corresponding positive model training data and negative model training data according to each topic of the dialogue, for the model for generating the sub-summary (Given a pair of positive and negative examples, denoted as PDco = (SD k , f SD k ), the contextual representations of each snippet can be obtained through the last layer of the Transformer encoder, denoted as ESD k , Ef SD k , individually, see section 2.2);
and taking all content of the dialogue as a model input, and taking a complete summary as a model output, for the model for generating the full-text summary (See algorithm 1).
Regarding claim 8 Liu teaches the method according to claim 2, wherein the constructing the mode for detecting the coherence of the dialogue sentences, to model a switching relationship between different topics in the dialogue by learning the semantic coherence of dialogue sentences to obtain topic segmentation information of the dialogue, comprises: calculating coherence scores of positive model training data and negative model training data of the mode for detecting the coherence of the dialogue sentences respectively (The contrastive margin-based coherence loss is then calculated, see section 2.2);
randomly selecting a second predetermined number of predetermined positive and negative pairs, and calculating a coherence loss based on contrastive learning in a training stage (For a dialogue D, there exist at least |D − k| contrastive snippet pairs, while, for simplicity, we randomly select Nco < |D−k| pairs for each epoch. The coherence loss can be used to update the parameters in the encoder, see section 2.2); and calculating an objective function of the mode for detecting the coherence of the dialogue sentences according to an edge contrastive loss (margin-based contrastive loss is calculated, see section 2.2).
Regarding claim 9 Liu teaches the method according to claim 2, wherein the constructing the model for generating the sub-summary, to generate the sub-summary corresponding to each of the topics of the dialogue, comprises:
modeling a sub-summary generation task as a sequence-to-sequence learning problem; determining the degree of irrelevance between a dialogue segment and a sub-summary in the sub-summary generation task (The normalized scores after the softmax layer can be regarded as the irrelevance score to show how irrelevant a snippet is to a sub-summary, see section 2.2);
randomly selecting a third predetermined number of predetermined positive and negative pairs for training in a training stage (The corresponding negative example is randomly picked from the rest snip pets in W, denoted as Si neg. Now, we have con structed the contrastive sub-summary generation pairs {(Si pos, ti), (Si neg, ti)}, see 2.2);
and determining an objective function of the model for generating the sub-summary according to a marginal loss function based on contrastive learning (margin-based contrastive loss is calculated, see section 2.2).
Regarding claim 10 Liu teaches the method according to claim 2, wherein the constructing a model for generating a full-text summary, to generate the summary of the full-text of the dialogue, comprises: modeling a full-text summary generation task as a sequence-to-sequence learning problem (The abstractive summarization task learns to generate summaries by rewriting the input document, which is a typical sequence-to-sequence learning problem. Sequence-to-sequence attentive LSTMs, see section 4);
setting a training goal of the model for generating the full-text summary to learn an optimal model parameter and minimize a negative logarithmic likelihood function value (the goal is to learn the optimal model parameters θ and to minimize the negative log-likelihood, see section 2.1);
Xu teaches and determining an objective function of the model for generating the full-text summary (determine a trained model 125, see par. [0051]).
Regarding claim 11 Xu teaches a method for training a model, comprising:
determine topic information of the dialogue and generate a model for generating the summary of the dialogue (after the pre-training of the speech dialogue coding model is completed, the speech dialogue coding model may be adjusted based on a downstream task model (e.g., a classification model, a summary extraction model, a translation model) and sample data with annotations (which correspond to a corresponding downstream task, for example, “category,” “summary,” “translation”), which may improve the processing effect of a downstream task, see par. [0112]);
and performing model training on the model for generating the summary of the dialogue by using an alternating parameter updating mode, so that the model for generating the summary of the dialogue outputs the summary of a target dialogue according to an input target dialogue (obtaining target speech dialogue data. The method may include obtaining a text vector representation sequence, a phonetic symbol vector representation sequence, and a role vector representation sequence by performing a vector transformation on the target speech dialogue data based on a text embedding model, a phonetic symbol embedding model, and a role embedding model, respectively. The method may include determining a representation vector corresponding to the target speech dialogue data by inputting the text vector representation sequence, the phonetic symbol vector representation sequence, and the role vector representation sequence into a trained speech dialogue coding model. The method may include determining a summary of the target speech dialogue data by inputting the representation vector into a classification model, see par. [0004]).
However Xu does not teach modeling semantic coherence of dialogue sentences by a contrastive learning mode to determine topic information of the dialogue and generate a model for generating the summary of the dialogue.
In the same field of endeavor Liu teaches To capture the various topic information of a conversation and outline salient facts for the captured topics, this work proposes two topic-aware contrastive learning objectives, namely coherence detection and sub-summary generation objectives, which are expected to implicitly model the topic change and handle information scattering challenges for the dialogue summarization task, see abstract.
It would have been obvious to one of ordinary skill in the art to combine the Xu invention with the teachings of Liu for the benefit of handling information scattering challenges for the dialogue summarization task, see abstract.
Regarding claim 12 The method according to claim 11, wherein the modeling the semantic coherence of dialogue sentences by the contrastive learning mode to determine the topic information of the dialogue and generate the model for generating the summary of the dialogue, comprises:
constructing a model for detecting the semantic coherence of the dialogue sentences, to model a switching relationship between different topics in the dialogue by learning the semantic coherence of dialogue sentences to obtain topic segmentation information of the dialogue (obtain the topical information of a dialogue by modeling the coherence change among utterances. The assumption behind this is that utterances within the same topic are more coherent than those spanning across different topics, see section 2.2);
constructing a model for generating a sub-summary, to generate a sub-summary corresponding to each of topics of the dialogue (we introduce the contrastive sub-summary generation objective, see section 2.2).
Regarding claim 13 Liu teaches the method according to claim 12, wherein the performing model training on the model for generating the summary of the dialogue by using an alternating parameter updating mode, comprises:
sequentially updating parameters by using an objective function of the model for detecting the semantic coherence of the dialogue sentences, an objective function of the model for generating the sub-summary and an objective function of the model for generating the full-text summary, to train the model for generating the summary of the dialogue, wherein the model for generating the summary of the dialogue comprises the model for detecting the semantic coherence of the dialogue sentences, the model for generating the sub-summary and the model for generating the full-text summary in a training process (The summary of a long dialogue always consists of multiple sentences each of which is regarded as a sub-summary. Considering the fact that one dialogue may contain more than one topics, we assume that each sub-summary is related to one topic. Hence, we introduce the contrastive sub-summary generation objective. The sub-summary objective can be used to update the parameters in the encoder and decoder, see section 2.2);
and taking the model for detecting the semantic coherence of the dialogue sentences and the model for generating the sub-summary as auxiliary tasks to improve the quality of generating the summary by the model for generating the full-text summary in the training process (our model significantly improves the quality of generated summaries when the dialogue summary comprises of more than one sub-summaries, see section 3.5).
Regarding claim 14 Liu teaches The method according to claim 12, wherein the modeling the semantic coherence of dialogue sentences by the contrastive learning mode to determine the topic information of the dialogue and generate the model for generating the summary of the dialogue, further comprises:
pre-processing the dialogue (pre-processing procedure, see section A.3) ;
constructing corresponding model training data according to requirements of a model for understanding the dialogue, the model for generating the sub-summary and the model for generating the full-text summary (design two contrastive objectives as auxiliary task, i.e., coherence detection and sub-summary generation objectives, working together with the primary summarization task during training, see section 5);
Xu teaches and constructing the model for understanding the dialogue, to semantically encode the dialogue, and training the model for understanding the dialogue (by merging text information, phonetic symbol information, and role information of target speech dialogue data, the accuracy of semantic understanding of the target speech dialogue data can be improved, see par. [0045]).
Regarding claim 15 Liu teaches the method according to claim 14, wherein the pre-processing the dialogue, comprises: adding speaker information to each of speech contents of different speakers in the dialogue, and splicing the speech contents of the different speakers together (employ hierarchical models to capture features from different turns of different speakers, see section 1); and tokenizing the spliced dialogue by using a tokenizer of a pre-training model, and reserving a first predetermined number of words as a model input (we add interlocutors information before concatenating utterances, and then truncate the dialogues to keep only first 1024 tokens as input., see section A.2).
Regarding claim 16 Liu teaches the method according to claim 14, wherein the constructing the corresponding model training data according to the requirements of the model for understanding the dialogue, the model for generating the sub-summary and the model for generating the full-text summary, comprises: taking a window constructed according to consecutive dialogue sentences in the dialogue as positive model training data, and taking data obtained by scrambling and splicing again dialogue sentences in window content as negative model training data, for the model for understanding the dialogue (We introduce a window comprising a subsequence of k (k < |D|) utterances of a dialogue D, as a snippet, denoted as SD k . For instance, (uj,uj+1,...,uj+k) is an example snippet for dialogue D where j ∈ [1,|D| − k] is an integer utterance index. Such a snippet is regarded as a positive example, while the corresponding negative snippet f SD k is constructed by shuffling the order of sentences inside SD k . Given a pair of positive and negative examples, denoted as PDco = (SD k , f SD k ), the contextual representations of each snippet can be obtained through the last layer of the Transformer encoder, denoted as ESD k , Ef SD k , individually, see section 2.2);
generating corresponding positive model training data and negative model training data according to each topic of the dialogue, for the model for generating the sub-summary (Given a pair of positive and negative examples, denoted as PDco = (SD k , f SD k ), the contextual representations of each snippet can be obtained through the last layer of the Transformer encoder, denoted as ESD k , Ef SD k , individually, see section 2.2);
and taking all content of the dialogue as a model input, and taking a complete summary as a model output, for the model for generating the full-text summary (See algorithm 1).
Regarding claim 17 Liu teaches the method according to claim 12, wherein the constructing the mode for detecting the coherence of the dialogue sentences, to model a switching relationship between different topics in the dialogue by learning the semantic coherence of dialogue sentences to obtain topic segmentation information of the dialogue, comprises:
calculating coherence scores of positive model training data and negative model training data of the mode for detecting the coherence of the dialogue sentences respectively (The contrastive margin-based coherence loss is then calculated, see section 2.2);
randomly selecting a second predetermined number of predetermined positive and negative pairs, and calculating a coherence loss based on contrastive learning in a training stage (For a dialogue D, there exist at least |D − k| contrastive snippet pairs, while, for simplicity, we randomly select Nco < |D−k| pairs for each epoch. The coherence loss can be used to update the parameters in the encoder, see section 2.2); and calculating an objective function of the mode for detecting the coherence of the dialogue sentences according to an edge contrastive loss (margin-based contrastive loss is calculated, see section 2.2).
Regarding claim 18 Liu teaches the method according to claim 12, wherein: the constructing the model for generating the sub-summary, to generate the sub-summary corresponding to each of the topics of the dialogue, comprises:
modeling a sub-summary generation task as a sequence-to-sequence learning problem; determining the degree of irrelevance between a dialogue segment and a sub-summary in the sub-summary generation task (The normalized scores after the softmax layer can be regarded as the irrelevance score to show how irrelevant a snippet is to a sub-summary, see section 2.2);
randomly selecting a third predetermined number of predetermined positive and negative pairs for training in a training stage (The corresponding negative example is randomly picked from the rest snip pets in W, denoted as Si neg. Now, we have con structed the contrastive sub-summary generation pairs {(Si pos, ti), (Si neg, ti)}, see 2.2);
and determining an objective function of the model for generating the sub-summary according to a marginal loss function based on contrastive learning (margin-based contrastive loss is calculated, see section 2.2).
modeling a full-text summary generation task as a sequence-to-sequence learning problem (The abstractive summarization task learns to generate summaries by rewriting the input document, which is a typical sequence-to-sequence learning problem. Sequence-to-sequence attentive LSTMs, see section 4);
setting a training goal of the model for generating the full-text summary to learn an optimal model parameter and minimize a negative logarithmic likelihood function value (the goal is to learn the optimal model parameters θ and to minimize the negative log-likelihood, see section 2.1);
Xu teaches and determining an objective function of the model for generating the full-text summary (determine a trained model 125, see par. [0051]).
Regarding claim 24 Xu teaches a computer device comprising: a memory configured to store instructions; and a processor configured to execute a method for performing the instructions ( a method for processing speech dialogue may be implemented on a computing device having one or more processors and one or more storage devices, see par. [0004]) comprising:
determine topic information of the dialogue and generate a model for generating the summary of the dialogue (after the pre-training of the speech dialogue coding model is completed, the speech dialogue coding model may be adjusted based on a downstream task model (e.g., a classification model, a summary extraction model, a translation model) and sample data with annotations (which correspond to a corresponding downstream task, for example, “category,” “summary,” “translation”), which may improve the processing effect of a downstream task, see par. [0112]);
and inputting a target dialogue into the model for generating the summary of the dialogue to obtain the summary of the target dialogue (obtaining target speech dialogue data. The method may include obtaining a text vector representation sequence, a phonetic symbol vector representation sequence, and a role vector representation sequence by performing a vector transformation on the target speech dialogue data based on a text embedding model, a phonetic symbol embedding model, and a role embedding model, respectively. The method may include determining a representation vector corresponding to the target speech dialogue data by inputting the text vector representation sequence, the phonetic symbol vector representation sequence, and the role vector representation sequence into a trained speech dialogue coding model. The method may include determining a summary of the target speech dialogue data by inputting the representation vector into a classification model, see par. [0004]).
However Xu does not teach modeling semantic coherence of dialogue sentences by a contrastive learning mode to determine topic information of the dialogue and generate a model for generating the summary of the dialogue.
In the same field of endeavor Liu teaches To capture the various topic information of a conversation and outline salient facts for the captured topics, this work proposes two topic-aware contrastive learning objectives, namely coherence detection and sub-summary generation objectives, which are expected to implicitly model the topic change and handle information scattering challenges for the dialogue summarization task, see abstract.
Regarding claim 25 Xu teaches a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium has computer instructions stored thereon that, when executed by a processor, implement the method according to claim 1 ( a non-transitory computer readable medium may include at least one set of instructions. When executed by at least one processor of a computing device, the at least one set of instructions may cause the at least one processor to effectuate a method, see par. [0023]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Gehrmann ‘391 teaches method for generating a summary. The method includes one or more processing devices performing operations including generating a set of word embeddings corresponding to each word of a text input, see abstract.
Paulus ‘400 teaches robust and coherent abstractive text summarization model addresses these issues of general coherence, flow and readability, as well as unnatural summaries with repeated phrases, see par. [0012].
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Ortiz-Sanchez whose telephone number is (571)270-3711. The examiner can normally be reached Monday- Friday 9AM-6PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL ORTIZ-SANCHEZ/Primary Examiner, Art Unit 2656