DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
The claims recite mathematical concepts and mental-process-like steps (probabilities, vector operations and grouping). The specification nor the claims recite an explicit machine/architecture and there is no stated improvement to computer/ASR technology. The claims merely recite analyzing and transforming information to output a transcript summary. There is no integration into a practical application. The claims are directed to the abstract idea of transcript topic segmentation, as explained in detail below.
The limitations, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “processor” nothing in the claim element precludes the steps from practically being performed by mental processing, human activity or mathematical concept. For example, the language, obtaining a communication transcript including a set of utterances, wherein the obtained communication transcript is a text data set (can be done by a user listening to someone speak and convert the speech to text); dividing the set of utterances into a plurality of utterance windows of a defined window size, wherein each utterance window of the plurality of utterance windows includes a different subset of utterances of the set of utterances, and wherein each utterance of the set of utterances is included in at least one utterance window of the plurality of utterance windows (can be done by a user segmenting the data heard); for each utterance window of the plurality of utterance windows, classifying each utterance in the utterance window as a topic boundary or a non-boundary using a deep learning model applied to the utterance window (can be done by the user categorizing the data using a mathematical concept); identifying topic segments of the communication transcript based on utterances of the set of utterances that are classified as topic boundaries (can be done by a user identifying themes of the conversation); and generating a communication transcript summary using the communication transcript and the identified topic segments (can be done by a user summarizing the data).
According to Step 1, it includes determining whether the claims fall within a statutory category. The claims include a system and method, therefore the claims fall within a statutory category. Step 2A Prong one, includes evaluating whether the claims recite a judicial exception. The claims recite a judicial exception, therefore an evaluation is done to determine if the claims fit into one of the categories. As explained above, the claims collectively and individually, fall within categories courts and USPTO guidance commonly treat as abstract ideas: mental processes (recognizing/extracting/organizing information), mathematical concepts (mapping text to numerical vectors) and fundamental data-processing/manipulation.
Prong 2B is used to evaluate whether the claims recite additional elements that integrate the exception into a practical application. The judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements which are recited at a high-level of generality such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
Furthermore, it is noted that a few of the dependent claims recite training a model, however, the claims do not explicitly recite how the trained data is used for a particular purpose. The claims are presented at a high level and do not meaningfully limit the claim to a specific, unconventional improvement in computer or vehicle technology. The claims do not recite a particular hardware architecture, specialized data structures, concrete signal-processing steps, defined latency or safety constraints, or a specific machine-learning architecture or training regime that produces a technological improvement. The mere mention of training is insufficient to transform the abstract idea into patent-eligible subject matter. The claims do not supply an inventive concept that amounts to significantly more than the judicial exception because the claimed elements are routine, conventional data-processing activities implemented on generic computing hardware.
Regarding claims 2 and 10, classifying each utterance in the utterance window as a topic boundary or a non-boundary using a model applied to the utterance window includes: generating output vectors for each utterance in the utterance window using an encoder based on the utterances of the utterance window (can be done by manipulating mathematical algorithm and generating vector data) and classifying the utterances of the utterance window as topic boundaries or non-boundaries using a binary classifier applied to the generated output vector (can be done by manipulating mathematical algorithm and classifying data).
Regarding claims 3, 11 and 18, classifying each utterance in the utterance window as a topic boundary or a non-boundary using a model applied to the utterance window includes: generating an utterance boundary score for a target utterance in the utterance window (can be done by a user generating a score regarding topic data); based on the generated utterance boundary score of the target utterance in the utterance window exceeding a boundary score threshold, classifying the target utterance as a topic boundary (can be done by a user classifying data based on a particular criteria); and based on the generated utterance boundary score of the target utterance in the utterance window not exceeding the boundary score threshold, classifying the target utterance as a non-boundary (can be done by a user classifying data based on a particular criteria).
Regarding claims 4, 12 and 19, generating an utterance boundary score for the target utterance in the utterance window further includes: generating a plurality of utterance boundary scores for the target utterance based on a plurality of utterance windows that include the target utterance (can be done by a user generating a score based on certain data) and identifying a highest utterance boundary score in the plurality of utterance boundary scores (can be done by a user identifying the highest score); and setting the identified highest utterance boundary score to be the utterance boundary score of the target utterance (can be done by a user setting the highest score to represent the target utterance).
Regarding claims 5 and 13, pre-training the deep learning model using general text training data; and training the pre-trained deep learning model continually using communication transcript training data including at least one of the following: labeled communication transcript training data and unlabeled communication transcript training data (mere instruction to perform the abstract idea on a generic “computing device” or to use conventional machine learning is insufficient to transform the abstract idea into patent-eligible subject matter. The claim does not supply an inventive concept that amounts to significantly more than the judicial exception because the claimed elements are routine, conventional data-processing activities implemented on generic computing hardware).
Regarding claims 6 and 14, training the pre-trained deep learning model continually includes: identifying topic segments in communication transcript training data and organizing the identified topic segments in a plurality of different orders to form a plurality of training transcripts; and training the pre-trained deep learning model using the plurality of training transcripts (mere instruction to perform the abstract idea on a generic “computing device” or to use conventional machine learning is insufficient to transform the abstract idea into patent-eligible subject matter. The claim does not supply an inventive concept that amounts to significantly more than the judicial exception because the claimed elements are routine, conventional data-processing activities implemented on generic computing hardware).
Regarding claims 7 and 15, training the pre-trained deep learning model continually includes: training a teacher model using the communication transcript training data; initialize a student model with parameters of the teacher model; and train the student model with output from the teacher model using a self-distillation process, wherein the trained student model is used as the deep learning model (mere instruction to perform the abstract idea on a generic “computing device” or to use conventional machine learning is insufficient to transform the abstract idea into patent-eligible subject matter. The claim does not supply an inventive concept that amounts to significantly more than the judicial exception because the claimed elements are routine, conventional data-processing activities implemented on generic computing hardware).
Regarding claims 8, 16 and 20, identifying a topic segment of the identified topic segments that has a time length that exceeds a segment length threshold (can be done by a user identifying a topic duration), identifying an utterance in the identified topic segment with a highest utterance boundary score (can be done by a user identifying the highest score) and classifying the identified utterance as a topic boundary, wherein the identified topic segments are updated based on classifying the identified utterance as a topic boundary (can be done by a user classifying the data).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4, 8-12 and 16-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Szymanski et al. (PGPUB 2021/0027783), hereinafter referenced as Szymanski in view of Joller et al. (PGPUB 2020/0105274), hereinafter referenced as Joller.
Regarding claims 1, 9 and 17, Szymanski discloses a system, method and media, hereinafter referenced as a system comprising:
at least one processor (fig. 1, element 900 with processor; p. 0099); and
at least one memory comprising computer program code, the at least one memory and the computer program code configured to (fig. 9, elements 906 or 908 with p. 0099), with the at least one processor, cause the at least one processor to:
obtain a communication transcript including a set of utterances, wherein the obtained communication transcript is a text data set (In the present example, it will be assumed that the electronic data packages received are raw audio files of a conversation between entities. Accordingly, in one embodiment, there is a speech to text processing module 206 that is operative to convert a raw audio data file to text. For example, the interaction engine 240 may use natural language processing (NLP) to process the raw natural language content of the conversational audio data. This natural language content may be provided as a text (e.g., short message system (SMS), e-mail, etc.,) or via voice. Regarding the latter, the speech to text processing module 206 of the interaction engine 240 can perform speech recognition to determine the textual representation thereof; figure 2, element 206 with p. 0036-0037);
divide the set of utterances into a plurality of utterance windows of a defined window size, wherein each utterance window of the plurality of utterance windows includes a different subset of utterances of the set of utterances, and wherein each utterance of the set of utterances is included in at least one utterance window of the plurality of utterance windows (The interaction engine applies a sequential RNN 540 that has a predetermined window 512 (having a length of four sentences in the present example). For example, the RNN used herein is a type of neural network where connections between social actions form a predetermined orderly cycle that is consistent with that of a topic; fig. 5, element 512 with p. 0046-0048);
for each utterance window of the plurality of utterance windows, classify each utterance in the utterance window as a topic boundary (fig. 5, element 520) or a non-boundary using a deep learning model applied to the utterance window; (An RNN is able to analyze the sequence of social actions to identify a dynamic temporal behavior. Unlike feedforward neural networks, where connections between the units do not form a cycle, RNNs discussed herein can use their memory to process sequences of inputs provided by social action classifiers. Thus, RNN's can leverage the dependence of a previous state of an input (i.e., social action) to determine its next state. During a training phase, the RNN can learn a probability distribution over a sequence by being trained to determine whether a next social action is consistent with a topic. In one embodiment, the neural network used is a many to one sequential RNN 540 that provides as output a probability of a boundary 520; p. 0046)
identify topic segments of the communication transcript based on utterances of the set of utterances that are classified as topic boundaries (The social action sequence in the window is used as inputs for RNN to predict if the window is a transition boundary. After all transition boundaries are detected, they can be used to divide the original conversation sentences 610 into segments. Accordingly, the length of segments is unrelated to, and different from, the window size. We analyze the segment of original conversation sentences (not social actions) to understand the topical meaning, and measure the similarity of adjacent segments (topic_sim(si, sj)). Social actions are used to help find the transition/division boundary, and calculate the strength (seg_div(S.sub.i,S.sub.j)) of dividing Si, Sj as two segments. For example, the conversation includes various utterances, which are converted to sequence of sentences 610. Each utterance, which was converted to a sentence, is classified to a corresponding social action 620. In the sequence of social actions 620, social actions in a customized window size (e.g. size 4) are used as input for RNN to detect if they can be transition boundaries. For example, the social actions in the dashed rectangle 630 indicate detected transition boundaries, which divide the sequence of sentences 610 into segments S.sub.i−1, S.sub.i, and S.sub.i+1 (640, 650, and 660, respectively). In one embodiment, the grouping of the segments into topics may be iterative. For example, the segmentation may be performed repeatedly between adjacent boundaries, such that similar topics are merged together if the similarity between the adjacent segments is above a predetermined threshold. This concept may be better understood in view of FIG. 7 discussed below; fig. 6, elements 610, 630 640, 650, 660 with p. 0047-0048, 0068); and
generate a communication transcript using the communication transcript and the identified topic segments (For example, such conversation may be received as actual text transcript or in the form of raw audio data, which is converted to text by the interaction engine; p. 0045-0046), but does not specifically teach a transcript summary.
Joller discloses a transcript summary (p. 0089-0091), to provide short-form highlights.
Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to provide an improved way to search and engage with content.
Regarding claims 2 and 10, Szymanski discloses a system, wherein classifying each utterance in the utterance window as a topic boundary or a non-boundary using a model applied to the utterance window includes:
generating output vectors for each utterance (probability which is a numerical representation) in the utterance window using an encoder based on the utterances of the utterance window (The interaction engine applies a sequential RNN 540 that has a predetermined window 512 (having a length of four sentences in the present example). For example, the RNN used herein is a type of neural network where connections between social actions form a predetermined orderly cycle that is consistent with that of a topic. An RNN is able to analyze the sequence of social actions to identify a dynamic temporal behavior. Unlike feedforward neural networks, where connections between the units do not form a cycle, RNNs discussed herein can use their memory to process sequences of inputs provided by social action classifiers. Thus, RNN's can leverage the dependence of a previous state of an input (i.e., social action) to determine its next state. During a training phase, the RNN can learn a probability distribution over a sequence by being trained to determine whether a next social action is consistent with a topic. In one embodiment, the neural network used is a many to one sequential RNN 540 that provides as output a probability of a boundary 520; p. 0044-0048); and
classifying the utterances of the utterance window as topic boundaries or non-boundaries using a binary classifier (In various embodiments, the machine learning discussed herein may be supervised or unsupervised. In supervised learning, during a training phase, the interaction engine 103 may be presented with historical data 113 from the historical data repository 112 as data that is related to various segments, respectively. Put differently, the historical data repository 112 acts as a teacher for the interaction engine 103. In unsupervised learning, the historical data repository 112 does not provide any labels as what is acceptable, rather, it simply provides historic data 113 to the collaboration server 116, which can be used to find its own structure among the data to identify the correct segmentation of a conversation. In various embodiments, the machine learning may make use of techniques such as supervised learning, unsupervised learning, semi-supervised learning, naive Bayes, Bayesian networks, decision trees, neural networks, fuzzy logic models, and/or probabilistic classification models.; p. 0041) applied to the generated output vectors (The social action sequence in the window is used as inputs for RNN to predict if the window is a transition boundary. After all transition boundaries are detected, they can be used to divide the original conversation sentences 610 into segments. Accordingly, the length of segments is unrelated to, and different from, the window size. We analyze the segment of original conversation sentences (not social actions) to understand the topical meaning, and measure the similarity of adjacent segments (topic_sim(si, sj)). Social actions are used to help find the transition/division boundary, and calculate the strength (seg_div(S.sub.i,S.sub.j)) of dividing Si, Sj as two segments. For example, the conversation includes various utterances, which are converted to sequence of sentences 610. Each utterance, which was converted to a sentence, is classified to a corresponding social action 620. In the sequence of social actions 620, social actions in a customized window size (e.g. size 4) are used as input for RNN to detect if they can be transition boundaries. For example, the social actions in the dashed rectangle 630 indicate detected transition boundaries, which divide the sequence of sentences 610 into segments S.sub.i−1, S.sub.i, and S.sub.i+1 (640, 650, and 660, respectively). In one embodiment, the grouping of the segments into topics may be iterative. For example, the segmentation may be performed repeatedly between adjacent boundaries, such that similar topics are merged together if the similarity between the adjacent segments is above a predetermined threshold; fig. 6, elements 610, 630 640, 650, 660 with p. 0047-0048, 0068).
Regarding claims 3, 11 and 18, Szymanski discloses a system wherein classifying each utterance in the utterance window as a topic boundary or a non-boundary using a model applied to the utterance window includes:
generating an utterance boundary score for a target utterance in the utterance window (segment/boundary score; p. 0048-0056);
based on the generated utterance boundary score of the target utterance in the utterance window exceeding a boundary score threshold (likelihood), classifying the target utterance as a topic boundary (p. 0047-0056); and
based on the generated utterance boundary score of the target utterance in the utterance window not exceeding the boundary score threshold, classifying the target utterance as a non-boundary (At block 808, the similarity of the topic between adjacent segments is determined. If the similarity is above a predetermine threshold (i.e., “YES” at decision block 818), then the process continues with block 820, where the adjacent segments having a topic similarity that is above the predetermined threshold, are merged into a common segment. The process then continues with block 826. However, if the similarity is not above the predetermined threshold (i.e., “NO” at decision block 818), the process goes directly to block 826, where a determination is made whether all segments in the sequence have been evaluated for similarity of topic. If not (i.e., “NO” at decision block 826), the process returns to block 808 until all segments have been evaluated in the present iteration. However, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626) in the present iteration, the process stops. In one embodiment, multiple iterations are performed. For example, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626), the process continues with block 830, where a determination is made whether any segments were merged in the present iteration. If not (i.e., “NO” at decision block 830), it is indicative that the segments were sufficiently merged, and the process stops. However, upon determining that one or more segments in the present iteration have been merged so (i.e., “YES” at decision block 830), it is indicative that the segments may be merged further, as discussed previously in the context of FIG. 6. Accordingly, the process returns to block 808 and the adjacent segments are evaluated for similarity. In one embodiment, the threshold for comparison is loosened between each iteration; p. 0044, 0064-0065).
Regarding claims 4, 12 and 19, Szymanski discloses a system wherein generating an utterance boundary score for the target utterance in the utterance window further includes:
generating a plurality of utterance boundary scores for the target utterance based on a plurality of utterance windows that include the target utterance (The interaction engine 240 includes an event bounded topic extractor 218 module that is operative to group the sequential topics into broader topics. For example, the interaction engine 240 may evaluate two adjacent segments in a sequence to determine the similarity between the topics. By way of continuing the previous example, if segment A (e.g., comprising social actions 1 to 10) relates to the focused topic of “how to improve yield in the DRAM portion of a new processor X,” and adjacent segment B (e.g., comprising social actions 11 to 20) relates to the focused topic of “the reasons why the performance of processor X deteriorates after 200 hours of operation,” the event bounded topic extractor module 218 is operative to determine how related the two focused topics are (e.g., by way of a similarity function). If the similarity is above a predetermined threshold, then the two segments A and B are merged into a common segment having a broader topic (e.g., “chip reliability.) In one embodiment, the process of merging adjacent segments in a sequence continues until the similarity between segments is not above a predetermined threshold. The resulting topics are then provided by the interaction engine 240 by way of an output module 220. At block 808, the similarity of the topic between adjacent segments is determined. If the similarity is above a predetermine threshold (i.e., “YES” at decision block 818), then the process continues with block 820, where the adjacent segments having a topic similarity that is above the predetermined threshold, are merged into a common segment. The process then continues with block 826. However, if the similarity is not above the predetermined threshold (i.e., “NO” at decision block 818), the process goes directly to block 826, where a determination is made whether all segments in the sequence have been evaluated for similarity of topic. If not (i.e., “NO” at decision block 826), the process returns to block 808 until all segments have been evaluated in the present iteration. However, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626) in the present iteration, the process stops. In one embodiment, multiple iterations are performed. For example, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626), the process continues with block 830, where a determination is made whether any segments were merged in the present iteration. If not (i.e., “NO” at decision block 830), it is indicative that the segments were sufficiently merged, and the process stops. However, upon determining that one or more segments in the present iteration have been merged so (i.e., “YES” at decision block 830), it is indicative that the segments may be merged further, as discussed previously in the context of FIG. 6. Accordingly, the process returns to block 808 and the adjacent segments are evaluated for similarity. In one embodiment, the threshold for comparison is loosened between each iteration; p. 0044-0048 and 0064-0065); and
identifying a highest utterance boundary score in the plurality of utterance boundary scores (The interaction engine 240 includes an event bounded topic extractor 218 module that is operative to group the sequential topics into broader topics. For example, the interaction engine 240 may evaluate two adjacent segments in a sequence to determine the similarity between the topics. By way of continuing the previous example, if segment A (e.g., comprising social actions 1 to 10) relates to the focused topic of “how to improve yield in the DRAM portion of a new processor X,” and adjacent segment B (e.g., comprising social actions 11 to 20) relates to the focused topic of “the reasons why the performance of processor X deteriorates after 200 hours of operation,” the event bounded topic extractor module 218 is operative to determine how related the two focused topics are (e.g., by way of a similarity function). If the similarity is above a predetermined threshold, then the two segments A and B are merged into a common segment having a broader topic (e.g., “chip reliability.) In one embodiment, the process of merging adjacent segments in a sequence continues until the similarity between segments is not above a predetermined threshold. The resulting topics are then provided by the interaction engine 240 by way of an output module 220. At block 808, the similarity of the topic between adjacent segments is determined. If the similarity is above a predetermine threshold (i.e., “YES” at decision block 818), then the process continues with block 820, where the adjacent segments having a topic similarity that is above the predetermined threshold, are merged into a common segment. The process then continues with block 826. However, if the similarity is not above the predetermined threshold (i.e., “NO” at decision block 818), the process goes directly to block 826, where a determination is made whether all segments in the sequence have been evaluated for similarity of topic. If not (i.e., “NO” at decision block 826), the process returns to block 808 until all segments have been evaluated in the present iteration. However, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626) in the present iteration, the process stops. In one embodiment, multiple iterations are performed. For example, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626), the process continues with block 830, where a determination is made whether any segments were merged in the present iteration. If not (i.e., “NO” at decision block 830), it is indicative that the segments were sufficiently merged, and the process stops. However, upon determining that one or more segments in the present iteration have been merged so (i.e., “YES” at decision block 830), it is indicative that the segments may be merged further, as discussed previously in the context of FIG. 6. Accordingly, the process returns to block 808 and the adjacent segments are evaluated for similarity. In one embodiment, the threshold for comparison is loosened between each iteration; p. 0044-0048 and 0064-0065); and
setting the identified highest utterance boundary score to be the utterance boundary score of the target utterance (The interaction engine 240 includes an event bounded topic extractor 218 module that is operative to group the sequential topics into broader topics. For example, the interaction engine 240 may evaluate two adjacent segments in a sequence to determine the similarity between the topics. By way of continuing the previous example, if segment A (e.g., comprising social actions 1 to 10) relates to the focused topic of “how to improve yield in the DRAM portion of a new processor X,” and adjacent segment B (e.g., comprising social actions 11 to 20) relates to the focused topic of “the reasons why the performance of processor X deteriorates after 200 hours of operation,” the event bounded topic extractor module 218 is operative to determine how related the two focused topics are (e.g., by way of a similarity function). If the similarity is above a predetermined threshold, then the two segments A and B are merged into a common segment having a broader topic (e.g., “chip reliability.) In one embodiment, the process of merging adjacent segments in a sequence continues until the similarity between segments is not above a predetermined threshold. The resulting topics are then provided by the interaction engine 240 by way of an output module 220. At block 808, the similarity of the topic between adjacent segments is determined. If the similarity is above a predetermine threshold (i.e., “YES” at decision block 818), then the process continues with block 820, where the adjacent segments having a topic similarity that is above the predetermined threshold, are merged into a common segment. The process then continues with block 826. However, if the similarity is not above the predetermined threshold (i.e., “NO” at decision block 818), the process goes directly to block 826, where a determination is made whether all segments in the sequence have been evaluated for similarity of topic. If not (i.e., “NO” at decision block 826), the process returns to block 808 until all segments have been evaluated in the present iteration. However, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626) in the present iteration, the process stops. In one embodiment, multiple iterations are performed. For example, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626), the process continues with block 830, where a determination is made whether any segments were merged in the present iteration. If not (i.e., “NO” at decision block 830), it is indicative that the segments were sufficiently merged, and the process stops. However, upon determining that one or more segments in the present iteration have been merged so (i.e., “YES” at decision block 830), it is indicative that the segments may be merged further, as discussed previously in the context of FIG. 6. Accordingly, the process returns to block 808 and the adjacent segments are evaluated for similarity. In one embodiment, the threshold for comparison is loosened between each iteration; p. 0044-0048 and 0064-0065).
Regarding claims 8, 16 and 20, Szymanski discloses a system wherein the at least one memory and the computer program code are configured to, with the at least one processor, further cause the at least one processor to:
identify an utterance in the identified topic segment with a highest utterance boundary score (The interaction engine 240 includes an event bounded topic extractor 218 module that is operative to group the sequential topics into broader topics. For example, the interaction engine 240 may evaluate two adjacent segments in a sequence to determine the similarity between the topics. By way of continuing the previous example, if segment A (e.g., comprising social actions 1 to 10) relates to the focused topic of “how to improve yield in the DRAM portion of a new processor X,” and adjacent segment B (e.g., comprising social actions 11 to 20) relates to the focused topic of “the reasons why the performance of processor X deteriorates after 200 hours of operation,” the event bounded topic extractor module 218 is operative to determine how related the two focused topics are (e.g., by way of a similarity function). If the similarity is above a predetermined threshold, then the two segments A and B are merged into a common segment having a broader topic (e.g., “chip reliability.) In one embodiment, the process of merging adjacent segments in a sequence continues until the similarity between segments is not above a predetermined threshold. The resulting topics are then provided by the interaction engine 240 by way of an output module 220.; p. 0044-0048); and
classify the identified utterance as a topic boundary, wherein the identified topic segments are updated based on classifying the identified utterance as a topic boundary (The interaction engine 240 includes an event bounded topic extractor 218 module that is operative to group the sequential topics into broader topics. For example, the interaction engine 240 may evaluate two adjacent segments in a sequence to determine the similarity between the topics. By way of continuing the previous example, if segment A (e.g., comprising social actions 1 to 10) relates to the focused topic of “how to improve yield in the DRAM portion of a new processor X,” and adjacent segment B (e.g., comprising social actions 11 to 20) relates to the focused topic of “the reasons why the performance of processor X deteriorates after 200 hours of operation,” the event bounded topic extractor module 218 is operative to determine how related the two focused topics are (e.g., by way of a similarity function). If the similarity is above a predetermined threshold, then the two segments A and B are merged into a common segment having a broader topic (e.g., “chip reliability.) In one embodiment, the process of merging adjacent segments in a sequence continues until the similarity between segments is not above a predetermined threshold. The resulting topics are then provided by the interaction engine 240 by way of an output module 220. At block 808, the similarity of the topic between adjacent segments is determined. If the similarity is above a predetermine threshold (i.e., “YES” at decision block 818), then the process continues with block 820, where the adjacent segments having a topic similarity that is above the predetermined threshold, are merged into a common segment. The process then continues with block 826. However, if the similarity is not above the predetermined threshold (i.e., “NO” at decision block 818), the process goes directly to block 826, where a determination is made whether all segments in the sequence have been evaluated for similarity of topic. If not (i.e., “NO” at decision block 826), the process returns to block 808 until all segments have been evaluated in the present iteration. However, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626) in the present iteration, the process stops. In one embodiment, multiple iterations are performed. For example, upon determining that all segments in the sequence have been evaluated for similarity of topic (i.e., “YES” at decision block 626), the process continues with block 830, where a determination is made whether any segments were merged in the present iteration. If not (i.e., “NO” at decision block 830), it is indicative that the segments were sufficiently merged, and the process stops. However, upon determining that one or more segments in the present iteration have been merged so (i.e., “YES” at decision block 830), it is indicative that the segments may be merged further, as discussed previously in the context of FIG. 6. Accordingly, the process returns to block 808 and the adjacent segments are evaluated for similarity. In one embodiment, the threshold for comparison is loosened between each iteration; p. 0044-0048 and 0064-0065). In addition, Joller discloses a system causing the processor to identify a topic segment of the identified topic segments that has a time length that exceeds a segment length threshold (FIG. 6 illustrates an example of content summarization consistent with certain embodiments of the present disclosure. As illustrated, a longer-form content file 600 and/or associated transcribed text may be analyzed to identify potentially less-informative content portions such as introductions, advertisements, and/or conclusions, as well as potentially more informative topics and/or associated content segments. Identified segments may be analyzed and/or scored consistent with segment relevance scoring processes described herein. Segments associated with a threshold score may be included in and/or otherwise combined into a shorter-form content summary 602. In some embodiments, segments associated with a threshold score may be included in the shorter-form content summary 602 regardless if other segments associated with the same topic are also included in the shorter-form content summary 602. In further embodiments, segments associated with the highest threshold score in each identified topic may be included in the shorter-form content summary 602. For example, in the illustrated content summarization process, the most relevant segment based on associated scoring from each topic included in the longer-form content file 600 may be included in the shorter-form content summary 602. Scoring used in connection with topic and/or segment identification for inclusion in a shorter-form content summary may take into account a variety of signals including, without limitation, one or more of a segment's semantic similarity to a topic and/or the associated longer-form and/or full length content; the freshness of content within a segment; a segment's relative importance as indicated by a number of important concepts, keywords, key phrases, and/or entities within a segment, time, and/or duration of the associated transcribed text; a segment's completeness as indicated by vocabulary overlap, co-references, speaker turns and/or identifies, lexical and/or grammatical patterns, and/or punctuation; a segment's overall quality based on cue phrases, time, duration, and/or the like of the associated transcribed text, speaker turns, audio features, etc.; and/or the like. In some embodiments, automated content summarization may be managed, at least in part, by one or more user specified conditions and/or parameters. For example, a user may set, among other things, a target size and/or format for shorter-form segments, highlights, and/or summaries, a number of topics included in generated summaries, a number of associated segments for each included topics, a threshold ranking, scoring, and/or relevance level for segment inclusion in a summary, and/or the like. Using various disclosed embodiments, the user may thus have shorter-form segments, highlights, and/or summaries “made to order” based on the target parameters.; p. 0044-0048, 0090-0092).
Claim(s) 5-7 and 13-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Szymanski in view of Joller and in further view of Li et al. (PGPUB 2016/0078339), hereinafter referenced as Li.
Regarding claims 5 and 13, Szymanski in view Joller disclose a system as described above, but does not specifically teach wherein the at least one memory and the computer program code are configured to, with the at least one processor, further cause the at least one processor to:
pre-train the deep learning model using general text training data; and
train the pre-trained deep learning model continually using communication transcript training data including at least one of the following: labeled communication transcript training data and unlabeled communication transcript training data.
Li discloses a system further cause the at least one processor to:
pre-train the deep learning model using general text training data (In one embodiment, initialization component 124 creates and initializes the untrained student DNN model by assigning random numbers to the weights of the nodes in the model (i.e. the weights of matrix W). In another embodiment, initialization component 124 receives from accessing component 122 data for pre-training the student DNN model, such as un-transcribed data that is used to establish initial node weights for the student DNN model. In some embodiments, initialization component 124 also initializes or creates the teacher DNN model. In particular, using labeled or transcribed data from data source(s) 108 provided by accessing component 122, initialization component 124 may create a teacher DNN model (which may be pre-trained), and provide the initialized but untrained teacher DNN model to training component 126 for training. Similarly, initialization component 124 may create an ensemble teacher model by determining a plurality of sub-DNN models (e.g., creating and handing off to training component 126 for training or identifying already existing DNN model(S)) to be included as members of the ensemble). In these embodiments, initialization component 124 may also determine the relationships between the output layer of the ensemble and the output layers of the member sub-DNN models (e.g. by taking a raw average of the member model outputs), or may provide the initialized but untrained ensemble teacher DNN to training component 126 for training. Training component 126 is generally responsible for training the student DNN based on the teacher. In particular training component 126 receives from initialization component 124 and/or accessing component 122, an untrained (or pre-trained) DNN model, which will be the student and trained DNN model, which will serve as the teacher. Training component 126 also receives un-labeled data for training the student DNN from accessing component 112.; p. 0031-0033, 0041-00043); and
train the pre-trained deep learning model continually using communication transcript training data including at least one of the following: labeled communication transcript training data and unlabeled communication transcript training data (As described above, an advantage of some embodiments described herein is that the student DNN model may be trained using un-labeled (or un-transcribed data) because its supervised signal (P.sub.L(s|x), as will be further described) is obtained by passing the un-labeled training data through the teacher DNN model. Because labeling (or transcribing) data for training costs time and money, a much smaller amount of labeled (or transcribed) data is available as compared to un-labeled data. Without the need for transcribed (or labeled) training data, much more data becomes available for training. With more training data available to cover a particular feature space, the accuracy of a deployed (student) DNN model is even further improved. This advantage is especially useful for industry scenarios with large amounts of un-labeled data available due to the deployment feed-back loop (wherein deployed models provide their usage data to application developers, who use the data to further tailor future versions of the application). For example, many search engines use such a deployment feed-back loop; p. 0016, 0019, 0022), to provide improvements to speech processing.
Therefore, it would have been obvious to one of ordinary skill of the art, before the effective filing date of the claimed invention, to modify the method as described above, to assist with further tailoring data.
Regarding claims 6 and 14, Szymanski in view of Joller disclose a system wherein training the pre-trained deep learning model continually includes:
identifying topic segments in communication transcript training data (Szymanski -the historical data repository 112 is configured to store and maintain a large set of historical data 113, sometimes referred to as massive data, which includes data related to prior conversations between various entities from which the interaction engine 103 can learn from. For example, the historical database 112 may provide training data related to conversations that have been successfully segmented and topics extracted therefrom. In one embodiment, the historical data 113 serves as a corpus of data from which the interaction engine 103 can learn from to create a sequential deep learning model that can then be used to evaluate various aspects of a conversation between one or more entities 102(1) to 102(N). During a training stage, the interaction engine 103 is configured to receive the historical data 113 over the network 106 to create models therefrom. These models can then be used by the interaction engine 103, during an active stage, to facilitate at least one of: (i) convert speech received in electronic data packages 105(1) to 105(N) to text, (ii) perform speech segmentation, (iii) classify each utterance in the speech into a social action, and (iv) group one or more social actions in a sequence into a segment. Each of these features is discussed in more detail below; p. 0031-0032, 0041, 0046); and
organizing the identified topic segments in a plurality of different orders to form a plurality of training transcripts (Szymanski - the historical data repository 112 is configured to store and maintain a large set of historical data 113, sometimes referred to as massive data, which includes data related to prior conversations between various entities from which the interaction engine 103 can learn from. For example, the historical database 112 may provide training data related to conversations that have been successfully segmented and topics extracted therefrom. In one embodiment, the historical data 113 serves as a corpus of data from which the interaction engine 103 can learn from to create a sequential deep learning model that can then be used to evaluate various aspects of a conversation between one or more entities 102(1) to 102(N). During a training stage, the interaction engine 103 is configured to receive the historical data 113 over the network 106 to create models therefrom. These models can then be used by the interaction engine 103, during an active stage, to facilitate at least one of: (i) convert speech received in electronic data packages 105(1) to 105(N) to text, (ii) perform speech segmentation, (iii) classify each utterance in the speech into a social action, and (iv) group one or more social actions in a sequence into a segment. Each of these features is discussed in more detail below; p. 0031-0032, 0041, 0046). In addition, Li discloses a system comprising training the pre-trained deep learning model using the plurality of training transcripts (e.g., creating and handing off to training component 126 for training or identifying already existing DNN model(S)) to be included as members of the ensemble). In these embodiments, initialization component 124 may also determine the relationships between the output layer of the ensemble and the output layers of the member sub-DNN models (e.g. by taking a raw average of the member model outputs), or may provide the initialized but untrained ensemble teacher DNN to training component 126 for training. Training component 126 is generally responsible for training the student DNN based on the teacher. In particular training component 126 receives from initialization component 124 and/or accessing component 122, an untrained (or pre-trained) DNN model, which will be the student and trained DNN model, which will serve as the teacher. Training component 126 also receives un-labeled data for training the student DNN from accessing component 112.; p. 0031-0033, 0041-00043).
Regarding claims 7 and 15, it is interpreted and rejected for similar reasons as set forth above. In addition, Li discloses a system wherein training the pre-trained deep learning model continually includes:
training a teacher model using the communication transcript training data (As described above, an advantage of some embodiments described herein is that the student DNN model may be trained using un-labeled (or un-transcribed data) because its supervised signal (P.sub.L(s|x), as will be further described) is obtained by passing the un-labeled training data through the teacher DNN model. Because labeling (or transcribing) data for training costs time and money, a much smaller amount of labeled (or transcribed) data is available as compared to un-labeled data. Without the need for transcribed (or labeled) training data, much more data becomes available for training. With more training data available to cover a particular feature space, the accuracy of a deployed (student) DNN model is even further improved. This advantage is especially useful for industry scenarios with large amounts of un-labeled data available due to the deployment feed-back loop (wherein deployed models provide their usage data to application developers, who use the data to further tailor future versions of the application). For example, many search engines use such a deployment feed-back loop; p. 0019-0022, 0026-0037);
initialize a student model with parameters of the teacher model (As described above, an advantage of some embodiments described herein is that the student DNN model may be trained using un-labeled (or un-transcribed data) because its supervised signal (P.sub.L(s|x), as will be further described) is obtained by passing the un-labeled training data through the teacher DNN model. Because labeling (or transcribing) data for training costs time and money, a much smaller amount of labeled (or transcribed) data is available as compared to un-labeled data. Without the need for transcribed (or labeled) training data, much more data becomes available for training. With more training data available to cover a particular feature space, the accuracy of a deployed (student) DNN model is even further improved. This advantage is especially useful for industry scenarios with large amounts of un-labeled data available due to the deployment feed-back loop (wherein deployed models provide their usage data to application developers, who use the data to further tailor future versions of the application). For example, many search engines use such a deployment feed-back loop; p. 0019-0022, 0026-0037); and
train the student model with output from the teacher model using a self-distillation process, wherein the trained student model is used as the deep learning model (Initially, student DNN 301 is untrained or may be pre-trained, but has not yet been trained by the teacher DNN. In an embodiment, system 300 may be used to learn student DNN 301 from teacher DNN 301 using an iterative process until the output distribution 351 of student DNN 301 converges (or otherwise approximates) output distribution 352 of teacher DNN 302. In particular, for each iteration, a small piece of unlabeled (or un-transcribed) data 310 is provided to both student DNN 301 and teacher DNN 302. Using forward propagation, the posterior distribution (output distribution 351 and 352) is determined. An error signal 360 is then determined from the distribution 351 and 352. The error signal may be calculated by determining the KL divergence between distributions 351 and 352, or by using regression, or other suitable technique, and may be determined using evaluating component 128 of FIG. 1. (The term signal as in “error signal” is a term of the art and does not mean that the error signal comprises a transitory signal such as a propagated communications signal. Rather, in some embodiments, error signal comprises a vector.) Embodiments that determine the KL divergence provide an advantage over other alternatives such as regression because minimizing the KL divergence is equivalent to minimizing the cross entropy of the distributions, as further described in method 500 of FIG. 5. If the output distribution 351 of student DNN 301 has converged with the output distribution 352 of teacher DNN 302, then the student DNN is deemed to be trained. However, if the output has not converged, and in some embodiments the output still appears to be converging, then the student DNN 301 is trained based on the error. For example, as shown at 370, using back propagation the weights of student DNN 301 are updated using the error signal.; p. 0043, 0057, wherein the self-distillation is the machine learning technique where a neural network acts as both teacher and student, using its own internal predictions or deeper layers to refine its learning).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. This information has been detailed in the PTO 892 attached (Notice of References Cited).
Saggi et al. (PGPUB 2020/0372066) teaches identifying transition in topics, summarizing data and identifying time lengths (p. 0010).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAKIEDA R JACKSON whose telephone number is (571)272-7619. The examiner can normally be reached Mon - Fri 6:30a-2:30p.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at 571.272.5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAKIEDA R JACKSON/Primary Examiner, Art Unit 2657