Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
2. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
3. Claims 1, 5 and 7, are rejected under 35 U.S.C. 103 as being unpatentable over Li (US 2021/0104228) in view of Wan (US 2022/0358288).
Regarding Claim 1:
Li discloses a model learning apparatus (Li: ¶7 discloses a system that includes data processing hardware and memory hardware) comprising:
processing circuitry configured train an utterance feature reconstruction model that is a neural network model (Li: ¶31-34 discloses unsupervised loss where an encoder receives as input a sequence of input audio features/vectors, the encoders include neural networks) that randomly selects some of utterance feature (Li: ¶29, discloses the training process trains the auditor speech recognition model using training data that includes un-transcribed speech and a set of transcribed speech utterances, ¶32 discloses the feature encoder receives as input audio features/vectors which correspond to the transcribed and transcribed utterances, ¶33 discloses the latent speech representations output from the feature encoder may be fed to a masking module where some of the latent speech representations are randomly chosen. Li therefore teaches training a neural network model that receives utterance features, masks selected feature representations and learns to predict the feature representation corresponding to each masked position) and replace the selected utterance feature (Li: ¶33 discloses a shared trained feature vector is predetermined masking information. Therefore Li teaches randomly selecting speech feature representations and replacing them with predetermined masking information), and
estimate utterance features of the masked utterance feature (Li ¶32 converts the utterance acoustic feature sequence into latent speech representations at respective time steps, ¶33 randomly masks some of those latent representations, ¶34 generates a contrastive context vector for each masked latent representation ¶35 discloses a quantizer receives the latent representation and produces a target quantized vector and target token index from the unmasked latent vectors to make them learnable training targets, ¶37-39 discloses the masked context vectors are used to generate a high level context vector representing a prediction of the target token index and train the model by comparing the prediction associated with the masked position against the target derived from the corresponding unmasked latent speech representation, therefore estimating the utterance feature corresponding to the masked position); and
output the trained utterance feature reconstruction model as an unsupervised pre-trained model (Li ¶17-19 discloses training using unlabeled training data and self-supervised techniques, i.e., training the model using an unsupervised/self-supervised technique and producing a pretrained checkpoint representing the model at the completion of that unsupervised training stage. The pretrained checkpoint teaches outputting the trained model as an unsupervised pretrained model).
Li does not explicitly disclose the randomly selected utterance feature is an entire sequence that are sequences of utterance features. Rather Li discloses utterances that are sequenced into multiple frames, wherein individual frames of each utterance are masked, rather than an entire utterance.
However, Wan discloses utterance feature sequences that are sequences of utterance features (Wan: teaches performing masked context regression by randomly selecting and masking an utterance within a dialogue and predicting an encoding vector for the masked utterance. Therefore, Wan teaches applying the masking and reconstruction process at an utterance level rather than at the individual per frame step within an individual utterance).
Li and Wan are combinable because they are from the same field of endeavor, autoencoding of natural language data, i.e., both disclose inputting speech or text data into an encoder, masking portions of the inputting and training a model to predict these masked portions. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to disclose pretraining process to randomly mask and reconstruct utterance level feature representations as taught in Wan in order to learn contextual representations of utterances within a dialogue or subsequent dialogue related tasks as disclosed in ¶61 of Wan: “Model training data is adjustable based on the eventual use of the model's output. Some non-limiting examples of eventual uses of the model's output are to perform masked language modeling (using context words surrounding a mask token, or blank to be filled in, to try to predict what word should replace the mask token), masked context regression (masking a randomly-selected utterance and predicting an encoding vector for the masked utterance).
Regarding Claim 5:
Claim 5 has been analyzed with regard to claims 5 (see rejection above) and is rejected for the same reasons of obviousness as used above.
Regarding Claim 7:
The proposed combination of Li and Wan further discloses a non-transitory computer readable medium storing a computer program for causing a computer to function as the model learning apparatus according to claim 1 (Li: ¶51 discloses memory 520 stores information non-transitorily within the computing device 500. The memory 520 may be a computer-readable medium, a volatile memory unit(s), or non-volatile memory unit(s)).
4. Claims 2 and 3, are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Wan and further in view of Ando (US 2020/0152178).
Regarding Claim 2:
The proposed combination of Li in view of Wan further discloses the model learning apparatus according to claim 1, comprising
by using parameters of the utterance feature reconstruction model as initial values of model parameters (Li: ¶30-31 discloses a unsupervised training method which may be used to train with masking techniques for predicting features; Wan: ¶330 teaches the training module trains the untrained encoder model using dialogue training corpus to produce a trained encoder model which then uses dialogue task specific training data to further train the model)
Li and Wan are from the same field of endeavor of autoencoding of natural language data, i.e., both disclose inputting speech or text data into an encoder, masking portions of the inputting and training a model to predict these masked portions. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the parameters learned through Li’s unsupervised masked-speech training, as modified by Wan’s utterance-level masked context regression, as the initial parameter values for subsequent task specific training. Wan explicitly teaches: “Using, as the base set of parameters, those of already-trained token embedding and token self-attention portions saves training time by starting from a partially trained state.” In ¶60 Therefore, using the trained parameters of the Li and Wan utterance feature reconstruction model as initial model parameters would have predictably reduced training time by beginning from already learned contextual speech and dialogue representations rather than randomly initialized parameters.
The proposed combination of Li in view of Wan does not explicitly disclose
a satisfaction level estimation model learning unit configured to perform supervised learning on a satisfaction level estimation model
that is a model for estimating an utterance satisfaction level and a conversation satisfaction level
and using utterance feature sequences and corresponding utterance satisfaction level labels and conversation satisfaction level labels as learning data
However Ando discloses:
a satisfaction level estimation model learning unit configured to perform supervised learning on a satisfaction level estimation model (Ando: ¶29, ¶36 and ¶39-41 discloses training the model by propagating the conversation satisfaction label and the speech satisfaction label corresponding to each speech and learns the satisfaction estimation model)
that is a model for estimating an utterance satisfaction level and a conversation satisfaction level (Ando: ¶17 discloses the invention is speech satisfaction of a model)
and using utterance feature sequences (Ando ¶31-32 and ¶36 discloses the ordered feature quantities corresponding to the respective speeches or utterances) and corresponding utterance satisfaction level labels and conversation satisfaction level labels as learning data (Ando: ¶29 discloses each speech satisfaction label corresponds to a respective speech or utterance in the conversation and therefore teaches the claimed utterance satisfaction level levels, further teaching using both the conversation satisfaction label and the speech satisfaction labels to learn the satisfaction estimation model).
Li, Wan and Ando are combinable because they are pertinent to fields. Each concern neural network processing of speech or dialogue sequences and the training of models to derive contextual information or predictions from utterance-level data. Ando is also reasonably pertinent to the problem addressed by Li and Wan because Ando uses features extracted from respective utterances of a target speaker to train a downstream model that estimates both utterance level and conversation level information. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination of Li and Wan’s parameters using Ando’s utterance-feature sequences and corresponding utterance and conversation satisfaction labels/ Ando Explicitly teaches: “Learning a model that simultaneously estimates the conversation satisfaction and the speech satisfaction as a single model simultaneously and integrally contributes to the improvement in the estimation accuracy” in ¶22. The combination therefore would have predictably used the pretrained con textual representations of LI and Wan to initialize Ando’s downstream satisfaction model while obtaining Ando’s stated benefit of improved utterance and conversation satisfaction estimation accuracy.
Regarding Claim 3:
The model learning apparatus according to claim 1, wherein the utterance features are any of (Wan: ¶45 teaches linguistic features by processing natural language dialogue tokens, Wan: ¶42, ¶132 also teaches conversation features through metadata describing the relationships among utterances in a dialogue, number of turns, number of conversational turns and the amount of time elapsed).
Li and Wan are from the same field of endeavor autoencoding of natural language data, i.e., both disclose inputting speech or text data into an encoder, masking portions of the inputting and training a model to predict these masked portions. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use Wan’s linguistic and conversation features, including natural language token embeddings and metadata representing conversation turns, speakers, timestamps, elapsed time and turn relationships. Wan recognizes that conventional transformer architectures do not explicitly account for relationships between tokens supplied by metadata and teaches incorporating such metadata into the model to represent dialogue context in ¶61.
The proposed combination of Li and Wan do not explicitly disclose:
prosody features.
However, Ando discloses:
prosody features (Ando: ¶32 discloses prosodic features).
Li, Wan and Ando are analogous art because each concerns neural-network analysis of speech or dialogue using feature extracted from respective utterances and Ando is reasonably pertinent to the problem of selecting utterance features useful for estimating speaker related information from a conversation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the Li and Wan combined model to include Ando’s prosodic features, such as speech speed and final phoneme duration, because such features provide information regarding how an utterance was spoken that is not necessarily represented by Wan’s linguistic and conversation features. Ando specifically teaches that conversation satisfaction may be estimated using “a feature of a speaking style such as a speech speed” and that speech satisfaction may be estimated using features such as “a voice tone” and a backchannel frequency in ¶3.
5. Claims 4 and 6, are rejected under 35 U.S.C. 103 as being unpatentable over Ando in view of Wan.
Regarding Claim 4:
Ando discloses a satisfaction estimation apparatus comprising processing circuitry configured to a satisfaction level estimation unit configured to estimate an utterance satisfaction level and a conversation satisfaction level corresponding to an utterance of a target speaker on the basis of a satisfaction level estimation model (Ando: ¶47-49 discloses a trained model may acquire target speech of a target speaker and extract a feature quantity for each speech and inputting the feature quantities followed by estimating the satisfaction of each speech and the satisfaction of the overall conversation)
and using utterance feature sequences and corresponding utterance satisfaction level labels and conversation satisfaction level labels as learning data.
Ando does not explicitly disclose:
trained by using, as initial values of model parameters, parameters of an utterance feature reconstruction model that is a neural network model that randomly selects some of utterance feature sequences that are sequences of utterance features corresponding to respective utterances of a target speaker and replaces the selected utterance feature sequences with predetermined masking information to mask the utterance feature sequences.
and estimate utterance features of the masked utterance feature sequences
However Wan discloses:
trained by using, as initial values of model parameters, parameters of an utterance feature reconstruction model that is a neural network model that randomly selects some of utterance feature sequences that are sequences of utterance features corresponding to respective utterances of a target speaker and replaces the selected utterance feature sequences with predetermined masking information to mask the utterance feature sequences (Wan: ¶60-62 discloses saving training time by starting from a partially trained state, and that it is a neural network model that randomly selects some of utterance feature sequences to mask the utterance features sequences. Wan ¶45-46 teaches this model is a transformer neural network and that it uses masked context regression to randomly select utterance and predicting an encoding vector for the masked utterance)
and estimate utterance features of the masked utterance feature sequences (Wan: ¶61 the predicted encoding vector for the masked utterance is the estimated feature representation of that masked utterance).
Ando and Wan constitute analogous art because they are in the same field of endeavor, i.e., both disclose systems and methods for training neural network models that process utterances within conversations to produce utterance or conversation level output it would have been obvious to one ordinary skill in the art before the effective filing date of the claimed invention to modify Ando’s satisfaction estimation model to use, as its initial model parameters, parameters learned by Wan’s pretrained masked utterance reconstruction model before training the model using Ando’s utterance and conversation satisfaction labels. Wan explicitly provides the motivation for doing this modification by teaching that “Using, as the base set of parameters, those of already-trained token embedding and token self-attention portions saves training time by starting from a partially trained state” in ¶60. Therefore the modification would have predictably reduced the time required to train Ando’s satisfaction estimation model by beginning with parameters that had already learned contextual representations of conversational utterances.
Regarding Claim 6:
Claim 6 has been analyzed with regard to claims 4 (see rejection above) and is rejected for the same reasons of obviousness as used above.
Conclusion
The following prior art is not used in the rejection but considered pertinent to Applicant’s disclosure:
Rosenberg et al. US 2023/0317059 discloses a method includes receiving training data that includes unspoken textual utterances, un-transcribed non-synthetic speech utterances, and transcribed non-synthetic speech utterances. Each unspoken textual utterance is not paired with any corresponding spoken utterance of non-synthetic speech. Each un-transcribed non-synthetic speech utterance not paired with a corresponding transcription. Each transcribed non-synthetic speech utterance paired with a corresponding transcription. The method also includes generating a corresponding alignment output for each unspoken textual utterance of the received training data using an alignment model. The method also includes pre-training an audio encoder on the alignment outputs generated for corresponding to the unspoken textual utterances, the un-transcribed non-synthetic speech utterances, and the transcribed non-synthetic speech utterances to teach the audio encoder to jointly learn shared speech and text representations.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IAN SCOTT MCLEAN whose telephone number is (703)756-4599. The examiner can normally be reached "Monday - Friday 8:00-5:00 EST, off Every 2nd Friday".
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IAN SCOTT MCLEAN/Examiner, Art Unit 2654
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654