Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
2. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
3. Claims 1-2 and 4 are rejected under 35 U.S.C. 103 as being unpatentable over Hsu et al., “HIERARCHICAL GENERATIVE MODELING FOR CONTROLLABLE SPEECH SYNTHESIS,” herein Hsu, in view of Kang et al., “MULTI-DISTRIBUTION DEEP BELIEF NETWORK FOR SPEECH SYNTHESIS,” herein Kang.
Regarding Claim 1:
Hsu discloses a method for generating information such as an image or speech (Hsu: Section 1 discloses a neural text-to-speech method that generates a sequence of speech frames and converts the generated frames into speech), comprising:
providing a storage unit storing a set of instructions, and a processing unit for executing said set of instructions (Hsu: Section 2.4 and Fig. 2 discloses executing the trained synthesizer, decoder and neural vocoder to predict speech features and convert the predicted features into a time-domain speech waveform),
said set of instructions including a generator routine having a decoder constituted by a statistically learned model (Hsu: Section 1 and Section 2 and Section 2.4 disclose a trained neural speech generator having an autoregressive speech decoder that predicts acoustic features frame by frame which is a statistically learned model),
said decoder being defined by an observable variable and a decoder hierarchy of a set of random variables, said decoder hierarchy constituted by layers having at least one random variable from said set of random variables in each layer (Hsu: Abstract, Section 2.1 and Fig. 2 discloses observable speech frames and two hierarchical latent random variables: top categorical variable yl and zl and lower continuous variable. Fig. 1 places the variables at respective levels of the generative hierarchy),
said observable variable, and said set of random variables being jointly distributed according to a prior probability distribution (Hsu: Section 2.1 and Eq(1) discloses factorizing the joint generative probability distribution into the speech frame distribution, the conditional distribution of the lower latent variable and the prior distribution of the top latent variable),
said prior probability distribution being factorized having:
- a first factor defined as a first probability distribution of said observable variable conditioned on at least one random variable from said set of random variables (Hsu: Section 2.1 and Eq. (1) discloses a probability distribution for observable speech frame sequence (X) conditioned on lower latent random variable zl.),
- a second factor defined as a second probability distribution of the random variable of the top layer of said decoder (Hsu: Section 2.1 and Eq.(1) first paragraph disclose a prior probability distribution for top layer categorical random variable yl. Because Hsu employs a two level hierarchy, the sequence below the top layer contains one conditional probability factor),
- a third factor defined as the product of sequence of the probability distributions for the random variables of said set of random variables, the random variable of each respective element in said product of sequence being conditioned on (Hsu: Section 2.1 and Eq.(1) first paragraph disclose lower latent random variable zl, conditioned on higher latent random variable yl, Hsu therefore discloses the claimed higher to lower conditional relationship, but does not explicitly disclose conditioning the lower random variable on at least two higher random variables),
said method further comprising sampling a value of the random variable of the top layer, and processing said value through said hierarchy such that said information is generated (Hsu: Section 2.1 and Eq. (1) disclose first sampling top categorical variable yi, using the sampled value to generate lower latent variable zl. Hsu: Section 2.4 discloses processing the resulting representation through the speech decoder and neural vocoder to generate a time-domain speech waveform).
Hsu does not explicitly disclose:
at least two.
However, Kang discloses:
at least two (Kang: Section 1 discloses a speech synthesis deep belief network containing multiple stochastic hidden variables, Section 2.1 and Eq (3) disclose that the probability assigned to each individual random variable in a lower layer is determined from the combined weighted contributions of the random variables in the higher layer. Section 4.1 discloses 2,000 random units in each hidden layer. Accordingly, each lower layer random variable is conditioned on substantially more than two higher layer random variables, Section 3.2 further discloses recursively passing the resulting probabilities downward through the hierarchy to generate the speech parameters used to produce the final speech signal).
Hsu and Kang are in the same field of endeavor, i.e., both disclose systems and generating speech using probabilistic, hierarchical, statistically learned models. Hsu conditions its lower latent representation on a single higher latent variable. Kang teaches a speech synthesis hierarchy in which each lower random variable receives the combined contributions of multiple of multiple higher layer random variable. A person of ordinary skill in the art would have been motivated to modify Hsu’s lower latent layer so that each lower random variable receives the combined contributions of multiple higher layer random variables as taught by Kang. Kang explicitly states that its deep belief network “model the speech parameters including spectrum and F0 simultaneously and generate these parameters from DBN for speech synthesis” in the Abstract.
Regarding Claim 2:
The proposed combination of Hsu and Kang discloses method according to claim 1, the random variable of each respective element in said product of sequence being conditioned on the random variables in the higher layers (Hsu: Section 2.1 and Eq (1) discloses lower latent random variable zl conditioned on higher latent random variable yi. Hsu therefore disclose condition a lower random variable on a random variable in a higher layer, but does not explicitly disclose conditioning on the random variables in the higher layer. Kang: ¶2.1 discloses that the probability assigned to each individual random variable in a lower layer is determined from the combined weighted contribution of all the random variables in the higher layer).
Hsu and Kang are in the same field of endeavor, i.e., both disclose systems and generating speech using probabilistic, hierarchical, statistically learned models. Hsu conditions its lower latent representation on a single higher latent variable. Kang teaches a speech synthesis hierarchy in which each lower random variable receives the combined contributions of multiple of multiple higher layer random variable. A person of ordinary skill in the art would have been motivated to modify Hsu’s lower latent layer so that each lower random variable receives the combined contributions of multiple higher layer random variables as taught by Kang. Kang explicitly states that its deep belief network “model the speech parameters including spectrum and F0 simultaneously and generate these parameters from DBN for speech synthesis” in the Abstract.
Regarding Claim 4:
Hsu and Kang further disclose the method according to claim 1, said decoder comprising a deterministic variable for summarizing information from random variables higher in said decoder hierarchy (Kang: Section 2.1 discloses calculating the probability of each lower layer random variable by combining the weighted values of the random variables in the higher layer and applying a deterministic activation function. The resulting calculated probability is a deterministic variable that summarizes the combined information received from the higher layer random variables).
Hsu and Kang are in the same field of endeavor, i.e., both disclose systems and generating speech using probabilistic, hierarchical, statistically learned models. Hsu conditions its lower latent representation on a single higher latent variable. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to disclose Kang’s deterministic summary calculation into Hsu’s decoder. Kang explicitly states that its deep belief network “model the speech parameters including spectrum and F0 simultaneously and generate these parameters from DBN for speech synthesis” in the Abstract.
4. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Hsu in view of Kang and further in view of Maaløe et al. “Auxiliary Deep Generative Models” herein Maaløe .
Regarding Claim 3:
The proposed combination of Hsu and Kange further discloses the method according to claim 1, except:
the random variables of said set of random variables being divided into a bottom-up path and a top-down path.
However, Maaløe, discloses:
the random variables of said set of random variables being divided into a bottom-up path and a top-down path (Maaløe: Section 4.3 and Eqs. 19-20 discloses dividing the multilayer variables into an upward auxiliary path in which each variable is generated based on the preceding lower layer variable and observable x, and a downward latent variable path in which each variable is generated based on the higher layer variable).
Hsu and Kang in view of Maaløe are combinable because they are in the same field of endeavor, i.e., both disclose systems or methods for deep generative modeling using hierarchies of stochastic latent variables. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Hsu to divide its hierarchical random variables into the upward auxiliary variable path and downward latent variable path taught by Maaløe. This creates a more expressive model that converges faster and produces improved results, Maaløe explicitly states that “more expressive and properly specified deep generative models converge faster with better results” in the Abstract.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IAN SCOTT MCLEAN whose telephone number is (703)756-4599. The examiner can normally be reached "Monday - Friday 8:00-5:00 EST, off Every 2nd Friday".
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IAN SCOTT MCLEAN/Examiner, Art Unit 2654
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654