DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-11 are pending and have been examined.
Priority
Acknowledgment is made of applicant's claim for foreign priority under 35 U.S.C. 119(a)-(d). The present application claims priority to Korean patent application KR 10-2024-0077449, filed June 14, 2024. The effective filing date of the claimed invention is therefore June 14, 2024 for all claims. Certified copies of the priority document have not been received; applicant is advised to ensure that the foreign priority claim is perfected under 37 CFR 1.55.
Specification
The disclosure is objected to because of the following informalities:
a. In paragraph [0004], "phenome expressions" should apparently read "phoneme expressions."
b. In paragraph [0006], "Montoreal Forced Aligner" should apparently read "Montreal Forced Aligner."
c. In paragraph [0009], "variational interference" should apparently read "variational inference."
d. In paragraphs [0009], [0024], and [0054], "banding" appears where "bending" is apparently intended.
e. In paragraph [0057], "MICI pitch" should apparently read "MIDI pitch."
f. In paragraph [0062], "in the at" should apparently read "in the art."
Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
This application includes one or more claim limitations that do not use the word "means," but are nonetheless being interpreted under 35 U.S.C. 112(f) because the claim limitation uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation is: "a monotonic alignment search module configured to perform monotonic alignment search by using information on MIDI duration to extract phoneme duration" in claims 1 and 11. The term "module" is a generic placeholder for "means"; it is coupled with the functional language "configured to perform monotonic alignment search by using information on MIDI duration to extract phoneme duration"; and the placeholder is not preceded by a structural modifier, because "monotonic alignment search" identifies the function performed rather than the structure that performs it.
Because this claim limitation is being interpreted under 35 U.S.C. 112(f), it is being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. The corresponding structure is the algorithm described at paragraphs [0044]-[0046] and [0052]-[0053] and illustrated in Figures 4 and 5: dividing the phoneme sections by using the information on MIDI duration, performing the monotonic alignment search within each divided section so as to maximize the likelihood between the prior probability distribution and the posterior probability distribution, and extracting the duration of each phoneme independently. Because the specification discloses this algorithm as the corresponding structure, no rejection under 35 U.S.C. 112(b) follows from this interpretation. If applicant does not intend to have this limitation interpreted under 35 U.S.C. 112(f), applicant may amend the claim so that it expressly recites sufficient structure to perform the claimed function, or present a sufficient showing that the claim limitation recites sufficient structure to perform the claimed function.
The terms "prior encoder," "posterior encoder," "flow," and "decoder" recited in claims 1, 10, and 11 are not being interpreted under 35 U.S.C. 112(f). These terms are understood in this art as names for classes of structures rather than as generic placeholders; consistent with that understanding, claim 7 and paragraph [0038] of the specification recite the structure of the prior encoder as a text encoder and a projection layer.
Claim construction: the term "acoustic features" is expressly defined by claim 8 and by paragraphs [0019] and [0042] of the specification as a linear spectrogram or a Mel-spectrogram. The term is given that definition throughout this Office action.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-11 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 (Statutory Category). Claim 1 is directed to a system (a machine), claim 10 is directed to a method (a process), and claim 11 is directed to a system (a machine). Each independent claim is therefore directed to a statutory category of invention.
Step 2A, Prong One (Recitation of a Judicial Exception). Independent claims 1, 10, and 11 recite, in pertinent part, receiving phonemes converted from a text and outputting a prior probability distribution, receiving acoustic features and outputting a posterior probability distribution, converting, by a flow, the probability distribution to simplify the posterior probability distribution, and performing monotonic alignment search by using information on MIDI duration to extract phoneme duration. Under their broadest reasonable interpretation, these limitations recite mathematical concepts, that is, mathematical relationships, mathematical formulas or equations, and mathematical calculations. The prior probability distribution and the posterior probability distribution are mathematical objects defined by their parameters (a mean and a variance). The flow is an invertible change-of-variables transformation applied to a probability distribution, which is a mathematical operation. Monotonic alignment search is a dynamic-programming optimization that selects, over a matrix of likelihood values, the monotonic path that maximizes a likelihood, and extracting each phoneme duration from the selected path is a mathematical calculation performed on the entries of the alignment matrix. Using information on MIDI duration in the search further specifies the numeric inputs of the same mathematical optimization. The claims therefore recite an abstract idea.
Step 2A, Prong Two (Integration into a Practical Application). The judicial exception is not integrated into a practical application. Beyond the mathematical concepts, the independent claims recite the prior encoder, the posterior encoder, the flow, and the monotonic alignment search module as components of a generically recited machine-learning architecture that serves as a tool to perform the mathematical concepts, and the receipt of the text phonemes, the acoustic features, and the MIDI information is insignificant extra-solution activity in the nature of gathering the data on which the mathematics operates. Claim 11 recites no output beyond the extracted phoneme duration: the claim receives data, performs the mathematical operations, and outputs the numeric result of those operations, and therefore recites no additional element that applies the mathematical concepts beyond generally linking them to a computing environment. Claims 1 and 10 additionally recite a decoder that outputs a voice digital signal reflecting the extracted phoneme duration. This additional element is expressly considered and does not integrate the mathematical concepts into a practical application: outputting a waveform from the duration-reflecting input is insignificant extra-solution activity in the nature of outputting the result of the mathematical alignment, performed by a generically recited decoder, and amounts to mere instructions to apply the mathematical concepts using a generic component. The claims do not recite any improvement to the functioning of a computer or to any other technology or technical field; the asserted advance lies in the mathematical alignment procedure itself, that is, constraining and partitioning the alignment optimization with additional numeric score inputs, and not in any improvement to the way the recited components operate. The encoder, flow, and decoder architecture surrounding the recited mathematics is acknowledged in the specification as related art (paragraph [0051], describing the sentence-level monotonic alignment search of the related art). Accordingly, the additional elements do not integrate the abstract idea into a practical application, and the claims are directed to the abstract idea.
Step 2B (Inventive Concept). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration, the encoders, the flow, the module, and the decoder are recited at a high level of generality and perform well-understood, routine, and conventional functions of receiving data, transforming representations, and outputting a synthesized waveform. The five-component arrangement of prior encoder, posterior encoder, flow, monotonic alignment search, and decoder within which the recited mathematics is performed is acknowledged as related art in the specification (paragraph [0051]) and is evidenced by Kim of record, discussed below, which discloses that arrangement. Considered individually and as an ordered combination, the additional elements amount to no more than mere instructions to apply the mathematical concepts using generic components and therefore do not provide an inventive concept. The claims are not patent eligible.
The dependent claims have been considered and do not cure the deficiencies of the independent claims. Each dependent claim is addressed separately below.
Claim 2 recites that the prior encoder is configured to additionally receive MIDI pitch and MIDI duration as input, in addition to phonemes converted from lyrics, to perform the monotonic alignment search by using the MIDI duration information. This limitation further specifies the numeric data supplied to the mathematical operations and is insignificant extra-solution activity in the nature of data gathering; it adds no additional element beyond those addressed above and is neither integrated into a practical application nor significantly more.
Claim 3 recites that the information inputted to the prior encoder is information in which a text, pitch, and duration of the MIDI corresponding to each phoneme are mapped. Organizing the numeric inputs at the phoneme level prepares the data for the mathematical alignment computation and further describes the data on which the mathematics operates; it is neither integrated into a practical application nor significantly more.
Claim 4 recites dividing phoneme sections by using the MIDI duration information and then performing the monotonic alignment search for each phoneme section. Partitioning a sequence into sections according to numeric duration values is an arithmetic operation, and performing the search for each section is the same mathematical optimization applied over each partition; the claim further narrows the abstract mathematical procedure and supplies no additional element.
Claim 5 recites performing the monotonic alignment search between the posterior probability distribution and the prior probability distribution in every phoneme section. This is the mathematical optimization itself, restated at the level of its operands, and supplies no additional element.
Claim 6 recites dividing the respective phoneme sections and independently extracting phoneme duration for all phonemes. Independent evaluation of each partition further describes the execution of the mathematical procedure and supplies no additional element.
Claim 7 recites that the prior encoder includes a text encoder and a projection layer. These are generically recited neural-network components employed as tools to perform the mathematical concepts; naming the components does not integrate the mathematics into a practical application and is not significantly more.
Claim 8 recites that the acoustic features are a linear spectrogram or a Mel-spectrogram. A spectrogram is itself a mathematical representation of a signal, and specifying the form of the input data limits only the data on which the mathematics operates; it is neither integrated into a practical application nor significantly more.
Claim 9 recites that the decoder receives the posterior probability distribution as input when learning and receives the prior probability distribution undergoing inverse transformation on the probability distribution as input when inferring, outputting a waveform in each case. The inverse transformation is the same change-of-variables mathematics performed in reverse, and routing inputs according to an operating mode is a generic implementation detail; the waveform output is addressed above with respect to claims 1 and 10; the claim is neither integrated into a practical application nor significantly more.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3 and 7-11 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al., "Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech," arXiv:2106.06103v1 (made publicly available June 11, 2021), hereinafter Kim, in view of Zhang et al., "VISinger: Variational Inference with Adversarial Learning for End-to-End Singing Voice Synthesis," arXiv:2110.08813v2 (made publicly available February 24, 2022), hereinafter Zhang.
Each of Kim and Zhang was publicly available before the June 14, 2024 effective filing date of the claimed invention and qualifies as prior art under 35 U.S.C. 102(a)(1). Kim, the first-named author of the primary reference, is Jaehyeon Kim of Kakao Enterprise and is not the named inventor Tae Woo Kim of the present application.
Regarding claim 1, Kim discloses:
A singing voice phoneme duration extraction system using a MIDI, comprising: (Kim discloses an end-to-end speech synthesis system that extracts phoneme duration: "The overall architecture of the proposed model consists of a posterior encoder, prior encoder, decoder, discriminator, and stochastic duration predictor" (Kim, Sec. 2.5; Fig. 1(a)/(b)); the phoneme duration extraction is quoted with the monotonic alignment search limitation below.)
a prior encoder configured to receive phonemes converted from a text as input, and to output a prior probability distribution; (Kim discloses that "The prior encoder consists of a text encoder that processes the input phonemes ctext and a normalizing flow," the phonemes being converted from text: "The input condition of the prior encoder c is composed of phonemes ctext extracted from text" and "We convert text sequences to IPA phoneme sequences using open-source software" (Kim, Sec. 2.5.2; Sec. 2.1.3; Sec. 3.2). Kim discloses that the prior encoder outputs the prior probability distribution: the hidden representation is obtained through the text encoder and "a linear projection layer above the text encoder that produces the mean and variance used for constructing the prior distribution" (Kim, Sec. 2.5.2; Sec. 2.1.3, Eq. (4)).)
a posterior encoder configured to receive acoustic features as input and to output a posterior probability distribution; (Kim discloses that "The linear projection layer above the blocks produces the mean and variance of the normal posterior distribution" (Kim, Sec. 2.5.1) and that "We use linear spectrograms which can be obtained from raw waveforms through the Short-time Fourier transform (STFT), as input of the posterior encoder" (Kim, Sec. 3.2; Sec. 2.1.3, Eq. (3)). Under the express definition of "acoustic features" set forth in claim 8 and paragraphs [0019] and [0042], stated in the Claim Interpretation section above, the linear-scale spectrogram input of Kim is the recited acoustic features.)
a flow configured to convert the probability distribution to simplify the posterior probability distribution; (Kim discloses "a normalizing flow ..., which allows an invertible transformation of a simple distribution into a more complex distribution following the rule of change-of-variables" (Kim, Sec. 2.1.3, Eq. (4)), the flow being "a stack of affine coupling layers" (Kim, Sec. 2.5.2). Per Eq. (4), the flow operates on the latent representation sampled from the posterior, which is the interpretation of "the probability distribution" stated in the Claim Interpretation section above.)
a monotonic alignment search module configured to perform monotonic alignment search by using information on MIDI duration to extract phoneme duration; and (Kim discloses "Monotonic Alignment Search (MAS) ..., a method to search an alignment that maximizes the likelihood of data parameterized by a normalizing flow," in which "the candidate alignments are restricted to be monotonic and non-skipping" (Kim, Sec. 2.2.1, Eq. (5)). Kim redefines MAS "to find an alignment that maximizes the ELBO, which reduces to finding an alignment that maximizes the log-likelihood of the latent variables z," that is, the search is performed between the flow-transformed posterior representation and the prior (Kim, Sec. 2.2.1, Eq. (6)). Kim extracts each phoneme duration from the search result: "We can calculate the duration of each input token di by summing all the columns in each row of the estimated alignment" (Kim, Sec. 2.2.2). This limitation is interpreted under 35 U.S.C. 112(f) as set forth in the Claim Interpretation section above.)
a decoder configured to output a waveform which is a voice digital signal, based on input reflecting a result of extracting the phoneme duration. (Kim discloses "We upsample the latent variables z to the waveform domain ... through a decoder" (Kim, Sec. 2.1.2), the decoder being "essentially the HiFi-GAN V1 generator" (Kim, Sec. 2.5.3), where the alignment reflected in the latent input represents "how long each input phoneme expands to be time-aligned with the target speech" (Kim, Sec. 2.1.3).)
Kim does not disclose that the phoneme duration extraction system is a singing voice phoneme duration extraction system using a MIDI (claim 1, preamble), and does not disclose performing the monotonic alignment search by using information on MIDI duration.
Zhang discloses information on MIDI duration provided at the phoneme level as an input to a singing voice synthesis system having the architecture of Kim: "The music score of a song mainly includes lyrics, note duration and note pitch. We first convert the lyrics to a phoneme sequence. Note duration is the number of frames corresponding to each note, and note pitch is converted to Pitch ID following the MIDI standard." (Zhang, Sec. 2.3.1.) "The note duration sequence and note pitch sequence are extended to the length of the phoneme sequence." (Zhang, Sec. 2.3.1.) Zhang states that "VISinger follows the main architecture of VITS" (Zhang, Abstract) and "we build upon VITS and propose VISinger, an end-to-end singing voice synthesis system" (Zhang, Sec. 1), VITS being the system of Kim.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the system of Kim to singing voice synthesis with the phoneme-level music-score inputs of Zhang and to modify the system such that the monotonic alignment search of Kim is performed by using the phoneme-level note duration information of Zhang, that is, the information on MIDI duration, to extract each phoneme duration. One of ordinary skill in the art would have been motivated to make this modification in order that the synthesized singing conform to the labels of the music score, as Zhang expressly teaches: "synthetic singing should not only be pronounced correctly according to the lyrics, but also conform to the labels of the music score" (Zhang, Sec. 118-29).
Regarding claim 10, Kim discloses:
A singing voice phoneme duration extraction method using a MIDI, comprising: (Kim discloses the method performed by its system: "Figures 1a and 1b show the training and inference procedures of our method, respectively." Kim, Sec. 2; Sec. 2.5.)
receiving, by a prior encoder, phonemes converted from a text as input, and outputting a prior probability distribution; (as quoted for claim 1: Kim, Sec. 2.5.2; Sec. 2.1.3, Eq. (4); Sec. 3.2.)
receiving, by a posterior encoder, acoustic features as input and outputting a posterior probability distribution; (as quoted for claim 1: Kim, Sec. 2.5.1; Sec. 3.2; Sec. 2.1.3, Eq. (3). Under the express definition of "acoustic features" in claim 8 and paragraphs [0019] and [0042], the linear-scale spectrogram input of Kim is the recited acoustic features.)
converting, by a flow, the probability distribution to simplify the posterior probability distribution; (as quoted for claim 1: Kim, Sec. 2.1.3, Eq. (4); Sec. 2.5.2.)
performing, by a monotonic alignment search module, monotonic alignment search by using information on MIDI duration; (as quoted for claim 1: Kim, Sec. 2.2.1, Eqs. (5)-(6).)
extracting, by the monotonic alignment search module, phoneme duration through a result of the monotonic alignment search; and (Kim extracts each phoneme duration from the result of the search: "We can calculate the duration of each input token di by summing all the columns in each row of the estimated alignment." Kim, Sec. 2.2.2.)
outputting, by a decoder, a waveform which is a voice digital signal, based on input reflecting a result of extracting the phoneme duration. (as quoted for claim 1: Kim, Sec. 2.1.2; Sec. 2.5.3; Sec. 2.1.3.)
Kim does not disclose that the phoneme duration extraction method is a singing voice phoneme duration extraction method using a MIDI (claim 10, preamble), and does not disclose performing the monotonic alignment search by using information on MIDI duration.
Zhang discloses information on MIDI duration provided at the phoneme level as an input to a singing voice synthesis system having the architecture of Kim, as quoted for claim 1 (Zhang, Sec. 2.3.1; Abstract; Sec. 1).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the method of Kim to singing voice synthesis with the phoneme-level music-score inputs of Zhang and to perform the monotonic alignment search of Kim by using the phoneme-level note duration information of Zhang to extract each phoneme duration. One of ordinary skill in the art would have been motivated to make this modification in order that the synthesized singing conform to the labels of the music score, as Zhang expressly teaches: "synthetic singing should not only be pronounced correctly according to the lyrics, but also conform to the labels of the music score" (Zhang, Sec. 118-29).
Regarding claim 11, Kim discloses:
A singing voice phoneme duration extraction system using a MIDI, comprising: (Kim discloses an end-to-end synthesis system comprising a posterior encoder, a prior encoder including a normalizing flow, and monotonic alignment search: "The overall architecture of the proposed model consists of a posterior encoder, prior encoder, decoder, discriminator, and stochastic duration predictor." Kim, Sec. 2.5; Fig. 1(a)/(b).)
a prior encoder configured to receive phonemes converted from a text, MIDI pitch, and MIDI duration as input, and to output a prior probability distribution; (as quoted for claim 1: Kim, Sec. 2.5.2; Sec. 2.1.3; Sec. 3.2; Eq. (4).)
a posterior encoder configured to receive acoustic features as input and to output a posterior probability distribution; (as quoted for claim 1: Kim, Sec. 2.5.1; Sec. 3.2; Sec. 2.1.3, Eq. (3). Under the express definition of "acoustic features" in claim 8 and paragraphs [0019] and [0042], the linear-scale spectrogram input of Kim is the recited acoustic features.)
a flow configured to convert the probability distribution to simplify the posterior probability distribution; and (as quoted for claim 1: Kim, Sec. 2.1.3, Eq. (4); Sec. 2.5.2.)
a monotonic alignment search module configured to perform monotonic alignment search by using information on MIDI duration to extract phoneme duration. (as quoted for claim 1: Kim, Sec. 2.2.1, Eqs. (5)-(6); Sec. 2.2.2. This limitation is interpreted under 35 U.S.C. 112(f) as set forth in the Claim Interpretation section above.)
Kim does not disclose that the phoneme duration extraction system is a singing voice phoneme duration extraction system using a MIDI (claim 11, preamble), does not disclose that the prior encoder receives MIDI pitch, and MIDI duration as input together with the phonemes, and does not disclose performing the monotonic alignment search by using information on MIDI duration.
Zhang discloses a prior-side text encoder of a system having the architecture of Kim that receives, for singing voice synthesis, the note pitch and note duration of the music score at the phoneme level: "The music score of a song mainly includes lyrics, note duration and note pitch. We first convert the lyrics to a phoneme sequence. Note duration is the number of frames corresponding to each note, and note pitch is converted to Pitch ID following the MIDI standard. The note duration sequence and note pitch sequence are extended to the length of the phoneme sequence. The text encoder which consists of multiple FFT blocks takes the above three sequences in phoneme-level as input and generates a phoneme-level representation of music score." (Zhang, Sec. 2.3.1; see also Zhang, Abstract; Sec. 1.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the system of Kim to singing voice synthesis and to provide the MIDI pitch and the MIDI duration of Zhang as inputs to the prior encoder of Kim together with the phonemes. One of ordinary skill in the art would have been motivated to make this modification in order that the synthesized singing conform to the music score, as Zhang expressly teaches: "synthetic singing should not only be pronounced correctly according to the lyrics, but also conform to the labels of the music score" (Zhang, Sec. 1). It would further have been obvious to perform the monotonic alignment search of Kim by using the phoneme-level note duration information of Zhang to extract each phoneme duration. One of ordinary skill in the art would have been motivated to make this modification in order that the extracted phoneme durations likewise conform to the labels of the music score, per the express teaching of Zhang quoted above (Zhang, Sec. 118-29).
Regarding claim 2: The singing voice phoneme duration extraction system of claim 1, wherein the prior encoder is configured to additionally receive MIDI pitch and MIDI duration, as input, in addition to phonemes converted from lyrics which is a text, to perform monotonic alignment search by using the MIDI duration information. (Kim in view of Zhang teaches claim 1 as set forth above. Kim does not disclose that the prior encoder additionally receives MIDI pitch and MIDI duration as input, in addition to the phonemes converted from lyrics, to perform the monotonic alignment search by using the MIDI duration information. Zhang discloses this limitation: "We first convert the lyrics to a phoneme sequence. Note duration is the number of frames corresponding to each note, and note pitch is converted to Pitch ID following the MIDI standard. The note duration sequence and note pitch sequence are extended to the length of the phoneme sequence. The text encoder which consists of multiple FFT blocks takes the above three sequences in phoneme-level as input" (Zhang, Sec. 2.3.1), and in the combination set forth for claim 1 the monotonic alignment search is performed by using that note duration information. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the prior encoder additionally receives the MIDI pitch and the MIDI duration of Zhang as input. One of ordinary skill in the art would have been motivated to make this modification in order that the synthesized singing conform to the music score, as Zhang expressly teaches: "synthetic singing should not only be pronounced correctly according to the lyrics, but also conform to the labels of the music score" (Zhang, Sec. 1).)
Regarding claim 3: The singing voice phoneme duration extraction system of claim 2, wherein information inputted to the prior encoder is information in which a text, pitch, and duration of the MIDI corresponding to each phoneme are mapped. (Kim in view of Zhang teaches claim 2 as set forth above. Kim does not disclose that the information inputted to the prior encoder is information in which a text, pitch, and duration of the MIDI corresponding to each phoneme are mapped. Zhang discloses this limitation: "The note duration sequence and note pitch sequence are extended to the length of the phoneme sequence. The text encoder which consists of multiple FFT blocks takes the above three sequences in phoneme-level as input and generates a phoneme-level representation of music score." (Zhang, Sec. 2.3.1.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the text, the pitch, and the duration of the MIDI are mapped for each phoneme in the manner of Zhang. One of ordinary skill in the art would have been motivated to make this modification in order to supply the score information in one-to-one correspondence with each phoneme so that a phoneme-level representation of the music score is generated, as Zhang expressly teaches (Zhang, Sec. 2.3.1).)
Regarding claim 7: The singing voice phoneme duration extraction system of claim 1, wherein the prior encoder comprises a text encoder and a projection layer. (Kim in view of Zhang teaches claim 1 as set forth above. Kim further discloses this limitation: the hidden representation is obtained "through the text encoder and a linear projection layer above the text encoder that produces the mean and variance used for constructing the prior distribution." Kim, Sec. 2.5.2.)
Regarding claim 8: The singing voice phoneme duration extraction system of claim 1, wherein the acoustic features are a linear spectrogram or a Mel-spectrogram. (Kim in view of Zhang teaches claim 1 as set forth above. Kim further discloses this limitation: "We use linear spectrograms which can be obtained from raw waveforms through the Short-time Fourier transform (STFT), as input of the posterior encoder" (Kim, Sec. 3.2) and "We, therefore, use the linear-scale spectrogram of target speech xlin as input rather than the mel-spectrogram" (Kim, Sec. 2.1.3), with the Mel-spectrogram disclosed as an alternative posterior input: "Replacing the linear-scale spectrogram for posterior input with the mel-spectrogram" (Kim, Sec. 4.1; Table 2, row "with Mel-spectrogram").)
Regarding claim 9: The singing voice phoneme duration extraction system of claim 1, wherein the decoder is configured to receive the posterior probability distribution as input when learning, and to output a waveform which is a voice digital signal, and to receive the prior probability distribution undergoing inverse transformation on the probability distribution as input when inferring, and to output a waveform which is a voice digital signal. (Kim in view of Zhang teaches claim 1 as set forth above. Kim further discloses this limitation: "The posterior encoder and discriminator are only used for training, not for inference" (Kim, Sec. 2.5); in training, the latent representation sampled from the posterior probability distribution is the input upsampled by the decoder to the waveform, while in inference the sample from the prior probability distribution undergoes the inverse transformation of the flow and is then input to the decoder (Kim, Fig. 1(a)/(b); Sec. 2.1.2), Kim disclosing synthesis "through the inverse transformation of the normalizing flow ... and decoder G" (Kim, App. D, Eq. (14)). "The probability distribution" is interpreted as set forth in the Claim Interpretation section above.)
Claims 4-6 are rejected under 35 U.S.C. 103 as being unpatentable over Kim in view of Zhang as applied to claim 2 above, and further in view of Hono et al., "Singing Voice Synthesis Based on a Musical Note Position-Aware Attention Mechanism," Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2023, hereinafter Hono.
Hono was published at ICASSP 2023, and a corresponding preprint, arXiv:2212.13703v1, was made publicly available December 28, 2022. Hono was therefore publicly available before the June 14, 2024 effective filing date of the claimed invention and qualifies as prior art under 35 U.S.C. 102(a)(1).
Regarding claim 4: The singing voice phoneme duration extraction system of claim 2, wherein the monotonic alignment search module is configured to divide phoneme sections by using the MIDI duration information, and then to perform monotonic alignment search for each phoneme section.
Kim in view of Zhang teaches claim 2 as set forth above, including performing the monotonic alignment search between the flow-transformed posterior representation and the prior (Kim, Sec. 2.2.1, Eqs. (5)-(6)) by using the phoneme-level note duration information of the music score (Zhang, Sec. 2.3.1).
The combination of Kim and Zhang does not disclose that the monotonic alignment search module divides phoneme sections by using the MIDI duration information, and then performs the monotonic alignment search for each phoneme section.
Hono discloses constraining singing-voice alignment by the note boundaries of the musical score and dividing each note duration among the sung units within the note. The note position representations of Hono are computed from the note start and end positions of the score, "where sn and en denote the start and end positions of the n-th musical note" (Hono, Sec. 3.1, Eqs. (5)-(7)). Hono generates "a penalty matrix based on the note boundaries," whose "pseudo-boundaries are obtained by equally dividing note duration in accordance with the number of morae in each note" (Hono, Sec. 3.3; Fig. 2). Hono states: "Since singing voices are generally sung to follow the rhythm of a musical score, it is natural to assume that the alignment of the singing voice should be close to the path determined by the note timing in the score." (Hono, Sec. 3.3.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combination such that the phoneme sections are divided by using the MIDI duration information of Zhang and the monotonic alignment search of Kim is then performed for each divided phoneme section, bounding the search by the note boundaries of the score in the manner taught by Hono. One of ordinary skill in the art would have been motivated to make this modification in order to keep the alignment of each phoneme within the note boundaries determined by the musical score, because "the alignment of the singing voice should be close to the path determined by the note timing in the score" (Hono, Sec. 3.3).
Regarding claim 5: The singing voice phoneme duration extraction system of claim 4, wherein the monotonic alignment search module is configured to perform monotonic alignment search between the posterior probability distribution and the prior probability distribution in every phoneme section. (The combination of Kim, Zhang, and Hono as set forth for claim 4 teaches this limitation: within each phoneme section, the monotonic alignment search of Kim is performed between the flow-transformed posterior representation and the prior, the search finding "an alignment that maximizes the log-likelihood of the latent variables z" evaluated under the prior (Kim, Sec. 2.2.1, Eq. (6)), and the penalty matrix of Hono confines the alignment of each section to the note boundaries of that section, whose "pseudo-boundaries are obtained by equally dividing note duration in accordance with the number of morae in each note" (Hono, Sec. 3.3; Fig. 2).The reason to combine is set forth with respect to claim 4.)
Regarding claim 6: The singing voice phoneme duration extraction system of claim 4, wherein the monotonic alignment search module is configured to divide the respective phoneme sections, and to independently extract phoneme duration for all phonemes. (The combination of Kim, Zhang, and Hono as set forth for claim 4 teaches this limitation: Hono divides each note duration among the sung units within that note, its "pseudo-boundaries are obtained by equally dividing note duration in accordance with the number of morae in each note" (Hono, Sec. 3.3; Fig. 2), so that the duration of each phoneme is extracted from the alignment of its own section, by summing the alignment entries of that section (Kim, Sec. 2.2.2), independently of the other sections. The reason to combine is set forth with respect to claim 4.)
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
a. WO 2021/101665 A1: singing voice synthesis in which a phoneme duration predictor is trained using a loss based on the note duration of the musical score.
b. Wang et al., "Singing-Tacotron: Global Duration Control Attention and Dynamic Filter for End-to-end Singing Voice Synthesis," arXiv:2202.07907v1: singing voice synthesis in which the alignment between the input phonemes and the output acoustic features is constrained by the duration information of the musical score.
c. Blaauw et al., "Sequence-to-sequence Singing Synthesis Using the Feed-forward Transformer," arXiv:1910.09989v2: singing synthesis in which the alignment is constrained by the note timings of the musical score.
d. Nakano et al. (US 8,244,546 B2): singing synthesis parameter data estimation describing per-syllable lyric alignment of singing audio.
e. Ren et al., "PortaSpeech: Portable and High-Quality Generative Text-to-Speech," Advances in Neural Information Processing Systems 34 (NeurIPS 2021), pp. 13963-13974: text-to-speech in which phoneme-level alignment is restricted within word-level segment boundaries by an attention mask.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUVAL H. LEVENTAL whose telephone number is (571) 270-3130. The examiner can normally be reached Monday-Friday, 8:00 AM - 5:00 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, PIERRE-LOUIS DESIR, can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YUVAL HAIM LEVENTAL/Examiner, Art Unit 2659
/PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659