DETAILED ACTION
This office action is in response to Applicant’s submission filed on 11/27/2024. Claims 1-15 are pending in the application. As such, claims 1-15 have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) was submitted on 5/01/2025 and 2/18/2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-15 are rejected under 35 U.S.C. 103 as being unpatentable over Griffin et al. (US 11715477 B1; hereinafter referred to as Griffin) in view of Lim et al. (US 5715365 A; hereinafter referred to as Lim).
Regarding claim 1, Griffin discloses: a method of estimating speech model parameters from a digitized speech signal, the method comprising: dividing the digitized speech signal into two or more frequency band signals ([col 2, lines 52-55] estimating speech model parameters from a digitized speech signal, includes dividing the digitized speech signal into two or more frequency band signals);
determining a first excitation parameter at a first time sample ([col 2, lines 33-38] The first element of the vector of excitation strength parameters may correspond to an associated frequency band and time interval, and the first weight may depend on an energy of the associated frequency band and time interval and an energy of at least one other frequency band or time interval), wherein determining the first excitation parameter comprises performing a nonlinear operation on at least two of the frequency band signals to produce at least first and second modified frequency band signals… ([col 2, lines 55-59] A first preliminary excitation parameter is determined using a first method that includes performing a nonlinear operation on at least two of the frequency band signals to produce at least two modified frequency band signals);
determining a first weight at the first time sample based on at least one of the first and second modified frequency band signals ([col 2, lines 59-63] weights to apply to the at least two modified frequency band signals are determined, and the first preliminary excitation parameter is determined using a first weighted combination of the at least two modified frequency band signals);
determining a second weight at the second time sample based on at least one of the third and fourth modified frequency band signals ([col 8-9, lines 62-2] The weight generation unit 410 divides the Fourier transform into bands and generates weights based on the energy in each band and parameters generated by the speech analysis unit 420. The vector quantizer unit 415 compares codebook entries to the input excitation strengths based on the weights from the weight generation unit 410 and the speech analysis parameters from the speech analysis unit 420 to determine the best codebook entry);
determining a third weight at a third time sample based on at least one of the first through fourth modified frequency band signals ([col 8-9, lines 62-2] The weight generation unit 410 divides the Fourier transform into bands and generates weights based on the energy in each band and parameters generated by the speech analysis unit 420. The vector quantizer unit 415 compares codebook entries to the input excitation strengths based on the weights from the weight generation unit 410 and the speech analysis parameters from the speech analysis unit 420 to determine the best codebook entry);
determining a third excitation parameter at the third time sample using the first and second excitation parameters ([col 3, lines 9-15] determining a third preliminary excitation parameter by comparing energy near a peak frequency to total energy and using the first, second and third preliminary excitation parameters to determine the excitation parameter for the digitized speech signal. The peak frequency may be determined after excluding frequencies below a threshold level) and the first, second, and third weights ([col 8, lines 57-2] The excitation parameter quantization system 400 jointly quantizes the voiced, unvoiced, and pulsed strengths to produce quantized strengths and the best codebook index. The window and Fourier transform unit 405 computes the Fourier transform of the windowed signal. The weight generation unit 410 divides the Fourier transform into bands and generates weights based on the energy in each band and parameters generated by the speech analysis unit 420. The vector quantizer unit 415 compares codebook entries to the input excitation strengths based on the weights from the weight generation unit 410 and the speech analysis parameters from the speech analysis unit 420 to determine the best codebook entry).
Griffin does not explicitly, but Lim teaches: determining a second excitation parameter at a second time sample ([col 1, lines 44-49] In typical approaches to determining excitation parameters, an analog speech signal s(t) is sampled to produce a speech signal s(n). Speech signal s(n) is then multiplied by a window w(n) to produce a windowed signal s.sub.w (n) that is commonly referred to as a speech segment or a speech frame), wherein determining the second excitation parameter comprises performing the nonlinear operation on the at least two of the frequency band signals to produce at least third and fourth modified frequency band signals… ([col 3, lines 21-26] excitation parameters for an input signal are generated by dividing the input signal into at least two frequency band signals. Thereafter, a nonlinear operation is performed on at least one of the frequency band signals to produce at least one modified frequency band signal. Multiple modified frequency band signal can be generated.).
Griffin and Lim are considered analogous in the field of speech processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Griffin to combine the teachings of Lim because doing so would allow for better excitation parameter estimation by improving the accuracy of fundamental frequency estimation using parabolic interpolation, leading to more accurate excitation parameters when the fundamental frequency serves as the excitation parameter (Lim [Fig.2; col 4, lines 7-17] the invention features using nonlinear operations to improve the accuracy of fundamental frequency estimation. A nonlinear operation is performed on the input signal to produce a modified signal from which the fundamental frequency is estimated. In another approach, the input signal is divided into at least two frequency band signals. Next, a nonlinear operation is performed on these frequency band signals to produce modified frequency band signals. Finally, the modified frequency band signals are combined to produce a combined signal from which a fundamental frequency is estimated).
Regarding claim 2, the combination of Griffin and Lim teaches: the method of claim 1. Griffin further teaches: wherein the third excitation parameter is a voiced error or a voiced strength ([col 2, lines 42-49] The vector of excitation strength parameters may include a voiced strength/pulsed strength pair, and the first weight may be selected such that the error between a high voiced strength/low pulsed strength pair and a quantized low voiced strength/high pulsed strength pair is less than the error between the high voiced strength/low pulsed strength pair and a quantized low voiced strength/low pulsed strength pair).
Regarding claim 3, the combination of Griffin and Lim teaches: the method of claim 1. Griffin further teaches: wherein the third excitation parameter is closer to the first excitation parameter when the third weight is closer to the first weight ([col 8-9, lines 57-2] The excitation parameter quantization system 400 jointly quantizes the voiced, unvoiced, and pulsed strengths to produce quantized strengths and the best codebook index. The window and Fourier transform unit 405 computes the Fourier transform of the windowed signal. The weight generation unit 410 divides the Fourier transform into bands and generates weights based on the energy in each band and parameters generated by the speech analysis unit 420. The vector quantizer unit 415 compares codebook entries to the input excitation strengths based on the weights from the weight generation unit 410 and the speech analysis parameters from the speech analysis unit 420 to determine the best codebook entry).
Regarding claim 4, Griffin teaches: a method of estimating speech model parameters from a digitized speech signal, the method comprising: dividing the digitized speech signal into two or more frequency band signals ([col 2, lines 52-55] estimating speech model parameters from a digitized speech signal, includes dividing the digitized speech signal into two or more frequency band signals);
determining a first excitation parameter at a first time sample ([col 2, lines 33-38] The first element of the vector of excitation strength parameters may correspond to an associated frequency band and time interval, and the first weight may depend on an energy of the associated frequency band and time interval and an energy of at least one other frequency band or time interval), wherein determining the first excitation parameter comprises performing a nonlinear operation on at least two of the frequency band signals to produce at least first and second modified frequency band signals… ([col 2, lines 55-59] A first preliminary excitation parameter is determined using a first method that includes performing a nonlinear operation on at least two of the frequency band signals to produce at least two modified frequency band signals);
determining a first weight at the first time sample based on at least the first and second modified frequency band signals ([col 2, lines 35-41] the first weight may depend on an energy of the associated frequency band and time interval and an energy of at least one other frequency band or time interval. The first weight may be increased when an excitation strength is different between the associated frequency band and time interval and the at least one other frequency band or time interval.) and corresponding voiced strengths ([col 5, lines 5-9] The voiced strength V(t,ω), unvoiced strength U(t,ω), and pulsed strength P(t,ω) parameters control the proportion of quasi-periodic, noise-like, and pulsed signals in each frequency band. These parameters are functions of time (t) and frequency (ω)), wherein determining the first weight comprises increasing the first weight when energy in at least one of the first and second modified frequency band signals near the first time sample increases ([col 2, lines 33-41] the vector of excitation strength parameters may correspond to an associated frequency band and time interval, and the first weight may depend on an energy of the associated frequency band and time interval and an energy of at least one other frequency band or time interval. The first weight may be increased when an excitation strength is different between the associated frequency band and time interval and the at least one other frequency band or time interval);
determining a second weight at the second time sample based on at least the third and fourth modified frequency band signals and corresponding voiced strengths ([col 8-9, lines 62-2] The weight generation unit 410 divides the Fourier transform into bands and generates weights based on the energy in each band and parameters generated by the speech analysis unit 420. The vector quantizer unit 415 compares codebook entries to the input excitation strengths based on the weights from the weight generation unit 410 and the speech analysis parameters from the speech analysis unit 420 to determine the best codebook entry), wherein determining the second weight comprises increasing the second weight when energy in at least one of the third and fourth modified frequency band signals near the second time sample increases ([col 11, lines 11-17] The weights may be computed by comparing the energy in a frequency band to an estimate of the background noise in that band to produce a signal to noise ratio (SNR). The weights may be determined from the estimated SNR so that weights are higher when the estimated SNR is higher);
determining a third excitation parameter at a third time sample using the first and second excitation parameters ([col 3, lines 9-13] The method also may include determining a third preliminary excitation parameter by comparing energy near a peak frequency to total energy and using the first, second and third preliminary excitation parameters to determine the excitation parameter for the digitized speech signal) and the first and second weights ([col 2, lines 61-5] the first preliminary excitation parameter is determined using a first weighted combination of the at least two modified frequency band signals. A second preliminary excitation parameter is determined by applying weights corresponding to the weights determined in the first method to the at least two of the frequency band signals to form a second weighted combination of at least two frequency band signals and using a second method different from the first method to determine the second preliminary excitation parameter from the second weighted combination. The first and second preliminary excitation parameters are used to determine an excitation parameter for the digitized speech signal).
Griffin does not explicitly, but Lim teaches: determining a second excitation parameter at a second time sample ([col 1, lines 44-49] In typical approaches to determining excitation parameters, an analog speech signal s(t) is sampled to produce a speech signal s(n). Speech signal s(n) is then multiplied by a window w(n) to produce a windowed signal s.sub.w (n) that is commonly referred to as a speech segment or a speech frame), wherein determining the second excitation parameter comprises performing the nonlinear operation on the at least two of the frequency band signals to produce at least third and fourth modified frequency band signals… ([col 3, lines 21-26] excitation parameters for an input signal are generated by dividing the input signal into at least two frequency band signals. Thereafter, a nonlinear operation is performed on at least one of the frequency band signals to produce at least one modified frequency band signal. Multiple modified frequency band signal can be generated.).
Regarding claim 5, the combination of Griffin and Lim teaches: the method of claim 4. Griffin further teaches: wherein the third excitation parameter is a fundamental frequency ([col 5, lines 17-22] The vector of parameters v(t,ω) associated with the voiced strength parameter V(t,ω) includes voiced excitation parameters and voiced system parameters. The voiced excitation parameters may include a time and frequency dependent fundamental frequency ω.sub.0(t,ω) (or equivalently a pitch period n.sub.0(t,ω))).
Regarding claim 6, the combination of Griffin and Lim teaches: the method of claim 1. Griffin further teaches: wherein dividing the digital speech signal into the two or more frequency band signals further comprises: applying a bandpass filter to generate two or more bandpass filter outputs ([col 4, lines 46-49] A speech signal s.sub.0(n) may be divided into multiple frequency bands using bandpass filters. Characteristics of these bandpass filters are allowed to change as a function of time and/or frequency);
applying the nonlinearity operation to each bandpass filter output… ([col 11, lines 4-6] A nonlinearity is applied to the output of each bandpass filter to emphasize the fundamental frequency).
Lim further teaches: and subsequent to applying the nonlinearity operation, applying a lowpass filter and a downsampling operation to each bandpass filter output ([col 6, lines 42-48] The output of nonlinear operation unit 36 is passed through a lowpass filtering and downsampling unit 38 to reduce the data rate and consequently reduce the computational requirements of later components of the system. Lowpass filtering and downsampling unit 38 uses a seven point FIR filter computed every other sample for a downsampling factor of two).
Regarding claim 7, the combination of Griffin and Lim teaches: the method of claim 6. Lim further teaches: wherein the bandpass filter comprises a Finite Impulse Response (FIR) filter or Infinite Impulse Response (IIR) filter ([col 6, lines 24-26] Bandpass filter 34 can be implemented as a Finite Impulse Response (FIR) or Infinite Impulse Response (IIR) filter).
Regarding claim 8, the combination of Griffin and Lim teaches: the method of claim 6. Griffin further teaches: wherein applying the bandpass filter ([col 4, lines 46-51] A speech signal s.sub.0(n) may be divided into multiple frequency bands using bandpass filters. Characteristics of these bandpass filters are allowed to change as a function of time and/or frequency. A speech signal may also be divided into multiple bands by applying frequency windows or weightings to the speech signal STFT S(t,ω)) comprises multiplying the digital speech signal by a time window to generate the two or more bandpass filter outputs ([col 6, lines 15-17, 27-29] The window and Fourier transform unit 305 multiplies the input speech signal s.sub.0(n) by a window w(t,n) centered at time t to obtain a windowed signal s(t,n)… The Fourier transform computed by unit 305 is divided into bands by unit 310 and the energy in each band is computed to generate weights for vector quantizer unit 315).
Regarding claim 9, the combination of Griffin and Lim teaches: the method of claim 8. Griffin further teaches: wherein the time window is a 32 point Kaiser window, and a 32 point Fast Fourier transform (FFT) is used to generate the two or more bandpass filter outputs ([col 6, lines 18-24] The window used is typically a Hamming window or Kaiser window and is typically constant as a function of t so that w(t,n)=w.sub.0(n−t). The length of the window w(t,n) typically ranges between 5 ms and 40 ms. The Fourier transform (FT) of the windowed signal S(t,ω) is typically computed using a fast Fourier transform (FFT) with a length greater than or equal to the number of samples in the window).
Regarding claim 10, the combination of Griffin and Lim teaches: the method of claim 1. Griffin further teaches: a speech encoder configured to perform the method of claim 1 ([col 17, lines 11-16] The digital speech is processed by a MBE speech encoder unit 1115 to produce a digital bit stream 1120 suitable for transmission or storage. The speech encoder processes the digital speech signal in short frames. Each frame of digital speech samples produces a corresponding frame of bits in the bit stream output of the encoder).
Regarding claim 11, the combination of Griffin and Lim teaches: the speech encoder of claim 10. Griffin further teaches: a handset or mobile radio comprising the speech encoder of claim 10 ([col 3, lines 39-41] The speech coder may be included in, for example, a handset, a mobile radio, a base station or a console).
Regarding claim 12, the combination of Griffin and Lim teaches: the speech encoder of claim 10. Griffin further teaches: a base station or console comprising the speech encoder of claim 10 ([col 3, lines 39-41] the speech coder may be included in, for example, a handset, a mobile radio, a base station or a console) .
Regarding claim 13, the combination of Griffin and Lim teaches: the method of claim 4. Griffin further teaches: a speech encoder configured to perform the method of claim 4 ([col 17, lines 11-16] The digital speech is processed by a MBE speech encoder unit 1115 to produce a digital bit stream 1120 suitable for transmission or storage. The speech encoder processes the digital speech signal in short frames. Each frame of digital speech samples produces a corresponding frame of bits in the bit stream output of the encoder).
Regarding claim 14, the combination of Griffin and Lim teaches: the speech encoder of claim 13. Griffin further teaches: a handset or mobile radio comprising the speech encoder of claim 13 ([col 3, lines 39-41] The speech coder may be included in, for example, a handset, a mobile radio, a base station or a console).
Regarding claim 15, the combination of Griffin and Lim teaches: the speech encoder of claim 13. Griffin further teaches: a base station or console comprising the speech encoder of claim 13 ([col 3, lines 39-41] the speech coder may be included in, for example, a handset, a mobile radio, a base station or a console) .
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Popa et al. (US 20080082320 A1) – discloses determining target encoding parameters, including excitation parameters, from conversion of source encoding parameters characterizing a source speech signal.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Nathan Tengbumroong whose telephone number is (703)756-1725. The examiner can normally be reached Monday - Friday, 11:30 am - 8:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NATHAN TENGBUMROONG/Examiner, Art Unit 2654
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654