DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 10/14/2025 has been entered.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Response to Amendment/Arguments
Applicant’s arguments with respect to claim(s) 9 – 12 and 14-15 have been considered but are moot because of the new ground of rejection below. Upon further search and consideration, the indication of allowable subject matter has been withdrawn and the claims are now rejected using newly discovered prior art and different interpretation of the claim’s language.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 9 -12, 14 and 15 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 9, line 8, the phrase “And is capable of synthesizing …” (and claim 10, line 4) is indefinite. “Capable of” doing something does not indicate a positive action being performed but simply the ability to do so.
Claims 14 and 15 are indefinite for the reason indicated above.
Dependent claims 11 -12 incorporate the deficiencies of claim 9 upon which they depend.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 14 and 15 are rejected under 35 U.S.C. 101 because the claimed subject matter is directed to an abstract idea and does not include additional elements that amount to significantly more than the abstract idea itself.
Claims 14 and 15 are directed to a process and a manufacture, respectively, and therefore fall within a statutory category of eligible subject matter.
Claim 14 recites steps of:
inputting designation of a conversion destination voice,
analyzing voice data of a conversion source voice and extracting time-series data including a phoneme and a pitch,
extracting an utterance section of each phoneme,
matching a height of the pitch to a height of the designated conversion destination voice,
compressing or expanding the pitches in a time direction in accordance with compression or expansion of the utterance section, and
inputting the phoneme and the pitch to a deep learning model and generating converted voice data.
Claim 15 recites substantially the same operations in the form of program instructions stored on a non-transitory recording medium.
These limitations, considered as a whole, are directed to collecting, analyzing, manipulating, and generating data. More specifically, the claims recite the abstract idea of processing speech-related information to produce a desired output using generic computational and machine-learning operations. The recited operations amount to a results-oriented data transformation workflow, which is an abstract idea similar to other recognized forms of information processing.
The claims do not recite any specific technological improvement to computer functionality, speech-recognition hardware, audio signal processing architecture, or deep-learning model architecture. Instead, they recite the desired result of converting source voice data into target voice data by performing high-level functional steps.
Accordingly, claims 14 and 15 recite an abstract idea (STEP 2A, prong 1).
The additional claim language does not integrate the abstract idea into a practical application. The claims merely apply the abstract idea in the context of voice conversion, using generic components and generic computer implementation.
The recited “computer” in claim 14 and “non-transitory recording medium” in claim 15 are merely nominal or generic computer-implementation limitations that do not impose a meaningful limitation on the claim scope. Likewise, the recited “deep learning model” is claimed at a high level of abstraction, as a functional tool for synthesizing voice data, without any specific improvement to the model itself or to the operation of a computer.
The claims do not recite:
a particular voice-conversion algorithm,
a particular neural-network architecture,
a specific pitch-alignment technique,
a specific signal-processing mechanism,
or any other technical implementation that improves the functioning of a computer or another technology.
Instead, the claims are directed to using conventional computing and machine-learning functionality to perform a desired information-processing result.
Therefore, the claims are not directed to a practical application (STEP 2A, prong 2).
The additional elements, considered individually and in combination, do not amount to significantly more than the abstract idea itself.
The claims merely recite:
generic voice data,
generic extraction and adjustment operations,
a generic deep learning model,
and conventional program/medium language.
These elements are described only at a functional level and do not add an inventive concept. The claims do not include any unconventional arrangement of components or any specific technical feature that confines the claims to a particular implementation. Rather, the claims preempt the use of basic data-processing techniques for voice conversion.
Thus, the claims amount to no more than instructions to apply the abstract idea using generic computer technology (STEP 2B).
Accordingly, claims 14 and 15 are ineligible under 35 U.S.C. §101.
Claim 9 is directed to a specific technological improvement in voice conversion systems. The claim recites a defined speech-processing pipeline that extracts phoneme and pitch time-series data, adjusts utterance duration and pitch timing in a coordinated manner, and supplies the processed data to a deep learning model for synthesis of a designated target voice. This is not merely the use of a computer to perform an abstract idea, but a particular technical solution to improving the conversion of source voice data into target voice data.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 9, 10-12, 13 -14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. PgPub No. 2021/0280202 (Wang) in view of Yu et al., US Patent No.: 11,4304,31(Yu), a pdf copy of the patent is provided as text form showing the paragraph numbers used below) .
As per claim, 9, wang discloses a voice conversion apparatus comprising:
an input unit that inputs designation of a conversion destination voice (Wang discloses acquiring a reference speech of a second user for use in voice conversion ¶¶0021–0024, 0041–0046);
an extraction unit that analyzes voice data of a conversion source voice and extracts time series data including a phoneme and a pitch (Wang discloses extracting first speech content information and a first acoustic feature from the source speech ¶¶ 0029–0034);
an adjustment unit that matches a height of the pitch to a height of the designated conversion destination voice (Wang discloses reconstructing a third acoustic feature using source and reference acoustic features, and synthesizing target speech based on the reconstructed feature ¶¶0034–0046, 0050–0067);
a generation unit that inputs the phoneme and the pitch to a deep learning model that learns voice data of many people and is capable of synthesizing a designated person’s voice in time-series order, and generates voice data obtained by synthesizing the designated conversion destination voice (Wang discloses a pre-trained voice conversion model that reconstructs acoustic features and synthesizes target speech ¶¶ 0034–0046, 0068–0080);
wherein the extraction unit extracts an utterance section of each of the phonemes, and inputs the utterance section that is compressed or expanded to the generation unit, and the adjustment unit compresses or expands the pitches in a time direction in accordance with the compression or expansion of the utterance section (¶0047]–[¶0049], [Wang ¶0053]–[¶0056], [Wang ¶0063]–[¶0067]).
Wang fails to explicitly teach the extract time series data includes a phoneme, a height of the pitch and extracting an utterance section of each phoneme or compressing or expanding pitches in a time direction in accordance with utterance-section compression or expansion.
However, these feature are well known in the art as evidenced by Yu which teaches:
receiving a phoneme sequence input and using F0 and other frame-level acoustic information in the voice conversion process ¶¶0021–0023, 0025–0027); use of F0 and speaker conditioning in singing voice conversion (height of the pitch ¶¶0022–0023, 0027);
Yu also teaches a deep neural network -based generation framework that receives phoneme-related inputs and speaker information to generate mel-spectrogram features recursively ( ¶¶0021–0023, 0025–0027).
Yu teaches a context associated with one or more phonemes corresponding to the singing voice of a first person is encoded, the one or more phonemes are aligned to one or more target acoustic frames based on the encoded context (¶¶ 0003, 0004, 0005, 0020, 0025). Yu also teaches encoding phoneme context, aligning phonemes to target acoustic frames based on encoded context, and generating mel-spectrogram features using phoneme duration, F0, and RMSE inputs, including state expansion based on duration ¶¶ 0021–0023, 0025–0027); duration-based state expansion, including expansion of hidden states according to phoneme duration, and use of F0 and frame-aligned hidden states in the conversion process (¶¶ 0022–0023, 0025–0027).
Therefore, it would have been obvious to one of ordinary skill in the art before the invention was made to modify Wang in view of Yu to improve temporal alignment and prosodic consistency in voice conversion.
As per claim 10, Wang in view of Yu disclose all the limitations of claim 9 upon which claim 10 depends. Wang further discloses training a pre-trained voice conversion model based on speech data of a third user, and then using that model for later conversion ¶¶ 0034–0035, 0068–0080). Wang fails to expressly teach extracting phonemes and pitches from many people’s voice data which become conversion destination voices. Yu teaches using phoneme sequence input, phoneme duration, F0, RMSE, and speaker data in a training and generation framework for singing voice conversion ¶¶ 0021–0023, 0025–0027). Therefore, it would have been obvious to of ordinary skill in the art at the time of the invention to modify Wang’s pre-trained model with Yu-A’s phoneme and pitch-based training inputs because both references address deep-learning-based voice conversion and speaker-conditioned synthesis. Incorporating Yu’s phoneme-duration and pitch control into Wang’s training framework would have been a predictable variation to improve multi-speaker synthesis control.
As per claim 11, Wang in view of Yu disclose all the limitations of claim 9 upon which claim 11 depends. Wang fails to expressly teach inputting, as text, same sentences as in an utterance content of the conversion source voice in combination with the voice data of the conversion source voice to analyze the sentence and extract phonemes. But Yu teaches phoneme-sequence-based processing for singing voice conversion ¶¶ 0021–0023, 0025–0027. It would have been obvious to one of ordinary skill in the art at the time of the invention to use sentence-based phoneme extraction in Wang’s system in view of Yu’s phoneme-centered conversion framework, because both references rely on linguistic content representations to drive acoustic conversion, and combining them would improve content control and alignment.
As per claim 12, Wang in view of Yu disclose all the limitations of claim 9 upon which claim 12 depends. Wang discloses extraction of acoustic features from source speech and reference speech and use of those features in a pre-trained voice conversion model ¶¶ 0029–0035, 0041–0046. Wang fails to expressly teach reading out pitches corresponding to phonemes from a storage device. Yu teaches use of F0 and duration-based frame alignment in conjunction with phoneme processing ¶¶ 0022–0023, 0025–0027. It would have been obvious to modify Wang with Yu-A so that pitch-related features are associated with phoneme timing because both references teach using acoustic and timing information to guide voice conversion. The combination would have yielded predictable improvement in alignment and synthesis accuracy.
Claims 14 and 15 are similar in scope and content to claim 9 rejected above, therefore claims 14 and 15 are rejected under similar rationale.
As per claims 14 and 15, Wang discloses A voice conversion method causing a computer ¶0021]–[¶0023], [Wang ¶0029]–[¶0035], [Wang ¶0041]–[¶0045], [Wang ¶0068]–[¶0080) to:
input designation of a conversion destination voice ( ¶0038]–[¶0040);
analyze voice data of a conversion source voice and extract time series data including a phoneme and a pitch (¶0029]–[¶0030], [Wang ¶0047]–[¶0049], [Wang ¶0053]–[¶0056]);
extract an utterance section of each of the phonemes wherein the utterance section of each of the phonemes is compressed or expanded (¶0047]–[¶0049], [Wang ¶0063]–[¶0067);
match a height of the pitch to a height of the designated conversion destination voice (¶0053]–[¶0056], [Wang ¶0063]–[¶0067);
compress or expand the pitches in a time direction in accordance with the compression or expansion of the utterance section (¶0053]–[¶0056], [Wang ¶0063]–[¶0067); and
input the phoneme and the pitch to a deep learning model that learns voice data of many people and is capable of synthesizing a designated person’s voice in time-series order, and generate voice data obtained by synthesizing the designated conversion destination voice (¶0034]–[¶0035], [Wang ¶0041]–[¶0045], [Wang ¶0063]–[¶0067], [Wang ¶0068]–[¶0080)
Wang fails to explicitly teach a phoneme sequence input, aligning the one or more phonemes to one or more target acoustic frames based on the encoded context. However, these features are well known in the art as evidenced by Yu which discloses:
phoneme sequence input (¶(21)], [Yu ¶(25)); aligning the one or more phonemes to one or more target acoustic frames based on the encoded context (Yu ¶(22)], [Yu ¶(26)); recursively generating one or more mel-spectrogram features (¶(23)], [Yu ¶(27)).
It would have been obvious to one of ordinary skill in art at the time of the invention to modify Wang’s method to include phoneme-based alignment and frame expansion from Yu in order to improve timing and synthesis quality in a neural voice conversion system. The combination would have been a predictable use of known techniques in the same art.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO-892.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RICHEMOND DORVIL whose telephone number is (571)272-7602. The examiner can normally be reached 8:30 - 5:30 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658