Prosecution Insights
Last updated: August 16, 2026
Application No. 18/277,857

LYRIC FILE GENERATION METHOD AND APPARATUS

Non-Final OA §101§102§103
Filed
Aug 18, 2023
Priority
Feb 19, 2021 — CN 202110192245.9 +1 more
Examiner
CRESPO FEBLES, HECTOR J
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Lemon Inc.
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
1 granted / 1 resolved
+38.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
10 currently pending
Career history
7
Total Applications
across all art units

Statute-Specific Performance

§101
8.3%
-31.7% vs TC avg
§103
70.8%
+30.8% vs TC avg
§102
16.7%
-23.3% vs TC avg
§112
4.2%
-35.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1 to 4, 17, 18 and 21 to 23 rejected under 35 U.S.C. 101 because Regarding claim 1, limitations include the actions of: obtaining a phoneme propagation sequence of a song and an audio frame sequence of the song, determining an audio frame corresponding to the text unit on the audio frame sequence by matching the phoneme, determine the time information and generating a lyric file of the song based on the information. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can listen to a song and obtain a set of phonemes related to the song, determine the phonemes of a section of the song and the time that this appear and write them down in a piece of paper alongside the lyrics of the song. The claim does not recite additional elements that integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 2, limitations include the actions of: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, obtaining the phoneme corresponding to the text unit from the phoneme set corresponding to the text unit based on a pronunciation of the text unit in the song; and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can obtain the lyrics of a song and using a dictionary obtain the pronunciation of the lyrics then modify them to match the pronunciation in the song and finally, write them down in a piece of paper. The claim recites additional elements such as a pronunciation lexicon which is a generic pronunciation dictionary used as a tool to implement the mental process and don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 3, limitations include the actions of: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, obtaining a target phoneme subset of the text unit from the phoneme set corresponding to the text unit based on a pronunciation of the text unit in the song; obtaining the phoneme corresponding to the text unit from the target phoneme subset of the text unit based on a pronunciation duration of the text unit in the song; and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can obtain the lyrics of a song and using a dictionary obtain multiple ways of pronunciation of the lyrics then modify them to match the pronunciation in the song and finally, write them down in a piece of paper. The claim recites additional elements such as a pronunciation lexicon which is a generic pronunciation dictionary used as a tool to implement the mental process and don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 4, limitations include the actions of: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, obtaining a target phoneme subset of the text unit from the phoneme set corresponding to the text unit based on a pronunciation thereof in the song; obtaining the phoneme corresponding to the text unit from the target phoneme subset of the text unit based on a tone conversion of the text unit in the song; and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can obtain the lyrics of a song and using a dictionary obtain multiple ways of pronunciation of the lyrics then modify them to match the pronunciation in the song and finally, write them down in a piece of paper. The claim recites additional elements such as a pronunciation lexicon which is a generic pronunciation dictionary used as a tool to implement the mental process and don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 17, limitations include the actions of: obtaining a phoneme propagation sequence of a song and an audio frame sequence of the song, determining an audio frame corresponding to the text unit on the audio frame sequence by matching the phoneme, determine the time information and generating a lyric file of the song based on the information. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can listen to a song and obtain a set of phonemes related to the song, determine the phonemes of a section of the song and the time that this appear and write them down in a piece of paper alongside the lyrics of the song. The claim recites additional elements such as an electronic device, a processor and a memory but those are generic computer components used as tools to implement the mental process and that don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 18, limitations include the actions of: obtaining a phoneme propagation sequence of a song and an audio frame sequence of the song, determining an audio frame corresponding to the text unit on the audio frame sequence by matching the phoneme, determine the time information and generating a lyric file of the song based on the information. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can listen to a song and obtain a set of phonemes related to the song, determine the phonemes of a section of the song and the time that this appear and write them down in a piece of paper alongside the lyrics of the song. The claim recites additional elements such as a processor and a non-transitory computer-readable medium but those are generic computer components used as tools to implement the mental process and that don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 21, limitations include the actions of: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, obtaining the phoneme corresponding to the text unit from the phoneme set corresponding to the text unit based on a pronunciation of the text unit in the song; and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can obtain the lyrics of a song and using a dictionary obtain the pronunciation of the lyrics then modify them to match the pronunciation in the song and finally, write them down in a piece of paper. The claim recites additional elements such as a pronunciation lexicon and an electronic device which are a generic pronunciation dictionary and generic computer components used as tools to implement the mental process and don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 22, limitations include the actions of: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, obtaining a target phoneme subset of the text unit from the phoneme set corresponding to the text unit based on a pronunciation of the text unit in the song; obtaining the phoneme corresponding to the text unit from the target phoneme subset of the text unit based on a pronunciation duration of the text unit in the song; and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can obtain the lyrics of a song and using a dictionary obtain multiple ways of pronunciation of the lyrics then modify them to match the pronunciation in the song and finally, write them down in a piece of paper. The claim recites additional elements such as a pronunciation lexicon and an electronic device which are a generic pronunciation dictionary and generic computer components used as a tool to implement the mental process and don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Regarding claim 23, limitations include the actions of: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, obtaining a target phoneme subset of the text unit from the phoneme set corresponding to the text unit based on a pronunciation thereof in the song; obtaining the phoneme corresponding to the text unit from the target phoneme subset of the text unit based on a tone conversion of the text unit in the song; and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. Those steps constitute an abstract idea directed to a mental process that can be executed by a human mentally or using pen and paper, which is a judicial exception to patent eligibility. A human can obtain the lyrics of a song and using a dictionary obtain multiple ways of pronunciation of the lyrics then modify them to match the pronunciation in the song and finally, write them down in a piece of paper. The claim recites additional elements such as a pronunciation lexicon and an electronic device which are a generic pronunciation dictionary and generic computer components used as a tool to implement the mental process and don’t integrate the mental process into a practical application and the sole purpose of the system is not significantly more than performing the mental steps listed. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 17 and 18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Goto; Masataka et al. (US 20080097754 A1) hereinafter GOTO. Regarding claim 1, GOTO teaches: A lyric file generation method, comprising: obtaining a phoneme propagation sequence of a song and an audio frame sequence of the song, the phoneme propagation sequence comprising a phoneme corresponding to a text unit in a lyric text of the song; “FIG. 10B shows an example of converting English lyrics into a sequence of phonemes for alignment (phoneme network). In this example, the English lyrics are represented by English phonemes. Most preferably, an English phone model may be used for English lyrics using English phonemes. However, a Japanese phone model may be used for English lyrics if English phonemes are converted into Japanese phonemes. In an example of FIG. 10B, first, text data A representing phrases of original lyrics are converted into a sequence of phonemes B. Then, the sequence is further converted into a sequence of phonemes for alignment C only including phonemes used to identify the English phonemes (N, AA, TH . . . ) and short pauses (sp) by applying the two rules described above to the English lyrics converted into the sequence of phonemes B.” (GOTO [0107]). “… Alternatively, lyrics tagged with time information is generated in advance by the system of the present invention may be stored in storage means such as a hard disc provided in a music audio signal reproducing apparatus, or may be stored in a server over the network. The lyrics tagged with time information that have been acquired from the storage means or the server over the network may be displayed on the screen in synchronization with music digital data reproduced by the music audio signal reproducing apparatus.” (GOTO [0060]). determining an audio frame corresponding to the text unit in the audio frame sequence, wherein the phoneme corresponding to the text unit matches with an audio feature of the audio frame; “Returning to FIG. 1, the temporal-alignment feature extraction means 11 extracts a temporal-alignment feature suitable to make temporal alignment between lyrics of the vocal and the music audio signal from the dominant sound audio signal at each time. Specifically, in this embodiment, the 25th order features such as a resonance property of the phoneme are extracted as temporal-alignment features. This step is a pre-processing necessary for the subsequent alignment. Details will be described later with reference to the analysis conditions for Viterbi alignment shown in FIG. 9. The 25th order features are extracted in this embodiment, including the 12th order MFCC, the 12th order .DELTA.MFCC, and .DELTA. power.” (GOTO [0100]). “In this example, the English lyrics A of "Nothing untaken. Nothing lost" are converted into a sequence of English phonemes B of "N AA TH IH NG AH N T EY K AH N N AA TH IH NG L A O S T". Then, short pauses (sp) are combined with the sequence of phonemes B to form a sequence of phonemes for alignment C. The sequence of phonemes for alignment C is a phoneme network SN.” (GOTO [0108]). determining time information of the text unit based on playback duration of the audio frame; “...Therefore, according to the present invention, lyric data tagged with time information that is synchronized with the music audio signal may automatically be generated using an output from the alignment means.” (GOTO [0044]). and generating a lyric file of the song based on the time information of the text unit, wherein the lyric file is used to indicate a display of the text unit when the song is played to a location indicated by the time information. “In a music audio signal reproducing apparatus which reproduces a music audio signal while displaying on a screen lyrics temporally aligned with the music audio signal to be reproduced, the computer program of the present invention can be run for temporal alignment between a music audio signals and lyrics. The lyrics are displayed on a screen after the lyrics have been tagged with time information. When the lyrics are displayed on the screen, a portion of the displayed lyrics is selected with a pointer. In this manner, the music audio signal may be reproduced from that point, based on the time information corresponding to the selected lyric portion. Alternatively, lyrics tagged with time information is generated in advance by the system of the present invention may be stored in storage means such as a hard disc provided in a music audio signal reproducing apparatus, or may be stored in a server over the network. The lyrics tagged with time information that have been acquired from the storage means or the server over the network may be displayed on the screen in synchronization with music digital data reproduced by the music audio signal reproducing apparatus.” (GOTO [0060]). Regarding claim 17, GOTO teaches: An electronic device, comprising: a memory for storing computer programs; a processor for performing, by invoking the computer programs, the lyric file generation method comprising: obtaining a phoneme propagation sequence of a song and an audio frame sequence of the song, the phoneme propagation sequence comprising a phoneme corresponding to a text unit in a lyric text of the song; “FIG. 10B shows an example of converting English lyrics into a sequence of phonemes for alignment (phoneme network). In this example, the English lyrics are represented by English phonemes. Most preferably, an English phone model may be used for English lyrics using English phonemes. However, a Japanese phone model may be used for English lyrics if English phonemes are converted into Japanese phonemes. In an example of FIG. 10B, first, text data A representing phrases of original lyrics are converted into a sequence of phonemes B. Then, the sequence is further converted into a sequence of phonemes for alignment C only including phonemes used to identify the English phonemes (N, AA, TH . . . ) and short pauses (sp) by applying the two rules described above to the English lyrics converted into the sequence of phonemes B.” (GOTO [0107]). “According to the present invention, when a computer is used to make temporal alignment between lyrics and a music audio signal of music including vocals and accompaniment sounds, the computer may be identified as a program which implements the dominant sound audio signal extraction means, the vocal-section feature extraction means, the vocal section estimation means, the temporal-alignment feature extraction means, the phoneme network storage means, and the alignment means. The computer program may be stored in a computer-readable recording medium.” (GOTO [0059]). “… Alternatively, lyrics tagged with time information is generated in advance by the system of the present invention may be stored in storage means such as a hard disc provided in a music audio signal reproducing apparatus, or may be stored in a server over the network. The lyrics tagged with time information that have been acquired from the storage means or the server over the network may be displayed on the screen in synchronization with music digital data reproduced by the music audio signal reproducing apparatus.” (GOTO [0060]). determining an audio frame corresponding to the text unit in the audio frame sequence, wherein the phoneme corresponding to the text unit matches with an audio feature of the audio frame; “Returning to FIG. 1, the temporal-alignment feature extraction means 11 extracts a temporal-alignment feature suitable to make temporal alignment between lyrics of the vocal and the music audio signal from the dominant sound audio signal at each time. Specifically, in this embodiment, the 25th order features such as a resonance property of the phoneme are extracted as temporal-alignment features. This step is a pre-processing necessary for the subsequent alignment. Details will be described later with reference to the analysis conditions for Viterbi alignment shown in FIG. 9. The 25th order features are extracted in this embodiment, including the 12th order MFCC, the 12th order .DELTA.MFCC, and .DELTA. power.” (GOTO [0100]). “In this example, the English lyrics A of "Nothing untaken. Nothing lost" are converted into a sequence of English phonemes B of "N AA TH IH NG AH N T EY K AH N N AA TH IH NG L A O S T". Then, short pauses (sp) are combined with the sequence of phonemes B to form a sequence of phonemes for alignment C. The sequence of phonemes for alignment C is a phoneme network SN.” (GOTO [0108]). determining time information of the text unit based on playback duration of the audio frame; “...Therefore, according to the present invention, lyric data tagged with time information that is synchronized with the music audio signal may automatically be generated using an output from the alignment means.” (GOTO [0044]). and generating a lyric file of the song based on the time information of the text unit, wherein the lyric file is used to indicate a display of the text unit when the song is played to a location indicated by the time information. “In a music audio signal reproducing apparatus which reproduces a music audio signal while displaying on a screen lyrics temporally aligned with the music audio signal to be reproduced, the computer program of the present invention can be run for temporal alignment between a music audio signals and lyrics. The lyrics are displayed on a screen after the lyrics have been tagged with time information. When the lyrics are displayed on the screen, a portion of the displayed lyrics is selected with a pointer. In this manner, the music audio signal may be reproduced from that point, based on the time information corresponding to the selected lyric portion. Alternatively, lyrics tagged with time information is generated in advance by the system of the present invention may be stored in storage means such as a hard disc provided in a music audio signal reproducing apparatus, or may be stored in a server over the network. The lyrics tagged with time information that have been acquired from the storage means or the server over the network may be displayed on the screen in synchronization with music digital data reproduced by the music audio signal reproducing apparatus.” (GOTO [0060]). Regarding claim 18, GOTO teaches: A non-transitory computer-readable storage medium stored thereon a computer program that, when executed by a processor, implements the lyric file generation method comprising: obtaining a phoneme propagation sequence of a song and an audio frame sequence of the song, the phoneme propagation sequence comprising a phoneme corresponding to a text unit in a lyric text of the song; “FIG. 10B shows an example of converting English lyrics into a sequence of phonemes for alignment (phoneme network). In this example, the English lyrics are represented by English phonemes. Most preferably, an English phone model may be used for English lyrics using English phonemes. However, a Japanese phone model may be used for English lyrics if English phonemes are converted into Japanese phonemes. In an example of FIG. 10B, first, text data A representing phrases of original lyrics are converted into a sequence of phonemes B. Then, the sequence is further converted into a sequence of phonemes for alignment C only including phonemes used to identify the English phonemes (N, AA, TH . . . ) and short pauses (sp) by applying the two rules described above to the English lyrics converted into the sequence of phonemes B.” (GOTO [0107]). “According to the present invention, when a computer is used to make temporal alignment between lyrics and a music audio signal of music including vocals and accompaniment sounds, the computer may be identified as a program which implements the dominant sound audio signal extraction means, the vocal-section feature extraction means, the vocal section estimation means, the temporal-alignment feature extraction means, the phoneme network storage means, and the alignment means. The computer program may be stored in a computer-readable recording medium.” (GOTO [0059]). “… Alternatively, lyrics tagged with time information is generated in advance by the system of the present invention may be stored in storage means such as a hard disc provided in a music audio signal reproducing apparatus, or may be stored in a server over the network. The lyrics tagged with time information that have been acquired from the storage means or the server over the network may be displayed on the screen in synchronization with music digital data reproduced by the music audio signal reproducing apparatus.” (GOTO [0060]). determining an audio frame corresponding to the text unit in the audio frame sequence, wherein the phoneme corresponding to the text unit matches with an audio feature of the audio frame; “Returning to FIG. 1, the temporal-alignment feature extraction means 11 extracts a temporal-alignment feature suitable to make temporal alignment between lyrics of the vocal and the music audio signal from the dominant sound audio signal at each time. Specifically, in this embodiment, the 25th order features such as a resonance property of the phoneme are extracted as temporal-alignment features. This step is a pre-processing necessary for the subsequent alignment. Details will be described later with reference to the analysis conditions for Viterbi alignment shown in FIG. 9. The 25th order features are extracted in this embodiment, including the 12th order MFCC, the 12th order .DELTA.MFCC, and .DELTA. power.” (GOTO [0100]). “In this example, the English lyrics A of "Nothing untaken. Nothing lost" are converted into a sequence of English phonemes B of "N AA TH IH NG AH N T EY K AH N N AA TH IH NG L A O S T". Then, short pauses (sp) are combined with the sequence of phonemes B to form a sequence of phonemes for alignment C. The sequence of phonemes for alignment C is a phoneme network SN.” (GOTO [0108]). determining time information of the text unit based on playback duration of the audio frame; “...Therefore, according to the present invention, lyric data tagged with time information that is synchronized with the music audio signal may automatically be generated using an output from the alignment means.” (GOTO [0044]). and generating a lyric file of the song based on the time information of the text unit, wherein the lyric file is used to indicate a display of the text unit when the song is played to a location indicated by the time information. “In a music audio signal reproducing apparatus which reproduces a music audio signal while displaying on a screen lyrics temporally aligned with the music audio signal to be reproduced, the computer program of the present invention can be run for temporal alignment between a music audio signals and lyrics. The lyrics are displayed on a screen after the lyrics have been tagged with time information. When the lyrics are displayed on the screen, a portion of the displayed lyrics is selected with a pointer. In this manner, the music audio signal may be reproduced from that point, based on the time information corresponding to the selected lyric portion. Alternatively, lyrics tagged with time information is generated in advance by the system of the present invention may be stored in storage means such as a hard disc provided in a music audio signal reproducing apparatus, or may be stored in a server over the network. The lyrics tagged with time information that have been acquired from the storage means or the server over the network may be displayed on the screen in synchronization with music digital data reproduced by the music audio signal reproducing apparatus.” (GOTO [0060]). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 2, 3, 4, 21, 22 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view of ZHENG, Gui-tao (CN 110148427 A) hereinafter ZHENG. Regarding claim 2, GOTO teaches the lyric file generation method according to claim 1, furthermore GOTO teaches: obtaining the phoneme corresponding to the text unit from the phoneme set corresponding to the text unit based on a pronunciation of the text unit in the song; “... A representing phrases of original lyrics are converted into a sequence of phonemes B. Then, the sequence is further converted into a sequence of phonemes for alignment C only including phonemes used to identify the English phonemes (N, AA, TH . . . ) and short pauses (sp) by applying the two rules described above to the English lyrics converted into the sequence of phonemes B.” (GOTO [0107]). “...Further, a phoneme network is stored in phoneme network storage means (in the storage step). The phoneme network is constituted from a plurality of phonemes corresponding to the music audio signal and temporal intervals between two adjacent phonemes are connected in such a manner that the temporal intervals can be adjusted. Then, alignment means is provided with a phone model for singing voice that estimates a phoneme corresponding to the temporal-alignment feature, based on the temporal-alignment feature, and performs an alignment operation that makes the temporal alignment between the plurality of phonemes in the phoneme network and the dominant sound audio signals (in the alignment step). ...” (GOTO [0058]). and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. “...This final selection, or finally selected assumed sequence of phonemes has been defined based on the features corresponding to the time. In other words, the finally selected sequence of phonemes is a sequence of phonemes synchronized with the music audio signal. Therefore, lyric data to be displayed based on the finally selected sequence of phonemes will be "lyrics tagged with time information" or lyrics having time information required for synchronization with the music audio signal.” (GOTO [0115]). GOTO does not teach, but ZHENG teaches: wherein the obtaining the phoneme propagation sequence of the song comprises: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, wherein the pronunciation lexicon comprises a correspondence between the text unit and the phoneme set, and the phoneme set corresponding to the text unit is a set of phoneme(s) corresponding to pronunciation(s) of the text unit; “When it is necessary to obtain the phonemes (i.e. pronunciation) of reference words, the audio processing device can call the corresponding pronunciation dictionary according to the application scenario, and then obtain the phoneme sequence corresponding to each reference word according to the pronunciation dictionary. For example, in an intelligent assisted teaching scenario for spoken English, the audio processing device can call up the corresponding English pronunciation dictionary and look up the phonemes of the words I, am, and OK.” (ZHENG [0081]). “S204. Synthesize the phoneme sequences corresponding to all reference words in the word sequence to form a reference audio.” (ZHENG [0082]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to obtain phonemes of the words by accessing a pronunciation/lexicon dictionary. The benefit and motivation of such modification is discussed by ZHENG in the following portion: “The audio processing device can find the phoneme sequence corresponding to each reference word in the word sequence by using a pronunciation dictionary. Specifically, the audio processing device can build different pronunciation dictionaries for different scenarios.” (ZHENG [0080]). Regarding claim 3, GOTO teaches the lyric file generation method according to claim 1, furthermore GOTO teaches: the phoneme set corresponding to the text unit being a set composed of phoneme subset(s) corresponding to pronunciation(s) of the text unit, “In this example, the English lyrics A of "Nothing untaken. Nothing lost" are converted into a sequence of English phonemes B of "N AA TH IH NG AH N T EY K AH N N AA TH IH NG L A O S T". Then, short pauses (sp) are combined with the sequence of phonemes B to form a sequence of phonemes for alignment C. The sequence of phonemes for alignment C is a phoneme network SN.” (GOTO [0108]). and a phoneme subset corresponding to any of the pronunciation(s) being a set comprising permutation and combinations of various pronunciation durations of the phoneme(s) corresponding to the pronunciation; “Instep ST103, loop 1 is performed on all of the assumed sequences of phonemes. Loop 1 is to calculate scores for each of the assumed sequences as of the time that the previous frame has been processed. For example, it is assumed that temporal alignment should be made in connection with a phoneme network of "a-i-sp-u-e . . . ". In this example, a possible assumed sequence of phonemes up to the sixth frame or the sixth phoneme may be "a a a a a a" or "a a a i i i" or "a a u u sp u" or others. In the process of the search, these possible assumed sequences are retained at the same time and calculation is performed on all of the assumed sequences. “ (GOTO [0112]). obtaining a target phoneme subset of the text unit from the phoneme set corresponding to the text unit based on a pronunciation of the text unit in the song; “In step ST103, loop 1 is performed on all of the currently assumed sequences of phonemes. Loop 1 is to calculate scores for each of the currently assumed sequences as of the point of time that the previous frame has been processed. For example, it is assumed that temporal alignment should be performed in connection with a phoneme network of "a-i-sp-u-e . . . ". In this example, a possible assumed sequence of phonemes up to the sixth frame or the sixth phoneme may be "a a a a a a" or "a a a i i i" or "a a u u sp u" or others. In the process of the search, these possible assumed sequences are retained at the same time and calculation is performed on all of the assumed sequences. These assumed sequences have respective scores. Assuming that there are six frames, the score is obtained from calculations of possibilities or log likelihoods that features of each frame up to the sixth frame may correspond to, for example, a sequence of phonemes of "a a a i i i" by comparing the features with a phone model. For example, once the sixth frame (t=6) has been processed and then processing of the seventh frame is started, calculations are done on all of the currently retained assumed sequences. The processing as described above is Loop 1.” (GOTO [0116]). obtaining the phoneme corresponding to the text unit from the target phoneme subset of the text unit based on a pronunciation duration of the text unit in the song; “The phoneme network storage means stores a phoneme network constituted from a plurality of phonemes and short pauses in respect of lyrics of the music corresponding to the music audio signal. For example, lyrics are converted into a sequence of phonemes, phrase boundaries are converted into a plurality of short pauses, and a word boundary is converted into one short pause. Thus, a phoneme network is constituted. Preferably, Japanese lyrics may be converted into a sequence of phonemes including only vowels and short pauses. Preferably, English lyrics may be converted into a sequence of phonemes including English phonemes and short pauses.” (GOTO [0042]). and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. “The phoneme network storage means stores a phoneme network constituted from a plurality of phonemes and short pauses in respect of lyrics of the music corresponding to the music audio signal. For example, lyrics are converted into a sequence of phonemes, phrase boundaries are converted into a plurality of short pauses, and a word boundary is converted into one short pause. Thus, a phoneme network is constituted. Preferably, Japanese lyrics may be converted into a sequence of phonemes including only vowels and short pauses. Preferably, English lyrics may be converted into a sequence of phonemes including English phonemes and short pauses.” (GOTO [0042]). “...This final selection, or finally selected assumed sequence of phonemes has been defined based on the features corresponding to the time. In other words, the finally selected sequence of phonemes is a sequence of phonemes synchronized with the music audio signal. Therefore, lyric data to be displayed based on the finally selected sequence of phonemes will be "lyrics tagged with time information" or lyrics having time information required for synchronization with the music audio signal.” GOTO does not teach, but ZHENG teaches: wherein the obtaining the phoneme propagation sequence of the song comprises: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, wherein the pronunciation lexicon comprises a correspondence between the text unit and the phoneme set, “When it is necessary to obtain the phonemes (i.e. pronunciation) of reference words, the audio processing device can call the corresponding pronunciation dictionary according to the application scenario, and then obtain the phoneme sequence corresponding to each reference word according to the pronunciation dictionary. For example, in an intelligent assisted teaching scenario for spoken English, the audio processing device can call up the corresponding English pronunciation dictionary and look up the phonemes of the words I, am, and OK.” (ZHENG [0081]). “S204. Synthesize the phoneme sequences corresponding to all reference words in the word sequence to form a reference audio.” (ZHENG [0082]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to obtain phonemes of the words by accessing a pronunciation/lexicon dictionary. The benefit and motivation of such modification is discussed by ZHENG in the following portion: “The audio processing device can find the phoneme sequence corresponding to each reference word in the word sequence by using a pronunciation dictionary. Specifically, the audio processing device can build different pronunciation dictionaries for different scenarios.” (ZHENG [0080]). Regarding claim 4, GOTO teaches the lyric file generation method according to claim 1, furthermore GOTO teaches: the phoneme set corresponding to the text unit being a set composed of phoneme subset(s) corresponding to pronunciation(s) of the text unit, “The phoneme network storage means stores a phoneme network constituted from a plurality of phonemes and short pauses in respect of lyrics of the music corresponding to the music audio signal. For example, lyrics are converted into a sequence of phonemes, phrase boundaries are converted into a plurality of short pauses, and a word boundary is converted into one short pause. Thus, a phoneme network is constituted. Preferably, Japanese lyrics may be converted into a sequence of phonemes including only vowels and short pauses. Preferably, English lyrics may be converted into a sequence of phonemes including English phonemes and short pauses.” (GOTO [0042]). and a phoneme subset corresponding to any of the pronunciation(s) being a set comprising permutation and combination of various phoneme(s) and various phoneme(s) with transposition phoneme corresponding to the pronunciation; “Instep ST103, loop 1 is performed on all of the assumed sequences of phonemes. Loop 1 is to calculate scores for each of the assumed sequences as of the time that the previous frame has been processed. For example, it is assumed that temporal alignment should be made in connection with a phoneme network of "a-i-sp-u-e . . . ". In this example, a possible assumed sequence of phonemes up to the sixth frame or the sixth phoneme may be "a a a a a a" or "a a a i i i" or "a a u u sp u" or others. In the process of the search, these possible assumed sequences are retained at the same time and calculation is performed on all of the assumed sequences. These assumed sequences have their own scores. Assuming that there are six frames, the score is obtained from calculations of possibilities or log likelihoods that features of each frame up to the sixth frame may be, for example, a sequence of phonemes of "a a a i i i" by comparing the features with a phone model. For example, when the sixth frame (t=6) has been processed and then processing of the seventh frame is started, calculations are done on all of the currently retained assumed sequences. The processing as described above is Loop 1.” (GOTO [0112]). “… Next, temporal-alignment feature extraction means extracts a temporal-alignment feature suitable to make temporal alignment between lyrics of the vocal and the music audio signal from the dominant sound audio signal at each time (in the temporal-alignment feature extraction step). Further, a phoneme network is stored in phoneme network storage means (in the storage step). The phoneme network is constituted from a plurality of phonemes corresponding to the music audio signal and temporal intervals between two adjacent phonemes are connected in such a manner that the temporal intervals can be adjusted. …” (GOTO [0058]). obtaining a target phoneme subset of the text unit from the phoneme set corresponding to the text unit based on a pronunciation thereof in the song; “Instep ST103, loop 1 is performed on all of the assumed sequences of phonemes. Loop 1 is to calculate scores for each of the assumed sequences as of the time that the previous frame has been processed. For example, it is assumed that temporal alignment should be made in connection with a phoneme network of "a-i-sp-u-e . . . ". In this example, a possible assumed sequence of phonemes up to the sixth frame or the sixth phoneme may be "a a a a a a" or "a a a i i i" or "a a u u sp u" or others. In the process of the search, these possible assumed sequences are retained at the same time and calculation is performed on all of the assumed sequences. These assumed sequences have their own scores. Assuming that there are six frames, the score is obtained from calculations of possibilities or log likelihoods that features of each frame up to the sixth frame may be, for example, a sequence of phonemes of "a a a i i i" by comparing the features with a phone model. For example, when the sixth frame (t=6) has been processed and then processing of the seventh frame is started, calculations are done on all of the currently retained assumed sequences. The processing as described above is Loop 1.” (GOTO [0112]). obtaining the phoneme corresponding to the text unit from the target phoneme subset of the text unit based on a tone conversion of the text unit in the song; “FIG. 10B shows an example of converting English lyrics into a sequence of phonemes for alignment (phoneme network). In this example, the English lyrics are represented by English phonemes. Most preferably, an English phone model may be used for English lyrics using English phonemes. However, a Japanese phone model may be used for English lyrics if English phonemes are converted into Japanese phonemes. In an example of FIG. 10B, first, text data A representing phrases of original lyrics are converted into a sequence of phonemes B. Then, the sequence is further converted into a sequence of phonemes for alignment C only including phonemes used to identify the English phonemes (N, AA, TH . . . ) and short pauses (sp) by applying the two rules described above to the English lyrics converted into the sequence of phonemes B.” (GOTO [0107]). and generating the phoneme propagation sequence based on the phoneme corresponding to the text unit. “...This final selection, or finally selected assumed sequence of phonemes has been defined based on the features corresponding to the time. In other words, the finally selected sequence of phonemes is a sequence of phonemes synchronized with the music audio signal. Therefore, lyric data to be displayed based on the finally selected sequence of phonemes will be "lyrics tagged with time information" or lyrics having time information required for synchronization with the music audio signal.” (GOTO [0115]). GOTO does not teach, but ZHENG teaches: wherein the obtaining the phoneme propagation sequence of the song comprises: obtaining a phoneme set corresponding to the text unit based on a pronunciation lexicon, wherein the pronunciation lexicon comprises a correspondence between the text unit and the phoneme set, “When it is necessary to obtain the phonemes (i.e. pronunciation) of reference words, the audio processing device can call the corresponding pronunciation dictionary according to the application scenario, and then obtain the phoneme sequence corresponding to each reference word according to the pronunciation dictionary. For example, in an intelligent assisted teaching scenario for spoken English, the audio processing device can call up the corresponding English pronunciation dictionary and look up the phonemes of the words I, am, and OK.” (ZHENG [0081]). “S204. Synthesize the phoneme sequences corresponding to all reference words in the word sequence to form a reference audio.” (ZHENG [0082]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to obtain phonemes of the words by accessing a pronunciation/lexicon dictionary. The benefit and motivation of such modification is discussed by ZHENG in the following portion: “The audio processing device can find the phoneme sequence corresponding to each reference word in the word sequence by using a pronunciation dictionary. Specifically, the audio processing device can build different pronunciation dictionaries for different scenarios.” (ZHENG [0080]). Regarding claim 21, the rejection of claim 17 is incorporated, furthermore arguments analogous to claim 2 are applicable. Regarding claim 22, the rejection of claim 17 is incorporated, furthermore arguments analogous to claim 3 are applicable. Regarding claim 23, the rejection of claim 17 is incorporated, furthermore arguments analogous to claim 4 are applicable. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view of ZHUANG, Xiao-bin (CN 111210850 A) hereinafter ZHUANG. Regarding claim 5, GOTO teaches the lyric file generation method according to claim 1, furthermore GOTO does not teach, but ZHUANG teaches: wherein the obtaining the audio frame sequence of the song comprises: sampling audio signals of the song based on a preset sampling frequency and a preset format to obtain a sampling sequence of the song; “102: The lyrics alignment device processes the human voice according to a preset time window to obtain N audio frames.” (ZHUANG [0046]). “Generally, the sampling frequency of the human voice separated from the song is 44.1KHz, and the sampling frequency of the target human voice obtained after downsampling is 16KHz. This reduces the amount of data required for subsequent lyric data matching, thereby improving the accuracy of lyric data matching. ” (ZHUANG [0050]). and generating the audio frame sequence of the song based on a duration of the audio frame and the sampling sequence. “The human voice is processed according to a preset time window to obtain N audio frames; ” (ZHUANG [0011]). “Specifically, the playback duration of the song is first divided according to the preset time window to obtain N playback time segments. That is, the playback duration is divided into N playback time segments according to the processing method of dividing the frequency domain signal into frames, so the N playback time segments correspond one-to-one with the N audio frames. ” (ZHUANG [0054]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to control the sampling frequency a song and generating the frames related to such sampling frequency. The benefit and motivation of such modification is discussed by ZHUANG in the following portion: “ …This reduces the amount of data required for subsequent lyric data matching, thereby improving the accuracy of lyric data matching. ” (ZHUANG [0050]). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view of ZHUANG in further view of ZHENG in further view of ZHAO; Weifeng ( US 20180349495 A1) hereinafter ZHAO. Regarding claim 5, GOTO in view of ZHUANG teaches the lyric file generation method according to claim 5, furthermore GOTO does not teach, but ZHUANG teaches: wherein: the preset sampling frequency is 16kHz; Zhuang [0050] Generally, the sampling frequency of the human voice separated from the song is 44.1KHz, and the sampling frequency of the target human voice obtained after downsampling is 16KHz. This reduces the amount of data required for subsequent lyric data matching, thereby improving the accuracy of lyric data matching. It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to control the sampling frequency a song and generating the frames related to such sampling frequency. The benefit and motivation of such modification is discussed by ZHUANG in the following portion: “ …This reduces the amount of data required for subsequent lyric data matching, thereby improving the accuracy of lyric data matching. ” (ZHUANG [0050]). GOTO in view of ZHUANG does not teach, but ZHENG teaches: the preset format is … mono Wave Pulse Code Modulation (PCM) format. ”To improve the accuracy and efficiency of audio recognition, audio processing devices can perform noise filtering on target speech data to obtain target audio. Specifically, the audio processing device can use a noise filtering algorithm to filter the target speech data to obtain the target audio. The noise filtering algorithms here include Voice Activity Detection (VAD), etc., and the target audio can be an audio file in Pulse Code Modulation (PCM) format.” (ZHENG [0074]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO in view of ZHUANG the capability to control the audio format. The benefit and motivation of such modification is discussed by ZHUANG in the following portion: ”To improve the accuracy and efficiency of audio recognition, audio processing devices can perform noise filtering on target speech data to obtain target audio. ...” (ZHENG [0074]). GOTO in view of ZHUANG in further view of ZHENG do not teach, but ZHAO teaches: and the preset format is a 16-bit “For example, it is assumed that both the accompaniment audio and the user audio (that is, the spliced audio data) are in a format of 44 k and 16 bit. …” (ZHAO [0113]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO in view of ZHUANG in further view of ZHENG the capability to control the audio format. The benefit and motivation of such modification is discussed by ZHAO in the following portion: “… As such, it will be understood that such steps and operations, which are at times referred to as being computer-executed, include the manipulation by the processing unit of the computer of electrical signals representing data in a structured form. This manipulation converts the data or maintains the location of the data in a memory system of the computer, which can be reconfigured, or otherwise a person skilled in this art changes the way of operation of the computer in a well-known manner.” Claims 7 is rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view of ZHUANG in further view of Todic; Ognjen (US 20110288862 A1) hereinafter TODIC. Regarding claim 7, GOTO teaches the lyric file generation method according to claim 1, furthermore GOTO does not teach, but ZHUANG teaches: further comprising: performing a Fourier transform on each audio frame in the audio frame sequence to obtain a Fourier transform spectrum for the each audio frame before determining the audio frames corresponding to the text units in the audio frame sequence; “Furthermore, a Fourier transform (including short-time Fourier transform and fast Fourier transform) is performed on the target voice to obtain the corresponding frequency domain signal. Then, a preset time window (window function) is used to perform frame processing on the frequency domain signal to obtain N audio frames. ” (ZHUANG [0051]). separating a human voice spectrum from an accompaniment spectrum in the Fourier transform spectrum of the each audio frame to obtain a human voice spectrum of the each audio frame; “201: The lyrics alignment device performs a Fourier transform on the song to obtain the first spectrum of the song.” (ZHUANG [0075]). “202: The lyrics alignment device inputs the first spectrogram into the neural network to obtain the second spectrogram of the human voice and the third spectrogram of the accompaniment. ” (ZHUANG [0080]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to obtain the frames of audio and to apply the Fourier transform to them. The benefit and motivation of such modification is discussed by ZHUANG in the following portion: “This application provides a lyrics alignment method and related products, which automatically aligns lyrics using multiple audio frames and a marked sequence for each lyric data, thereby improving the efficiency and intelligence of lyrics alignment.” (ZHUANG [0008]). GOTO in view of ZHUANG do not teach, but TODIC teaches: and converting the human voice spectrum of the each audio frame into a Mel Frequency Cepstrum Coefficient (MFCC) feature to obtain an audio feature of the each audio frame. “The audio engine 102 may analyze the suppressed audio signal by extracting feature vectors about every 10 ms (e.g., using Mel Frequency Cepstral Coefficients or (MFCC)). …” (TODIC [0035]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to compute Mel Frequency Cepstral Coefficients to extract audio features. The benefit and motivation of such modification is discussed by ZHUANG in the following portion: “In still another example aspect, a system is provided that comprises a Hidden Markov Model (HMM) database that may include statistical modeling of phonemes in a multidimensional feature space (e.g. using Mel Frequency Cepstral Coefficients), an optional expected grammar that defines words that a speech decoder can recognize, a pronunciation dictionary database that maps words to the phonemes, and a speech decoder. …” (TODIC [0009]). Claims 8 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view of TODIC. Regarding claim 8, GOTO teaches the lyric file generation method according to claim 1, furthermore GOTO does not teach, but TODIC teaches: wherein the determining the audio frame corresponding to the text unit in the audio frame sequence comprises: obtaining matching relationships between audio features and phonemes in the phoneme propagation sequence based on a matching model, “The audio engine 102 may analyze the suppressed audio signal by extracting feature vectors about every 10 ms (e.g., using Mel Frequency Cepstral Coefficients or (MFCC)). The ASR decoder 104 may then map the sequence of feature vectors to the expected response defined in the grammar. The ASR decoder 104 will expand the word grammar created by the grammar processor 108 into a phonetic grammar by using the dictionary database 110 to expand words into phonemes. The ASR decoder 104 may use a Hidden Markov Model (HMM) database 112 that statistically describes each phoneme in the features space (e.g., using MFCC) to obtain an optimal sequence of words from the phonemes that matches the grammar of the audio signal and corresponding feature vector. Although the HMM database 112 is illustrated separate from the system 100, in other examples, the HMM database 112 may be a component of the system 100 or may be contained within components of the system 100.” (TODIC [0035]). wherein the matching model is a model obtained by training a neural network model based on training samples comprising matched audio features and phonemes; “The ASR decoder 104 may use an HMM from the database 112 to decode the audio signal using a Viterbi decoding algorithm that determines an optimal sequence of text given the audio signal, expected grammar, and a set of HMMs that are trained on a large set of data, for example. Thus, the ASR decoder 104 uses the HMM database 112 of phonemes to map spoken words to a phonetic description, and uses the dictionary database 110 to map words to the phonetic description, for example. ” (TODIC [0037]). and determining the audio frame corresponding to the text unit in the audio frame sequence based on the matching relationships between various audio features and phonemes in the phoneme propagation sequence, as well as the phoneme corresponding to the text unit. “The ASR decoder 104 may determine timing information as a by-product of performing speech recognition. For example, a Viterbi decoder determines an optimal path through a matrix in which a vertical dimension represents HMI states and a horizontal dimension represents frames of speech (e.g., 10 ms). When an optimal sequence of HMM states is determined, an optimal sequence of corresponding phonemes and words is available. Because each pass through the HMM state consumes a frame of speech, the timing information at the state/phoneme/word level is available as the output of the automated speech recognition. ” (TODIC [0045]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to incorporate a trained model in the invention to do the matching between lyric, phonemes and audio. The benefit and motivation of such modification is discussed by TODIC in the following portion: “The ASR decoder 104 may use an HMM from the database 112 to decode the audio signal using a Viterbi decoding algorithm that determines an optimal sequence of text given the audio signal, expected grammar, and a set of HMMs that are trained on a large set of data, for example. Thus, the ASR decoder 104 uses the HMM database 112 of phonemes to map spoken words to a phonetic description, and uses the dictionary database 110 to map words to the phonetic description, for example. ” (TODIC [0045]). Regarding claim 10, GOTO in view of TODIC teaches the lyric file generation method according to claim 8, furthermore GOTO teaches: wherein the determining the time information of the text unit based on the playback duration of the audio frame comprises: obtaining time information of the various phonemes in the phoneme propagation sequence based on playback durations of the audio frames and the matching relationships between the various phonemes in the phoneme propagation sequence and the audio features; “...Next, the temporal-alignment feature extraction means 11 extracts a temporal-alignment feature suitable to make temporal alignment between lyrics of the vocal and the music audio signal from the dominant sound audio signal S2 at each time (in the temporal-alignment feature extraction step). Further, a phoneme network SN is stored in phoneme network storage means 13 (in the storage step). The phoneme network SN is constituted from a plurality of phonemes corresponding to the music audio signal S1 and temporal intervals between two adjacent phonemes are connected in such a manner that the temporal intervals can be adjusted. Then, alignment means 17 is provided with the phone model 15 for singing voice that estimates a phoneme corresponding to the temporal-alignment feature, based on the temporal-alignment feature, and performs an alignment operation that makes the temporal alignment between the plurality of phonemes in the phoneme network SN and the dominant sound audio signals S2 (in the alignment step). In the alignment step, the alignment means 17 receives the temporal-alignment feature obtained in the step of extracting the temporal-alignment feature, the information on the vocal section and the non-vocal section, and the phoneme network SN, and performs the alignment operation using the phone model 15 for singing voice on condition that no phoneme exists at least in the non-vocal section. ” (GOTO [0132]). and determining the time information of the text unit based on the time information of the various phonemes in the phoneme propagation sequence and the phoneme corresponding to the text unit. “In a music audio signal reproducing apparatus which reproduces a music audio signal while displaying on a screen lyrics temporally aligned with the music audio signal to be reproduced, the computer program of the present invention can be run for temporal alignment between a music audio signals and lyrics. The lyrics are displayed on a screen after the lyrics have been tagged with time information. When the lyrics are displayed on the screen, a portion of the displayed lyrics is selected with a pointer. In this manner, the music audio signal may be reproduced from that point, based on the time information corresponding to the selected lyric portion. Alternatively, lyrics tagged with time information is generated in advance by the system of the present invention may be stored in storage means such as a hard disc provided in a music audio signal reproducing apparatus, or may be stored in a server over the network. The lyrics tagged with time information that have been acquired from the storage means or the server over the network may be displayed on the screen in synchronization with music digital data reproduced by the music audio signal reproducing apparatus. ” (GOTO [0060]). “...This final selection, or finally selected assumed sequence of phonemes has been defined based on the features corresponding to the time. In other words, the finally selected sequence of phonemes is a sequence of phonemes synchronized with the music audio signal. Therefore, lyric data to be displayed based on the finally selected sequence of phonemes will be "lyrics tagged with time information" or lyrics having time information required for synchronization with the music audio signal.” (GOTO [0115]). Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view TODIC of in further view of ZHUANG. Regarding claim 9, GOTO in view of TODIC teaches the lyric file generation method according to claim 8, furthermore GOTO does not teach, but TODIC teaches: obtaining a genre of the song based on the accompaniment spectrum of the each audio frame; “A result of a one-time training process is a database of different Hidden Markov Models each of which may include metadata specifying a specific genre, tempo, gender of the trained data, for example.” (TODIC [0098]). obtaining a target matching model, the target matching model being a matching model corresponding to the genre of the song; “In another example embodiment, methods described herein may be used to train data-specific HMMs to be used to recognize corresponding audio signals. For example, rather than using a general HMM for a given song, selection of a most appropriate model for a given song can be made. Multiple Hidden Markov models can be trained on subsets of training data using song metadata information (e.g., genre, singer, gender, tempo, etc.) as a selection criteria. FIG. 9 is a block diagram illustrating a hierarchical HMM training and model selection. An initial HMM training set 902 may be further adapted using genre information to generate separate models trained for a hip-hop genre 904, a pop genre 906, a rock genre 908, and a dance genre 910. The genre HMMs may be further adapted to a specific tempo, such as slow hip-hop songs 912, fast hip-hop songs 914, slow dance songs 916, and fast dance songs 918. Still further, these HMMs may be adapted based on a gender of a performer, such as a slow dance song with female performer 920 and slow dance song with male performer 922. Corresponding reverse models could also be trained using the training sets with reversed audio, for example. ” (TODIC [0097]). and obtaining the matching relationships between the various audio features and the phonemes in the phoneme propagation sequence based on the target matching model. “FIG. 11 is a block diagram illustrating a parallel audio and lyric synchronization system 1100. The system 1100 includes a number of aligners (1, 2, . . . , N), each of which receives a copy of an input audio signal and corresponding lyrics text. The aligners operate to output time-annotated synchronized audio and lyrics, and may be or include any of the components as described above in system 100 of FIG. 1 or system 200 of FIG. 2. Each of the aligners may operate using different HMMs models (such as the different HMMs described in FIG. 9), and there may a number of aligners equal to a number of different possible HMMs. ” (TODIC [0102]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to distinguish between different genre of music to select the most appropriate model. The benefit and motivation of such modification is discussed by TODIC in the following portion: “… Multiple Hidden Markov models can be trained on subsets of training data using song metadata information (e.g., genre, singer, gender, tempo, etc.) as a selection criteria. FIG. 9 is a block diagram illustrating a hierarchical HMM training and model selection. An initial HMM training set 902 may be further adapted using genre information to generate separate models trained for a hip-hop genre 904, a pop genre 906, a rock genre 908, and a dance genre 910. The genre HMMs may be further adapted to a specific tempo, such as slow hip-hop songs 912, fast hip-hop songs 914, slow dance songs 916, and fast dance songs 918. …”(TODIC [0097]). GOTO in view of TODIC do not teach, but ZHUANG teaches: wherein the obtaining the matching relationships between the various audio features and the phonemes in the phoneme propagation sequence based on the matching model comprises: performing a Fourier transform on each audio frame in the audio frame sequence to obtain a Fourier transform spectrum for the each audio frame; “Furthermore, a Fourier transform (including short-time Fourier transform and fast Fourier transform) is performed on the target voice to obtain the corresponding frequency domain signal. Then, a preset time window (window function) is used to perform frame processing on the frequency domain signal to obtain N audio frames. ” (ZHUANG [0051]). separating a human voice spectrum and an accompaniment spectrum in the Fourier transform spectrum of the each audio frame to obtain the accompaniment spectrum of the each audio frame; “201: The lyrics alignment device performs a Fourier transform on the song to obtain the first spectrum of the song. ” (ZHUANG [0075]). “202: The lyrics alignment device inputs the first spectrogram into the neural network to obtain the second spectrogram of the human voice and the third spectrogram of the accompaniment. ” (ZHUANG [0080]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to obtain the frames of audio and to apply the Fourier transform to them and separating the human voice and the accompaniment. The benefit and motivation of such modification is discussed by ZHUANG in the following portion: “This application provides a lyrics alignment method and related products, which automatically aligns lyrics using multiple audio frames and a marked sequence for each lyric data, thereby improving the efficiency and intelligence of lyrics alignment.” (ZHUANG [0008]). Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view of TODIC in further view of ZHAO. Regarding claim 11, GOTO in view of TODIC teaches the lyric file generation method according to claim 10, furthermore GOTO in view of TODIC does not teach, but ZHAO teaches: wherein: the time information of each phoneme of the phonemes in the phoneme propagation sequence comprises a start time and a duration of the each phoneme; “a segmentation subunit 3043, configured to: when it is determined that the lyric word count is the same as the word count of the text data, segment one or more phonemes of a word indicated by the text data, and determine time information that corresponds to the word, where the time information includes start time information and duration information.” (ZHAO [0131]). "In this application, the time information of each word may be specifically time information of corresponding phonemes of each word that includes, for example, start time information and duration information of a corresponding initial consonant and vowel. ” (ZHAO [0045]). and the time information of the text unit comprises a start time and a duration of the text unit. “a segmentation subunit 3043, configured to: when it is determined that the lyric word count is the same as the word count of the text data, segment one or more phonemes of a word indicated by the text data, and determine time information that corresponds to the word, where the time information includes start time information and duration information.” (ZHAO [0131]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO in view of TODIC the capability to indicate the start and duration times of words and phonemes. The benefit and motivation of such modification is discussed by ZHAO in the following portion: “Specifically, because the lyric file includes the start time and the duration that correspond to each word, the music score file includes the start time and the duration that correspond to each musical note, and the pitch of each musical note, and because each word may correspond to one or more musical notes, when a word corresponds to one musical note, a start time, duration, and pitch information that correspond to each word may be obtained from the music score file; or when a word corresponds to multiple musical notes, a start time, duration, and pitch of the word may be correspondingly obtained according to a start time, duration, and pitch of the multiple musical notes. However, the rap part of the song is not content that is to be sung but to be spoken, and therefore there is no pitch information. Therefore, after comparison is performed on the lyric file and the music score file by alignment, pitch that corresponds to each word may be obtained. If some of the words do not have pitch, these words may be determined as the rap part of the song.” (ZHAO [0040]). Claims 12 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over GOTO in view of TODIC in further view of Mahar; Matthew Dominick et al. (US 11245950 B1) hereinafter MAHAR Regarding claim 12, GOTO in view of TODIC teaches the lyric file generation method according to claim 8, furthermore GOTO in view of TODIC does not teach, but ZHAO teaches: further comprising: obtaining a confidence sequence, wherein various confidence scores in the confidence sequence are used to represent matching degrees between the audio features and the phonemes; “In still another example aspect, a system is provided that comprises a Hidden Markov Model (HMM) database that may include statistical modeling of phonemes in a multidimensional feature space (e.g. using Mel Frequency Cepstral Coefficients), an optional expected grammar that defines words that a speech decoder can recognize, a pronunciation dictionary database that maps words to the phonemes, and a speech decoder. The speech decoder receives an audio signal and accesses the HMM, expected grammars, and a dictionary to map vocal elements in the audio signal to words. The speech decoder further performs an alignment of the audio signal with corresponding textual transcriptions of the vocal elements, and determines timing boundary information associated with an elapsed amount of time for a duration of a portion of the vocal elements. The speech decoder further determines a confidence metric indicating a level of certainty for the timing boundary information for the duration of the portion of the vocal elements. ” (TODIC [0009]). determining whether the matching degrees between the phonemes and the audio features corresponding to the text unit are all less than a preset threshold based on the confidence sequence; “To determine a mismatch between the forward and reverse alignment, the confidence score engine 216 compares a difference between the forward and reverse timing information to a predefined threshold, and marks the line as a low or high confidence line in accordance with the comparison. Line timing information may be defined as T.sub.n.sup.BP where n is the line index, B defines a boundary type (S for start time, E end time) and P defines pass type (F for forward, R for reverse), then a start mismatch for line n is defined as: The mismatch metrics can then be compared to a predefined threshold to determine if the line should be flagged as a low or high confidence line. ” (TODIC [0065]). It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO the capability to have a confidence measurement and adjust the synchronization based on it. The benefit and motivation of such modification is discussed by TODIC in the following portion: “Methods and systems for performing audio synchronization with corresponding textual transcription and determining confidence values of the timing-synchronization are provided. Audio and a corresponding text (e.g., transcript) may be synchronized in a forward and reverse direction using speech recognition to output a time-annotated audio-lyrics synchronized data. Metrics can be computed to quantify and/or qualify a confidence of the synchronization. Based on the metrics, example embodiments describe methods for enhancing an automated synchronization process to possibly adapted Hidden Markov Models (HMMs) to the synchronized audio for use during the speech recognition. Other examples describe methods for selecting an appropriate HMM for use.” (TODIC [Abstract]). GOTO in view of TODIC do not teach, but MAHAR teaches: and in a case where the matching degrees of the phonemes and the audio features corresponding to a text unit are all less than the preset threshold value, adjusting the start times of all the text units after the text unit in the lyrics text forward by a preset time length. “As described herein, the service provider computers may also identify language or text errors included in the lyrics file 206 by comparing the transcribed words included in the generated timestamped file 212 to the words in the lyrics file 206. For example, the third party lyrics file may include misspelled words, include missing words that were detected and transcribed by the automated speech recognition 210, include incorrect words, or include phrases that are meant as shorthand for repeating choruses or other phrases such as “repeat x 10.” The language or text errors can be corrected by the service provider computer modifying the text of the lyrics included in the lyrics file 206 using the text or words included in the generated timestamped file 212. In embodiments, the service provider computer may apply a fixed offset to the time stamps of the lyrics file 206 to correct the synchronization error. The lyrics provided by in the lyrics file 206 may be represented as X={(x.sub.i, a.sub.i)}, where x.sub.i is an n-gram and a.sub.i is a corresponding time stamp. In embodiments, the lyrics file 206 may be provided by a third party or maintained by the service provider computers implementing the automated synchronization features described herein. The transcribed lyrics included in the generated timestamped file 212 may be represented as Y={y.sub.i, b.sub.i)}, where y.sub.i is an n-gram and b.sub.i is a corresponding time offset from the beginning of the audio for the audio file 202.” (MAHAR [Column 7 line 45 to column 8 line 2]) It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of GOTO in view of TODIC the capability to adjust the preset time between the audio and the lyric (text) file in order to synchronize them. The benefit and motivation of such modification is discussed by MAHAR in the following portion: “Some digital music content may include lyric sheets that list the lyrics corresponding to the music content. When played, a user can read the lyric sheet while listening to the content. Some music content may be presented via other channels such as via a streaming media device. In such cases, lyric sheets may not be offered. Current media content may include lyric files which can be used in an attempt to present synchronized lyrics. However, there may be various different versions of the music content and the corresponding lyric files can include synchronization errors.” (MAHAR [Column 1 line 5 to line 14]). Regarding claim 13, GOTO in view of TODIC in further view of MAHAR teaches the lyric file generation method according to claim 12, furthermore GOTO in view of TODIC does not tech, but MAHAR teaches: The lyric file generation method according to claim 12, wherein the preset time length is 0.3 seconds. “In response to identifying an offset between the time stamps 108 and 110 of the lyric file 100 and generated file 102, the service provider computer may modify the lyric file 100 by applying the determined offset to the time stamps 108 thereby resolving the synchronization error between the lyric file 100 and the corresponding media file. If no offset is detected then the lyric file 100 may be flagged or otherwise marked as synchronized with the corresponding media file. In embodiments, the service provider computer may maintain a threshold of time that represents the offset that a content streamer, content creator, or user would find acceptable for applying an offset correction. For example, the service provider computer may only apply a correction using a time offset when the detected offset 112 is greater than 500 milliseconds. In embodiments, the service provider computer, content streamer, content creator, or user may specify the threshold of time for applying offsets.” (MAHAR [Column 5 line 64 to column 6 line 15]). At the time of invention, it is known that there needs to be an offset to synchronize the lyrics. There are only finite values that can be applied which yields the predictable solution for the identified problem which is the non-alignment of the file. Mahar allows the user to select the specify the offset and user may select 0.3 seconds for this offset. One of the ordinary skills in the art before effective filing date could have pursued using the 0.3 second as a preset time since both the solution of user providing the preset time and using the 0.3 second present time would result in synchronization. Further motivation to include the preset time mentioned is discussed by MAHAR in: “Some digital music content may include lyric sheets that list the lyrics corresponding to the music content. When played, a user can read the lyric sheet while listening to the content. Some music content may be presented via other channels such as via a streaming media device. In such cases, lyric sheets may not be offered. Current media content may include lyric files which can be used in an attempt to present synchronized lyrics. However, there may be various different versions of the music content and the corresponding lyric files can include synchronization errors.” (MAHAR [Column 1 line 5 to line 14]). Allowable Subject Matter Claims 14 and 15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to HECTOR J. CRESPO FEBLES whose telephone number is (571)272-4512. The examiner can normally be reached Mon - Fri 7:30 - 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /H.J.C./ Examiner, Art Unit 2657 /DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Aug 18, 2023
Application Filed
Jun 16, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 1m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month