DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-10 and 12-21 are rejected under rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 12 and 13 each initially recite “a music generation control” whose trigger operations causes display of a music generation interface, and subsequently recite that the displayed music generation interface includes “a music generation control”, thereby introducing two potential antecedents for the later recitation of “the music generation control”. Accordingly, it is unclear which music generation control is subsequently triggered to generate the voice and music.
For purposes of examination, Examiner interprets the subsequently recited “the music generation control” as referring to the music generation control included in the displayed music generation interface.
Claims 2-10 and 14-21 depend directly or indirectly therefrom and inherit the ambiguity.
Claims 7, 10, and 21 recite, in pertinent part, “generating a song [or music] with the music format based on a music format corresponding to the music format configuration item.” The recitation is unclear because “the music format” is recited prior to the introduction of “a music format,” and it is therefore unclear what particular music format is referenced by “the music format” and its relationship to the subsequently recited “a music format corresponding to the music format configuration item”
For the purposes of examination, the Examiner interprets “the music format” as referring to the subsequently recited “a music format corresponding to the music format configuration item.”
Claims 8, 9, 19 and 20, lines 7-8, 1-2 ,7-8 and 1-2, respectively recite “generating a music including the voice for the slogan based on the voice for the slogan...” The recitation is unclear because it is unclear what is intended by generating music “including the voice for the slogan based on the voice for the slogan”, particularly with respect to the relationship between the custom text input in the slogan input box, the generated voice for the slogan, and the generated music.
For the purposes of examination, the Examiner interprets the recitation “generating a music including the voice for the slogan based on the voice for the slogan” as “generating a music including the voice for the slogan based on the custom text input in the slogan input box”.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1,3,5,8,12,13,15,17 and 19 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US20200372896 (Cui), hereinafter US’896.
Regarding claim 1, US’896 discloses ‘A music generation method (US’896, [Abstract]:”The method includes: obtaining a target text; determining a target song … synthesizing a self-made song using the target text and tune information … the target text being used as the lyrics of the self-made song), comprising:
in response to a trigger operation of a music generation control, displaying a music generation interface (US’US’896,Fig. 3 depicts the user interface containing “User-created lyrics” and “Select music”; ¶[0062]:a video recording application is run “according to a user instruction” and that the terminal continues to enter a “background audio configuration interface (for example, the interface shown in FIG. 3) according to the user instruction”),
the music generation interface (US including a text input box, a music generation control (US’896, Fig. 3; ¶[0062]: a background-audio configuration interface (FIG. 3) through which the terminal obtains target text input by the user and a target song selected by the user, “synthesize the self-made song by using the target text as the lyrics and in combination with the determined tune information”;¶¶[0034]-[0042]) and at least one music configuration item (US’896, ¶[0041]: Fig. 3 includes candidate-song list 330. User selects song 331 as the target song, and the target song provides the tune for the self-made song);
in response to a trigger operation of the text input box, obtaining a custom text input by a user (US’896, ¶[0062], [0064]-[0065]: inputting the target text and tune information into the tune-control model and outputting the self-made song after “speech synthesis is performed on each character in the target text according to the corresponding tune feature”),
and in response to a trigger operation of the at least one music configuration item, determining configuration information corresponding to the at least one music configuration item (US’896, ¶[0039]: the interface terminal monitors a “predetermined trigger operation” on displayed candidate songs; trigger may be a touch operation or cursor clicking operation; ¶[0040]:the resulting selected candidate song becomes the target song, ¶[0042]:and tune information is obtained);
and in response to a trigger operation of the music generation control, generating a voice based on the custom text (US’896, ¶[0064]:”inputting the target text and the tune information into the tune control model”; “the self-made song obtained after speech synthesis is performed on each character in the target text according to the corresponding tune feature; ¶[0087]:”determining a phoneme sequence corresponding to the target text; synthesizing a self-made speech according to the phoneme sequence by using the timbre control model”),
and determining a music melody based on the configuration information corresponding to the at least one music configuration item (US’896, ¶[0037]:”is used for providing a tune for the to-be-synthesized self-made song”; ¶[0065]:”searches for the tune information of the target song”… “determine the tune feature corresponding to each character in the target text”),
generating a music including the voice corresponding to the custom text based on the voice and the music melody (US’896, ¶[0064]:”the self-made song obtained after speech synthesis is performed on each character in the target text according to the corresponding tune feature”).
Regarding claim 3, US’896 discloses ‘The method according to claim 1, as discussed above.
US’896 further discloses ‘wherein, the music configuration items can be preset with configuration information (US’896, ¶[0038]-[0041], [0062]: interface provides candidate songs from which a target song is selected; the selected song supplies the tune information used in generation; alternative selection mechanism which the terminal can determine a target song without an affirmative user selection), and the method further comprising:
in case that the user does not trigger the at least one music configuration item (US’896, ¶[0040]:an alternative to the user’s selection instruction “receive a selection instruction of the user for a candidate song, to obtain the candidate song selected by the selection instruction as the target song”), in response to the trigger operation of the music generation control (US’896, ¶[0062]: entering the configuration interface “according to user instruction”, obtaining the targe text/target song synthesizing the self-made song using the target text and “determined tune information”), generating a voice based on the custom text (US’896, ¶¶[0064]-[0065]:”inputting the target text and the tune information into the tune control model” and performing ”speech synthesis is performed on each character in the target text according to the corresponding tune feature”).
and determining a music melody based on the preset configuration information of the at least one music configuration item (US’896, ¶[0040]:”the terminal may alternatively randomly select a song from the candidate songs as the target song…”;¶[0047]: the tune information of the target song can be extracted from the song file, such as a MIDI file containing pitches and corresponding time information,¶[0065]:”… the terminal searches for the tune information of the target song obtained in advance, inputs the target text and the tune information into the trained tune control model, performs tune matching”),
generating a music including the voice corresponding to the custom text based on the voice and the music melody (US’896, ¶[0042]:”obtains a self-made song synthesized according to a tune control model, the target text, and tune information of the target song, the target text being used as the lyrics”).
Regarding claim 5, US’896 discloses ‘The method according to claim 1, as discussed above.
US’896 further discloses ‘wherein, the music generation control is a first control; the music generation interface is a first interface (US’896, ¶[0062]:the application enters a background audio configuration interface according to a user instruction, and through that interface obtains the target text and selected target song, determines the tune, and synthesizes the self-made song);
the text input box is a lyrics input box (US’896, FIG. 3 identifies text box 310 under “User-created lyrics”; ¶¶[0034]-[0035]:obtaining target text entered by the user through text input box 310; The target text subsequently serves as lyrics of the self-made song);
and the at least one music configuration item includes a song melody configuration item (US’896, FIG. 3, provides a “Select music” interface/candidate song list 330; ¶[0041]: the user selects a song from candidate song list 330 and the terminal obtains selected song 331 as the target song);
and the method further comprising: in response to the trigger operation of the music generation control, generating a voice singing the lyrics based on the custom text input in the lyrics input box (US’896, ¶[0064]:inputting the target and tune information into the tune-control model and outputting the self-made song after “speech synthesis is performed on each character in the target text according to the corresponding tune feature”; ¶[0065]: repeats this operation),
and determining a song melody based on configuration information corresponding to the song melody configuration item (US’896, ¶[0047]: tune information of the target song can be extracted from the target-song file, including a MIDI file containing pitches and corresponding timing information; [0065]: searching for the target song’s tune information and determining corresponding tune features),
performing synthesis based on the voice singing the lyrics and the song melody (US’896, ¶[0064]:inputs both target text and tune information into the tune-control model and performs speech synthesis on the target-text characters according to their corresponding tune features), generating a song including the voice singing the lyrics (US’896, ¶[0042]: obtaining a “self-made song” synthesized according to the target text and tune information, “the target text being used as the lyrics of the self-made song”; ¶[0065]:the resulting self-made song is produced after speech synthesis of the target-text characters according to the tune features).
Regarding claim 8, US’896 discloses ‘The method according to claim 1, as discussed above.
US’896 further discloses ‘wherein, the music generation control is a second control (US’896, ¶[0039]: FIG. 3, terminal displays candidate songs in the interface and detects a “predefined trigger operation”, such as “touch operation or a cursor clicking operation”, which triggers the corresponding selection instruction);
the music generation interface is a second interface (US’896, FIG.3, ¶[0062]: terminal enters a “background audio configuration interface (for example, the interface shown in FIG. 3) according to the user instruction”; target text and target song are obtained and used to synthesize the song”);
the text input box is a slogan input box (US’896, ¶¶[0034]-[0035]/ FIG. 3: obtaining target text input by the user through the text input box 310;¶[0086]:”the terminal obtains the self-made audio synthesized according to the timbre control model and the target text, the timbre control model matching the target speaking object”);
and the at least one music configuration item includes a music melody configuration item (US’896, ¶¶[0038]-[0041]: Fig. 3 display candidate songs from which the user selects target song 331. The selected target song supplies the tune information used for synthesis; the tune information reflects; ¶[0047]:tune information reflects song’s pitch over time and may be obtained from MIDI information);
and the method further comprising: in response to the trigger operation of the music generation control (US’896, ¶¶[0039]-[0042]: a selection instruction triggered according to a user operation, including detecting a predefined trigger operation on displayed candidate songs/objects; once the target song is selected, the system obtains the synthesized self-made song), generating a voice for a slogan based on a custom text input in the slogan input box (US’896, ¶[0083]”self-made audio synthesized according to a timbre control model and the target text”; ¶[0086]:this enables the user to have what the user wants to say spoken through a virtual or real character “a user selects to speak in the sound of a virtual character or a real character”),
and determining a music melody based on configuration information corresponding to the music melody configuration item (US’896, ¶¶[0064]-[0065]:determining the target song selected by the selection instruction, searching for “the tune information matching the target song” and using that information for synthesis),
generating a music including the voice for the slogan based on the voice for the slogan and the music melody (US’896, ¶¶[0061]-[0065]: synthesizing the self-made song according to the target text and tune information; ¶[0081]:”the server mixes the self-made song or the self-made speech with an accompaniment and delivers the self-made song or the self-made speech with the accompaniment to the terminal”).
Regarding claim 12, US’896 discloses ‘A system (US’896, ¶[0024]: FIG. 1, audio synthesis system; ¶[0128]: Fig. 10) comprising at least one computing apparatus (US’896, Fig. 10, ¶[0128]: includes a processor with memory) and at least one storage apparatus storing instructions (US’896, Fig. 10, ¶[0128]: storage of the program/instructions used), wherein the instructions, when executed by the at least one computing apparatus (US’896, Fig. 10, ¶[0128]: execution of the stored computer program), cause the at least one computing apparatus to perform a music generation method (US’896, Fig. 10, with ¶¶[0034]-[0042], [0062]-[0065], [0075]-[0078]: implements audio synthesizing method) comprising:
The remaining limitations of claim 12 are taught by US’086 for the same reasons set forth above with respect to claim 1.
in response to a trigger operation of a music generation control, displaying a music generation interface,
the music generation interface including a text input box, a music generation control and at least one music configuration item;
in response to a trigger operation of the text input box, obtaining a custom text input by a user,
and in response to a trigger operation of the at least one music configuration item, determining configuration information corresponding to the at least one music configuration item;
and in response to a trigger operation of the music generation control, generating a voice based on the custom text,
and determining a music melody based on the configuration information corresponding to the at least one music configuration item, generating a music including the voice corresponding to the custom text based on the voice and the music melody.
Regarding claim 13, US’896 discloses ‘A non-transitory computer-readable storage medium that stores programs or instructions that, (US’896, Fig. 10, ¶[0128]: non-transitory computer-readable storage medium storing computer programs) when executed by at least one computing apparatus, cause at least one computing apparatus to execute a music generation method (US’896, Fig. 10, ¶[0128]: may cause the processor to perform the audio synthesis method) comprising:
The remaining limitations of claim 13 are taught by US’086 for the same reasons set forth above with respect to claim 1.
in response to a trigger operation of a music generation control, displaying a music generation interface,
the music generation interface including a text input box, a music generation control and at least one music configuration item;
in response to a trigger operation of the text input box, obtaining a custom text input by a user,
and in response to a trigger operation of the at least one music configuration item, determining configuration information corresponding to the at least one music configuration item;
and in response to a trigger operation of the music generation control, generating a voice based on the custom text,
and determining a music melody based on the configuration information corresponding to the at least one music configuration item, generating a music including the voice corresponding to the custom text based on the voice and the music melody.
Regarding claim 15, US’896 discloses ‘The system according to claim 12, as discussed above.
The remaining limitations of claim 15 are taught by US’086 for the same reasons set forth above with respect to claim 3.
wherein, the music configuration items can be preset with configuration information,
and the method further comprises: in case that the user does not trigger the at least one music configuration item, in response to the trigger operation of the music generation control, generating a voice based on the custom text,
and determining a music melody based on the preset configuration information of the at least one music configuration item, generating a music including the voice corresponding to the custom text based on the voice and the music melody.
Regarding claim 17, US’896 discloses ‘The system according to claim 12, as discussed above.
The remaining limitations of claim 17 are taught by US’086 for the same reasons set forth above with respect to claim 5.
US’896 further discloses ‘wherein, the music generation control is a first control;
the music generation interface is a first interface;
the text input box is a lyrics input box;
and the at least one music configuration item includes a song melody configuration item;
and the music generation method further comprises: in response to the trigger operation of the music generation control, generating a voice singing the lyrics based on the custom text input in the lyrics input box,
and determining a song melody based on configuration information corresponding to the song melody configuration item,
performing synthesis based on the voice singing the lyrics and the song melody), generating a song including the voice singing the lyrics.
Regarding claim 19, US’896 discloses ‘The system according to claim 12, as discussed above.
The remaining limitations of claim 19 are taught by US’086 for the same reasons set forth above with respect to claim 8.
wherein, the music generation control is a second control;
the music generation interface is a second interface;
the text input box is a slogan input box;
and the at least one music configuration item includes a music melody configuration item;
and the music generation method further comprises: in response to the trigger operation of the music generation control, generating a voice for a slogan based on a custom text input in the slogan input box,
and determining a music melody based on configuration information corresponding to the music melody configuration item,
generating a music including the voice for the slogan based on the voice for the slogan and the music melody.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2,4,14 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over US’896, in view of US20030159566 (Sater), hereinafter US’566.
Regarding claim 2, US’896 discloses ‘The method according to claim 1, as discussed above.
US’896 further discloses ‘and the method further comprising: ‘…, in response to the trigger operation of the music generation control (US’896,¶[0062]: the interface/user instruction and synthesis functionality: target text and target song are obtained, tune information determined, and the self-made song synthesized), ,
and determining a music melody based on the configuration information corresponding to the at least one music configuration item (US’896, ¶[0041]:selects target song from candidate-song list 330; target song provides the tune; searches tune information and determines corresponding tune features),
generating a music including the voice corresponding to the sample text based on the voice and the music melody (US’896, ¶[0042]: obtains a self-made song synthesized according to target text and tune information of the target song, with the target text used as the lyrics; ¶¶[0064]-[0065]: speech synthesis performed according to the corresponding tune features)
US’896 does not expressly disclose ‘wherein, a sample text is displayed in the text input box,
[and]
in case that the custom text input by the user is not obtained
[and]
generating a voice based on the sample text.
However, US’566 discloses ‘wherein, a sample text is displayed in the text input box (US’566, ¶[0024]: FIG. 2, presenting the user with a “lyric sheet template” displaying “base lyrics” and “default placeholders” for the “custom lyric fields”; the user may replace/customize those fields by entering desired words),
[and]
in case that the custom text input by the user is not obtained (US’566, ¶[0024]: FIG. 2: the lyric template already contains base/default lyric content prior to the customization and that custom fields may be populated by the user; ¶¶[0038]-[0039]: pre-populating lyric templates with existing information)
[and]
generating a voice based on the sample text (US’566, ¶[0031]:a producer may use artificial intelligence to digitally simulate/synthesize a human voice from the lyric content; recording the song with default lyrics and reconstructing the customized song using default/non-customized and custom phrases).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify US’896’s text-input/template interface to display default/sample lyric text, as taught by US’566, so that music generation may proceed using available lyric content when the user does not provide custom text, thereby reducing required user input and facilitating song generation.
Regarding claim 4, US’896 discloses ‘The method according to claim 1, as discussed above.
US’896 further discloses ‘the music configuration items can be preset with configuration information (US’896, ¶[0047]: tune information may be extracted from the target-song file, such as MIDI information containing pitches and corresponding time information), and the method further comprising:
and the user does not trigger the at least one music configuration item (US’896, ¶[0040]: non-user-selection route: the terminal may “randomly select a song from the candidate songs as the target song”), in response to the trigger operation of the music generation control (US’896, ¶[0062]:user-instructed-based operation of the music-generation/configuration interface followed by synthesis of the “self-made song” from “target text” and “tune information”),
and determining a music melody based on configuration information preset by the at least one music configuration item (US’896, ¶[0040]:the terminal itself to determine the target song; ¶[0047]:associated tune information; ¶[0065]: searching for “the tune information of the target song obtained in advance” and using that information to determine the corresponding tune features),
generating a music including the voice corresponding to the sample text based on the voice and the music melody (US’896, ¶[0042]: obtains a self-made song synthesized according to target text and tune information of the target song, with the target text used as the lyrics; ¶¶[0064]-[0065]: speech synthesis performed according to the corresponding tune features).
US’896 does not expressly disclose ‘wherein, a sample text is displayed in the text input box,
[and]
in case that the custom text input by the user is not obtained,
[and]
generating a voice based on the sample text.
However, US’566 discloses ‘wherein, a sample text is displayed in the text input box (US’566, ¶[0024]: FIG. 2, presenting the user with a “lyric sheet template” displaying “base lyrics” and “default placeholders” for the “custom lyric fields”; the user may replace/customize those fields by entering desired words),
in case that the custom text input by the user is not obtained (US’566, ¶[0024]: FIG. 2: the lyric template already contains base/default lyric content prior to the customization and that custom fields may be populated by the user; ¶¶[0038]-[0039]: pre-populating lyric templates with existing information),
generating a voice based on the sample text (US’566, ¶[0031]:a producer may use artificial intelligence to digitally simulate/synthesize a human voice from the lyric content; recording the song with default lyrics and reconstructing the customized song using default/non-customized and custom phrases)
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to provide US’896’s text-input/template interface to display default/sample lyric text, as taught by US’566, to allow song generation when the user does not enter custom text and thereby reduce required user interaction.
Regarding claim 14, US’896 discloses ‘The system according to claim 12, as discussed above.
The remaining limitations of claim 14 are taught by US’086 for the same reasons set forth above with respect to claim 2.
wherein, a sample text is displayed in the text input box, and the music generation method further comprises:
in case that the custom text input by the user is not obtained, in response to the trigger operation of the music generation control, generating a voice based on the sample text,
and determining a music melody based on the configuration information corresponding to the at least one music configuration item, generating a music including the voice corresponding to the sample text based on the voice and the music melody.
Regarding claim 16, US’896 discloses ‘The system according to claim 12, as discussed above.
The remaining limitations of claim 16 are taught by US’086 for the same reasons set forth above with respect to claim 4.
wherein, a sample text is displayed in the text input box, the music configuration items can be preset with configuration information,
and the music generation method further comprises: in case that the custom text input by the user is not obtained and the user does not trigger the at least one music configuration item,
in response to the trigger operation of the music generation control, generating a voice based on the sample text,
and determining a music melody based on configuration information preset by the at least one music configuration item, generating a music including the voice corresponding to the sample text based on the voice and the music melody.
Claims 6 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over US’896, in view of US20190318715 (Danjyo), hereinafter US’715.
Regarding claim 6, US’896 discloses ‘The method according to claim 5, as discussed above.
US’896 further discloses ‘wherein, the generating a voice singing the lyrics based on the custom text input in the lyrics input box (US’896, ¶¶[0064]-[0065]: inputting target text and tune information into the trained tune-control model and performing speech synthesis on each character according to its corresponding tune feature) comprises:
aligning the custom text with the song melody corresponding to the song melody configuration item (US’896, ¶[0065]:inputting the target text and target-song tune information into the trained tune-control model performing tune matching on the characters, and outputting the self-made song after speech synthesis),
and generating the voice singing the lyrics from the aligned custom text (US’896, ¶[0065]: outputting the self-made song after speech synthesis is performed on each character in the target text according to the corresponding tune feature).
US’896 does not expressly disclose ‘and determining the correspondence between text units in the custom text and notes in the song melody.
However, US’715 discloses ‘and determining the correspondence between text units in the custom text and notes in the song melody (US’715, ¶[0147]:MusicXML music data contains “lyric strings (characters) and a melody (notes)” and “a melody corresponding to a lyric string”; FIG. 13, illustrates the relationship within a structure containing a pitch and associated lyric text);
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to implement US’896’s character-by-character tune matching with US’715’s known correspondence between lyric text units and musical notes to properly synchronized the lyrics with the melody and generate the singing voice according to the corresponding notes.
Regarding claim 18, US’896 further discloses ‘The system according to claim 17, as discussed above.
The remaining limitations of claim 18 are taught by US’086 for the same reasons set forth above with respect to claim 6.
wherein, the generating a voice singing the lyrics based on the custom text input in the lyrics input box comprises:
aligning the custom text with the song melody corresponding to the song melody configuration item,
and determining the correspondence between text units in the custom text and notes in the song melody;
and generating the voice singing the lyrics from the aligned custom text.
Claims 7,10 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over US’896, in view of US20190237051 (Silverstein), hereinafter US’051.
Regarding claim 7, US’896 discloses ‘The method according to claim 5, as discussed above.
US’896 further discloses ‘wherein, the at least one music configuration item further comprises: a timbre configuration item (US’896, ¶¶[0080]-[0088], [0088]: the resulting user self-made song has a timbre consistent with selected target speaking object; ¶[0082]:user can independently select lyrics, tune, and timbre);
and the method further comprising: generating a voice singing the lyrics from the custom text input in the lyrics input box based on a timbre corresponding to the timbre configuration item (US’896, ¶[0072]:“The target speaking object is an object to which a target timbre belongs. The target timbre is a voice feature of a self-made song or a self-made speech that the user intends to synthesize”; ¶[0077]:”the target text is first converted by using a timbre control model into a self-made speech with the timbre conforming to the target speaking object, and the self-made speech is then converted by using the tune control model into the self-made song with the tune conforming to the target song”),
the timbre of the voice singing the lyrics being the timbre corresponding to the timbre configuration item (US’896, ¶[0088]:”The timbre control model matching the target speaking object is a timbre control model obtained through training according to audio data of the target speaking object”);
and generating a song … based on the voice singing the lyrics and the song melody (US’896, ¶[0077]:”the target text is first converted by using a timbre control model into a self-made speech with the timbre conforming to the target speaking object, and the self-made speech is then converted by using the tune control model into the self-made song with the tune conforming to the target song”).
US’896 does not expressly disclose ‘and a music format configuration item;
[and]
and generating a song with the music format based on a music format corresponding to the music format configuration item.
However, US’051 discloses ‘and a music format configuration item (US’051, ¶[0862]:”The Piece Format Translator subsystem B50 analyzes the audio and text representation of the digital piece and creates new formats of the piece as requested by the system user or system including. Such new formats may include, but are not limited to, MIDI, Video, Alternate Audio, Image, and/or Alternate Text format”);
and generating a song with the music format based on a music format corresponding to the music format configuration item (US’051, FIG. 27001, ¶[0862]:”creates new formats of the piece as requested”; ”Subsystem B50 translates the completed music piece into desired alterative formats requested during the automated music composition and generation process of the present invention”)
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify US’896 to provide user-selectable music-format configuration, as taught by US’051, to enable the generated song to be provided in a desired output format selected by the user, thereby facilitating use of the generated song in different applications and playback environments.
Regarding claim 10, US’896 discloses ‘The method according to claim 8, as discussed above.
US’896 further discloses ‘wherein, the at least one music configuration item further comprises: a timbre configuration item (US’896, ¶¶[0097]-[0098]: a self-made song having target text as lyrics, tune consistent with the target song, and timbre consistent with the target speaking object; the user can independently select the lyrics, tune, and timbre);
and the method further comprising: based on a timbre corresponding to the timbre configuration item (US’896, ¶¶[0075]-[0078]:the target text is first converted using a timbre control model into self-made speech having timbre conforming to the target speaking object), generating a voice for a slogan from a custom text input in the slogan input box (US’896, ¶¶[0034]-[0035], FIG. 3; ¶¶[0075]-[0078]: the target text can be input by the user through text input box 310; converting that target text using the timbre-control model into self-made speech), the timbre of the voice for the slogan being the timbre corresponding to the timbre configuration item (US’896, ¶¶[0075]-[0078]; ¶¶[0097]-[0098]:converting target text into self-made speech having the timbre conforming to the target speaking object; a resulting song having “the timbre consistent with that of the target speaking object” and permits the user to independently select the timbre);
and based on the voice for the slogan and the music melody (US’896, ¶¶[0075]-[0078]: converting the target text using the timbre-control model into self-made speech, followed by conversion of that speech using the tune-control model into the self-made song having the tune corresponding to the target song; ¶¶[0064]-[0065]: determining target-song tune information and using the target text/tune information to generate the self-made song).
US’896 does not expressly disclose‘ …and a music format configuration item
[and]
‘ …and generating a music including the music format based on a music format corresponding to the music format configuration item.
However, US’051 discloses‘ …and a music format configuration item (US’051, FIG. 27OO1, ¶[0862]:the Piece Format Translator subsystem B50 “creates new formats of the piece as requested by the system user or system” including “MIDI, Video, Alternate Audio, Image, and/or Alternate Text format” and translates the completed music piece into the desired requested alternative format)
[and]
‘ …and generating a music including the music format based on a music format corresponding to the music format configuration item (US’051, FIG. 27OO1, ¶[0862]: “creates new formats of the piece as requested by the system user or system” and translates the completed music piece into “desired alterative formats requested: during automated music composition and generation ¶[0863]: transmission of the resulting “formatted digital audio file(s)” to the user).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify US’896 to provide user-selectable music-format configuration, as taught by US’051, to enable the generated music to be provided in a format desired by the user, thereby facilitating compatibility with different playback and processing applications.
Regarding claim 21, US’896 discloses ‘The system according to claim 19 as discussed above.
The remaining limitations of claim 21 are taught by US’086 for the same reasons set forth above with respect to claim 10.
wherein, the at least one music configuration item further comprises: a timbre configuration item and a music format configuration item;
and the music generation method further comprises: based on a timbre corresponding to the timbre configuration item,
generating a voice for a slogan from a custom text input in the slogan input box,
the timbre of the voice for the slogan being the timbre corresponding to the timbre configuration item;
and generating a music including the music format based on a music format corresponding to the music format configuration item and based on the voice for the slogan and the music melody.
Claims 9 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over US’896, in view of US20220059063 (Balassanian), hereinafter US’063.
Regarding claim 9, US’896 discloses ‘The method according to claim 8, as discussed above.
US’896 further discloses ’wherein, the generating a music including the voice for the slogan based on the voice for the slogan and the music melody (US’896, ¶¶[0061]-[0065]: synthesizing the self-made song according to the target text and tune information; ¶[0081]:”the server mixes the self-made song or the self-made speech with an accompaniment and delivers the self-made song or the self-made speech with the accompaniment to the terminal”) comprises:
‘…, performing synthesis of the voice for the slogan with the music melody to obtain a synthesized music (US’896, ¶¶[0077]-[0078]: the user-provided target text is converted into synthesized voice/self-made speech, and that voice is combined/processed with the selected target-song tune information to produce the self-made song);
US’896 does not expressly disclose ‘determining a musical key point in the music melody, the music melody having a mutation at the position of the musical key point;
[and]
‘…and based on the position of the musical key point;
in the synthesized music, the voice for the slogan appears at the position of the musical key point of the music melody.
However, US’063 discloses ‘determining a musical key point in the music melody (US’063, ¶[0071]: algorithmically defined “musical sections and transitions” ;¶[0072]: identifies a “climax point marking the end of a section”), the music melody having a mutation at the position of the musical key point (US’063, ¶[0071]:identifies abrupt musical changes ”ending a repeated phrase halfway through a repetition cycle, suddenly changing key, introducing a new texture… [and] sudden increases in volume”;¶[0072]: a “climax point marking the end of a section”);
[and]
‘…and based on the position of the musical key point (US’063, ¶[0157]: selection of a loop for inclusion in the generated music is a function of “the current position in a section or subsection”; an arrangement rule that may “chop a vocal loop to create a melody”’; techniques specifying the order in which vocals are added and placing chopped vocals “on every second beat”);
in the synthesized music, the voice for the slogan appears at the position of the musical key point of the music melody (US’063, FIG. 7, depicts Vocal loop A positioned within a buildup section having transition-in, main-content, and transition-out subsections; ¶[0157]: makes loop selection dependent upon “the current position in a section or subsection).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify US’896 to position the generated slogan voice at a musically significant transition or change point, as taught by US’063, in order to coordinate the vocal content with changes in the musical structure and thereby provide a more musically coherent synthesized output.
Regarding claim 20, US’896 discloses ‘The system according to claim 19, as discussed above.
The remaining limitations of claim 20 are taught by US’086 for the same reasons set forth above with respect to claim 9.
wherein, the generating a music including the voice for the slogan based on the voice for the slogan and the music melody comprises:
determining a musical key point in the music melody, the music melody having a mutation at the position of the musical key point;
and based on the position of the musical key point, performing synthesis of the voice for the slogan with the music melody to obtain a synthesized music;
in the synthesized music, the voice for the slogan appears at the position of the musical key point of the music melody.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US20150148927 (Georges) teaches a music-generation system that receives text through a text-input interface, processes the text into phonetic information, and generates a vocal track for output as music.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICOLE K GILLESPIE whose telephone number is (571)482-4187. The examiner can normally be reached Monday-Friday 7:30-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei K Hammond can be reached at (571)270-3819. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NICOLE K GILLESPIE/Examiner, Art Unit 2837
/DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837