DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement(s) (IDS(s)) submitted on 5/19/2023, 7/3/2023, 9/23/2024, 2/9/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) are being considered by the examiner.
Specification
The disclosure is objected to because of the following informalities:
In ¶0015, “analogue-to-digital converter for producing an analog singing sound signal V” is logically inconsistent. Analog-to-digital converters produce digital signals.
In ¶0016, “digital-to-analogue converter for producing a digital reproduction signal Z” is logically inconsistent. Digital-to-analog converters produce analog signals.
In ¶0053, “The pitch Fy1 represents a fundamental frequency of a pitch of a singing sound” should read, “The pitch Fy1 represents a fundamental frequency of a pitch of a musical instrument sound.”
In ¶0116, “14: sound emitting device, 15: sound receiving device” should read, “14: sound receiving device, 15: sound emitting device”
Appropriate correction is required.
Claim Objections
Claim 15 is objected to because of the following informality: "duration of sound output" should read, "duration of the singing sound." Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5 and 9-14 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “known sound data” in claim 5 is a relative term which renders the claim indefinite. The term “known sound data” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Paragraph 0110 of the instant description also recites “known sound data” but provides no indication as to what is “known,” by whom or in what capacity.
Claim 9 recites the limitation "singing data” in line 4. There is insufficient antecedent basis for this limitation in the claim. Amending this phrase to “singing sound data” congruent with the recitation in claim 1 would resolve this rejection.
Claims 11-14 are likewise rejected for depending, directly or indirectly, on claim 9.
Claim 10 recites the limitation "singing data” in line 4. There is insufficient antecedent basis for this limitation in the claim. Amending this phrase to “singing sound data” congruent with the recitation in claim 1 would resolve this rejection.
The term “known sound data” in claim 13 is a relative term which renders the claim indefinite. The term “known sound data” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Paragraph 0110 of the instant description also recites “known sound data” but provides no indication as to what is “known,” by whom or in what capacity.
The term “known sound data” in claim 14 is a relative term which renders the claim indefinite. The term “known sound data” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Paragraph 0110 of the instant description also recites “known sound data” but provides no indication as to what is “known,” by whom or in what capacity.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 6, and 16 are rejected under 35 U.S.C. 102(a)(1) as anticipated by Cardinaux et al. (US 20150278686 A1, October 1, 2015), hereinafter Cardinaux.
Regarding claim 1, Cardinaux discloses a computer-implemented sound processing method (Cardinaux ¶0040: "The term circuitry as used above comprises one or more programmable processors which are configured to run software.")comprising: outputting singing sound data based on a sound signal (Cardinaux ¶0061: "This time-to-frequency transformation unit 103 implements a Short-Time Fourier Transform (STFT) to determine a frequency spectrum of local sections of an input audio signal as it changes over time.") representing singing sound (Cardinaux ¶0024: "In a Karaoke system, the target instrument is typically the human voice of a singer."); and outputting sound data representing musical instrument sound (Cardinaux ¶0044: "Morphing a vocal track into a violin track may for example enable a violin player to e.g. learn to play the vocal track on the violin.") that correlates with musical elements of the singing sound (Cardinaux ¶0079: "The artificial neural network is thus trained to produce the morphing spectrum 305 whenever a corresponding mixture spectrum is fed to the inputs of the artificial neural network. The artificial neural network is thus trained to replace notes and articulations played by the target instrument or phonemes in singing/talking by corresponding notes and articulations played by the replacement instrument."), by inputting input data that includes the singing sound data to a trained model (Cardinaux ¶0062: "The spectral amplitudes |s(ωk)| are fed to the input nodes of an artificial neural network 107. In this embodiment, the artificial neural network 107 is a Deep Neural Network (DNN).") that has learned, by machine learning (Cardinaux ¶0020: "The term artificial neural network refers to any computational model that is capable of machine learning and pattern recognition, in particular to those computational models inspired by the human or animals central nervous systems (in particular the brain)."), a relationship between singing sound for training and musical instrument sound for training (Cardinaux ¶0081: "An artificial neural network may thus be trained to output the spectral difference between the target instrument and the desired replacement instrument using training samples where both instruments have played the same notes. If multiple instruments should be processed, the artificial neural network is trained for each combination to be morphed (e.g. voice to flute, guitar to strings). Each combination of instruments will result in a specific set of parameters for the artificial neural network.").
Regarding claim 2, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux further discloses that the outputting outputs the sound data in parallel with progress of the singing sound (Cardinaux ¶0084: "The disclosed approach may be realized by neural network forward passes in the frequency domain and may therefore be processed in real-time.").
Regarding claim 6, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux further discloses that the input data includes musical instrument data that specifies a first musical instrument from among a plurality of musical instruments (Cardinaux ¶0055: "According to some embodiments a system is provided which comprises a database for storing parameters of neural network, a user interface for selecting at least a target instrument, and circuitry, the circuitry implementing an artificial neural network which is configured according to parameters retrieved from the database and selected via the user interface, the artificial neural network being further configured to process a mixture spectrum which corresponds to input music in order to obtain an output spectrum based on the parameters selected via the user interface."), and the sound data represents musical instrument sound of the first musical instrument specified by the musical instrument data (Cardinaux ¶0059: "When a particular target instrument is selected, e.g. by means of a user interface it is selected extraction of a violin, the parameters corresponding to this target instrument, the violin, may be retrieved from the database and the artificial neural network may be configured according to these retrieved parameters. The artificial neural network is thus configured to extract the selected target instrument.").
Regarding claim 16, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux further discloses providing the trained model, wherein the trained model includes a plurality of musical instrument sound models, each corresponding to a different musical instrument (Cardinaux ¶0059: "For example, an artificial neural network may be trained for each of multiple target instruments and the parameters resulting from training may be stored in a database."), the input data is input to a musical instrument sound model that corresponds to a musical instrument selected from among the plurality of musical instrument sound models (Cardinaux ¶0059: "When a particular target instrument is selected, e.g. by means of a user interface it is selected extraction of a violin, the parameters corresponding to this target instrument, the violin, may be retrieved from the database and the artificial neural network may be configured according to these retrieved parameters."), and the sound data represents musical instrument sound of the selected musical instrument (Cardinaux ¶0059: "If the user selects morphing vocals into a violin, then corresponding parameters which were obtained in a previous vocal-to-violin morphing training are obtained from the database and the artificial neural network is configured according to these parameters to morph a vocal track intro a violin track.").
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 3-4 are rejected under 35 U.S.C. 103 as unpatentable over Cardinaux in view of Jameson (US 20030066414 A1, April 10, 2003).
Regarding claim 3, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux does not explicitly disclose that the sound data represents a pitch of the musical instrument sound that changes in accordance with a pitch of the singing sound.
However, Jameson teaches that the sound data represents a pitch of the musical instrument sound that changes in accordance with a pitch of the singing sound (Jameson ¶0047: "In response, the Vocolo produces the sound at the output 13 of a musical instrument that closely follows in both pitch and volume the nuances of the player's voice.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux by adding the changing pitch of Jameson to give the player the impression of playing the actual instrument and controlling it intimately with the fine nuances of his voice (Jameson ¶0024).
Regarding claim 4, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux does not explicitly disclose that the sound data represents a pitch of the musical instrument sound that satisfies a relationship where a predetermined pitch difference exists between the pitch of the musical instrument sound and a pitch of the singing sound.
However, Jameson teaches that the sound data represents a pitch of the musical instrument sound that satisfies a relationship where a predetermined pitch difference exists between the pitch of the musical instrument sound and a pitch of the singing sound (Jameson ¶0025: "The chosen instrument is synthesized at the pitch determined by the FDM or at an offset from that pitch as desired by the player.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux by adding the predetermined pitch relationship of Jameson to give the player the impression of playing the actual instrument and controlling it intimately with the fine nuances of his voice (Jameson ¶0024).
Claim 5 is rejected under 35 U.S.C. 103 as unpatentable over Cardinaux in view of Browne (US 6297439 B1, October 2, 2001), to the extent understood.
Regarding claim 5, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux does not explicitly disclose that the input data includes known sound data output by the trained model.
However, Browne teaches that the input data includes known sound data output by the trained model (Browne col. 6, lines 30-32: "In yet other embodiments, the sets of previous output values can be used as additional input vectors, with or without previous sets of hidden layer values.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux by adding the input data of Browne to learn and emulate music of a given style or by a specific composer (Browne col. 1, lines 11-12).
Claim 7 is rejected under 35 U.S.C. 103 as unpatentable over Cardinaux in view of Morris et al. (US 20090064851 A1, March 12, 2009), hereinafter Morris.
Regarding claim 7, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux further teaches adding a time series signal of the sound data (Cardinaux ¶0047: "The system may further comprise a frequency-to-time transformation unit which is configured to convert the output spectrum from a frequency domain to a time domain to obtain an output audio signal.").
Cardinaux does not explicitly disclose adding the sound signal representing the singing sound; and a signal representing musical instrument sound of a second musical instrument that differs from the first musical instrument.
However, Morris teaches adding the sound signal representing the singing sound (Morris ¶0016: "An exemplary system allows a person to realize any of these examples by singing a melody into a microphone and clicking “accompany me” on a user interface of the system. In response, the system outputs what the person just sang plus a backing track (i.e., accompaniment). The backing track can include chords, instruments and all."); and a signal representing musical instrument sound of a second musical instrument that differs from the first musical instrument (Morris ¶0123: "For example, chords can be broken into root notes and thirds and sevenths where a bass instrument plays the root notes and a piano instrument plays the thirds and sevenths of the chords as voicings.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux by adding the signal representing the singing sound and signal representing musical instrument sound of Morris to output what the person just sang plus a backing track (Morris ¶ 0016).
Claim 8 and 15 are rejected under 35 U.S.C. 103 as unpatentable over Cardinaux in view of Tolonen et al. (US 20020035915 A1, March 28, 2002), hereinafter Tolonen.
Regarding claim 8, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux does not explicitly disclose that the singing sound data includes a plurality of features relating to the singing sound, and the plurality of features include: a pitch of the singing sound; and an onset of the singing sound.
However, Tolonen teaches that the singing sound data includes a plurality of features relating to the singing sound, and the plurality of features include: a pitch of the singing sound (Tolonen ¶0010: "The audio-to-notes conversion method according to the invention comprises estimating fundamental frequencies of the audio signal for obtaining a sequence of fundamental frequencies and detecting note events on the basis of the sequence of fundamental frequencies for obtaining the note-based code."); and an onset of the singing sound (Tolonen ¶0012: "Any transition from zero to a non-zero value is assigned to a note-on event and a pitch corresponding to the current fundamental frequency.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux by adding the pitch and onset of Tolonen to provide a means for generating real time accompaniment to a musical presentation (Tolonen ¶0014).
Regarding claim 15, Cardinaux (in view of Tolonen) teaches a sound processing method comprising the features of claim 8 as discussed above.
Tolonen further teaches that the plurality of features further include at least one of: an error at an onset of the singing sound; a duration of sound output; an inflection of the singing sound; or a timbre change of the singing sound (Tolonen ¶0052: "The above described method for estimating the fundamental frequency is quite reliable in detecting the fundamental frequency of a sound signal with a single prominent harmonic source (for example voiced speech, singing, musical instruments that provide harmonic sound). Furthermore, the method derives a time trajectory of the estimated fundamental frequencies such that it follows the changes in the fundamental frequency of the sound signal. However, as was stated before, the time trajectory of the fundamental frequencies needs to be further processed for obtaining a note based code. Specifically, the time trajectory needs to be analyzed into a sequence of event pairs indicating the start, pitch and end of a note, which is referred to as note detection. In other words, the note detection refers to forming note events from the fundamental frequency trajectory. A note event comprises for example a starting position (note-on event), pitch, and ending position (note-off event) of a note. For example, the time trajectory may be transformed into a sequence of single length units, such as quavers according to a user-determined tempo.").
Claims 9-10 and 14 are rejected under 35 U.S.C. 103 as unpatentable over Cardinaux in view of Zhao et al. (CN 111653256 A, September 11, 2020), hereinafter Zhao, and further in view of
Janer et al. (Mapping Phonetic Features for Voice-Driven Sound Synthesis, 2008, retrieved August 20, 2026 from https://link.springer.com/chapter/10.1007/978-3-540-88653-2_23), hereinafter Janer, to the extent understood.
Regarding claim 9, Cardinaux discloses a sound processing method comprising the features of claim 1 as discussed above.
Cardinaux does not explicitly disclose providing the trained model, wherein the singing data includes: first data including: a pitch of the singing sound; and an onset of the singing sound; and second data including a feature that relates to the singing sound and differs from the pitch and onset of the singing sound, and wherein the trained model includes: a first model that outputs third data in response to receipt of first intermediate data that includes the first data, the third data including: a pitch of the musical instrument sound; and an onset of the musical instrument sound, and a second model that outputs the sound data in response to receipt of second intermediate data that includes the second data and the third data.
Zhao further teaches providing the trained model (Zhou ¶0015: "The decoding is repeated N times, and the output results of all decoders are compared with the accompaniment track. The encoding neural network and the decoding neural network are trained until a well-trained encoder-decoder network model is obtained."), wherein the trained model includes: a first model (Zhou ¶0012: "Establish an encoder-decoder network structure, including an encoder neural network and a decoder neural network. The encoder neural network consists of multiple encoders, and the decoder neural network consists of multiple decoders.") that outputs third data in response to receipt of first intermediate data that includes the first data (Zhou ¶0016: "For the unaccompanied music to be processed, obtain the embedding vector of the main music track, use the embedding vector as the input of the trained encoder-decoder network model, and output the music accompaniment; combine the obtained music accompaniment with the unaccompanied music to synthesize music containing the complete accompaniment."), the third data including: a pitch of the musical instrument sound (Zhou ¶0043: "Taking MIDI-Like encoding as an example, MIDI-Like encoding encodes music sequences into encoding sequences composed of SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_OFF, etc., for example: SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_ON, TIME_SHIFT, NOTE_OFF, NOTE_OFF…"); and an onset of the musical instrument sound (Zhou ¶0043: "Taking MIDI-Like encoding as an example, MIDI-Like encoding encodes music sequences into encoding sequences composed of SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_OFF, etc., for example: SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_ON, TIME_SHIFT, NOTE_OFF, NOTE_OFF…").
Janer further teaches that the singing data includes: first data including: a pitch of the singing sound (Janer Table 2 includes pitch as an acoustic feature); and an onset of the singing sound (Janer §2.1: "Finally, in order to increase robustness when determining the phonetic class of each segment in a sequence of segments, we use a state transition model. The underlying idea is that a note consists of an onset, a nucleus (vowel) and a coda."); and second data including a feature that relates to the singing sound and differs from the pitch and onset of the singing sound (Janer § 2: "For both, traditional singing and syllabling, principal musical information involves pitch, dynamics and timing; and those are independent of the phonetics. In vocal imitation, though, the role of phonetics is reserved for determining articulation and timbre aspects."), and a second model that outputs the sound data in response to receipt of second intermediate data (Janer § 4: "From the output of the modules described in sections 2 and 3, the system generates corresponding synthesis parameters for the sound synthesizer. We re-use the ideas of the concatenative sample-based saxophone synthesizer described in [2]. Synthesis parameters include note duration, note MIDI-equivalent pitch, note dynamics, and note-to-note articulation type. Sound samples are retrieved from the database taking into account similarity and the transformations that need to be applied, by computing a distance measure.") that includes the second data and the third data (Janer § 4: "Synthesis parameters include note duration, note MIDI-equivalent pitch, note dynamics, and note-to-note articulation type. Sound samples are retrieved from the database taking into account similarity and the transformations that need to be applied, by computing a distance measure. Selected samples are first transformed to fit the synthesis parameters, and concatenated by applying some timbre interpolation around resulting note transitions.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux by adding the trained model of Zhao and the second model and singing data of Janer to apply singing mappings to instrument output controls (Janer § 1.1).
Regarding claim 10, Cardinaux (in view of Zhao and further in view of Janer) teaches a sound processing method comprising the features of claim 9 as discussed above.
Zhao further teaches providing the trained model (Zhou ¶0015: "The decoding is repeated N times, and the output results of all decoders are compared with the accompaniment track. The encoding neural network and the decoding neural network are trained until a well-trained encoder-decoder network model is obtained."), wherein the trained model includes: a first model (Zhou ¶0012: "Establish an encoder-decoder network structure, including an encoder neural network and a decoder neural network. The encoder neural network consists of multiple encoders, and the decoder neural network consists of multiple decoders.") that outputs third data in response to receipt of first intermediate data that includes the first data (Zhou ¶0016: "For the unaccompanied music to be processed, obtain the embedding vector of the main music track, use the embedding vector as the input of the trained encoder-decoder network model, and output the music accompaniment; combine the obtained music accompaniment with the unaccompanied music to synthesize music containing the complete accompaniment."), the third data including: a pitch of the musical instrument sound (Zhou ¶0043: "Taking MIDI-Like encoding as an example, MIDI-Like encoding encodes music sequences into encoding sequences composed of SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_OFF, etc., for example: SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_ON, TIME_SHIFT, NOTE_OFF, NOTE_OFF…"); and an onset of the musical instrument sound (Zhou ¶0043: "Taking MIDI-Like encoding as an example, MIDI-Like encoding encodes music sequences into encoding sequences composed of SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_OFF, etc., for example: SET_VELOCITY, NOTE_ON, TIME_SHIFT, NOTE_ON, TIME_SHIFT, NOTE_OFF, NOTE_OFF…").
Janer further teaches that the singing data includes: first data including: a pitch of the singing sound (Janer Table 2 includes pitch as an acoustic feature); and an onset of the singing sound (Janer §2.1: "Finally, in order to increase robustness when determining the phonetic class of each segment in a sequence of segments, we use a state transition model. The underlying idea is that a note consists of an onset, a nucleus (vowel) and a coda."); second data including a feature that relates to the singing sound and differs from the pitch and onset of the singing sound (Janer § 2: "For both, traditional singing and syllabling, principal musical information involves pitch, dynamics and timing; and those are independent of the phonetics. In vocal imitation, though, the role of phonetics is reserved for determining articulation and timbre aspects."), and a second model that outputs fourth data in response to receipt of second intermediate data that includes the second data and the third data (Janer § 4: "From the output of the modules described in sections 2 and 3, the system generates corresponding synthesis parameters for the sound synthesizer. We re-use the ideas of the concatenative sample-based saxophone synthesizer described in [2]. Synthesis parameters include note duration, note MIDI-equivalent pitch, note dynamics, and note-to-note articulation type. Sound samples are retrieved from the database taking into account similarity and the transformations that need to be applied, by computing a distance measure."), the fourth data including a feature that relates to the musical instrument sound and differs from the pitch and onset of the singing sound (Janer § 4: "Selected samples are first transformed to fit the synthesis parameters, and concatenated by applying some timbre interpolation around resulting note transitions." Janer's timbre teaches a fourth data.), and the sound data includes the third data and the fourth data (Janer § 4: "Synthesis parameters include note duration, note MIDI-equivalent pitch, note dynamics, and note-to-note articulation type. Sound samples are retrieved from the database taking into account similarity and the transformations that need to be applied, by computing a distance measure. Selected samples are first transformed to fit the synthesis parameters, and concatenated by applying some timbre interpolation around resulting note transitions." The resulting instrument sound is produced from note pitch and timing together with dynamics, articulation, and timbre processing. This teaches sound data that includes contributions from third and fourth data.).
Regarding claim 14, Cardinaux (in view of Zhao and further in view of Janer) teaches a sound processing method comprising the features of claim 9 as discussed above.
Janer further teaches that the second intermediate data includes known sound data (Janer § 4: "Sound samples are retrieved from the database taking into account similarity and the transformations that need to be applied, by computing a distance measure. Selected samples are first transformed to fit the synthesis parameters, and concatenated by applying some timbre interpolation around resulting note transitions.").
Claims 11-12 are rejected under 35 U.S.C. 103 as unpatentable over Cardinaux in view of Zhao, and further in view of Janer and Kim et al. (Neural Music Synthesis for Flexible Timbre Control, November 1, 2018, retrieved August 20, 2026 from https://arxiv.org/pdf/1811.00223), hereinafter Kim, to the extent understood.
Regarding claim 11, Cardinaux (in view of Zhao and further in view of Janer) teaches a sound processing method comprising the features of claim 9 as discussed above.
Cardinaux (in view of Zhao and further in view of Janer) does not explicitly disclose that the first intermediate data includes musical instrument data that specifies a musical instrument.
However, Kim teaches that the first intermediate data includes musical instrument data that specifies a musical instrument (Kim § 2.1: "For each instrument to model, its timbre is represented in an embedding vector t, implemented as a learned matrix multiplication on one-hot encoded instrument labels.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux (as modified by Zhao and Janer) by adding the musical instrument data of Kim to apply a temporal dynamics according to a timbre-dependent input (Kim § 2.1).
Regarding claim 12, Cardinaux (in view of Zhao and further in view of Janer and Kim) teaches a sound processing method comprising the features of claim 11 as discussed above.
Kim further teaches that the second intermediate data includes the musical instrument data (Kim § 2: "The input goes through a linear 1x1 convolution layer, which is essentially a time-distributed fully connected layer, followed by a FiLM layer, to be described in the following subsection, which takes the timbre embedding vector and transforms the features accordingly. After a bidirectional LSTM layer and another FiLM layer for timbre conditioning, another linear 1x1 convolution layer produces the Mel spectrogram prediction." Kim fig. 1 teaches the instrument embedding feeding both FiLM blocks.).
Claim 13 is rejected under 35 U.S.C. 103 as unpatentable over Cardinaux in view of Zhao, and further in view of Janer and Browne, to the extent understood.
Regarding claim 13, Cardinaux (in view of Zhao and further in view of Janer) teaches a sound processing method comprising the features of claim 9 as discussed above.
Cardinaux (in view of Zhao and further in view of Janer) does not explicitly disclose that the second intermediate data includes known sound data.
However, Browne teaches that the second intermediate data includes known sound data (Browne col. 6, lines 30-32: "In yet other embodiments, the sets of previous output values can be used as additional input vectors, with or without previous sets of hidden layer values.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound processing method of Cardinaux (as modified by Zhao and Janer) by adding the input data of Browne to learn and emulate music of a given style or by a specific composer (Browne col. 1, lines 11-12).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHILIP SCOLES whose telephone number is (703)756-1831. The examiner can normally be reached Monday-Friday 8:30-4:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei Hammond can be reached on 571-270-7938. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHILIP G SCOLES/
Examiner, Art Unit 2837
/DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837