Prosecution Insights
Last updated: October 02, 2026
Application No. 18/472,119

SIGNAL PROCESSING METHOD, SIGNAL PROCESSING DEVICE, AND SOUND GENERATION METHOD USING MACHINE LEARNING MODEL

Non-Final OA §102§103§112
Filed
Sep 21, 2023
Priority
Mar 25, 2021 — JP 2021-051091 +1 more
Examiner
SCOLES, PHILIP GRANT
Art Unit
Tech Center
Assignee
Yamaha Corporation
OA Round
1 (Non-Final)
57%
Grant Probability
Moderate
1-2
OA Rounds
6m
Est. Remaining
72%
With Interview

Examiner Intelligence

Grants 57% of resolved cases
57%
Career Allowance Rate
40 granted / 70 resolved
-2.9% vs TC avg
Moderate +14% lift
Without
With
+14.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
33 currently pending
Career history
100
Total Applications
across all art units

Statute-Specific Performance

§101
1.5%
-38.5% vs TC avg
§103
59.3%
+19.3% vs TC avg
§102
19.6%
-20.4% vs TC avg
§112
16.5%
-23.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 70 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statement(s) (IDS(s)) submitted on 9/21/2023 and 6/5/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) are being considered by the examiner. Claim Objections Claim 13 is objected to because of the following informality: "at least one processor configured to execute a receiving unit configured to receive a control value" should read "at least one processor configured to execute instructions implementing: a receiving unit configured to receive a control value," or another acceptable formulation of Applicant’s choosing. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claim 5 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. The first and second reference acoustic features established in the preceding claims from which claim 5 depends must inherently be the same or different. Therefore, claim 5 recites nothing that further limits its inherited limitations. Applicant may cancel the claim, amend the claim to place the claim in proper dependent form, rewrite the claim in independent form, or present a sufficient showing that the dependent claim complies with the statutory requirements. Claim 7 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. The first and second reference acoustic features established in the preceding claims from which claim 7 depends must inherently be the same or different. Therefore, claim 7 recites nothing that further limits its inherited limitations. Applicant may cancel the claim, amend the claim to place the claim in proper dependent form, rewrite the claim in independent form, or present a sufficient showing that the dependent claim complies with the statutory requirements. Claim 17 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. The first and second reference acoustic features established in the preceding claims from which claim 17 depends must inherently be the same or different. Therefore, claim 7 recites nothing that further limits its inherited limitations. Applicant may cancel the claim, amend the claim to place the claim in proper dependent form, rewrite the claim in independent form, or present a sufficient showing that the dependent claim complies with the statutory requirements. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless –(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 8-10, and 13 are rejected under 35 U.S.C. 102(a)(1) as anticipated by Kochanski et al. (US 20030078780 A1, April 24, 2003), hereinafter Kochanski. Regarding claim 1, Kochanski discloses a signal processing method realized by a computer (Kochanski ¶0023: "Therefore, in accordance with an illustrative embodiment of the present invention, a computer model may be built to mimic a particular style by advantageously including processes that simulate each of the steps above with precise instructions at every step."), the signal processing method comprising: receiving a control value representing a musical feature (Kochanski ¶0057: "One of the phrase level tags, referred to herein as step_to, advantageously moves f 0 to a specified value which remains effective until the next step_to tag is encountered. When described by a sequence of step_to tags, the phrase curve is essentially treated as a piece-wise differentiable function. (This method is illustratively used below to describe Martin Luther King's phrase curve and Dinah Shore's music notes.)"); receiving a selection signal for selecting either a first degree of enforcement or a second degree of enforcement that is lower than the first degree of enforcement (Kochanski ¶0060: "The errors may be advantageously weighted by the strength S 1 of the tag which indicates how important it is to satisfy the specifications of the tag. If the strength of a tag is weak, the physiological constraint takes over and in those cases, smoothness becomes more important than accuracy. The strength S1 controls the interaction of accent tags with their neighbors by way of the smoothness requirement, G—stronger tags exert more influence on their neighbors."); and generating (Kochanski ¶0067: "In accordance with illustrative embodiments of the present invention, after one or more tags are generated they are fed into a prosody evaluation module such as prosody evaluation module 55 of FIG. 5. This module advantageously produces the final time series of features."), by using a trained model (Kochanski ¶0070: "In accordance with principles of the present invention as illustratively shown in FIG. 5, tag selection module 52 advantageously selects which of a given voice's tag templates to use at each syllable. In accordance with one illustrative embodiment of the present invention, this subsystem consists of a classification and regression (CART) tree trained on human-classified data."), in accordance with the selection signal, either an acoustic feature amount sequence that reflects the control value in accordance with the first degree of enforcement, or an acoustic feature amount sequence that reflects the control value in accordance with the second degree of enforcement (Kochanski ¶0123: "Note that the strength specification of the musical step_to is very strong (i.e., strength=8). This helps to maintain the specified frequency as the tags pass through the prosody evaluation component."). Regarding claim 8, Kochanski discloses a signal processing method comprising the features of claim 1 as discussed above. Kochanski further discloses that the acoustic feature amount sequence generated at the first degree of enforcement changes over time following the control value (Kochanski ¶0123: "Note that the strength specification of the musical step_to is very strong (i.e., strength=8). This helps to maintain the specified frequency as the tags pass through the prosody evaluation component."), and the acoustic feature amount sequence generated at the second degree of enforcement changes independently of the control value (Kochanski ¶¶0150-0153: "ACname=DROOP; pos=3.43; strength=0.00; wscale=0.21. # your. ACname=DROOP; pos=3.64; strength=0.00; wscale=0.21. Finally, the prosody evaluation module produces a time series of amplitude vs. time." Kochanski produces an amplitude time series with strength = 0.00 tags, which equates to zero enforcement weight in the scheme of ¶0060. Therefore, the resulting sequence is independent of the control target.). Regarding claim 9, Kochanski discloses a signal processing method comprising the features of claim 1 as discussed above. Kochanski further discloses that the acoustic feature amount sequence generated at the first degree of enforcement changes over time following the control value, and the acoustic feature amount sequence generated at the second degree of enforcement changes over time following the control value more loosely than the acoustic feature amount sequence generated at the first degree of enforcement (Kochanski ¶0060: "The errors may be advantageously weighted by the strength S 1 of the tag which indicates how important it is to satisfy the specifications of the tag. If the strength of a tag is weak, the physiological constraint takes over and in those cases, smoothness becomes more important than accuracy. The strength S1 controls the interaction of accent tags with their neighbors by way of the smoothness requirement, G—stronger tags exert more influence on their neighbors." Weaker strength tags render smoothness more important than accuracy, whereas stronger tags exert more influence.). Regarding claim 10, Kochanski discloses a signal processing method comprising the features of claim 1 as discussed above. Kochanski further discloses generating a sound signal from the acoustic feature amount sequence generated at the first degree of enforcement or the second degree of enforcement (Kochanski ¶0039: "Next, prosody evaluation module 55 converts the tags into a time series of prosodic features (or the equivalent) which can be used to directly control the synthesizer. The result of prosody evaluation module 55 may be referred to as a “stylized voice control information stream,” since it provides voice control information adjusted for a particular style. And finally, text-to-speech synthesis module 56 generates the voice (e.g., speech or song) waveform, based on the marked-up text and the time series of prosodic features or equivalent (i.e., based on the stylized voice control information stream)."). Regarding claim 13, Kochanski discloses a signal processing device comprising: at least one processor (Kochanski ¶0169: "The functions of the various elements shown in the figures, including functional blocks labeled as 'processor' or 'modules' may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software.") configured to execute a receiving unit configured to receive a control value representing a musical feature (Kochanski ¶0057: "One of the phrase level tags, referred to herein as step_to, advantageously moves f 0 to a specified value which remains effective until the next step_to tag is encountered. When described by a sequence of step_to tags, the phrase curve is essentially treated as a piece-wise differentiable function. (This method is illustratively used below to describe Martin Luther King's phrase curve and Dinah Shore's music notes.)"), and receive a selection signal for selecting either a first degree of enforcement or a second degree of enforcement that is lower than the first degree of enforcement (Kochanski ¶0060: "The errors may be advantageously weighted by the strength S 1 of the tag which indicates how important it is to satisfy the specifications of the tag. If the strength of a tag is weak, the physiological constraint takes over and in those cases, smoothness becomes more important than accuracy. The strength S1 controls the interaction of accent tags with their neighbors by way of the smoothness requirement, G—stronger tags exert more influence on their neighbors."), and an audio generation unit configured to generate (Kochanski ¶0067: "In accordance with illustrative embodiments of the present invention, after one or more tags are generated they are fed into a prosody evaluation module such as prosody evaluation module 55 of FIG. 5. This module advantageously produces the final time series of features."), by using a trained model (Kochanski ¶0070: "In accordance with principles of the present invention as illustratively shown in FIG. 5, tag selection module 52 advantageously selects which of a given voice's tag templates to use at each syllable. In accordance with one illustrative embodiment of the present invention, this subsystem consists of a classification and regression (CART) tree trained on human-classified data."), in accordance with the selection signal, either an acoustic feature amount sequence that reflects the control value in accordance with the first degree of enforcement or an acoustic feature amount sequence that reflects the control value in accordance with the second degree of enforcement (Kochanski ¶0123: "Note that the strength specification of the musical step_to is very strong (i.e., strength=8). This helps to maintain the specified frequency as the tags pass through the prosody evaluation component."). Claim 19 is rejected under 35 U.S.C. 102(a)(1) as anticipated by Danjyo et al. (US 20190392799 A1, December 26, 2019), hereinafter Danjyo. Regarding claim 19, Danjyo discloses a sound generation method comprising: in a system configured to generate sound of a musical piece corresponding to a given sequence of notes (Danjyo ¶¶0026-0027: "In addition to the aforementioned control program and various kinds of permanent data, the ROM 202 stores musical piece data including lyric data and accompaniment data. The ROM 202 (memory) is also pre-stored with melody pitch data (215 d) indicating operation elements that a user is to operate, singing voice output timing data (215 c) indicating output timings at which respective singing voices for pitches indicated by the melody pitch data (215 d) are to be output, and lyric data (215 a) corresponding to the melody pitch data (215 d)."), receiving from a user an instruction on a control value representing a musical feature (Danjyo ¶0031: "Either the melody pitch data 215 d pre-stored in the ROM 202 or pitch data 215 b for a note number obtained in real time due to a user key press operation is input to the voice synthesis LSI 205 as pitch data."); generating, by using a trained model (Danjyo ¶¶0056-0057: "The text analysis unit 307 performs this analysis and outputs a linguistic feature sequence 316 expressing, inter alia, phonemes, parts of speech, and words corresponding to the singing voice data 215. As described in Non-Patent Document 1, the trained acoustic model 306 is input with the linguistic feature sequence 316, and using this, the trained acoustic model 306 estimates and outputs an acoustic feature sequence 317 (acoustic feature data 317) corresponding thereto."), sound reflecting the instruction (Danjyo ¶0058: "The vocalization model unit 308 is input with the acoustic feature sequence 317. With this, the vocalization model unit 308 generates output data 321 corresponding to the singing voice data 215 including lyric text specified by the CPU 201. An acoustic effect is applied to the output data 321 in the acoustic effect application section 320, described later, and the output data 321 is converted into the final inferred singing voice data 217.") in accordance with a first degree of enforcement, in response to receiving from the user the instruction on the control value at the first degree of enforcement (Danjyo ¶0032: "In other words, when there is a user key press operation at a prescribed timing, an inferred singing voice is produced at a pitch corresponding to the key on which there was a key press operation, and when there is no user key press operation at a prescribed timing, an inferred singing voice is produced at a pitch indicated by the melody pitch data 215 d stored in the ROM 202."); and generating, by using the trained model (Danjyo ¶¶0056-0057: "The text analysis unit 307 performs this analysis and outputs a linguistic feature sequence 316 expressing, inter alia, phonemes, parts of speech, and words corresponding to the singing voice data 215. As described in Non-Patent Document 1, the trained acoustic model 306 is input with the linguistic feature sequence 316, and using this, the trained acoustic model 306 estimates and outputs an acoustic feature sequence 317 (acoustic feature data 317) corresponding thereto."), sound reflecting the instruction (Danjyo ¶0058: "The vocalization model unit 308 is input with the acoustic feature sequence 317. With this, the vocalization model unit 308 generates output data 321 corresponding to the singing voice data 215 including lyric text specified by the CPU 201. An acoustic effect is applied to the output data 321 in the acoustic effect application section 320, described later, and the output data 321 is converted into the final inferred singing voice data 217.") at a lower degree of enforcement lower than the first degree of enforcement, in response to receiving from the user the instruction on the control value at a second degree of enforcement (Danjyo ¶0069: "In this case, the user is able to vary the degree of the pitch effect in the acoustic effect application section 320 by, with respect to the pitch of the first key specifying a singing voice, specifying the second key that is repeatedly struck such that the difference in pitch between the second key and the first key is a desired difference."). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 2-7 and 14-18 are rejected under 35 U.S.C. 103 as unpatentable over Kochanski in view of Shechtman (US 20210082408 A1, March 18, 2021). Regarding claim 2, Kochanski discloses a signal processing method comprising the features of claim 1 as discussed above. Kochanski does not explicitly disclose that the trained model has already been trained by machine-learning a relationship between a reference acoustic feature amount sequence and a reference control value sequence indicating a musical feature at each of the first degree of enforcement and the second degree of enforcement. However, Shechtman teaches that the trained model has already been trained by machine-learning a relationship between a reference acoustic feature amount sequence and a reference control value sequence indicating a musical feature (Shechtman ¶¶0037-0038: "The observed prosody info can be temporally aligned and combined, for example by using concatenation or summation, to obtain a sequence of observed combined prosody info. In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences.")at each of the first degree of enforcement and the second degree of enforcement (Shechtman ¶0020: "For example, a pitch and energy trajectory may be calculated using pitch and energy estimators, and an automatic phonetic alignment is applied to divide the time signal to phoneme, syllable, word, and phrase segments. The pitch, duration, and energy observations may then be derived for various time spans. The observations may then be aligned and combined with each other to generate combined prosody info vector sequence."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing method of Kochanski by adding the finer temporal control of Shechtman to adjust prosody of the final acoustic sequence (Shechtman ¶0024). Regarding claim 3, Kochanski (in view of Shechtman) teaches a signal processing method comprising the features of claim 2 as discussed above. Shechtman further teaches that the trained model has already been trained by machine-learning, with respect to reference data representing sound waveforms (Shechtman ¶0020: "In various examples, a prosodic info vector sequence is automatically calculated for a training set of input utterances. The utterances may include both recordings and transcriptions for the recordings."), a first relationship between a first reference control value sequence indicating a musical feature at the first degree of enforcement as an input and a first reference acoustic feature amount sequence of the reference data as an output (Shechtman ¶¶0037-0038: "The observed prosody info can be temporally aligned and combined, for example by using concatenation or summation, to obtain a sequence of observed combined prosody info. In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences."), and a second relationship between a second reference control value sequence indicating a musical feature at the second degree of enforcement as an input and the first reference acoustic feature amount sequence as an output (Shechtman ¶¶0037-0038: "The observed prosody info can be temporally aligned and combined, for example by using concatenation or summation, to obtain a sequence of observed combined prosody info. In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences." A coarser hierarchical observation from the same utterance can serve as the second control value sequence.). Regarding claim 4, Kochanski (in view of Shechtman) teaches a signal processing method comprising the features of claim 3 as discussed above. Shechtman further teaches that the first reference control value sequence changes over time at a first fineness in accordance with a second reference acoustic feature amount sequence (Shechtman ¶0020: "For example, a pitch and energy trajectory may be calculated using pitch and energy estimators, and an automatic phonetic alignment is applied to divide the time signal to phoneme, syllable, word, and phrase segments. The pitch, duration, and energy observations may then be derived for various time spans. The observations may then be aligned and combined with each other to generate combined prosody info vector sequence."), and the second reference control value sequence changes over time at a second fineness in accordance with the second reference acoustic feature amount sequence (Shechtman ¶0020: "For example, a pitch and energy trajectory may be calculated using pitch and energy estimators, and an automatic phonetic alignment is applied to divide the time signal to phoneme, syllable, word, and phrase segments. The pitch, duration, and energy observations may then be derived for various time spans. The observations may then be aligned and combined with each other to generate combined prosody info vector sequence." Shechtman's coarser observation sequence from the same pitch-energy trajectory teaches a second fineness in accordance with the same acoustic feature amount sequence. ). Regarding claim 5, Kochanski (in view of Shechtman) teaches a signal processing method comprising the features of claim 4 as discussed above. Shechtman further teaches that the first reference acoustic feature and the second reference acoustic feature are same acoustic features or different acoustic features (Shechtman ¶¶0037-0038: "In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences. For example, the neural network may include a prosody info encoder, a linguistic encoder and an acoustic decoder. As one example, the embedded prosody info and embedded linguistic sequence are fed into the acoustic decoder outputting mel-spectrogram sequence."). Regarding claim 6, Kochanski (in view of Shechtman) teaches a signal processing method comprising the features of claim 4 as discussed above. Shechtman further teaches that the first reference control value at each time point (Shechtman ¶0032: "In various examples, the aligner and combiner 212 can align and combine the hierarchical observations 206, 208, and 210. For example, the aligner and combiner 212 can align and combine the hierarchical observations 206, 208, and 210 by summation or concatenation to produce combined prosody info, which may include a sequence of observation vectors, synchronized with input linguistic sequence.") is a representative value of the second reference acoustic feature amount sequence of the reference data (Shechtman ¶0023: "In various examples, the observation set may include at least a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof.") within a first time interval that includes each time point (Shechtman ¶0017: "Prosody info, as used herein, refers to a set of interpretable temporal observations. For example, the observations may be evaluated globally and/or locally and hierarchically at different temporal spans. Each observation is a linear combination or set of linear combinations of statistical measures, evaluating a prosodic component over a predetermined period of time."), and the second reference control value at each time point is a representative value of the second reference acoustic feature amount sequence within a second time interval that includes each time point (Shechtman ¶0023: "In various examples, the observation set may include at least a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof.") and that is longer than the first time interval (Shechtman ¶0022: "The observations include linear combinations of statistical measures evaluating a prosodic component over a predetermined period of time. In various examples, the observations may be evaluated globally or locally and hierarchically at different temporal spans. For example, a global observation may be at the utterance level. The hierarchical locally evaluated observations may be at the level of each paragraph, sentence, phrase, word, syllable, or phoneme segment. As used herein, a segment refers to a time span within this hierarchical temporal structure of paragraph/sentence/phrase/word/syllable/phone."). Regarding claim 7, Kochanski (in view of Shechtman) teaches a signal processing method comprising the features of claim 6 as discussed above. Shechtman further teaches that the first reference acoustic feature and the second reference acoustic feature are same acoustic features or different acoustic features (Shechtman ¶¶0037-0038: "In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences. For example, the neural network may include a prosody info encoder, a linguistic encoder and an acoustic decoder. As one example, the embedded prosody info and embedded linguistic sequence are fed into the acoustic decoder outputting mel-spectrogram sequence."). Regarding claim 14, Kochanski discloses a signal processing device comprising the features of claim 13 as discussed above. Kochanski does not explicitly disclose that the trained model has already been trained by machine-learning a relationship between a reference acoustic feature amount sequence and a reference control value sequence indicating a musical feature at each of the first degree of enforcement and the second degree of enforcement. However, Shechtman teaches that the trained model has already been trained by machine-learning a relationship between a reference acoustic feature amount sequence and a reference control value sequence indicating a musical feature (Shechtman ¶¶0037-0038: "The observed prosody info can be temporally aligned and combined, for example by using concatenation or summation, to obtain a sequence of observed combined prosody info. In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences.")at each of the first degree of enforcement and the second degree of enforcement (Shechtman ¶0020: "For example, a pitch and energy trajectory may be calculated using pitch and energy estimators, and an automatic phonetic alignment is applied to divide the time signal to phoneme, syllable, word, and phrase segments. The pitch, duration, and energy observations may then be derived for various time spans. The observations may then be aligned and combined with each other to generate combined prosody info vector sequence."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing device of Kochanski by adding the finer temporal control of Shechtman to adjust prosody of the final acoustic sequence (Shechtman ¶0024). Regarding claim 15, Kochanski (in view of Shechtman) teaches a signal processing device comprising the features of claim 14 as discussed above. Shechtman further teaches that the trained model has already been trained by machine-learning, with respect to reference data representing sound waveforms (Shechtman ¶0020: "In various examples, a prosodic info vector sequence is automatically calculated for a training set of input utterances. The utterances may include both recordings and transcriptions for the recordings."), a first relationship between a first reference control value sequence indicating a musical feature at the first degree of enforcement as an input and a first reference acoustic feature amount sequence of the reference data as an output (Shechtman ¶¶0037-0038: "The observed prosody info can be temporally aligned and combined, for example by using concatenation or summation, to obtain a sequence of observed combined prosody info. In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences."), and a second relationship between a second reference control value sequence indicating a musical feature at the second degree of enforcement as an input and the first reference acoustic feature amount sequence as an output (Shechtman ¶¶0037-0038: "The observed prosody info can be temporally aligned and combined, for example by using concatenation or summation, to obtain a sequence of observed combined prosody info. In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences." A coarser hierarchical observation from the same utterance can serve as the second control value sequence.). Regarding claim 16, Kochanski (in view of Shechtman) teaches a signal processing device comprising the features of claim 15 as discussed above. Shechtman further teaches that the first reference control value sequence changes over time at a first fineness in accordance with a second reference acoustic feature amount sequence (Shechtman ¶0020: "For example, a pitch and energy trajectory may be calculated using pitch and energy estimators, and an automatic phonetic alignment is applied to divide the time signal to phoneme, syllable, word, and phrase segments. The pitch, duration, and energy observations may then be derived for various time spans. The observations may then be aligned and combined with each other to generate combined prosody info vector sequence."), and the second reference control value sequence changes over time at a second fineness in accordance with the second reference acoustic feature amount sequence (Shechtman ¶0020: "For example, a pitch and energy trajectory may be calculated using pitch and energy estimators, and an automatic phonetic alignment is applied to divide the time signal to phoneme, syllable, word, and phrase segments. The pitch, duration, and energy observations may then be derived for various time spans. The observations may then be aligned and combined with each other to generate combined prosody info vector sequence." Shechtman's coarser observation sequence from the same pitch-energy trajectory teaches a second fineness in accordance with the same acoustic feature amount sequence. ). Regarding claim 17, Kochanski (in view of Shechtman) teaches a signal processing device comprising the features of claim 16 as discussed above. Shechtman further teaches that the first reference acoustic feature and the second reference acoustic feature are same acoustic features or different acoustic features (Shechtman ¶¶0037-0038: "In various examples, the observed prosody info may include any combination of observations including statistic measures associated with the input utterances, such as a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof. At block 306, the observed combined prosody info together with the linguistic and the acoustic sequences are used to train a neural network to predict acoustic sequences. For example, the neural network may include a prosody info encoder, a linguistic encoder and an acoustic decoder. As one example, the embedded prosody info and embedded linguistic sequence are fed into the acoustic decoder outputting mel-spectrogram sequence."). Regarding claim 18, Kochanski (in view of Shechtman) teaches a signal processing device comprising the features of claim 16 as discussed above. Shechtman further teaches that the first reference control value at each time point (Shechtman ¶0032: "In various examples, the aligner and combiner 212 can align and combine the hierarchical observations 206, 208, and 210. For example, the aligner and combiner 212 can align and combine the hierarchical observations 206, 208, and 210 by summation or concatenation to produce combined prosody info, which may include a sequence of observation vectors, synchronized with input linguistic sequence.") is a representative value of the second reference acoustic feature amount sequence of the reference data (Shechtman ¶0023: "In various examples, the observation set may include at least a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof.") within a first time interval that includes each time point (Shechtman ¶0017: "Prosody info, as used herein, refers to a set of interpretable temporal observations. For example, the observations may be evaluated globally and/or locally and hierarchically at different temporal spans. Each observation is a linear combination or set of linear combinations of statistical measures, evaluating a prosodic component over a predetermined period of time."), and the second reference control value at each time point is a representative value of the second reference acoustic feature amount sequence within a second time interval that includes each time point (Shechtman ¶0023: "In various examples, the observation set may include at least a log-pitch observation within a segment, a sub-segment log-duration observation within a segment, a log-energy observation within a segment, or any combination thereof.") and that is longer than the first time interval (Shechtman ¶0022: "The observations include linear combinations of statistical measures evaluating a prosodic component over a predetermined period of time. In various examples, the observations may be evaluated globally or locally and hierarchically at different temporal spans. For example, a global observation may be at the utterance level. The hierarchical locally evaluated observations may be at the level of each paragraph, sentence, phrase, word, syllable, or phoneme segment. As used herein, a segment refers to a time span within this hierarchical temporal structure of paragraph/sentence/phrase/word/syllable/phone."). Claims 11-12 are rejected under 35 U.S.C. 103 as unpatentable over Kochanski in view of Harada et al. (US 20190096374 A1, March 28, 2019), hereinafter Harada. Regarding claim 11, Kochanski discloses a signal processing method comprising the features of claim 1 as discussed above. Kochanski does not explicitly disclose detecting a position of a detection target in a first direction and a second direction by a sensor, wherein the control value is received based on the position of the detection target in the first direction, and the selection signal is received based on the position of the detection target in the second direction. However, Harada teaches detecting a position of a detection target in a first direction and a second direction by a sensor (Harada ¶0072: "Each of the touchpads 32 and 34 is an assembly constituted by a plurality of electrodes EL of touch sensors arrayed in a grid form in two directions perpendicular to each other. In the present embodiment, a plurality of electrodes EL of touch sensors is arrayed at regular intervals in each of the longitudinal direction (the X direction) and the lateral direction (the Y direction) of the musical instrument body 102."), wherein the control value is received based on the position of the detection target in the first direction (Harada ¶0076: "For example, by the above-described method for effect setting, when PITCH BEND is set for the X-axis of an arbitrary touchpad 32 or 34 (here, 34) and MODULATION is set for the Y-axis thereof as effects to be given to a musical sound as shown in FIG. 8A, control is performed such that a pitch is bent up when the thumb 204 is moved in the +X direction of the touchpad 34, and a pitch is bent down when the thumb 204 is moved in the −X direction of the touchpad 34."), and the selection signal is received based on the position of the detection target in the second direction (Harada ¶0076: "Furthermore, the cycle of change in tone is controlled to be larger (longer) when the thumb 204 is moved in the +Y direction of the touchpad 34, and the cycle of change in tone is controlled to be smaller (shorter) when the thumb 204 is moved in the −Y direction of the touchpad 34."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing method of Kochanski by adding the control touchpad of Harada to achieve a plurality of effects by simple operations (Harada ¶0079). Regarding claim 12, Kochanski discloses a signal processing method comprising the features of claim 1 as discussed above. Kochanski does not explicitly disclose that the control value is received by an operation of a first user operable input, and the selection signal is received by an operation of a second user operable input. However, Harada teaches that the control value is received by an operation of a first user operable input (Harada ¶0031: "In other words, the operation of the touchpad 32 is performed by changing, in a specific direction, the position (area) to be contacted with the left thumb of the instrument player, and the operation of the touchpad 34 is performed by changing, in a specific direction, the position (area) to be contacted with the right thumb of the instrument player."), and the selection signal is received by an operation of a second user operable input (Harada ¶0031: "In other words, the operation of the touchpad 32 is performed by changing, in a specific direction, the position (area) to be contacted with the left thumb of the instrument player, and the operation of the touchpad 34 is performed by changing, in a specific direction, the position (area) to be contacted with the right thumb of the instrument player."). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing method of Kochanski by adding the control touchpad of Harada to achieve a plurality of effects by simple operations (Harada ¶0079). Claim 20 is rejected under 35 U.S.C. 103 as unpatentable over Danjyo in view of Kochanski. Regarding claim 20, Danjyo discloses a sound generation method comprising the features of claim 19 as discussed above. Danjyo does not explicitly disclose that the generating of the sound that reflects the instruction at the lower degree of enforcement includes generating sound that does not reflect the instruction. However, Kochanski teaches that the generating of the sound that reflects the instruction at the lower degree of enforcement includes generating sound that does not reflect the instruction (Kochanski ¶¶0150-0153: "ACname=DROOP; pos=3.43; strength=0.00; wscale=0.21. # your. ACname=DROOP; pos=3.64; strength=0.00; wscale=0.21. Finally, the prosody evaluation module produces a time series of amplitude vs. time." Kochanski produces an amplitude time series with strength = 0.00 tags, which equates to zero enforcement weight in the scheme of ¶0060. Therefore, the resulting sequence is independent of the control target.). It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the sound generation method of Danjyo by adding the degrees of enforcement Kochanski to weaken the degree of the acoustic effect for a lesser difference in pitch (Danjyo ¶0069). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHILIP SCOLES whose telephone number is (703)756-1831. The examiner can normally be reached Monday-Friday 8:30-4:30 ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei Hammond can be reached on 571-270-7938. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PHILIP G SCOLES/ Examiner, Art Unit 2837 /DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837
Read full office action

Prosecution Timeline

Sep 21, 2023
Application Filed
Sep 01, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731566
Electronic Cymbal
4y 3m to grant Granted Sep 08, 2026
Patent 12711933
KEY FOR KEYBOARD DEVICE
3y 11m to grant Granted Aug 18, 2026
Patent 12670888
ELECTRONIC MUSICAL INSTRUMENT, ELECTRONIC MUSICAL INSTRUMENT CONTROLLING METHOD AND NON-TRANSITORY COMPUTER-READABLE STORAGE MEDIUM
4y 2m to grant Granted Jun 30, 2026
Patent 12646494
Attaching Hand-Actuated Music Controllers to a Saxophone
4y 0m to grant Granted Jun 02, 2026
Patent 12620381
SWITCH LOCK APPARATUS AND METHOD THEREOF
1y 9m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
57%
Grant Probability
72%
With Interview (+14.5%)
3y 7m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 70 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month