DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement(s) (IDS(s)) submitted on 9/7/2023, 9/12/2024, 12/12/2024, 5/20/2026 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) is/are being considered by the examiner.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 3-4, 15, 17, and 19-20 are rejected under 35 U.S.C. 102(a)(2) as anticipated by Echeveste et al. (US 20230082086 A1, filed February 1, 2021), hereinafter Echeveste.
Regarding claim 1, Echeveste discloses a signal processing system for causing a reproduction device to reproduce a time series signal (Echeveste ¶0060: "Music accompaniment data are therefore read by the processor PROC so as to drive the output interface OUT to feed at least one loudspeaker SPK (a baffle or an earphone) with an output acoustic signal based on the pre-recorded music accompaniment data.") that follows reproduction of a musical piece (Echeveste ¶0054: "The present disclosure proposes to solve the problem of synchronizing a pre-recorded accompaniment to a musician in real-time. To this aim, a device DIS (as shown in the example of FIG. 1 which is described hereafter) is used."), the signal processing system comprising: an electronic controller including at least one processor, the electronic controller being configured to execute a plurality of units (Echeveste ¶0055: "The device DIS comprises in an embodiment, at least:
An input interface INP, A processing unit PU, including a storage memory MEM and a processor PROC cooperating with memory MEM, and An output interface OUT.") including an acquisition unit configured to acquire an indicated position indicated by a user in the reproduction of the musical piece (Echeveste ¶0071: "The lag at time t between the musician's real-time musical position on the music score and that of the accompaniment track on the same score (both in beats) is denoted as diff. Therefore, parameter diff reflects exactly the difference between the position on the music score in beats of the detected musician's event in real-time and the position on the music score (in beats) of the accompaniment music that is to be synchronized."), and a control unit configured to execute time stretching of the time series signal in accordance with the indicated position (Echeveste ¶0083: "In the musician time-map, in step S7 the tempo of the output signal which is played on the basis of the pre-recorded accompaniment data can be corrected (from slope e1 to slope e2 of FIG. 3 b ) so as to adjust smoothly the position on the music score of the output signal to the position of the input signal at a future next synchronization time tsync as shown on FIG. 3 b ").
Regarding claim 3, Echeveste discloses a signal processing system comprising the features of claim 1 as discussed above.
Echeveste further discloses that the reproduction of the musical piece is the user's performance of the musical piece (Echeveste ¶0062: "A user US can hear the accompaniment music played by the loudspeaker SPK and can play with a music instrument on the accompaniment music, emitting thus a sound captured by a microphone MIC connected to the input interface INP. ").
Regarding claim 4, Echeveste discloses a signal processing system comprising the features of claim 1 as discussed above.
Echeveste further discloses that the control unit includes an identification unit configured to identify a reproduction position of the time series signal, and the reproduction position corresponds to the indicated position (Echeveste ¶0071: "The lag at time t between the musician's real-time musical position on the music score and that of the accompaniment track on the same score (both in beats) is denoted as diff. Therefore, parameter diff reflects exactly the difference between the position on the music score in beats of the detected musician's event in real-time and the position on the music score (in beats) of the accompaniment music that is to be synchronized."), and a reproduction unit configured to cause the reproduction device to reproduce a portion of the time series signal to execute the time stretching, and the portion corresponds to the reproduction position (Echeveste ¶0083: "In the musician time-map, in step S7 the tempo of the output signal which is played on the basis of the pre-recorded accompaniment data can be corrected (from slope e1 to slope e2 of FIG. 3 b ) so as to adjust smoothly the position on the music score of the output signal to the position of the input signal at a future next synchronization time tsync as shown on FIG. 3 b .").
Regarding claim 15, Echeveste discloses a signal processing system comprising the features of claim 4 as discussed above.
Echeveste further discloses that the indicated position is a performance position estimated by the acquisition unit (Echeveste ¶0079: "Where p is the current position of the musician playing on the score at current time t0.") analyzing the user's performance of the musical piece (Echeveste ¶0081: "Referring now to FIG. 2 , step S1 starts with receiving the input signal related to the musician playing. In step S2, acoustic features are extracted from the input signal so as to identify musical events in the musician playing which are related to events in the music score defined in the pre-recorded music accompaniment data.").
Regarding claim 17, Echeveste discloses a signal processing method, realized by a computer, for causing a reproduction device to reproduce a time series signal (Echeveste ¶0060: "Music accompaniment data are therefore read by the processor PROC so as to drive the output interface OUT to feed at least one loudspeaker SPK (a baffle or an earphone) with an output acoustic signal based on the pre-recorded music accompaniment data.") that follows reproduction of a musical piece (Echeveste ¶0054: "The present disclosure proposes to solve the problem of synchronizing a pre-recorded accompaniment to a musician in real-time. To this aim, a device DIS (as shown in the example of FIG. 1 which is described hereafter) is used."), the signal processing method comprising: acquiring an indicated position indicated by a user in the reproduction of the musical piece (Echeveste ¶0071: "The lag at time t between the musician's real-time musical position on the music score and that of the accompaniment track on the same score (both in beats) is denoted as diff. Therefore, parameter diff reflects exactly the difference between the position on the music score in beats of the detected musician's event in real-time and the position on the music score (in beats) of the accompaniment music that is to be synchronized."); and executing time stretching of the time series signal in accordance with the indicated position (Echeveste ¶0083: "In the musician time-map, in step S7 the tempo of the output signal which is played on the basis of the pre-recorded accompaniment data can be corrected (from slope e1 to slope e2 of FIG. 3 b ) so as to adjust smoothly the position on the music score of the output signal to the position of the input signal at a future next synchronization time tsync as shown on FIG. 3 b ").
Regarding claim 19, Echeveste discloses a signal processing method comprising the features of claim 17 as discussed above.
Echeveste further discloses that the reproduction of the musical piece is the user's performance of the musical piece (Echeveste ¶0062: "A user US can hear the accompaniment music played by the loudspeaker SPK and can play with a music instrument on the accompaniment music, emitting thus a sound captured by a microphone MIC connected to the input interface INP. ").
Regarding claim 20, Echeveste discloses a non-transitory computer-readable medium storing a program for causing a reproduction device to reproduce a time series signal that follows reproduction of a musical piece (Echeveste ¶0060: "Music accompaniment data are therefore read by the processor PROC so as to drive the output interface OUT to feed at least one loudspeaker SPK (a baffle or an earphone) with an output acoustic signal based on the pre-recorded music accompaniment data."), the program causing a computer to execute a process (Echeveste ¶0054: "The present disclosure proposes to solve the problem of synchronizing a pre-recorded accompaniment to a musician in real-time. To this aim, a device DIS (as shown in the example of FIG. 1 which is described hereafter) is used.") comprising: acquiring (Echeveste ¶0055: "The device DIS comprises in an embodiment, at least: An input interface INP, A processing unit PU, including a storage memory MEM and a processor PROC cooperating with memory MEM, and An output interface OUT.") an indicated position indicated by a user in the reproduction of the musical piece (Echeveste ¶0071: "The lag at time t between the musician's real-time musical position on the music score and that of the accompaniment track on the same score (both in beats) is denoted as diff. Therefore, parameter diff reflects exactly the difference between the position on the music score in beats of the detected musician's event in real-time and the position on the music score (in beats) of the accompaniment music that is to be synchronized."); and executing time stretching of the time series signal in accordance with the indicated position (Echeveste ¶0083: "In the musician time-map, in step S7 the tempo of the output signal which is played on the basis of the pre-recorded accompaniment data can be corrected (from slope e1 to slope e2 of FIG. 3 b ) so as to adjust smoothly the position on the music score of the output signal to the position of the input signal at a future next synchronization time tsync as shown on FIG. 3 b ").
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2, 5-7, and 18 are rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Miro et al. (US 20110230987 A1, September 22, 2011), hereinafter Miro.
Regarding claim 2, Echeveste discloses a signal processing system comprising the features of claim 1 as discussed above.
Echeveste further teaches that the time series signal is a signal representing audio or video (Echeveste ¶0095: "In fact, the method applies to any “continuous” media, including for example audio and video."), and the acquisition unit is configured to acquire a plurality of indicated positions over time (Echeveste ¶0077: "This Time-Map is constantly re-evaluated at each interaction of the system with a human musician.").
Echeveste does not explicitly disclose that the control unit is configured to execute the time stretching by a path search to which two or more different indicated positions from among the plurality of indicated positions, and a search condition in accordance with characteristics of the time series signal are applied.
However, Miro teaches that the control unit is configured to execute the time stretching by a path search to which two or more different indicated positions from among the plurality of indicated positions (Miro ¶0034: "Real-time online alignment: Continue computing feature vectors and follow an incremental DTW guided by a forward path selection, ensuring that both audio signals remain aligned during the whole duration of the audio track. That is, the alignment is done block-by-block, first with the initial buffered signals but during the processing the signals continue to arrive and being aligned."), and a search condition in accordance with characteristics of the time series signal are applied (Miro ¶0050: "The resulting spectrum is mapped onto a 12-dimensional normalized chroma representation. The 12 dimensions of the chroma bins correspond to the 12 notes found in western music. The effect of this mapping is to reduce the audio to that of a single octave. Chroma features are typically used in music alignment as they are robust to variations in how the music is played. Finally, the different costs between these chroma frames are calculated using the inner normalized product.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste by adding the control unit configuration of Miro to provide smooth playback in real time (Miro ¶0033).
Regarding claim 5, Echeveste discloses a signal processing system comprising the features of claim 4 as discussed above.
Echeveste further teaches that the acquisition unit is configured to sequentially identify the indicated position for each of a plurality of time points on a time axis (Echeveste ¶0077: "This Time-Map is constantly re-evaluated at each interaction of the system with a human musician.").
Echeveste does not explicitly disclose that this is in each of a plurality of processing time intervals on the time axis, the identification unit is configured to execute a path search and thereby identify a time series of two or more reproduction positions that correspond to different time points within at least a part of the processing time interval, two or more indicated positions respectively identified for two or more time points within the processing time interval from among the plurality of time points, and a search condition corresponding to characteristics of the time series signal are applied to the path search, and the reproduction unit is configured to cause the reproduction device to reproduce portions of the time series signal respectively corresponding to the two or more reproduction positions.
However, Miro teaches that this is in each of a plurality of processing time intervals on the time axis, the identification unit is configured to execute a path search and thereby identify a time series of two or more reproduction positions that correspond to different time points within at least a part of the processing time interval (Miro ¶0034: "Real-time online alignment: Continue computing feature vectors and follow an incremental DTW guided by a forward path selection, ensuring that both audio signals remain aligned during the whole duration of the audio track. That is, the alignment is done block-by-block, first with the initial buffered signals but during the processing the signals continue to arrive and being aligned."), two or more indicated positions respectively identified for two or more time points within the processing time interval from among the plurality of time points (Miro ¶0069: "To avoid any quantisation effects, the final path is smoothed by extrapolating its points so that for any point during the music there is a corresponding time (in milliseconds) of where the video should be."), and a search condition corresponding to characteristics of the time series signal are applied to the path search (Miro ¶0050: "The resulting spectrum is mapped onto a 12-dimensional normalized chroma representation. The 12 dimensions of the chroma bins correspond to the 12 notes found in western music. The effect of this mapping is to reduce the audio to that of a single octave. Chroma features are typically used in music alignment as they are robust to variations in how the music is played. Finally, the different costs between these chroma frames are calculated using the inner normalized product."), and the reproduction unit is configured to cause the reproduction device to reproduce portions of the time series signal respectively corresponding to the two or more reproduction positions (Miro ¶0070: "If the average difference (where the video should be in relation to the audio) differs from the video's actual difference (as known by the media player) by more than a certain threshold, for example, 35 ms (or one frame), video frames are skipped or replayed until the correct difference between the video and audio is reached.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste by adding the control unit configuration of Miro to provide smooth playback in real time (Miro ¶0033).
Regarding claim 6, Echeveste (in view of Miro) teaches a signal processing system comprising the features of claim 5 as discussed above.
Miro further teaches that the processing time interval is a time interval between a first time point and a second time point located after the first time point, from among the plurality of time points (Miro ¶0065: "In the first step, a forward path Pf is found using the same local constraint as explained before until L matching elements are found. In our experiments, L is set to 50 frames (5 seconds). The obtained path is a sub-optimal alignment between both signals but it is useful to obtain a good estimate for the end position at distance L."), and at least the part of the processing time interval is an analysis time interval from the first time point to a third time point between the first time point and the second time point (Miro ¶0012: "An standard DTW algorithm in which a path with minimizes a defined global cost is found, is applied with starting point pfL and final point pf1. The first half of this path is appended to the optimum path and W=W+L/2").
Regarding claim 7, Echeveste (in view of Miro) teaches a signal processing system comprising the features of claim 6 as discussed above.
Miro further teaches that the search condition includes a condition for fixing the reproduction position at the first time point to the indicated position at the first time point (Miro ¶0062: "In this case the starting point is fixed and only one path is computed forward with length L."), and fixing the reproduction position at the second time point to the indicated position at the second time point (Miro ¶0012: "An standard DTW algorithm in which a path with minimizes a defined global cost is found, is applied with starting point pfL and final point pf1.").
Regarding claim 18, Echeveste discloses a signal processing method comprising the features of claim 17 as discussed above.
Echeveste further teaches that the time series signal is a signal representing audio or video (Echeveste ¶0095: "In fact, the method applies to any “continuous” media, including for example audio and video."), and the acquisition unit is configured to acquire a plurality of indicated positions over time (Echeveste ¶0077: "This Time-Map is constantly re-evaluated at each interaction of the system with a human musician.").
Echeveste does not explicitly disclose that the control unit is configured to execute the time stretching by a path search to which two or more different indicated positions from among the plurality of indicated positions, and a search condition in accordance with characteristics of the time series signal are applied.
However, Miro teaches that the control unit is configured to execute the time stretching by a path search to which two or more different indicated positions from among the plurality of indicated positions (Miro ¶0034: "Real-time online alignment: Continue computing feature vectors and follow an incremental DTW guided by a forward path selection, ensuring that both audio signals remain aligned during the whole duration of the audio track. That is, the alignment is done block-by-block, first with the initial buffered signals but during the processing the signals continue to arrive and being aligned."), and a search condition in accordance with characteristics of the time series signal are applied (Miro ¶0050: "The resulting spectrum is mapped onto a 12-dimensional normalized chroma representation. The 12 dimensions of the chroma bins correspond to the 12 notes found in western music. The effect of this mapping is to reduce the audio to that of a single octave. Chroma features are typically used in music alignment as they are robust to variations in how the music is played. Finally, the different costs between these chroma frames are calculated using the inner normalized product.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing method of Echeveste by adding the control unit configuration of Miro to provide smooth playback in real time (Miro ¶0033).
Claims 8 is rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Miro, and further in view of Maezawa (US 20190156806 A1, May 23, 2019), hereinafter Maezawa '806.
Regarding claim 8, Echeveste (in view of Miro) teaches a signal processing system comprising the features of claim 5 as discussed above.
Echeveste (in view of Miro) does not explicitly disclose that the search condition includes an observation likelihood at each of the plurality of time points, the observation likelihood is a probability that each of a plurality of unit time intervals, obtained by dividing the time series signal on a time axis, corresponds to the reproduction position at each of the time points, and a probability distribution of the observation likelihood is defined by a mean corresponding to the indicated positions.
However, Maezawa '806 teaches that the search condition includes an observation likelihood at each of the plurality of time points (Maezawa '806 ¶0054: "The likelihood calculator 82 calculates a likelihood of observation L at each of multiple time points t within a piece for playback in conjunction with the performance of the piece for playback by performers P. That is, the distribution of likelihood of observation L across the multiple time points t within the piece for playback (hereafter, 'observation likelihood distribution') is calculated."), the observation likelihood is a probability that each of a plurality of unit time intervals, obtained by dividing the time series signal on a time axis, corresponds to the reproduction position at each of the time points (Maezawa '806 ¶0054: "An observation likelihood distribution is calculated for each unit segment (frame) obtained by dividing an audio signal A on the time axis. For an observation likelihood distribution calculated for a single unit segment of the audio signal A, a likelihood of observation L at a freely selected time point t is an index of probability that a sound represented by the audio signal A of the unit segment is output at the time point t within the piece for playback."), and a probability distribution of the observation likelihood is defined by a mean corresponding to the indicated positions (Maezawa '806 ¶0099: "The musical ensemble engine receives, from the score follower and at a point that is several frames after a position where a note switches to a new note in the music score, a normal distribution approximating an estimated current position or tempo distribution. Upon detecting the switch to the n-th note (hereafter, “onset event”) in the music data, the music score follower engine reports, to a musical ensemble timing generator, the time stamp tn indicating a time at which the onset event is detected, an estimated average position μn in the music score, and its variance σn 2.").It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste (as modified by Miro) by adding the observation likelihoods of Maezawa '806 to create a highly accurate estimate of playback position (Maezawa '806 ¶0005).
Claim 9 is rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Miro, and further in view of Maezawa '806, Meazawa (US 20140260911 A1, September 18, 2014), hereinafter Maezawa '911, and Cetinturk (US 20160042734 A1, February 11, 2016).
Regarding claim 9, Echeveste (in view of Miro and further in view of Maezawa '806) teaches a signal processing system comprising the features of claim 8 as discussed above.
Maezawa '806 further teaches that the time series signal is an audio signal representing performance sound of the musical piece (Maezawa '806: "Specifically, the performance analyzer 54 estimates each playback position T by analyzing a sound received by each of the sound receivers 224. As shown in FIG. 1, the performance analyzer 54 according to the present embodiment includes an audio mixer 542 and an analysis processor 544. The audio mixer 542 generates an audio signal A by mixing audio signals A0 generated by the sound receivers 224. Thus, the audio signal A is a signal representative of a mixture of multiple types of sounds represented by different audio signals A0.").
Echeveste (in view of Miro and further in view of Maezawa) does not explicitly disclose that the probability distribution of the observation likelihood at time points at which the indicated positions correspond to sound generation points of the audio signal, from among the plurality of time points, is defined by a first variance, and the probability distribution of the observation likelihood at time points at which the indicated positions do not correspond to the sound generation points of the audio signal, from among the plurality of time points, is defined by a second variance that is greater than the first variance.
However, Maezawa '911 teaches that the probability distribution of the observation likelihood at time points at which the indicated positions correspond to sound generation points of the audio signal, from among the plurality of time points, is defined by a first variance (Maezawa '911 ¶0082: "Assume that if the value of the number n of frames between the next beat is “0”, the onset feature values XO are distributed in accordance with the first normal distribution with a mean value of “3” and a variance of “1”. In other words, the value obtained by assigning the onset feature value XO(ti) as a random variable of the first normal distribution is the likelihood P (XO(ti)|Zb,n=0 (ti))."), and the probability distribution of the observation likelihood at time points at which the indicated positions do not correspond to the sound generation points of the audio signal, from among the plurality of time points, is defined by a second variance (Maezawa '911 ¶0082: "Furthermore, assume that if the value of the beat period b is “β”, with the value of the number n of frames between the next beat being “β/2”, the onset feature values XO are distributed in accordance with the second normal distribution with a mean value of “1” and a variance of “1”. In other words, the value obtained by assigning the onset feature value XO(ti) as a random variable of the second normal distribution is the likelihood P (XO(ti)|Zb=β3,n=β/2 (ti)).").
Furthermore, Cetinturk teaches a second variance that is greater than the first variance (Cetinturk ¶0116: "The non-signal (or weak-signal) and noise dominated spectral region coefficients will have high variance or deviation due to the infinite diversity of noise, in contrast, the speech-signal region coefficients will have relatively low variance or deviation due to the acoustic limits of the pronunciation of the same sound.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste (as modified by Miro and Maezawa '806) by adding the probability distribution and variances of Meazawa '911 and Cetinturk to improve the reliability and stability of the tempo (Maezawa '911 ¶0020).
Claim 10 is rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Miro, and further in view of Maezawa '806, Maezawa '911, Cetinturk, and Pachet et al. (US 20060074649 A1, April 6, 2006), hereinafter Pachet.
Regarding claim 10, Echeveste (in view of Miro and further in view of Maezawa '806, Maezawa '911, and Cetinturk) teaches a signal processing system comprising the features of claim 9 as discussed above.
Cetinturk further teaches that a variance of the probability distribution of the observation likelihood is set in accordance with the variation index (Cetinturk ¶0116: "The non-signal (or weak-signal) and noise dominated spectral region coefficients will have high variance or deviation due to the infinite diversity of noise, in contrast, the speech-signal region coefficients will have relatively low variance or deviation due to the acoustic limits of the pronunciation of the same sound.").
Echeveste (in view of Miro and further in view of Maezawa '806, Maezawa '911, and Cetinturk) does not explicitly disclose that the search condition includes a variation index representing a degree of variation of characteristics of the time series signal.
However, Pachet teaches that the search condition includes a variation index representing a degree of variation of characteristics of the time series signal (Pachet ¶0097: "When a stable zone has been identified, the detector 74 generates a stability score indicative of the level of spectral stability of this zone. In general, the stability score will be based on the value(s) of variation of the factor(s) taken into account when detecting the stable zones." Cetinturk links greater acoustic feature variation to greater model variance, and lower variation to lower variance. Pachet's variation-derived stability score to set the variance is a predictable parameterization.).
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste (as modified by Miro, Maezawa '806, Maezawa '911, and Cetinturk) by adding the variation index of Cetinturk and the parameterization of Pachet to set observation-likelihood variance so that more highly variable characteristics receive less probabilistic weight (Cetinturk ¶0116).
Claims 11 and 14 are rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Miro, and further in view of Maezawa '911 and Maezawa '806.
Regarding claim 11, Echeveste (in view of Miro) teaches a signal processing system comprising the features of claim 5 as discussed above.
Does not explicitly disclose that the search condition includes a transition probability that is set for each combination of two unit time intervals, from among the plurality of unit time intervals obtained by dividing the time series signal on the time axis, and that represents a probability of the reproduction position transitioning between the two unit time intervals.
However, Maezawa '911 teaches that the search condition includes a transition probability that is set for each combination of two unit time intervals, from among the plurality of unit time intervals obtained by dividing the time series signal on the time axis (Maezawa '911 ¶0088: "In this example, furthermore, the values of log transition probability T from a state where the value of the beat period b is 'βs' with the value of the number n of frames 'ηs' to a state where the value of the beat cycle b is 'βe' with the value of the number n of frames 'ηe' are set as follows: if 'ηe=0', 'βe=βs', and 'ηe=βe−1', the value of log transition probability T is '−0.2'.").
Furthermore, Maezawa '806 teaches that the search condition includes a transition probability that represents a probability of the reproduction position transitioning between the two unit time intervals (Maezawa '806 ¶0094: "First, a piece of music is divided into R segments, and each segment is treated as consisting of a single state. The r-th segment has n number of frames, and also has for each n the currently passing frame 0≤l<n as a state variable. Thus, n corresponds to a tempo within a given segment, and the combination of r and l corresponds to a position in a music score. Such a transition in a state space is expressed in the form of a Markov process such as follows.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste (as modified by Miro) by adding the transition probabilities of Maezawa '911 and Maezawa '806 to enhance accuracy of tempo estimation (Maezawa '911 ¶0018).
Regarding claim 14, Echeveste (in view of Miro, and further in view of Maezawa '911 and Maezawa '806) teaches a signal processing system comprising the features of claim 11 as discussed above.
Maezawa '806 further teaches that the transition probability of the reproduction position remaining at a last time point of a first inter-sounding time interval, from among a plurality of inter-sounding time intervals obtained by dividing the audio signal on the time axis by a plurality of sound generation points (Maezawa '806 ¶0095: "This means the selection of n enables the system to decide an approximate duration within a segment, and thus the self transition probability p can absorb subtle variations in tempo within the segment. The length of the segment or the self transition probability is obtained by analyzing the music data. Specifically, the system uses tempo indications or annotation information such as fermata.").
Furthermore, Maezawa '911 further teaches: that it is greater than the transition probability of the reproduction position transitioning from the last time point to a time point within a second inter-sounding time interval immediately following the first inter-sounding time interval (Maezawa '911 ¶0088: "In this example, furthermore, the values of log transition probability T from a state where the value of the beat period b is 'βs' with the value of the number n of frames 'ηs' to a state where the value of the beat cycle b is 'βe' with the value of the number n of frames 'ηe' are set as follows: if 'ηe=0', 'βe=βs', and 'ηe=βe−1', the value of log transition probability T is '−0.2'.").
Claim 12 is rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Miro, and further in view of Maezawa '911, Maezawa '806, and Blair et al. (US 20030165327 A1, September 4, 2003), hereinafter Blair.
Regarding claim 12, Echeveste (in view of Miro, and further in view of Maezawa '911 and Maezawa '806) teaches a signal processing system comprising the features of claim 11 as discussed above.
Maezawa '806 further teaches that the time series signal is an audio signal representing performance sound of the musical piece (Maezawa '806 ¶0033: "The audio mixer 542 generates an audio signal A by mixing audio signals A0 generated by the sound receivers 224. Thus, the audio signal A is a signal representative of a mixture of multiple types of sounds represented by different audio signals A0.").
Miro further teaches that the transition probability when the audio signal is silent in both of the two unit time intervals (Miro ¶0057: "It is worth noting that this selective process needs to be suspended during silent frames. Otherwise the noise of these frames would make the selection process random.").
Echeveste (in view of Miro, and further in view of Maezawa '911 and Maezawa '806) does not explicitly disclose that the transition probability is greater than the transition probability when the audio signal contains sound in one or both of the two unit time intervals.
However, Blair teaches that the transition probability is greater than the transition probability when the audio signal contains sound in one or both of the two unit time intervals (Blair ¶0009: "The removing step can include selectively removing a percentage portion of the periods of relative silence based on a selected video trick mode playback speed. According to one aspect of the invention, the removing step can further include determining an optimized percentage portion of each period of silence that must be removed in order to synchronize the audio portion and the video portion for playback after the concatenating step. The removing step can also include increasing the percentage portion of the periods of silence that are removed in order to achieve a faster video trick mode playback speed.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste (as modified by Miro, Maezawa '911, and Maezawa '806) by adding the silence removal of Blair to avoid having the noise of silent frames make the selection process random (Miro ¶0057).
Claim 13 is rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Miro, and further in view of Maezawa '911, Maezawa '806, Blair, Cetinturk, and Pachet.
Regarding claim 13, Echeveste (in view of Miro, and further in view of Maezawa '911, Maezawa '806, and Blair) teaches a signal processing system comprising the features of claim 11 as discussed above.
Maezawa '806 further teaches that the probability distribution of the transition probability when the audio signal contains the sound in one or both of the two unit time intervals is defined by a mean that is set to a prescribed value (Maezawa '806 ¶0102: "Accordingly, given that the matrix generated by adding σn (p) to τn,0,0 (p) is Σn (p), εn (p)˜N(0,Σn (p)) is derived. N(a, b) means the normal distribution of the average a and variance b.").
Echeveste (in view of Miro, and further in view of Maezawa '911, Maezawa '806, and Blair) does not explicitly disclose: by a variance corresponding to a variation index representing a degree of variation of acoustic characteristics of the audio signal.
However, Cetinturk teaches: by a variance corresponding to a variation index (Cetinturk ¶0116: "The non-signal (or weak-signal) and noise dominated spectral region coefficients will have high variance or deviation due to the infinite diversity of noise, in contrast, the speech-signal region coefficients will have relatively low variance or deviation due to the acoustic limits of the pronunciation of the same sound.").
Furthermore, Pachet teaches: representing a degree of variation of acoustic characteristics of the audio signal (Pachet ¶0097: "When a stable zone has been identified, the detector 74 generates a stability score indicative of the level of spectral stability of this zone. In general, the stability score will be based on the value(s) of variation of the factor(s) taken into account when detecting the stable zones.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste (as modified by Miro, Maezawa '911, Maezawa '806, and Blair) by adding the variation index of Cetinturk and parametrization of Pachet to use stability data during time-stretching (Pachet ¶0097).
Claim 16 is rejected under 35 U.S.C. 103 as unpatentable over Echeveste in view of Pachet and further in view of Little (US 20040114769 A1, June 17, 2004).
Regarding claim 16, Echeveste discloses a signal processing system comprising the features of claim 15 as discussed above.
Echeveste does not explicitly disclose that the reproduction unit is configured to select, in response to a first operation occurring at a first time point in the user's performance and a second operation occurring at a second time point after the first time point has elapsed in the user's performance, as an operating intensity at the second time point, a larger one of a first intensity obtained by reducing an intensity of the first operation over time from the first time point to the second time point, and a second intensity of the second operation, and control a volume of reproduction sound of the time series signal in accordance with the operating intensity.
However, Pachet teaches that the reproduction unit is configured to select, in response to a first operation occurring at a first time point in the user's performance and a second operation occurring at a second time point after the first time point has elapsed in the user's performance (Pachet ¶0122: "Certain play-modes are interesting because they select audio samples for output according to their original context in the original audio file, e.g. their position within the audio file (fourth sample, twentieth sample, etc.). This context is indicated by the meta-data associated with the audio sample. For instance, notes triggered from the user's operation of playable keys can, when he plays the next key, automatically be followed by playback of a sample representing a close event in the original music stream."), and control a volume of reproduction sound of the time series signal in accordance with the operating intensity (Pachet ¶0004: "When a key on the MIDI keyboard is depressed, a pre-stored audio sample is played back at a pitch corresponding to the depressed key and with a volume corresponding to the velocity of depression of the key.").
Furthermore, Little teaches: as an operating intensity at the second time point, a larger one of a first intensity obtained by reducing an intensity of the first operation over time from the first time point to the second time point, and a second intensity of the second operation (Little ¶0040: "This peak-hold circuit functions by efficiently calculating the L∞ norm of the control signal (28) whose samples have been weighted according to an exponentially decaying envelope. The circuit function is described by the following operation: y(z)=∥x(t−n)|K p|n∥∞").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the signal processing system of Echeveste by adding the reproduction unit of Pachet and Little to smoothly track the peaks of the control signal as they decay (Little ¶0041).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHILIP SCOLES whose telephone number is (703)756-1831. The examiner can normally be reached Monday-Friday 8:30-4:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei Hammond can be reached on 571-270-7938. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHILIP G SCOLES/
Examiner, Art Unit 2837
/DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837