DETAILED ACTION
This communication is in response to the Application filed on 08/31/2022 (foreign priority). Claims 1-15 have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Korean Patent Application No. KR10-2022-0110100, filed on August 31, 2022.
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/10/2025 was filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: ELECTRONIC DEVICE AND CONTROL METHOD FOR PITCH SHIFTING.
The abstract of the disclosure is objected to because it is longer than 150 words. A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
Claim Objections
Claim(s) 2-8 objected to because they claim dependence of claim 1 being a method when claim 1 recites an apparatus:
Regarding claims 2, 5, and 6 “The method of claim 1, wherein at least one processor, individually and/or collectively, is configured to: ...” should read:
The apparatus of claim 1, wherein at least one processor, individually and/or collectively, is configured to: ...
Regarding claim 3, “The method of claim 2, wherein at least one processor, individually and/or collectively, is configured to: ...” should read:
The apparatus of claim 2, wherein at least one processor, individually and/or collectively, is configured to: ...
Regarding claim 4, “The method of claim 3, wherein the first pitch shift value is an integer value and the second pitch shift value is a decimal value.” should read:
The apparatus of claim 3, wherein the first pitch shift value is an integer value and the second pitch shift value is a decimal value.
Regarding claim 7, “The method of claim 6, wherein at least one processor, individually and/or collectively, is configured to: ...” should read:
The apparatus of claim 6, wherein at least one processor, individually and/or collectively, is configured to: ...
Regarding claim 8, “The method of claim 6, wherein the feature information of the first voice data includes cepstrum, the pitch, or correlation.” should read:
The apparatus of claim 6, wherein the feature information of the first voice data includes cepstrum, the pitch, or correlation.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-5, and 9-13 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
With respect to claim(s) 1 and 9, the limitation(s) “identifying a target pitch shift value for shifting a pitch of voice data,” “dividing the identified target pitch shift value into a first pitch shift value and a second pitch shift value,” “identifying a pitch shift embedding value based on the first pitch shift value,” “updating a feature of a pitch of first voice data based on the second pitch shift value and obtaining second voice data,” “identifying a pitch embedding value based on a pitch of the obtained second voice data,” and “obtaining third voice data to which the pitch of the first voice data is shifted based on the pitch shift embedding value and the pitch embedding value,” as drafted, are processes that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “an electronic device,” “memory storing at least one instruction,” and “at least one processor, comprising processing circuitry, individually and/or collectively,” and specifically for claim 1, nothing in the claim’s elements preclude the steps from being practically performed in the mind. More specifically, but not including the generic computer components, the mental processes of a human singer mentally setting a desired key change, splitting it into a coarse interval and a fine micro-adjustment, humming the coarse pitch, sensing that hummed pitch as an internal reference, then bending their voice to the exact target pitch by ear. If a limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim(s) recites(s) an abstract idea.
This judicial exception is not integrated into a practical application. In particular, only claim 1 recites 3 additional elements — an electronic device, memory storing at least one instruction, and at least one processor, comprising processing circuitry, individually and/or collectively.
The electronic device is recited with a high level of generality (see [0046], where an electronic device 100 may be a vocoder. The vocoder may refer, for example, to a device comprising circuitry configured to generate a voice signal by utilizing feature information of voice data. A function of shifting or controlling a feature about a pitch among feature information of the voice data may be added to the vocoder. The electronic device 100 is not limited to the vocoder and may be various electronic devices 100 by which a pitch of the voice data may be shifted).
Furthermore, the memory storing at least one instruction is recited with a high-level of generality (see [0050], where memory 110 may store temporarily or non-temporarily various programs or data and transmit the stored information to a processor 120 according to a call of the processor 120. Also, the memory 110 may store various information required for an operation, processing, a control operation, or the like of the processor 120 as an electronic format. Also see [0051], where the memory 110 may include at least one of, for example, a main memory unit and an auxiliary memory unit. The main memory unit may be implemented using a semiconductor storage medium such as ROM and/or RAM. ROM may include, for example, a general ROM, EPROM, EEPROM, and/or MASK-ROM. RAM may include, for example, DRAM and/or SRAM. The auxiliary memory may be implemented using at least one storage medium that may permanently or semi-permanently store data such as a flash memory 110 device, a secure digital (SD) card, a solid-state drive (SSD), a hard disc drive (HDD), a magnetic drum, a compact disk (CD), an optical medium such as a DVD or a laser disk, a magnetic tape, a magneto-optical disk and/or a floppy disk).
Likewise, the at least one processor, comprising processing circuitry, individually and/or collectively is recited with a high-level of generality (see [0060], where the processor 120 may be implemented in various forms. For example, one or more processors 120 may include one or more of a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a many integrated core (MIC), a digital signal processor (DSP), a neural processing unit (NPU), a hardware accelerator, or a machine learning accelerator. The one or more processors 120 may control one or any combination of other components of the electronic device 100 and may perform an operation related to a communication or data processing. The one or more processors 120 may execute one or more programs or instructions stored in the memory 110. For example, the one or more processors 120 may execute one or more instructions stored in the memory 110, thereby performing a method according to an embodiment of the disclosure. The processor 120 may include various processing circuitry and/or multiple processors. For example, as used herein, including the claims, the term "processor" may include various processing circuitry, including at least one processor, wherein one or more of at least one processor, individually and/or collectively in a distributed manner, may be configured to perform various functions described herein. As used herein, when "a processor", "at least one processor", and "one or more processors" are described as being configured to perform numerous functions, these terms cover situations, for example and without limitation, in which one processor performs some of recited functions and another processor(s) performs other of recited functions, and also situations in which a single processor may perform all recited functions. Additionally, the at least one processor may include a combination of processors performing various of the recited /disclosed functions, e.g., in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions).
Accordingly, the additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim(s) is/are directed to an abstract idea.
The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, and concerning claim 1 alone, the additional elements of an electronic device, memory storing at least one instruction, and at least one processor, comprising processing circuitry, individually and/or collectively to perform the identifying, dividing, identifying, updating, identifying and obtaining steps amount to no more than mere instructions to apply an exception using generic computer components cannot provide an inventive concept. The claim(s) is/are not patent eligible.
With respect to claim(s) 2, and 10, the claim(s) recite(s) “dividing the target pitch shift value into the first pitch shift value and the second pitch shift value based on a pitch shift embedding table,” which reads on a human musician recalling a mental scale of discrete pitch steps, picking the nearest step to split the shift into a table-derived coarse part and a residual fine part. No additional limitations are present.
With respect to claim(s) 3, and 11, the claim(s) recite(s) “identifying an index value closest to the target pitch shift value among at least one index value included in the pitch shift embedding table as the first pitch shift value; and identifying a difference between the identified first pitch shift value and the target pitch shift value as the second pitch shift value,” which reads on a human mentally scanning their internalized list of familiar intervals (semitones), selecting the nearest matching interval, and subtracting it from the desired shift to compute the leftover difference. No additional limitations are present.
With respect to claim(s) 4, and 12, the claim(s) recite(s) “wherein the first pitch shift value is an integer value and the second pitch shift value is a decimal value,” which reads on a human representing a pitch change as a whole number of semitones (integer) plus a fractional cent adjustment (decimal) in their conscious auditory planning. No additional limitations are present.
With respect to claim(s) 5, and 13, the claim(s) recite(s) “wherein obtaining the third voice data includes: obtaining metadata based on the pitch shift embedding value and the pitch embedding value; encoding the metadata in a frame unit; and decoding the encoded metadata in a sample unit and obtaining the third voice data.,” which reads on a human grouping a target pitch and a current pitch into a mental phrase-level package, holding that package in working memory, then sequentially articulating each individual sound unit to produce the final pitched voice utterance. No additional limitations are present.
These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 5-10, and 13-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Morrison (US 11915714 B2) in view of Matsumoto (JP 2017111274 A).
Regarding claim 1, Morrison teaches:
An electronic device, comprising: memory storing at least one instruction (see column 18, lines 20-30, where the memory device 904 includes any suitable non-transitory computer-readable medium for storing data, program code, or both. A computer-readable medium can include any electronic, optical, magnetic, or other storage device capable of providing a processor with computer-readable instructions or other program code. Non-limiting examples of a computer-readable medium include a magnetic disk, a memory chip, a ROM, a RAM, an ASIC, optical storage, magnetic tape or other magnetic storage, or any other medium from which a processing device can read instructions);
and at least one processor, comprising processing circuitry, individually and/or collectively, configured to execute the at least one instruction (see column 18, lines 10-19, where the processor 902 communicatively is coupled to one or more memory devices 904. The processor 902 executes computer-executable program code stored in a memory device 904, accesses information stored in the memory device 904, or both. Examples of the processor 902 include a microprocessor, an application-specific integrated circuit ("ASIC"), a field-programmable gate array ("FPGA"), or any other suitable processing device. The processor 902 can include any number of processing devices, including a single processing device), wherein at least one processor, individually and/or collectively, is configured to:
identify a target pitch shift value for shifting a pitch of voice data (see column 1, lines 42-49, where the vocoder system is configured to apply a target prosody to audio data, where the target prosody indicates phoneme durations, a pitch contour, or a combination of both. In some examples, the target prosody has been determined based on a larger context of audio data around and including the audio data to which the target prosody is to be applied, so as to correct the prosody of the audio data);
identify a pitch shift embedding value based on the first pitch shift value (see column 8, lines 30-32, "the neural vocoder 175 inputs twenty acoustic features, including a pitch feature, a periodicity feature, and eighteen cepstral coefficients. Also see column 8, lines 35-39, "The frame-rate network 276 outputs an embedding, which can be a 128-dimensional embedding representing the current audio frame. The embedding may be held constant for the duration of processing the current frame, which may include multiple samples);
obtain second voice data by updating a feature of a pitch of first voice data based on the second pitch shift value (see column 1, lines 65-67, "The neural vocoder performs pitch-shifting and time-stretching to modify the audio data toward the target prosody. Also see column 7, lines 42-45, "The neural vocoder 175 outputs an updated version of that audio data having frequency and amplitude of the modulator signal but with the timbre of the carrier signal);
identify a pitch embedding value based on a pitch of the obtained second voice data (see column 8, lines 30-32, "the neural vocoder 175 inputs twenty acoustic features, including a pitch feature, a periodicity feature, and eighteen cepstral coefficients. Also see column 8, lines 35-39, "The frame-rate network 276 outputs an embedding, which can be a 128-dimensional embedding representing the current audio frame. The embedding may be held constant for the duration of processing the current frame, which may include multiple samples); and
obtain third voice data to which the pitch of the first voice data is shifted based on the pitch shift embedding value and the pitch embedding value (see column 1, lines 65-67, and column 2, lines 1-2, where the neural vocoder performs pitch-shifting and time-stretching to modify the audio data toward the target prosody, thus mapping the acoustic features to an updated version of the audio data whose pitch and rhythm now match or at least more closely match the target prosody. Also see column 8, lines 54-57, where the neural vocoder 175 outputs a synthesized output sample sₜ, which is raw audio data and which is a combination of the excitation value and the prediction value).
For purposes of this rejection, the claim term “pitch shift embedding value” and “pitch embedding value” are interpreted under broadest reasonable interpretation as a vector representation of pitch-related acoustic featured in a latent space or such representation. The term “embedding” is a known term in machine learning and signal processing that refers to a vectorized representation of input data learned by a neural network. While the claim language uses the term “embedding,” Morrison does not use the same terminology when referring to acoustic features being input into the neural vocoder. However, this mere difference in terminology does not create a substantive distinction and serve the same function of converting pitch-related information into a format that can be processed by a neural network.
Morrison fails to teach wherein at least one processor, individually and/or collectively, is configured to: divide the identified target pitch shift value into a first pitch shift value and a second pitch shift value. However, Matsumoto does teach:
at least one processor (see Page 3, where the control unit 31 includes an arithmetic processing circuit such as a CPU. The control unit 31 executes a control program stored in the storage unit 33 by the CPU, and realizes various functions in the data processing device 3), individually and/or collectively, is configured to:
divide the identified target pitch shift value into a first pitch shift value and a second pitch shift value (see Page 2, where the feature amount corresponds to the pitch of the singing voice. Also see Page 5, where the specifying unit 305 is based on the feature amount data (referred to as singing feature amount data) generated in the feature amount calculating unit 303 and the model feature amount data registered in the database 5 in association with the sung song. Also see Page 7, where the singing feature quantity data Ss3 is divided into a plurality of sections, the model feature quantity data Sd is compared with each section, and the singing feature quantity data Ss3 is expanded or contracted for each section (FIG. 5; step S109). This process is called a division expansion / contraction process).
Morrison and Matsumoto are both considered to be analogous to the claimed invention because they are in the same field of electric digital data processing and speech analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Morrison to incorporate the teachings of Matsumoto to divide the identified target pitch shift value into a first pitch shift value and a second pitch shift value in order to adjust the reproduction timing of the singing data; and therefore, effectively changing the pitch of a singing audio according to a target pitch (see Page 1, where in order to easily utilize sound data to adjust a reproduction timing of acquired sound data, a data processor comprises; a sound data acquisition part for acquiring first sound data including time information; a featured value calculation part for generating first featured value data associated with the time information; a specification part for, on the basis of second featured value data and the first featured value data, specifying a first data position of the first featured value data corresponding to a second data position of the second featured value data; a moving part for, on the basis of a temporal positional relation between the first data position and the second data position, changing the time information such that reproduction timings move in parallel; and an expansion/contraction part for comparing the first featured value data and the second featured value data associated with the changed time information to detect a correspondence relation between respective data portions, and for changing the time information such that an interval between the reproduction timings expands/contracts on the basis of the correspondence relation).
Regarding claim 2, which depends on claim 1, Morrison in view of Matsumoto teaches all of the limitations in claim 1. Furthermore, Morrison teaches at least one processor, individually and/or collectively, configured to:
divide the target pitch shift value into the first pitch shift value and the second pitch shift value based on a pitch shift embedding table (see Abstract, where a pitch feature as a quantized pitch value of the sample by assigning a pitch value, of the target prosody or the audio data, to at least one of a set of pitch bins having equal widths in cents. Also see column 16, lines 6-11, where the vocoder system 170 utilizes a quantization of the frequency range 50-550 Hz, in which pitch bins to which the pitch values are assigned are equally spaced in base-2 log scale. This results in equal-width pitch bins in terms of cents, specifically, with each pitch bin being 16.3 cents wide).
Regarding claim 3, which depends on claim 2, Morrison in view of Matsumoto teaches all of the limitations in claim 2. Furthermore, Morrison teaches wherein at least one processor, individually and/or collectively, is configured to:
identify an index value closest to the target pitch shift value among at least one index value included in the pitch shift embedding table as the first pitch shift value (see Abstract, where computing respective acoustic features for a sample includes computing a pitch feature as a quantized pitch value of the sample by assigning a pitch value, of the target prosody or the audio data, to at least one of a set of pitch bins having equal widths in cents. Also see column 16, lines 6-11, where the vocoder system 170 utilizes a quantization of the frequency range 50-550 Hz, in which pitch bins to which the pitch values are assigned are equally spaced in base-2 log scale. This results in equal-width pitch bins in terms of cents, specifically, with each pitch bin being 16.3 cents wide).
identify a difference between the identified first pitch shift value and the target pitch shift value as the second pitch shift value (see column 16, lines 38-43, where to compute the pitch features, the feature extraction subsystem 270 may dither the extracted pitch with random noise drawn from a triangular distribution, which may be centered at zero and may have a width equal to two Crepe pitch bins (i.e., 40 cents). This can reduce the quantization error without increasing the noise floor).
Regarding claim 5, which depends on claim 1, Morrison in view of Matsumoto teaches all of the limitations in claim 1. Furthermore, Morrison teaches wherein at least one processor, individually and/or collectively, is configured to:
obtain metadata based on the pitch shift embedding value and the pitch embedding value (see column 8, lines 35-37, where the frame-rate network 276 outputs an embedding, which can be a 128-dimensional embedding representing the current audio frame);
encode the metadata in a frame unit (see column 8, lines 22-33, where as shown in FIG. 2, in some embodiments, the frame-rate network 276 includes two one-dimensional (1D) convolution layers with tanh activations and two dense fully connected (FC) layers with tanh activations. As part of the frame-rate network 276, acoustic features first go through the two 1D convolution layers with a filter size of 3, thus labeled 1×3 in FIG. 2. This results in a receptive field of five frames, including two frames ahead and two frames behind a current frame. In some examples, the neural vocoder 175 inputs twenty acoustic features, including a pitch feature, a periodicity feature, and eighteen cepstral coefficients, such as BFCCs. Also see column 16, lines 54-62, where at block 720, the process 700 involves computing other acoustic features to represent the audio data as needed. To this end, in some embodiments, the feature extraction subsystem 270 also determines cepstral coefficients, such as Bark-frequency cepstral coefficients (BFCCs), to represent and encode the audio data. For instance, feature extraction subsystem 270 generates eighteen BFCCs, such as through the use of one or more techniques known in the art, to use as input acoustic features for the neural vocoder 175); and
obtain the third voice data by decoding the encoded metadata in a sample unit (see column 8, lines 40-57, where the sample-rate network 278 includes an embedding layer, labeled “concat” in FIG. 2, which combines (e.g., concatenates) its inputs. The sample-rate network 278 may take the following four inputs: the prior excitation value (i.e., from the last input sample); the prior synthesized output sample; a prediction value; and the embedding output by the frame-rate network 276, optionally after nearest neighbor upsampling. Following the embedding layer, the sample-rate network 278 includes two GRU layers and a dual fully connected layer. The output of the dual fully connected layer is used with a softmax activation to compute the probability P(e.sub.t) of each possible excitation value e.sub.t for the current sample. The sample-rate network 278 samples the probability distribution of P(e.sub.t) to determine and output an excitation value corresponding to the sample. The neural vocoder 175 outputs a synthesized output sample s.sub.t, which is raw audio data and which is a combination of the excitation value and the prediction value).
Regarding claim 6, which depends on claim 1, Morrison in view of Matsumoto teaches all of the limitations in claim 1. Furthermore, Morrison teaches wherein at least one processor, individually and/or collectively, is configured to:
identify input data of a pitch shift embedding model based on information about the feature of the first voice data and the target pitch shift value (see column 11, lines 21-32, where at block 330, the process 300 involves applying the target prosody to the subject audio data to match, or at least more closely match, the target prosody determined at blocks 315-325. To this end, in some embodiments, the vocoder system 170 of the operations system 120 extracts acoustic features of the subject audio data and the target prosody, as described in more detail below, and then utilizes the neural vocoder 175 to perform pitch-shifting and time-stretching, as is also described in more detail below);
identify output data of the pitch shift embedding model based on the third voice data (see column 11, lines 29-32, where the vocoder system 170 outputs an audio signal (e.g., a waveform) that is a modified version of the subject audio data having been pitch-shifted and time-stretched by the neural vocoder 175);
identify a loss of the pitch shift embedding model based on the input data and the output data; and learn the pitch shift embedding model and the pitch shift embedding table based on the input data, the output data, and the loss (see column 17, lines 55-61, where at block 820, the process involves training the neural vocoder 175 using the acoustic features extracted in block 815, thereby teaching the neural vocoder 175 to minimize the error between its output and the desired output audio data for each utterance used for training. As such, the neural vocoder 175 can learn to map audio data to modified audio data having shifted pitch or stretched rhythm).
Regarding claim 7, which depends on claim 6, Morrison in view of Matsumoto teaches all of the limitations in claim 6. Furthermore, Morrison teaches wherein at least one processor, individually and/or collectively, is configured to:
extract feature information from the first voice data (see column 1, lines 53-58, where the vocoder system computes acoustic features representing samples of the target prosody and the audio data, where respective acoustic features for each sample include a pitch feature and a periodicity feature representing the target prosody as well as cepstral coefficients representing the audio data).
obtain fourth voice data by augmenting the pitch of the first voice data based on the target pitch shift value (see column 1, lines 65-67, where the neural vocoder performs pitch-shifting and time-stretching to modify the audio data toward the target prosody. Also see column 1, lines 50-51, where the vocoder system applies the target prosody to the audio data);
extract feature information from the obtained fourth voice data (see column 1, lines 51-52, where the vocoder system extracts acoustic features from the target prosody and the audio data. Also see column 8, lines 25-37, where as part of the frame-rate network 276, acoustic features first go through the two ID convolution layers with a filter size of 3, thus labeled lx3 in FIG. 2. This results in a receptive field of five frames, including two frames ahead and two frames behind a current frame. In some examples, the neural vocoder 175 inputs twenty acoustic features, including a pitch feature, a periodicity feature, and eighteen cepstral coefficients, such as BFCCs. The output of the two convolution layers is added to a residual connection and then goes through the two fully connected layers. The frame-rate network 276 outputs an embedding, which can be a 128-dimensional embedding representing the current audio frame); and
identify input data of the pitch shift embedding model based on the feature information extracted from the first voice data and the feature information extracted from the fourth voice data (see column 1, lines 45-49, where the target prosody has been determined based on a larger context of audio data around and including the audio data to which the target prosody is to be applied, so as to correct the prosody of the audio data. Also see column 2, lines 64-67, where the correction system extracts acoustic features from the target prosody and the subject audio data, including utilizing a prediction model to predict some of those acoustic features. Also see column 3, lines 62-65, where the neural vocoder takes as input acoustic features of the subject audio data and the target prosody and performs pitch-shifting and time-stretching to modify the subject audio data toward the target prosody).
Regarding claim 8, which depends on claim 6, Morrison in view of Matsumoto teaches all of the limitations in claim 6. Furthermore, Morrison teaches wherein the feature information of the first voice data includes cepstrum, the pitch, or correlation (see column 1, lines 55-58, where respective acoustic features for each sample include a pitch feature and a periodicity feature representing the target prosody as well as cepstral coefficients representing the audio data).
Regarding claim 9, which recites a method, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 1. As detailed in the rejection of claim 1, the disclosed apparatus teaches each step of the method recited in claim 9. Accordingly, claim 9 is rejected for the same reasons set forth in the rejection of claim 1.
Regarding claim 10, which recites a method, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 2. As detailed in the rejection of claim 2, the disclosed apparatus teaches each step of the method recited in claim 10. Accordingly, claim 10 is rejected for the same reasons set forth in the rejection of claim 2.
Regarding claim 11, which recites a method, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 3. As detailed in the rejection of claim 3, the disclosed apparatus teaches each step of the method recited in claim 11. Accordingly, claim 11 is rejected for the same reasons set forth in the rejection of claim 3.
Regarding claim 13, which recites a method, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 5. As detailed in the rejection of claim 5, the disclosed apparatus teaches each step of the method recited in claim 13. Accordingly, claim 13 is rejected for the same reasons set forth in the rejection of claim 5.
Regarding claim 14, which recites a method, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 6. As detailed in the rejection of claim 6, the disclosed apparatus teaches each step of the method recited in claim 14. Accordingly, claim 14 is rejected for the same reasons set forth in the rejection of claim 6.
Regarding claim 15, which recites a method, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 7. As detailed in the rejection of claim 7, the disclosed apparatus teaches each step of the method recited in claim 15. Accordingly, claim 15 is rejected for the same reasons set forth in the rejection of claim 7.
Claim(s) 4, and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Morrison (US 11915714 B2) in view of Matsumoto (JP 2017111274 A), and further in view of Morsy (WO 2021175460 A1).
Regarding claim 4, which depends on claim 3, Morrison in view of Matsumoto teaches all of the limitations in claim 3, but Morrison and Matsumoto fail to teach wherein the first pitch shift value is an integer value and the second pitch shift value is a decimal value.
However, Morsy does teach:
wherein the first pitch shift value is an integer value (see Page 4, where pitch shift value (e.g. number of semitones or cents up/down). Also see Page 5, where the pitch scaling effect may shift the pitch of the audio data of the first audio track up or down by a predetermined number of semitones. Also see Page 9, where the pitch shift value may be a number of semitones by which the first key needs to be shifted up or down in order to assume a key that differs from the second key by a fifth); and
the second pitch shift value is a decimal value (see Page 4, where pitch shift value (e.g. number of semitones or cents up/down. Also see Page 8, wherein the first effect unit is a pitch scaling unit adapted to shift the pitch of the first audio track by the pitch shift value).
For purposes of this rejection, the claim term “integer” and “decimal” are interpreted under broadest reasonable interpretation as a semitone and cent. The term “integer” and “decimal” are known terms in audio and signal processing that refer to a whole-number and fractional values used to represent pitch shift amounts. A semitone is the smallest standard interval in Western music, and pitch shifts are conventionally expressed as a number of semitones (e.g., +2, -4, +5), which are integer values. Conversely, a cent is a logarithmic unit of measure used for musical interval, where one semitone equals 100 cents. As a result, pitch shifts expressed in cents are fractional values of a semitone and constitute decimal pitch shift values. Therefore, this mere difference in terminology does not create a substantive distinction and serve the same function of converting pitch-related information into a format that can be processed by a neural network.
Morrison, Matsumoto, and Morsy are considered to be analogous to the claimed invention because they are in the same field of electric digital data processing, and audio and speech analysis. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date to have modified Morrison and Matsumoto to incorporate the teachings of Morsy to include a first pitch shift value to be an integer in order to maintain the audio quality without distortion and perform coarse/fine pitch control (see Page 3, where for example, the audio effect may be a pitch scaling effect changing the pitch of audio data while maintaining its playback duration, which might be desired by DJs to match the key of one song to that of another song such as to smoothly crossfade between the two songs (without the clashing of different keys). Conventional pitch scaling will lead to an unnatural distortion of the music, when the pitch is shifted by more than one or two semitones. This results in a limitation of the creative freedom of the DJ).
Regarding claim 12, which recites a method, this claim is rejected as unpatentable over the same prior art and reasoning applied against claim 4. As detailed in the rejection of claim 4, the disclosed apparatus teaches each step of the method recited in claim 12. Accordingly, claim 12 is rejected for the same reasons set forth in the rejection of claim 4.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure:
Endo et al. (U.S. PG Pub. No. 2008/0091417): Pitch Conversion Method And Device.
Kang et al. (U.S. PG Pub. No. 2022/0165250): METHOD FOR CHANGING SPEED AND PITCH OF SPEECH AND SPEECH SYNTHESIS SYSTEM.
Morrison et al. (U.S. Patent No. 11830481): Context-aware Prosody Correction Of Edited Speech.
Hantrakul et al. (U.S. PG Pub. No. 2023/0154451): DIFFERENTIABLE WAVETABLE SYNTHESIZER.
Clark et al. (U.S. Patent No. 10923107): Clockwork hierarchical variational autoencoder.
Kayama et al. (U.S. Patent No. 10789937): Speech Synthesis device and method.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHN HONG FANG-WU whose telephone number is (571)270-0607. The examiner can normally be reached Monday - Friday, 9AM to 5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras Shah can be reached at (571)-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JOHN HONG FANG-WU/ Examiner, Art Unit 2653
/DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657