DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant’s election without traverse of claims 1-10 in the reply filed on 7/11/2026 is acknowledged.
Applicant is reminded that upon the cancelation of claims to a non-elected invention, the inventorship must be corrected in compliance with 37 CFR 1.48(a) if one or more of the currently named inventors is no longer an inventor of at least one claim remaining in the application. A request to correct inventorship under 37 CFR 1.48(a) must be accompanied by an application data sheet in accordance with 37 CFR 1.76 that identifies each inventor by his or her legal name and by the processing fee required under 37 CFR 1.17(i).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-10 and 21-30 are rejected under 35 U.S.C. 103 as being unpatentable over Cheng; Qiang et al. (US 6892175 B1) hereinafter CHENG, in view of Sharma; Ravi K. et al. (US 9305559 B2) hereinafter SHARMA, in further view of CHA JAE UK (KR 20220064817 A), hereinafter CHA.
Regarding claim 1, CHENG teaches:
(Original) A computer-implemented method for embedding watermarks in audio signals, comprising: obtaining, by a computer, a watermarked audio signal comprising a speech signal and a watermark signal embedded at the speech signal;
CHENG [Abstract] “Methods and apparatus for encoding an arbitrary digital message, e.g., a watermark, into a speech signal are provided. In one aspect of the invention, a method of embedding digital information in a speech signal comprises the steps of: (i) generating a spread spectrum signal, wherein the spread spectrum signal is representative of the digital information and further wherein the spread spectrum signal is within a frequency bandwidth corresponding to speech; and (ii) embedding the spread spectrum signal in the speech signal. By making use of spread spectrum technology and speech analysis techniques in the signal generation and embedding operations, respectively, significantly higher bit rates can be embedded into the speech signal without effecting the perceived quality of the recording. The invention also provides methods and apparatus for recovering the digital information embedded in the speech signal.”
determining, by the computer, a strength of the watermark signal
CHENG (Col 5 line 64- Col 6 line 10) “In accordance with the invention, the instantaneous watermark gain is dynamically determined to match the characteristics of the speech signal. In the simplest case, when little speech energy is present (i.e., during silence), the watermark may be added using a fixed gain threshold. This is selected so that the watermark becomes the effective noise floor of the recording. Perceptually, a small amount of noise is always expected in a recording and the watermark signal is not atypical of such recording noise. In many applications, silence may not be transmitted or might be by coded using extreme compression. In these circumstances, designers may preferably choose an error correcting code (such as a convolutional code) with the proper characteristics so that the message may be recovered despite these losses.
CHENG (Col 9 lines 34-47) “The gain calculation module 714 is used to dynamically determined the instantaneous watermark gain in order to match the characteristics of the speech signal. The operations of the gain calculation module are described in detail above in Section III. As mentioned therein, the watermark gain in each frame of the speech signal can be determined by the linear combination of the gains for silence, normalized per sample residual energy E, and normalized per sample speech energy E, as specified in equation (2). The gain calculation operation is designed to maximize the strength of the watermark signals without incurring perceptual degradations. The gain calculation module outputs a gain control signal for each frame of the speech signal in the manner described above with respect to equation (2).”
determining, by the computer, a watermark strength
CHENG (Col 9 lines 34-47) “The gain calculation module 714 is used to dynamically determined the instantaneous watermark gain in order to match the characteristics of the speech signal. The operations of the gain calculation module are described in detail above in Section III. As mentioned therein, the watermark gain in each frame of the speech signal can be determined by the linear combination of the gains for silence, normalized per sample residual energy E, and normalized per sample speech energy E, as specified in equation (2). The gain calculation operation is designed to maximize the strength of the watermark signals without incurring perceptual degradations. The gain calculation module outputs a gain control signal for each frame of the speech signal in the manner described above with respect to equation (2).”
CHENG (Col 9 lines 48-56) “The output of the vocal tract filter 708 is provided to the signal multiplier 716 along with the gain control signal generated by the gain calculation module 714. The signal multiplier adjusts the gain of the watermark signal in accordance with the gain control signal generated in accordance with the computation performed by the gain calculation module 714. Lastly, the output of the signal multiplier 716 is added to the speech signal 710 in the signal adder 718 to yield the watermarked speech signal 720.”
updating, by the computer, the strength of the watermark signal of the speech signal
CHENG (Col 10 line 56- Col 11 line 33) “As mentioned above, the watermark gain in each frame of the speech signal can be determined by the linear combination of the gains for silence, normalized per sample linear predictor residual energy E, and normalized per sample speech energy E.sub.s as specified in equation (2). As is evident from FIG. 9, the energy detector 904 and the weight factor unit 908 yield the gain contribution associated with normalized per sample speech energy E.sub.s from the speech signal 710; the residual energy predictor 906 and the weight factor unit 910 yield the gain contribution associated with the normalized per sample residual energy E from the output of the LPC analysis module 712; and the weight factor unit 912 yields a gain contribution representing silence (i.e., a fixed threshold generated by applying a unity input to the weight factor unit 912). The gain contribution outputs of all the weight factor units are then linearly combined in signal adder 914 to yield the watermark signal gain for the current frame of the speech signal. In this manner, the gain calculation operation is designed to maximize the strength of the watermark signal without incurring perceptual degradations. As noted in the description of FIG. 7, the gain control signal output by the signal adder 914 representing the current watermark signal gain is applied to the watermark signal before the watermark signal is embedded into the speech signal.”
and generating, by the computer, a revised watermarked audio signal comprising the watermark signal having the strength as updated using the watermark strength
CHENG (Col 10 line 56- Col 11 line 33) “As mentioned above, the watermark gain in each frame of the speech signal can be determined by the linear combination of the gains for silence, normalized per sample linear predictor residual energy E, and normalized per sample speech energy E.sub.s as specified in equation (2). As is evident from FIG. 9, the energy detector 904 and the weight factor unit 908 yield the gain contribution associated with normalized per sample speech energy E.sub.s from the speech signal 710; the residual energy predictor 906 and the weight factor unit 910 yield the gain contribution associated with the normalized per sample residual energy E from the output of the LPC analysis module 712; and the weight factor unit 912 yields a gain contribution representing silence (i.e., a fixed threshold generated by applying a unity input to the weight factor unit 912). The gain contribution outputs of all the weight factor units are then linearly combined in signal adder 914 to yield the watermark signal gain for the current frame of the speech signal. In this manner, the gain calculation operation is designed to maximize the strength of the watermark signal without incurring perceptual degradations. As noted in the description of FIG. 7, the gain control signal output by the signal adder 914 representing the current watermark signal gain is applied to the watermark signal before the watermark signal is embedded into the speech signal.”
CHENG does not teach, but SHARMA teaches:
Watermark strength reduction
SHARMA (Col 26 lines 52-57) “When spectral shaping models are used for shaping the spectrum of the watermark signal to appear similar to the host signal spectrum, large spectral peaks in the host signal can lead to correspondingly large spectral peaks in the watermark signal spectrum. These large peaks can adversely affect audio quality.”
Col 26 line 58- Col 27 line 4) “Audio quality can be improved by adaptively reducing the strength of such large peaks. For example, the largest frequency peak in the spectrum of an audio segment of interest is identified. A threshold is then set at say 10% of the value of this largest peak. All spectral values that are above this threshold are clipped to the threshold value. Since the value of the threshold is based on the spectrum in any given segment, the thresholding operation is adaptive. Further, the percentage at which to base the threshold can itself be adaptively set based on other statistics in the spectrum. For example if the spectrum is relatively flat (i.e., not peaky), then a higher percentage threshold can be set, thereby resulting in fewer frequency bins being clipped.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG the capability to adjust that gain to distinguish a reduction in the watermark signal explicitly, since CHENG only teaches adjusting a gain in order to have the maximum power possible for the watermark without distorting. The combination would allow that gain adjustment to explicitly be smaller than 1 therefore reducing the strength of the watermark. The benefit and motivation of such modification is discussed by SHARMA in the following portion: SHARMA (Col 26 lines 52-57) “When spectral shaping models are used for shaping the spectrum of the watermark signal to appear similar to the host signal spectrum, large spectral peaks in the host signal can lead to correspondingly large spectral peaks in the watermark signal spectrum. These large peaks can adversely affect audio quality.”
CHENG in view of SHARMA does not teach, but CHA teaches:
At the formant peak
CHA [0039] “The formant section extraction unit (120) can convert the input audio data into frequency spectrum data using the Fast Fourier Transform (FFT). The formant section extraction unit (120) can detect peak sections having power greater than a preset value through linear predictive coding (LPC) from the converted frequency spectrum data.”
CHA [0042] “Referring to Figure 2, in the frequency spectrum converted from audio data, the power (energy) has different values depending on the frequency. Here, the part of the frequency spectrum that has a peak shape is called a formant. In the frequency spectrum shown in Figure 2, three formants F1, F2, and F3 are formed in order of decreasing frequency.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG in view of SHARMA the capability to analyze the addition of the watermark using the energy at the formant peaks. The benefits and motivation of such modification is discussed by CHA in the following portion: CHA [0033] “The data processing device (100) can perform watermarking processing based on a filter that changes the frequency band of the formant section for audio data. The data processing device (100) can perform watermarking processing based on the inserted data while minimizing changes to the original sound of the audio data.”
Regarding claim 2, the rejection of claim 1 is incorporated, furthermore CHENG does not teach but SHARMA teaches:
(Original) The method according to claim 1, wherein the computer generates the revised watermarked audio signal by embedding the revised watermarked audio signal in a transform domain of the watermarked audio signal at the formant peak of the speech signal.
SHARMA [Abstract] Audio watermark encoding methods employ reversing polarity and pairwise embedding. A watermark signal is generated, mapped to pairs of embedding locations and inserted in the members of the pair with reverse polarity. The pairs of embedding locations correspond to adjacent regions or frames in time or frequency domains. A method for controlling an audio output device captures audio from the device, extracts its identity from a watermark, and modifies the operation of the device, e.g. varying watermark strength or audio loudness.
SHARMA (Col 14 line 58- col 15 line 14) “The embedder 406 takes a selected watermark type and protocol for the audio class and constructs the watermark signal of this selected type from auxiliary data As depicted in FIG. 4, the watermark type specifies a domain or “feature space” (422) in which the watermark signal is defined, along with the watermark signal structure and audio feature or features that are modified to convey the watermark. Examples of features include the amplitude or magnitude of discrete values in the feature space, such as amplitudes of discrete samples of the audio in a time domain, or magnitudes of transform domain coefficients in a transform domain of the audio signal. Additional examples of features include peaks or impulse functions (424), phase component adjustments (426), or other audio attributes, like an echo (428). From these examples, it is apparent that they can be represented in different domains. For instance, a frequency domain peak corresponds to a time domain sinusoid function. An echo corresponds to a peak in the autocorrelation domain. Phase, likewise has a representation of a time shift in the time domain, phase angle in a frequency domain. The watermark signal structure defines the structure of feature changes made to insert the watermark signal: e.g., signal patterns such as changes to insert a peak or collection of peaks, a set of amplitude changes, a collection of phase shifts or echoes, etc.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG the capability to perform the watermark embeddings in the transform domain, doing such analysis in different feature spaces allows for improvements in quality and robustness of the watermark, as taught by SHARMA. The benefit and motivation of such modification is discussed by SHARMA in the following portion: SHARMA (Col 17 lines 8-22) “While we believe that defining the watermark type from the perspective of the detector is most useful, one can see that there are other useful perspectives. Another perspective of watermark type is that of the embedder. While it is common to embed and detect a watermark in the same feature set, it is possible to represent a watermarks signal in different domains for embedding and detecting, and even different domains for processing stages within the embedding and detecting processes themselves. Indeed, as watermarking methods become more sophisticated, it is increasingly important to address watermark design in terms of many different feature spaces. In particular, optimizing watermarking for the design constraints of audio quality, watermark robustness and capacity dictate watermark design based an analysis in different feature spaces of the audio.”
Regarding claim 3, the rejection of claim 1 is incorporated, furthermore CHENG in view of SHARMA does not teach, but CHA teaches:
(Original) The method according to claim 1, wherein obtaining the watermarked audio signal includes parsing, by the computer, the watermarked audio signal into a plurality of frames,
CHA [0052] ”Fig. 4 illustrates an exemplary formant section extracted according to one embodiment of the present invention. Referring to FIG. 4, the formant section may include a plurality of frame sections (401 to 405).”
each frame having a preconfigured frame-length for speech,
CHA [0053] “For example, the total length of the formant interval may be 25 msec, and the length of each frame interval (401 to 405) may be 5 msec. Depending on the time domain window, the application rate of the filter may vary for each frame interval (401 to 405) included in the formant interval.”
and wherein the watermark signal having a frame-length is embedded by the computer at the frame of the speech signal containing the formant peak.
CHA [0051] “The extracted formant segment can be divided into multiple frame segments by a filter. For example, a formant segment of 25 msec length can be divided into multiple frame segments of 5 msec length. The watermarking processing unit (130) can determine the application ratio of a filter for each frame interval of a formant interval based on a time domain window or a frequency domain window.”
CHA [0019] “Another embodiment of the present invention may include a data processing method for processing audio data, comprising the steps of receiving audio data, extracting a formant section including a peak portion of a frequency spectrum from the received audio data, and performing watermarking processing based on a filter that changes the frequency band of the extracted formant section.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG in view of the capability to divide the portion that contains the formant of the speech signal into predetermined sections, thereby allowing controlled application of the watermarking operation across the formant section. The benefits and motivation of such modification is discussed by CHA in the following portion: CHA [0056] “The data processing device (100) according to the present invention has the effect of preventing the original sound of audio data from being unnecessarily altered and maintaining the naturalness of the voice by adjusting the degree to which a filter is applied over time.”
Regarding claim 4, the rejection of claim 1 is incorporated, furthermore CHENG does not teach but SHARMA teaches:
(Original) The method according to claim 1, wherein obtaining the watermarked audio signal having the watermark signal includes: executing, by the computer, a transform function to generate a transformed representation of a watermark-free audio signal in a transform domain;
SHARMA (Col 17 lines 23-37) “A related consideration that plays a role in watermark design is that well-developed implementations of signal transforms enable a discrete watermark signal, as well as sampled version of the host audio, to be represented in different domains. For example, time domain signals can be transformed into a variety of transform domains and back again (at least to some close approximation). These techniques, for example, allow a watermark that is detected based on analysis of frequency domain features to be embedded in the time domain. These techniques also allow sophisticated watermarks that have time, frequency and phase components. Further, the embedding and detecting of such components can include analysis of the host signal in each of these feature spaces, or in a subset of the feature space, by exploiting equivalence of the signal in different domains.”
and generating, by the computer, the watermarked audio signal comprising the watermark signal by embedding the watermark signal in the transform domain of the speech signal of the watermark-free audio signal.
SHARMA (Col 26 lines 24-30) “One approach to make the watermark signal components have the same spectral shape as the host audio is to multiply the frequency domain watermark signal components (e.g. +/− bumps or other patterns of the DWM structure as described above) with the host spectrum. The resulting signal can then be added to the host audio (either in the spectral domain or the time domain) after multiplying with a gain factor.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG the capability to perform the watermark and host audio signal combination in the transform domain, since the combination can be done either in the time domain or in the spectral domain, as taught by SHARMA. The benefit and motivation of such modification is discussed by SHARMA in the following portion: SHARMA (Col 17 lines 22 to 27) “A related consideration that plays a role in watermark design is that well-developed implementations of signal transforms enable a discrete watermark signal, as well as sampled version of the host audio, to be represented in different domains. …”
Regarding claim 5, the rejection of claim 1 is incorporated, furthermore CHENG does not teach but SHARMA teaches:
(Original) The method according to claim 4, wherein obtaining the watermarked audio signal having the watermark signal includes generating, by the computer, a watermark sequence of the watermark signal comprising one or more watermark values in the transform domain.
SHARMA (Col 19 lines 41-48) “The watermark protocol specifies signal communication techniques employed, such as a type of data modulation to encode data using a signal carrier. One such example is direct sequence spread spectrum (DSSS) where a pseudo random carrier is modulated with data. There are a variety of other types of modulation, phase modulation, phase shift keying, frequency modulation, etc. that can be applied to generate a watermark signal.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG the capability to use watermark elements mapped to the transform domain, allowing them to analyze and optimize parameters of the host signal and the watermark in any domain. The benefit and motivation of such modification is discussed by SHARMA in the following portion: SHARMA (Col 16 line 53 – Col 17 line 7) “Having introduced the concepts of watermark structure and audio features for conveying it, one can now appreciate finer aspects in watermark design and insertion methodology. The watermark structure is inserted into audio by altering audio features according to watermark signal elements that make up the structure. Watermarking algorithms are often classified in terms of signal domains, namely signal domains where the signal is embedded or detected, such as “time domain,” “frequency domain,” “transform domain,” “echo or autocorrelation” domain. For discrete audio signal processing, these signal domains are essentially a vector of audio features corresponding to units for an audio frame: e.g., audio amplitude at a discrete time values within a frame, frequency magnitude for a frequency within a frequency transform of a frame, phase for a frequency transform of a frame, echo delay pattern or auto-correlation feature within a frame, etc. For background, see watermarking types in U.S. Pat. Nos. 6,614,914 and 6,674,876, and Published Applications 20120214515 and 20120214544, which are hereby incorporated by reference. The domain of the signal is essentially a way of referring to the audio features that carry watermark signal elements, and likewise, a coordinate space of such features where one can define watermark structure.”
Regarding claim 6, the rejection of claim 5 is incorporated, furthermore CHENG does not teach but SHARMA teaches:
(Original) The method according to claim 1, wherein obtaining the watermarked audio signal having the watermark signal includes receiving, by the computer, the watermarked audio signal via one or more networks.
SHARMA (Col 31 lines 12-18) “For real time embedding applications, the evaluations of quality and robustness need to be computationally efficient and applicable to relatively small audio segments so as not to introduce latency in the transmission of the audio signal. Examples of real time operation include embedding with a payload at the point of distribution (e.g., terrestrial or satellite broadcast, or network delivery).”
SHARMA (Col 44 lines 44-51) “Ambient detection refers to detection of an audio watermark from audio captured from the ambient environment through a sensor (i.e. microphone). In addition to distortions that occur in electromagnetic wave transmission of the watermarked audio over a wire or wireless (e.g., RF signaling) transmission, the ambient audio is converted to sound waves via a loudspeaker into a space, where it can be reflected from surfaces, attenuated and mixed with background noise. ...”
SHARMA (Col 59 line 54 – Col 60 line 6) “In another embodiment, the identification information can be used to control or modify at least one attribute of the audio watermark signal output by the identified speaker(s). For example, the receiving device can be configured to directly or indirectly control or modify an attribute of the audio watermark signal output by the identified speaker(s) (e.g., similar to the manner exemplarily discussed above with respect to modification of the host audio signal). In such an example, the watermark embedder is located at the receiving device. In another example, the watermark embedder is remote from the receiving device, but is coupled to (e.g., via wired or wireless connection, either directly or indirectly via any network) or otherwise integrated into one or more of the aforementioned audio output control devices. One attribute of the audio watermark signal that may be adjusted is the strength of the watermark signal relative to the host audio signal. For example, the strength of the audio watermark signal can be adjusted (e.g., raised or lowered) to enhance ambient detection of the audio watermark signal, to reduce human perceptibility of the audio watermark signal, or the like or a combination thereof.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG the capability to receive via the network as through by SHARMA, using network-connected audio systems that stream audio and receiving devices capture and process the resulting watermark audio. The benefit and motivation of such modification is discussed by SHARMA in the following portion: SHARMA (Col 58 line 59 – Col 59 line 30) “In one aspect, the data channel provided by an audio watermark signal can be used to identify an audio output device (e.g., a loudspeaker, also referred to herein as a “speaker”) or a group or set of speakers (e.g., of the type found in public address systems, radio and television receivers, portable digital media players, smartphones, tablet computers, laptop computers, desktop computers, mobile phones, sound reinforcement systems for theaters and concerts, etc.). Generally, a speaker is configured to generate sound in response to receiving an electronic signal, wherein the sound produced corresponds to the electronic signal applied. The speaker or set may be communicatively coupled (e.g., via wired or wireless connection, either directly or indirectly via any network) to one or more audio output control devices configured to apply various electronic signals to the speaker(s), thereby controlling the manner in which audio signals are output by the speaker(s)) as sound, a watermark embedder as exemplarily described above, or any combination thereof. An exemplary audio output control device may include one or more devices such as remote servers configured to stream music or other audio information—including an audio watermark—to be output by the speaker(s), radio receivers, television receivers, portable digital media players, smartphone or other mobile phones, tablet computers, laptop computers, desktop computers, etc., each of which is generically referred to herein as a “audio output control device”). A microphone-equipped receiving device (e.g., a portable digital media player, a smartphone or other mobile phone, a tablet computer, a laptop computer, etc.) may be used to capture audio signals output by the speaker(s) and perform ambient detection on the captured audio signals (e.g., in the manner exemplarily described above). In the event that an embedded audio watermark is detected within the audio signal output by the speaker(s), the receiving device can extract from the watermark, information identifying the speaker or set thereof. As discussed in greater detail below, this identification information can then be used control or modify one or more audio signals (e.g., the host audio signal, the audio watermark signal, or both) output by the speaker(s).”
Regarding claim 7, the rejection of claim 1 is incorporated, furthermore CHENG in view of SHARMA does not teach, but CHA teaches:
(Original) The method according to claim 1, wherein determining the strength of the watermark signal at the formant peak includes identifying, by the computer, the formant peak in the watermarked audio signal, the formant peak of the speech signal of the watermarked audio signal containing a relatively higher amount of power satisfying a peak-detection threshold and indicative of the formant peak at a portion of the speech signal of the watermarked audio signal.
CHA [0039] The formant section extraction unit (120) can convert the input audio data into frequency spectrum data using the Fast Fourier Transform (FFT). The formant section extraction unit (120) can detect peak sections having power greater than a preset value through linear predictive coding (LPC) from the converted frequency spectrum data.
CHA [0042] Referring to Figure 2, in the frequency spectrum converted from audio data, the power (energy) has different values depending on the frequency. Here, the part of the frequency spectrum that has a peak shape is called a formant. In the frequency spectrum shown in Figure 2, three formants F1, F2, and F3 are formed in order of decreasing frequency.
CHA [0046] The formant segment extracted by the formant segment extraction unit (120) may correspond to a vowel utterance segment among the audio data. Generally, vowel articulation segments have more energy than consonant articulation segments. Therefore, the formants of the frequency spectrum have the characteristic of appearing more clearly in the vowel utterance intervals than in the consonant utterance intervals of the speech data.
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG in view of SHARMA the capability to identify and utilize the high-power peaks corresponding to a formant to prevent the detection and removal of the watermarking. The benefits and motivation of such modification is discussed by CHA in the following portion: CHA [0037] “In addition, the data processing device (100) can prevent the removal of watermarking through simple frequency filtering or deletion of some sections by performing watermarking processing on vowel utterance sections among voice section data.”
Regarding claim 8, the rejection of claim 7 is incorporated, furthermore CHENG in view of SHARMA does not teach, but CHA teaches:
(Original) The method according to claim 7, wherein the computer executes a Linear Predictive Coding (LPC) analysis for identifying the formant peak in a transform domain of the watermarked audio signal of the speech signal.
CHA [0039] The formant section extraction unit (120) can convert the input audio data into frequency spectrum data using the Fast Fourier Transform (FFT). The formant section extraction unit (120) can detect peak sections having power greater than a preset value through linear predictive coding (LPC) from the converted frequency spectrum data.
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG in view of SHARMA the capability to utilize FFT and LPC in the watermarking process to identify formant peaks and efficiently apply the watermarks to the speech signal. The benefits and motivation of such modification is discussed by CHA in the following portion: CHA [0047] “The data processing device (100) according to the present invention can efficiently perform watermarking processing by using a formant formed in a vowel phonation section. By doing so, it is possible to prevent the watermark from being removed by simple frequency filtering or the deletion of some sections.”
Regarding claim 9, the rejection of claim 1 is incorporated, furthermore CHENG teaches:
(Original) The method according to claim 1, wherein the computer determines the watermark strength reduction weighting based upon a preconfigured penalty parameter and the amount of power of the speech signal
CHENG (Col 6 lines 18-21) “The watermark gain in each frame can be determined by the linear combination of the gains for silence, normalized per sample residual energy E, and normalized per sample speech energy E.sub.s :”
CHENG (Col 6 line 23) “g(t)=.lambda..sub.0 +.lambda..sub.1 E+.lambda..sub.2 E.sub.s (2)”
CHENG (Col 6 lines 24-38) “which is designed to maximize the strength of the watermark signals without incurring perceptual degradations. It is to be appreciated that the parameters .lambda..sub.0, .lambda..sub.1 and .lambda..sub.2 are empirically chosen parameters that serve to trade off noise versus watermark signal strength. The designer may choose these parameters depending on the particular application. FIG. 3A shows a segment of speech and FIG. 3B shows the resulting watermarked speech. A listening test demonstrates that the watermarked speech is indistinguishable from the original speech with this watermark gain. If the gain is increased further, there may be "hoarseness" in the watermarked speech. Though it hardly affects the naturalness of the voice, the difference with the original speech may indeed be perceptible.”
CHENG in view of SHARMA does not teach, but CHA teaches:
At the formant peak
CHA [0039] “The formant section extraction unit (120) can convert the input audio data into frequency spectrum data using the Fast Fourier Transform (FFT). The formant section extraction unit (120) can detect peak sections having power greater than a preset value through linear predictive coding (LPC) from the converted frequency spectrum data.”
CHA [0042] “Referring to Figure 2, in the frequency spectrum converted from audio data, the power (energy) has different values depending on the frequency. Here, the part of the frequency spectrum that has a peak shape is called a formant. In the frequency spectrum shown in Figure 2, three formants F1, F2, and F3 are formed in order of decreasing frequency.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG in view of SHARMA the capability to analyze the addition of the watermark using the energy at the formant peaks. The benefits and motivation of such modification is discussed by CHA in the following portion: CHA [0033] “The data processing device (100) can perform watermarking processing based on a filter that changes the frequency band of the formant section for audio data. The data processing device (100) can perform watermarking processing based on the inserted data while minimizing changes to the original sound of the audio data.”
Regarding claim 10, the rejection of claim 1 is incorporated, furthermore CHENG does not teach but SHARMA teaches:
(Original) The method according to claim 1, further comprising executing, by the computer, a transform function on the revised watermarked audio signal in a transform domain to generate an audible representation of the revised watermarked audio signal in a time domain.
SHARMA (Col 17 lines 23-37) “A related consideration that plays a role in watermark design is that well-developed implementations of signal transforms enable a discrete watermark signal, as well as sampled version of the host audio, to be represented in different domains. For example, time domain signals can be transformed into a variety of transform domains and back again (at least to some close approximation). These techniques, for example, allow a watermark that is detected based on analysis of frequency domain features to be embedded in the time domain. These techniques also allow sophisticated watermarks that have time, frequency and phase components. Further, the embedding and detecting of such components can include analysis of the host signal in each of these feature spaces, or in a subset of the feature space, by exploiting equivalence of the signal in different domains.”
SHARMA (Col 26 lines 24-30) “One approach to make the watermark signal components have the same spectral shape as the host audio is to multiply the frequency domain watermark signal components (e.g. +/− bumps or other patterns of the DWM structure as described above) with the host spectrum. The resulting signal can then be added to the host audio (either in the spectral domain or the time domain) after multiplying with a gain factor. “
SHARMA (Col 37 lines 44-56) “The perceptual adaptation module 808 is a software function that transforms the watermark signal elements to changes to corresponding features of the host audio segment according to the perceptual masking envelope. The envelope specifies limits on a change in terms of magnitude, time and frequency dimensions. Perceptual adaptation takes into account these limits, the value of the watermark element, and host feature values to compute a detail gain factor that adjust watermark signal strength for a watermark signal element (e.g., a bump) while staying within the envelope. A global gain factor may also be used to scale the energy up or down, e.g., depending on feedback from iterative embedding, or user adjustable watermark settings.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of CHENG the capability to transform the modified signal between the transform domain and the time domain, such transformation would allow the spectrally processed watermarked audio to be mapped to a time domain signal for audio output. The benefit and motivation of such modification is discussed by SHARMA in the following portion: SHARMA (Col 37 lines 57-67) “Insertion function 810 makes the changes to embed a watermark signal element determined by perceptual adaptation. These can be a combination of changes in multiple domains (e.g., time and frequency). Equivalent changes from one domain can be transformed to another domain, where they are combined and applied to the host signal. An example is where parameters for frequency domain based feature masking are computed in the frequency domain and converted to the time domain for application of additional temporal masking (e.g., removal of pre-echoes) and insertion of a time domain change.”
Regarding claim 21, arguments analogous to claim 1 are applicable. Furthermore, CHENG teaches:
(New) A system for embedding watermarks in audio signals, the system comprising:
CHENG (Col 8 lines 18-28) “Referring now to FIG. 7, a block diagram illustrating a speech watermarking system according to an embodiment of the invention is shown. Generally, the system 700 inputs a digital message 702 and a speech signal 710 and embeds the digital message into the speech signal, as explained above and as will be further described below, to yield a watermarked speech signal 720. As shown in FIG. 7, the system 700 includes an error control coder 704, a spread spectrum modulator 706, a vocal tract filter 708, an LPC analysis module 712, a gain calculation module 714, a signal multiplier 716 and a signal adder 718.”
Regarding claim 22, the rejection of claim 21 is incorporated, furthermore arguments analogous to claim 2 are applicable.
Regarding claim 23, the rejection of claim 21 is incorporated, furthermore arguments analogous to claim 3 are applicable.
Regarding claim 24, the rejection of claim 21 is incorporated, furthermore arguments analogous to claim 4 are applicable.
Regarding claim 25, the rejection of claim 24 is incorporated, furthermore arguments analogous to claim 5 are applicable.
Regarding claim 26, the rejection of claim 21 is incorporated, furthermore arguments analogous to claim 6 are applicable.
Regarding claim 27, the rejection of claim 21 is incorporated, furthermore arguments analogous to claim 7 are applicable.
Regarding claim 28, the rejection of claim 27 is incorporated, furthermore arguments analogous to claim 8 are applicable.
Regarding claim 29, the rejection of claim 21 is incorporated, furthermore arguments analogous to claim 9 are applicable.
Regarding claim 30, the rejection of claim 21 is incorporated, furthermore arguments analogous to claim 10 are applicable.
Summary of References used as prior art
CHENG Is used as a main reference, teaches speech watermarking, Linear predictive coding and gain control of the watermark signal to maintain the watermark imperceptible and not create any noticeable distortion.
SHARMA is used as a secondary reference because it provides a great discussion about spectral domain watermarking, as well as a discussion about sequences and audio frames, and the transform processing for audio signals
CHA is used as an additional reference because it provides an explicit explanation of formant peaks and the identification of the peaks in a signal. Also provides a power analysis framework, with another LPC framework as well.
Pertinent art not cited in this Office Action
Agaskar; Ameya et al. (US 12136428 B1) is considered pertinent for teaching frame-based audio watermark embedding, for scaling factors to control watermark amplitude and the scaling based on spectral energy measurement.
FAUBEL; Friedrich et al. (US 20240038249 A1) is considered pertinent for teaching transform-domain speech watermarking, fixed-length frames and watermark sequences embedded into the speech signal.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HECTOR J. CRESPO FEBLES whose telephone number is (571)272-4512 and email is hcrespofebles@uspto.gov. The examiner can normally be reached Mon - Fri 7:30 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HECTOR J. CRESPO FEBLES/Examiner, Art Unit 2657
/DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657