DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 12 is objected to because of the following informalities:
Claim 12 recites the limitation "enabling a system to neutralize the voices …". The claim is a method claim and introduces a step of enabling a system to neutralize the voices within the claim, but it isn’t clear as to which system the claim is referring to, such as a system that comprises the processing unit described in claim 1, or any other possible system. A recommended correction would be “enabling the processing unit”.
Note: claims 13 and 14 appear to reference the system introduced in claim 12, so, based on the amendment to claim 12, they would likely require amendments as well.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 4 to 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “common speech anomalies” in claim 4, 15 and 16 is a relative term which renders the claim indefinite. The term “common speech anomalies” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. No objective boundary has been provided to what qualifies as a common anomaly in the speech, it is not clear in the claims or the specification what is included, or excluded, as a common speech anomaly.
The term “efficacy” in claim 8 is a relative term which renders the claim indefinite. The term “efficacy” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. It is not clear based on the language of the claim what are the limits in “maintaining efficacy against unauthorized devices”, specification does not disclose if it means that a certain attenuation needs to be archived, a measurable loss in quality or a rate of failure for a speech recognition models needs to be achieved.
Claim 12 recites the limitation "enabling a system to neutralize the voices of multiple target speakers…". There is insufficient antecedent basis for this limitation in the claim.
Claim 15 and 16 recites the limitation "the target speaker’s voice”. There is insufficient antecedent basis for this limitation in the claim. Claim 15 and 16 are independent claims and the term target speaker voice was not introduced before the claim mentioned “the target speaker’s voice”. A possible correction is to change the initial mention of “the target speaker’s voice” to “a target speaker’s voice” with subsequent references using “the target speaker’s voice”.
The term “perceptible delay” in claim 15 is a relative term which renders the claim indefinite. The term “perceptible delay” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. It is not clear if that perceptible delay would have to be perceptible to the listener, the speaker or a recording device. Furthermore, a quantitative description of a perceptible delay for a noise cancelling signal is not provided in the claims or the specification.
Regarding claims 5 to 7, 9 to 14 and 16 to 20, they inherit and do not correct the indefinite issues from the claims from which they depend, therefore rendering them indefinite.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1 is/are rejected under 35 U.S.C. 103 as being unpatentable over Alava; Galo M. et al. (US 12424192 B1), hereinafter ALAVA, in view of Narayanan; Ganesh et al. (US 20220417659 A1), hereinafter NARAYANAN, in view of Liu; Jing et al. (US 12531056 B1), hereinafter LIU, in further view of SANTOS; Alexandre (US 20190147902 A1), hereinafter SANTOS.
Regarding claim 1, ALAVA teaches:
A method of preventing unauthorized audio capture by neutralizing electronic personal digital assistants, comprising: " capturing audio signals from an environment using a microphone configured to detect ambient audio within a predefined range;
ALAVA (Col 3 lines 25 to 34): “In certain embodiments, the audio protection system may include a microphone that detects and/or receives the communication of the user. For example, the microphone may be in a communication device of the user (e.g., a mobile device, a desktop computer, a microphone communicatively coupled to the mobile device or desktop computer). The audio protection system may also include an audio modification mechanism (e.g., a device, a speaker, a filter) that provides the generated audio data in the second portion of the control area.”
generating, using an audio signal generator operably connected to the processing unit, a counteracting audio signal that is 180 degrees out of phase with the speech of the target speaker, the counteracting audio signal being configured to neutralize the target speaker's voice at point of detection by the electronic personal digital assistants;
ALAVA (Col 6 lines 16 to 36): “In response to determining that the communication of the user 104 includes confidential information, the audio modification system 102 may determine generated audio data (e.g., generated sound; active noise-cancelling audio data or voice-cancelling audio data) that renders the communication at least partially inaudible in the second portion 132 of the control area 106. The audio modification system 102 may determine the generated audio data based on parameters, such as a frequency and/or an amplitude, of sound waves of the communication of the user 104. For example, the generated audio data may include another frequency and/or another amplitude of sound waves that is generally opposite or otherwise configured to cancel or at least partially cancel the sound waves of the communication of the user 104. More specifically, the frequency and amplitude of the generated audio data may be 180 degrees out of phase relative to the frequency and amplitude of the communication of the user 104 (e.g., a phase shift of 180 degrees). In other embodiments, a phase shift between the communication of the user 104 and the generated audio data may be between 0 degrees and 180 degrees.”
ALAVA (Col 14 lines 1 to 22): “Accordingly, the audio protection systems 100 and 300 described herein may affect the audio of the user 104 in the control areas 106 and 302, respectively, such that other people cannot clearly hear the communication of the users and/or recording devices cannot clearly record the communication of the user. For example, the audio modification system 102 may receive audio data indicative of the communication of the user 104 and determine generated audio data that renders the communication of the user 104 at least partially inaudible, such that the other people and/or the recording devices cannot determine the confidential information in the communication of the user. The audio modification system 102 may automatically and/or iteratively perform the process, thereby enabling efficient protection of the user's communication in real-time (e.g., within seconds of detecting the user's voice and/or the confidential nature of the user's communication). Accordingly, the audio protection systems 100 and 300 may render the communication of the user at least partially inaudible, or otherwise modify the communication of the user, in the control area to protect the user 104 and/or others affected by the communication of the user 104.”
synchronizing, by the processing unit, timing of the counteracting audio signal with a target speaker's speech cadence to ensure real-time neutralization without perceptible delay;
ALAVA (Col 14 lines 1 to 22): “Accordingly, the audio protection systems 100 and 300 described herein may affect the audio of the user 104 in the control areas 106 and 302, respectively, such that other people cannot clearly hear the communication of the users and/or recording devices cannot clearly record the communication of the user. For example, the audio modification system 102 may receive audio data indicative of the communication of the user 104 and determine generated audio data that renders the communication of the user 104 at least partially inaudible, such that the other people and/or the recording devices cannot determine the confidential information in the communication of the user. The audio modification system 102 may automatically and/or iteratively perform the process, thereby enabling efficient protection of the user's communication in real-time (e.g., within seconds of detecting the user's voice and/or the confidential nature of the user's communication). Accordingly, the audio protection systems 100 and 300 may render the communication of the user at least partially inaudible, or otherwise modify the communication of the user, in the control area to protect the user 104 and/or others affected by the communication of the user 104.”
adapting, by the processing unit, the counteracting audio signal dynamically in response to changes in a target speaker's voice characteristics or environmental audio conditions, including variations in tone, volume, and background noise;
ALAVA (Col 5 line 59 to Col 6 line 3): “… In some embodiments, the audio modification system 102 may determine that the communication of the user 104 is confidential based on a tone, a volume, a pitch, and/or an inflection in the communication of the user 104. For example, if the communication becomes quieter (e.g., if a change in volume of the communication exceeds a threshold volume change) and/or if the pitch changes dramatically (e.g., a change in the pitch of the communication exceeds a threshold pitch change), the audio modification system 102 may determine that the communication is more likely to include confidential information or includes confidential information.”
ALAVA does not teach but NARAYANAN teaches:
analyzing the captured audio signals using a processing unit to isolate speech patterns associated with a target speaker based on unique vocal characteristics, including pitch, tone, cadence, and frequency patterns;
NARAYANAN [0075]: “At step 506, a voice profile is determined based on the first spoken audio content. The voice profile may be additionally or alternatively determined based on the first one or more words. The voice profile may be associated with a speaker of the first one or more words. The voice profile may be generally representative of the speaker's speech such that the voice profile may be used to generate spoken audio content. The voice profile may indicate various characteristics of the speaker's speech, such as audio, vocal, and/or linguistic characteristics. Example audio, vocal, and/or linguistic characteristics may include audio spectral patterns, vocal frequency, vocal pitch, speaking speed, intonation, loudness, amplitude, speech patterns, speech cadence, the number and/or duration of utterances, the number and/or duration of breaks between utterances, etc. Accordingly, the voice profile may be determined based on similar characteristics of the spoken audio content and/or the first one or more words spoken by the speaker.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA the capability to identify a speaker by their voice characteristics from a spoken audio content. The benefit and motivation of such modification is discussed by NARAYANAN in the following portion: NARAYANAN [0052] “Conversely, the speech recognized in a portion of the input audio content 402 by the ASR 414 may be used by the voice profile module 418 to aid in determining a voice profile associated with a speaker in the portion of the input audio content 402. For example, the speech recognized in the portion may reveal audio, vocal, and linguistic characteristics relating to word choices/preferences, phrase construction, word patterns, prevalence of filler words (“um” or “uh”), etc. These characteristics may be used to select a voice profile from a plurality of stored voice profiles, determine a new voice profile, or update an existing voice profile.”
and storing, in a local memory accessible to the processing unit, data representing the target speaker's voice characteristics for use in subsequent operations of the method, wherein the stored data facilitates faster identification and neutralization of the target speaker's voice in future interactions.
NARAYANAN [0024]: “The voice profile module 106 may store a plurality of voice profiles associated with various speakers (some of whom may or may not speak in the particular content). In this case, determining the voice profile associated with the speaker may comprise selecting the voice profile, out of the plurality of voice profiles, that is associated with the speaker. The voice profile module 106 may additionally or alternatively generate a voice profile associated with the speaker based on spoken audio content associated with (e.g., spoken by) the speaker. The spoken audio content used to determine the speaker's voice profile may be from the instant content and/or other content in which the speaker speaks. The spoken audio content used to determine the speaker's voice profile may be from sample recordings of the speaker made for the purpose of determining the speaker's voice profile.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA the capability to store a profile of the speaker based on their speech characteristics. The benefit and motivation of such modification is discussed by NARAYANAN in the following portion: NARAYANAN [0023]: “... A voice profile may be determined that is associated with the speaker of the incorrect word. For example, the voice profile module 106 may determine the voice profile based on the spoken audio content in the content portion that indicates the incorrect word (as opposed to background audio content in the content portion). The voice profile may be determined based on audio, vocal, and/or linguistic characteristics of the spoken audio content, such as audio spectral patterns, vocal frequency, vocal pitch, speaking speed, intonation, loudness, amplitude, speech patterns, speech cadence, the number and/or duration of utterances, the number and/or duration of breaks between utterances, etc. As such, the voice profile may characterize the speaker's speech according to these same (at least in part) audio, vocal, and/or linguistic characteristics. The voice profile may be used to generate spoken audio content with speech that resembles the speaker's other speech in the content. That is, spoken audio content generated based on a speaker's voice profile may include speech that sounds as it were actually spoken by the speaker.”
ALAVA in view of NARAYANAN does not teach but LIU teaches:
predicting, by a machine learning model implemented in the processing unit, subsequent words or phrases likely to be spoken by the target speaker, the prediction being based on the speech patterns and contextual data;
LIU (Col 10 line 60 to Col 11 line 6): “The LM 160 may process the prompt 230 and generate the LM output 162 which may include words 235. Based on the prompt 230 instructing the LM 160 to determine relevant words/context data, the LM output 162 may include the words 235 determined by the LM 160 as being relevant. The LM 160 may determine the words 235 as being relevant for performing speech recognition for the current user input (i.e. the audio 107). In other words, the LM 160 may determine/predict, based on the information in the prompt 230, a subsequent/future user input or the words 235 that may be included in the subsequent user input. For example, given the words 232 and/or topic 234, the LM 160 may predict words likely to be included in a user input, which can be based on a dialog topic, the rare/unique words, etc.”
LIU (Col 16 line 10 to 17): “In some embodiments, the system may include a machine learning model(s) other than a LM or in addition to a LM. Such machine learning model(s) may receive text and/or other types of data as inputs, and may output text and/or other types of data (e.g., context data relevant for performing ASR processing). Such model(s) may be neural network based models, deep learning models, classifier models, autoregressive models, seq2seq models, etc.
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN the capability to predict the subsequent word from a target speaker. The benefit and motivation of such modification is discussed by LIU in the following portion: LIU (Col 16 line 31 to 44): “In some embodiments, the LM 160 may be fine-tuned to perform relevant context determination. Fine-tuning of the LM 160 may be performed using one or more techniques. One example fine-tuning technique is transfer learning that involves reusing a pre-trained model's weights and architecture for a new task. The pre-trained model may be trained on a large, general dataset, and the transfer learning approach allows for efficient and effective adaptation to specific tasks. Another example fine-tuning technique is sequential fine-tuning where a pre-trained model is fine-tuned on multiple related tasks sequentially. This allows the model to learn more nuanced and complex language patterns across different tasks, leading to better generalization and performance.”
ALAVA in view of NARAYANAN in view of LIU does not teach but SANTOS teaches:
frequency-shifting, by the processing unit, the counteracting audio signal into an inaudible range , such that the signal is inaudible to humans but detectable by unauthorized audio capture devices;
SANTOS [0021] “In an embodiment, the randomly mixed voice recording is emitted at an audible level. In an embodiment, the randomly mixed voice recording is emitted at an inaudible level. In an embodiment, the randomly mixed voice recording is emitted at both an audible level and an inaudible level. In an embodiment, the inaudible level is selected from the group consisting of an infrasound and an ultrasound.”
SANTOS [0064] “The resulting spliced mixed recoding 34 contains the voice of the subject in mixed gibberish. As the mixed gibberish is emitted during the subject's speaking the audio surveillance recorder will record both the sound of the subject's voice in tandem with the gibberish thereby masking the subject's voice. The foregoing prevents the intelligible recording of the subject's voice. Moreover, if this unintelligible surveillance recording is analyzed in order to peel the gibberish masking from the voice recording of the subject, it will be difficult to distinguish between a layer of a gibberish segment and a layer of an intelligible voice portion of the subject's actual speech.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN in view of LIU the capability to include an audio disruption in the inaudible range. The benefit and motivation of such modification is discussed by SANTOS in the following portion: SANTOS [0064] “The resulting spliced mixed recoding 34 contains the voice of the subject in mixed gibberish. As the mixed gibberish is emitted during the subject's speaking the audio surveillance recorder will record both the sound of the subject's voice in tandem with the gibberish thereby masking the subject's voice. The foregoing prevents the intelligible recording of the subject's voice. Moreover, if this unintelligible surveillance recording is analyzed in order to peel the gibberish masking from the voice recording of the subject, it will be difficult to distinguish between a layer of a gibberish segment and a layer of an intelligible voice portion of the subject's actual speech.”
Claim(s) 2-6 is/are rejected under 35 U.S.C. 103 as being unpatentable over ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in further view of McElveen; James Keith et al. (US 20210092548 A1) hereinafter MCELVEEN.
Regarding claim 2, the rejection of claim 1 is incorporated, furthermore ALAVA in view of NARAYANAN in view of LIU in view of SANTOS does not teach but MCELVEEN teaches:
The method of claim 1, wherein the microphone is configured to capture audio signals within a 360-degree range, employing multiple directional sensors to detect the target speaker's voice regardless of their position relative to the microphone or presence of physical barriers.
MCELVEEN [0049] “Following below are more detailed descriptions of various concepts related to, and embodiments of, inventive methods, devices, systems and non-transitory computer-readable media having instructions stored thereon to enable one or more said systems, devices and methods for receiving an audio data input associated with an acoustic location; processing the audio data according to a linear framework configured to define one or more boundary conditions for the acoustic location to generate an acoustic propagation model; processing the audio data to determine at least one spatial or spectral characteristic of the audio data; identifying a three-dimensional spatial location corresponding to the at least one spatial or spectral characteristic, the three-dimensional spatial location defining a point source within the acoustic location; processing the audio data according to the acoustic propagation model to extract a subject audio signal associated with the point source; processing the audio data to suppress audio signals that are not associated with the point source; and rendering a digital audio output comprising the subject audio signal.”
MCELVEEN [0069] “Embodiments of the present disclosure are configured to accommodate for suboptimal acoustic propagation environments (e.g., large reflective surfaces, objects located between the target acoustic location and the transducers that interfere with the free-space propagation, and the like) by processing audio input data according to a data processing framework in which one or more boundary conditions are estimates within a Green's Function algorithm to derive an acoustic propagation model for a target acoustic location.”
MCELVEEN [0084] “Multiple channels of audio can be combined to create patterns of constructive and destructive interference across the frequency band of interest that will discriminate between sound waves arriving from different directions. …”
Wherein the 360-degree range microphone detection of the speaker is mapped to a microphone array that operates in a three-dimensional spatial location and has the capability to determine the location of the speaker.
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN in view of LIU in view of SANTOS the capability to record all around the surrounding area of the microphone and the capability to locate a target speaker within the range of the microphone. The benefit and motivation of such modification is discussed by MCELVEEN in the following portion: MCELVEEN [0026] “Further objects of the present disclosure provide for a spatial audio processing system to overcome deficiencies associated with prior art multi-channel techniques, such as beamforming and signal separation, that fail to accommodate real-world acoustic conditions, such as large reflective surfaces, inanimate and animate objects situated or moving in-between the target acoustic location and the transducers, and other factors that interfere with the ideal, free-space propagation of acoustics.”
Regarding claim 3, the rejection of claim 2 is incorporated, furthermore ALAVA in view of NARAYANAN in view of LIU in view of SANTOS does not teach but MCELVEEN teaches:
The method of claim 2, wherein the processing unit employs a Fast Fourier Transform (FFT) to analyze frequency components of the captured audio signals, using a frequency-domain filtering technique to isolate the target speaker's voice from overlapping background noise, echoes, and ambient sounds.
MCELVEEN[0031]: ”In accordance with certain aspects of the present disclosure, the at least one transform function is selected from the group consisting of Fourier transform, Fast Fourier transform, Short Time Fourier transform and modulated complex lapped transform. In certain embodiments, the one or more spatial audio processing operations may further comprise applying a spectral subtraction noise reduction filter to the at least one separated audio output signal. ...”
MCELVEEN [0058] “As used herein the term “noise” refers to anything that interferes with the intelligibility of a signal, including but not limited to background noise, competing speech, non-speech acoustic events, resonance reverberation (of both target speech and other sounds), and/or echo.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN in view of LIU in view of SANTOS the capability to employ FFT in order to identify target speaker and isolate their speech from the ambient sounds or other noise. The benefit and motivation of such modification is discussed by MCELVEEN in the following portion: MCELVEEN [0018] “An object of the present disclosure is to provide for a spatial audio processing system that provides for significant separation of target sources even when there are fewer microphones than real and virtual (i.e., reflected images of) noise sources (i.e., the under-determined mathematically case). In accordance with certain embodiments, the spatial audio processing system provides enhancement to target sounds emanating from a point source and reduction of non-desired sounds emanating from elsewhere than the targeted point source location, rather than filtering an audio input solely based on a sound wave's direction of arrival (i.e., along or within a “beam” as a conventional beamformer does). In accordance with certain embodiments, the system may provide for 15 dB (decibels) or more of additional signal-to-noise ratio (SNR) improvement compared to prior art beamforming techniques, while using far fewer transducers.”
Regarding claim 4, the rejection of claim 3 is incorporated, furthermore ALAVA does not teach but NARAYANAN teaches:
The method of claim 3, wherein the machine learning model is pre-trained on a dataset comprising varied linguistic patterns, regional accents, and common speech anomalies, enabling accurate prediction of subsequent words or phrases spoken by the target speaker under diverse conversational contexts.
NARAYANAN [0043] “Based on the input audio content 402, such as the spoken audio content (speech) thereof, a voice profile module 418 (e.g., the voice profile module 106 of FIG. 1) may determine a voice profile associated with a speaker in the input audio content 402 or portion thereof (e.g. associated with the spoken audio content from the speaker). The voice profile may be determined based on audio, vocal, and/or linguistic characteristics of the spoken audio content, such as audio spectral patterns, vocal frequency, vocal pitch, speaking speed, intonation, loudness, amplitude, speech patterns, speech cadence, the number and/or duration of utterances, the number and/or duration of breaks between utterances, etc. The voice profile may be preferably representative of the speaker's voice and mode of speech such that the voice profile may be used to generate spoken audio content that imitates the speaker's actual speech (or a best approximation thereof).”
NARAYANAN [0079] “A machine learning algorithm (e.g., a machine learning model or other machine learning techniques) may be used to determine (e.g., update) a voice profile (e.g., the voice profile associated with the speaker in the first spoken audio content). The feedback loop may comprise the machine learning algorithm. The feedback loop may be trained generally with spoken audio content associated with a speaker (e.g., any spoken audio content in any content and associated with any speaker). The feedback loop may additionally or alternatively comprise a real-time feedback loop trained using the first spoken audio content and/or other spoken audio content in other portions of the content that is associated with the speaker. In the case of a machine learning algorithm, which may be determined based on training data, an input component of training data may comprise spoken audio content associated with a speaker and a corresponding output component of the training data may comprise the determined (e.g., by the machine learning algorithm) voice profile associated with the speaker. The output component of the training data may additionally or alternatively include performance metrics associated with the determined voice profile, such as whether the voice profile is an accurate representation of the speaker's speech and/or whether the transitions in the corrected output content between corrected spoken audio content determined based on the voice profile and original spoken audio content are smooth and most likely inconspicuous to viewers.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA the capability to use a machine learning algorithm and train it on the target speaker voice characteristics. The benefit and motivation of such modification is discussed by NARAYANAN in the following portion: NARAYANAN [0055] “... As another example, a television news program may wish to normalize the accents spoken by the various news casters and reporters, in which case a word pronounced with a non-standard accent may be considered an incorrect word (with the same word pronounced with the normalized accent being considered the associated correct word). ...”
Regarding claim 5, the rejection of claim 4 is incorporated, furthermore ALAVA in view of NARAYANAN does not teach but LIU teaches:
The method of claim 4, wherein the machine learning model further incorporates contextual analysis by identifying semantic and syntactic relationships within captured speech, leveraging a natural language processing (NLP) framework to refine the prediction of likely future phrases.
LIU (Col 21 line 4 to 11): “The speech processing system 692 may further include a NLU component 660. The NLU component 660 may receive the ASR data from the ASR component 150. The NLU component 660 may attempts to make a semantic interpretation of the phrase(s) or statement(s) represented in the text data input therein by determining one or more meanings associated with the phrase(s) or statement(s) represented in the text data. ...”
LIU (Col 16 line 10 to 17): “In some embodiments, the system may include a machine learning model(s) other than a LM or in addition to a LM. Such machine learning model(s) may receive text and/or other types of data as inputs, and may output text and/or other types of data (e.g., context data relevant for performing ASR processing). Such model(s) may be neural network based models, deep learning models, classifier models, autoregressive models, seq2seq models, etc.”
LIU (Col 13 line 16 to 21): “At a step 310, of the process 300, the LM 160 may process the prompt to determine the relevant context data, for example, the LM output 162, including a second plurality of words, for example, the words 235. The LM output 162 may include the words 235 that may be relevant for recognizing the words in the spoken input. The LM 160 may predict the words 235 as being the likely words to be included in the spoken input. ...”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN the capability to include Natural Language Processing/Understanding to fine tune their word prediction. The benefit and motivation of such modification is discussed by LIU in the following portion:
LIU (Col 3 line 13 to 25): “In some embodiments, the system uses, among other things, interaction history information (e.g., past user inputs, past system actions, time/date of past interactions, visual content, context data related to the past interaction, etc.), user preference data, user profile data (e.g., personalized contacts, personalized device names, location, etc.), dialog data (e.g., past inputs of a current dialog), etc. to prompt the LM to generate words (e.g., text, tokens, etc.) that are contextually relevant for performing ASR for a future user input(s). The contextual information generated by the LM is then processed using an attention component (e.g., a multi-head attention model) to provide relevant features to the ASR model (e.g., a joiner network of the ASR model).”
Regarding claim 6, the rejection of claim 5 is incorporated, furthermore ALAVA teaches:
The method of claim 5, wherein the processing unit generates the counteracting audio signal by applying an adaptive phase-inversion algorithm that accounts for fluctuations in pitch, amplitude, and inflection in the target speaker's vocal characteristics, ensuring effective cancellation even during dynamic speech patterns.
ALAVA (Col 6 line 16 to 36): “In response to determining that the communication of the user 104 includes confidential information, the audio modification system 102 may determine generated audio data (e.g., generated sound; active noise-cancelling audio data or voice-cancelling audio data) that renders the communication at least partially inaudible in the second portion 132 of the control area 106. The audio modification system 102 may determine the generated audio data based on parameters, such as a frequency and/or an amplitude, of sound waves of the communication of the user 104. For example, the generated audio data may include another frequency and/or another amplitude of sound waves that is generally opposite or otherwise configured to cancel or at least partially cancel the sound waves of the communication of the user 104. More specifically, the frequency and amplitude of the generated audio data may be 180 degrees out of phase relative to the frequency and amplitude of the communication of the user 104 (e.g., a phase shift of 180 degrees). In other embodiments, a phase shift between the communication of the user 104 and the generated audio data may be between 0 degrees and 180 degrees.”
ALAVA (Col 5 line 59 to Col 6 line 3): “... In some embodiments, the audio modification system 102 may determine that the communication of the user 104 is confidential based on a tone, a volume, a pitch, and/or an inflection in the communication of the user 104. For example, if the communication becomes quieter (e.g., if a change in volume of the communication exceeds a threshold volume change) and/or if the pitch changes dramatically (e.g., a change in the pitch of the communication exceeds a threshold pitch change), the audio modification system 102 may determine that the communication is more likely to include confidential information or includes confidential information.”
Claim(s) 7-8 is/are rejected under 35 U.S.C. 103 as being unpatentable over ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in view of MCELVEEN in further view of York; Stanley J. et al. (US 6654467 B1) hereinafter YORK.
Regarding claim 7, the rejection of claim 6 is incorporated, furthermore ALAVA teaches:
target speaker vocal range
ALAVA (Col 6 line 16 to 36): “In response to determining that the communication of the user 104 includes confidential information, the audio modification system 102 may determine generated audio data (e.g., generated sound; active noise-cancelling audio data or voice-cancelling audio data) that renders the communication at least partially inaudible in the second portion 132 of the control area 106. The audio modification system 102 may determine the generated audio data based on parameters, such as a frequency and/or an amplitude, of sound waves of the communication of the user 104.”
ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in further view of MCELVEEN does not teach but YORK teaches:
The method of claim 6, wherein the counteracting audio signal is generated in multiple overlapping frequency bands, each tailored to neutralize distinct components of
YORK (Col 5 lines 19 to 31) “From the output stage of the preamplifier 22 the noise signal passes to the digital filter section 23 where the noise signal is analyzed in the frequency domain. The digital filters 23 are clock driven, and an adjustable clock determines the frequency at which the filters reach cutoff. There are a number of filters in this section, each one tuned to the fundamental frequency, or a harmonic of the fundamental frequency in question. In order for the circuit to respond favorably with a wide frequency range, it is necessary for the harmonics to be separated this way. For a squarewave, the harmonic content can be determined by taking the sum of the fundamental frequency and all the harmonics or overtones. …”
YORK (Col 5 lines 38 to 44) “Complex periodic waveforms associated with noise have some attributes of all three of these waveforms, and can be broken down into the fundamental frequency and harmonics in the same way. The anti-noise device described herein contains a number of active filters that can be individually tuned to the fundamental frequency and any harmonic of a frequency that is a significant contributor to the noise. …”
YORK (Col 6 lines 5 to 16) “The signals are then passed on to the mixer circuit 26 that takes the individual signals and mixes them to become one signal before passing it to the power amplifier 27, where it is boosted to a level where it can drive the loudspeaker 28 which yields an output noise which is shifted one-half cycle to the input noise. The level of the output from the loudspeaker 28 is now adjusted to give the same acoustic level as the noise coming from the noise source to produce a nullity where the noise and anti-noise signal wavefronts meet. Feedback 29 from the power amplifier 27 to the preamplifier 22 prevents the device from trying to nullify its own output.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in further view of MCELVEEN the capability to modify the counteracting audio signal so it generates in multiple frequency bands, in order to neutralize components of incoming sound including fundamental frequencies and harmonics, further including the target speaker’s voice as identified in ALAVA. The benefit and motivation of such modification is discussed by YORK in the following portion: YORK (Col 3 lines 28 to 39) “The present invention accomplishes the above and other objects by providing an apparatus that cancels ambient noise by having an input sensor means, an input amplifier, means for analyzing and dividing the signal, means for introducing a half cycle phase delay, means for recombining the signals into one signal with an output amplifier for passing the signal to an output loudspeaker to effectively cancel the ambient noise. The sensor means picks up ambient noise and converts it into an electrical input signal containing amplitude and temporal information corresponding to the frequency wave of the ambient noise. The input signal is then fed through an input amplifier to a series of digital filters which divide the signal into a fundamental signal and a series of separate harmonic signals of different frequency ranges. Then each of the harmonic signals is fed through delay lines which introduce a half cycle phase delay to the signals. Then the signals are combined into an output signal which is then amplified and passed to a transducer to yield an output noise having a frequency wave which is shifted one half cycle from the frequency of the ambient noise such that the ambient noise is canceled. …”
Regarding claim 8, the rejection of claim 7 is incorporated, furthermore ALAVA does not teach but NARAYANAN teaches:
a detected spectral profile of a target speaker's speech
NARAYANAN [0023] “The audio correction module 104 may generally determine that content comprises one or more incorrect words (e.g., spoken by a speaker) and automatically initiate steps to replace the incorrect(s) word in the content with corresponding correct word(s). For example, the audio correction module 104 may determine, such as via automatic content recognition (ACR), that a portion of content comprises an incorrect word (or multiple). A voice profile may be determined that is associated with the speaker of the incorrect word. For example, the voice profile module 106 may determine the voice profile based on the spoken audio content in the content portion that indicates the incorrect word (as opposed to background audio content in the content portion). The voice profile may be determined based on audio, vocal, and/or linguistic characteristics of the spoken audio content, such as audio spectral patterns, vocal frequency, vocal pitch, speaking speed, intonation, loudness, amplitude, speech patterns, speech cadence, the number and/or duration of utterances, the number and/or duration of breaks between utterances, etc. As such, the voice profile may characterize the speaker's speech according to these same (at least in part) audio, vocal, and/or linguistic characteristics. The voice profile may be used to generate spoken audio content with speech that resembles the speaker's other speech in the content. That is, spoken audio content generated based on a speaker's voice profile may include speech that sounds as it were actually spoken by the speaker.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA the capability to detect a spectral profile of a target speaker's speech. The benefit and motivation of such modification is discussed by NARAYANAN in the following portion: NARAYANAN [0023] “… As such, the voice profile may characterize the speaker's speech according to these same (at least in part) audio, vocal, and/or linguistic characteristics. The voice profile may be used to generate spoken audio content with speech that resembles the speaker's other speech in the content. That is, spoken audio content generated based on a speaker's voice profile may include speech that sounds as it were actually spoken by the speaker.”
ALAVA in view of NARAYANAN in view of LIU does not teach but SANTOS teaches:
the inaudible counteracting audio signal
SANTOS [0021] “In an embodiment, the randomly mixed voice recording is emitted at an audible level. In an embodiment, the randomly mixed voice recording is emitted at an inaudible level. In an embodiment, the randomly mixed voice recording is emitted at both an audible level and an inaudible level. In an embodiment, the inaudible level is selected from the group consisting of an infrasound and an ultrasound.”
maintaining its efficacy against unauthorized devices equipped with adaptive noise cancellation technologies.
SANTOS [0023] “In an embodiment, preventing the intelligible recording of the conversation comprises emitting the randomly mixed voice recording during the real-time audio recording of the conversation causing the real-time audio recording of the randomly mixed voice recording thereby masking the conversation.”
SANTOS [0039] “FIG. 1 shows an arrangement for preventing intelligible voice recordings also known as a voice recording jammer 10. The jammer 10 is shown here in the form of a device. Of course, the arrangement of FIG. 1 can also be a kit or a system. Device 10 prevents the intelligible recording of a voice. In an embodiment, device 10 prevents the intelligible recording of a conversation between at least two interlocutors.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN in view of LIU the capability to include an inaudible counteracting audio signal. The benefit and motivation of such modification is discussed by SANTOS in the following portion: SANTOS [0023] “In an embodiment, preventing the intelligible recording of the conversation comprises emitting the randomly mixed voice recording during the real-time audio recording of the conversation causing the real-time audio recording of the randomly mixed voice recording thereby masking the conversation.”
ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in further view of MCELVEEN does not teach but YORK teaches:
The method of claim 7, wherein the processing unit implements a frequency modulation algorithm to adjust the
YORK (Col 5 lines 19 to 31) “From the output stage of the preamplifier 22 the noise signal passes to the digital filter section 23 where the noise signal is analyzed in the frequency domain. The digital filters 23 are clock driven, and an adjustable clock determines the frequency at which the filters reach cutoff. There are a number of filters in this section, each one tuned to the fundamental frequency, or a harmonic of the fundamental frequency in question. In order for the circuit to respond favorably with a wide frequency range, it is necessary for the harmonics to be separated this way. For a squarewave, the harmonic content can be determined by taking the sum of the fundamental frequency and all the harmonics or overtones. ..."
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in further view of MCELVEEN the capability to implement a frequency modulation algorithm to adjust SANTOS inaudible counteracting audio signal and adjusted in response the detection of a spectral profile as discussed by NARAYANAN. The benefit and motivation of such modification is discussed by YORK in the following portion: YORK (Col 3 lines 28 to 39) “The present invention accomplishes the above and other objects by providing an apparatus that cancels ambient noise by having an input sensor means, an input amplifier, means for analyzing and dividing the signal, means for introducing a half cycle phase delay, means for recombining the signals into one signal with an output amplifier for passing the signal to an output loudspeaker to effectively cancel the ambient noise. The sensor means picks up ambient noise and converts it into an electrical input signal containing amplitude and temporal information corresponding to the frequency wave of the ambient noise. The input signal is then fed through an input amplifier to a series of digital filters which divide the signal into a fundamental signal and a series of separate harmonic signals of different frequency ranges. Then each of the harmonic signals is fed through delay lines which introduce a half cycle phase delay to the signals. Then the signals are combined into an output signal which is then amplified and passed to a transducer to yield an output noise having a frequency wave which is shifted one half cycle from the frequency of the ambient noise such that the ambient noise is canceled. …”
Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in view of MCELVEEN in further view of YORK in further view of Swerdlow; Nick (US 20240212666 A1) hereinafter SWERDLOW.
Regarding claim 9, the rejection of claim 8 is incorporated, furthermore ALAVA teaches:
counteracting audio signal
ALAVA (Col 6 line 16 to 36): “In response to determining that the communication of the user 104 includes confidential information, the audio modification system 102 may determine generated audio data (e.g., generated sound; active noise-cancelling audio data or voice-cancelling audio data) that renders the communication at least partially inaudible in the second portion 132 of the control area 106. The audio modification system 102 may determine the generated audio data based on parameters, such as a frequency and/or an amplitude, of sound waves of the communication of the user 104. For example, the generated audio data may include another frequency and/or another amplitude of sound waves that is generally opposite or otherwise configured to cancel or at least partially cancel the sound waves of the communication of the user 104. More specifically, the frequency and amplitude of the generated audio data may be 180 degrees out of phase relative to the frequency and amplitude of the communication of the user 104 (e.g., a phase shift of 180 degrees). In other embodiments, a phase shift between the communication of the user 104 and the generated audio data may be between 0 degrees and 180 degrees.”
ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in view of MCELVEEN in further view of YORK does not teach but SWERDLOW teaches:
The method of claim 8, wherein the synchronization of the counteracting audio signal is achieved by continuously monitoring the target speaker's speech cadence and applying a predictive timing adjustment algorithm that pre-aligns a signal's emission with anticipated speech segments.
SWERDLOW [0059] “... The vocal model can be trained using user accent data, speech cadence data, speech tone data, and/or speech inflection data of the real-time communication session participant. The vocal model can be used to extrapolate audio representations of the one or more predicted words in the voice of the real-time communication session participant. If there are multiple real-time communication session participants sharing one user device for the real-time communication session, the server 604 may use voice recognition software to identify the particular real-time communication session participant that is speaking, for example using voice fingerprinting, and synthesize the one or more predicted words in the identified real-time communication session participant's voice.”
SWERDLOW [0061] “... In some implementations, the audio stream combination software 414 may be configured to synthesize a predicted word. The predicted word may be synthesized in the voice of the user and combined with the audio stream to generate the combined audio stream. The audio stream combination software 414 may use the user speech profile to synthesize the predicted word in the voice of the user. The server 404 may transmit the combined audio stream via the communication software 406 to the user devices 402B-402N for output at those devices during the real-time communication session.”
It would have been obvious to someone of ordinary skill in the art before the effective filling date of the claimed invention to include in the teachings of ALAVA in view of NARAYANAN in view of LIU in view of SANTOS in view of MCELVEEN in further view of YORK the capability to synchronize the counteracting audio signal by continuously monitoring the target speaker's speech cadence and applying a predictive timing adjustment in real-time. The benefit and motivation of such modification is discussed by SWERDLOW in the following portion: SWERDLOW [Abstract] “…The server synthesizes a predicted word for replacing the missing word in a voice of a user of the user device and combines the synthesized word with the first audio stream to generate the continuous audio stream. The server transmits the continuous audio stream to other user devices connected to the real-time communication session.”
Allowable Subject Matter
Claims 10 to 14 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
Claims 15 to 20 would be allowable if rewritten or amended to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action.
The following is an examiner’s statement of reasons for allowance: Claims 10 to 14 are listed as allowable subject matter over prior art by merit of the combination of the limitations listed from claims 1 through 10. It would require elements from too many references in order to cover all the limitations leading to those claims, making it difficult to stablish a prima facia case of obviousness. Claims 15 to 20 are allowable subject matter over prior art for similar reasons regarding the combination of limitations.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HECTOR J. CRESPO FEBLES whose telephone number is (571)272-4512 and email is hcrespofebles@uspto.gov. The examiner can normally be reached Mon - Fri 7:30 - 5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HECTOR J. CRESPO FEBLES/Examiner, Art Unit 2657
/DANIEL C WASHBURN/Supervisory Patent Examiner, Art Unit 2657