DETAILED ACTION
This action is a First Action on the Merits (FAOM) for the claim set submitted on 04/28/2025. Claims 1-20 are pending and have been considered.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119
(a)-(d). The certified copy has been filed for the parent Application No. CN202410702965.9, filed on 05/31/2024.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1-20 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 3, 8-12 of copending Application No. 19/171,582 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other, as seen in the table below.
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented.
19/192,021 (instant app)
19/171,582 (reference app)
Notes
An audio processing method, comprising:
An audio processing method, comprising:
acquiring a plurality of first audio signals in a space of a mobile terminal;
acquiring a first audio signal in a space of a mobile terminal;
A plurality of first audio signals includes at least a first audio signal
determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal;
performing separation processing on the first audio signal to obtain at least two paths of second audio signals;
Separation processing will indicate determined positions associated with audio signals, wherein the examiner asserts that a path of an audio signal indicates a determined position with respect to the source, tracking to an audio signal having a position in space
performing audio mixing based on at least one second audio signal to obtain a third audio signal.
performing sound effect processing on the at least two paths of second audio signals respectively to correspondingly obtain at least two paths of third audio signals.
Sound effect processing and audio mixing are synonymous terms in the art under the BRI. Further, performing audio mixing “based on at least one second audio signal” indicates the mixing could be applied to an additional second audio signal, i.e. second path, resulting in two paths/positions of third audio signals
As the table demonstrates, each limitation of claim 1 of the present application is found in claim 1 of the copending Application 19/171,582, thus claim 1 of the instant application is anticipated by claim 1 of the copending reference application. Dependent claims 2-3, 8-12 of the instant application are substantially similar to dependent claims 2, 7-11 of the reference application and are, therefore, also anticipated by the copending reference application as shown below.
19/192,021 (instant app)
19/171,582 (reference app)
Notes
2. The method according to claim 1, wherein the determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal comprises: performing separation processing on the plurality of first audio signals to obtain a plurality of fourth audio signals; and determining the at least one second audio signal based on the plurality of fourth audio signals.
(as previously disclosed in the independent claim, see above). Any of the “paths” from the paths of audio signals track to fourth audio signals, wherein any of those can also be a second audio signal if selected (suggested by the processing to output third audio signals). Sound effect processing can be the form of determining
The differences between “second”, “third”, “fourth”, etc. is merely convention and does not affect functionality
3. The method according to claim 2, wherein the performing separation processing on the plurality of first audio signals to obtain a plurality of fourth audio signals comprises: inputting the plurality of first audio signals into a first neural network model, and outputting the plurality of fourth audio signals respectively through a plurality of output channels of the first neural network model.
2. The method according to claim 1, wherein the performing separation processing on the first audio signal to obtain at least two paths of second audio signals, comprises: inputting the first audio signal into a first neural network model, and outputting the at least two paths of second audio signals respectively by at least two output channels of the first neural network model.
The differences between “second” and “fourth” audio signals is negligible and strictly related to naming convention
8. The method according to claim 1, wherein the performing audio mixing based on the at least one second audio signal to obtain a third audio signal comprises: performing signal superposition on the at least one second audio signal to obtain a fifth audio signal; and performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal.
7. The method according to claim 6, wherein the performing audio mixing based on the at least two paths of third audio signals to obtain a fourth audio signal, comprises: executing signal superposition on the at least two paths of third audio signals to obtain a fifth audio signal; and performing audio mixing processing on the fifth audio signal and a preset signal to obtain the fourth audio signal.
Again, differences between the numbering of signals is convention, not related to functionality
9. The method according to claim 8, wherein before the performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal, the method further comprises: performing audio effect processing on the fifth audio signal to obtain a sixth audio signal; and the performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal comprises: performing audio mixing processing on the sixth audio signal and the preset signal to obtain the third audio signal.
9. The method according to claim 6 wherein the acquiring a first audio signal in a space of a mobile terminal, comprises: acquiring a sixth audio signal in the space of the mobile terminal; and eliminating an interference signal in the sixth audio signal to obtain the first audio signal.
10. The method according to claim 9, wherein the eliminating an interference signal in the sixth audio signal to obtain the first audio signal, comprises: performing interference signal elimination processing on the sixth audio signal based on a reference signal determined based on the fourth audio signal to obtain the first audio signal.
Again, differences between the numbering of signals is convention, not related to functionality. Eliminating interference tracks to a method of audio mixing
A reference signal tracks to a preset signal
10. The method according to claim 8, wherein the acquiring a plurality of first audio signals in a space of a mobile terminal comprises: acquiring sound signals at a plurality of positions in the space of the mobile terminal through a plurality of transducers to obtain the plurality of first audio signals.
8. The method according to claim 6, wherein the acquiring a first audio signal in a space of a mobile terminal, comprises: acquiring sound signals at a plurality of locations in the space of the mobile terminal by a plurality of acoustic sensors to obtain the first audio signal.
11. The method according to claim 8, wherein the acquiring a plurality of first audio signals in a space of a mobile terminal comprises: acquiring a plurality of seventh audio signals in the space of the mobile terminal; and eliminating interference signals in the plurality of seventh audio signals, respectively, to obtain the plurality of first audio signals.
9. The method according to claim 6 wherein the acquiring a first audio signal in a space of a mobile terminal, comprises: acquiring a sixth audio signal in the space of the mobile terminal; and eliminating an interference signal in the sixth audio signal to obtain the first audio signal.
Again, differences between the numbering of signals is convention, not related to functionality
12. The method according to claim 1, further comprising: playing the third audio signal inside the space of the mobile terminal and/or outside the space of the mobile terminal.
11. The method according to claim 6, further comprising: playing the fourth audio signal inside the space of the mobile terminal, and/or outside the space of the mobile terminal.
Again, differences between the numbering of signals is convention, not related to functionality
Similar rationale is extended to independent claim 13 and its associated dependent claims (14-15, 19-20) which are not patentably distinct from the claims above; therefore, claims 13-15, 19-20 are also rejected as being anticipated by the reference app.
Regarding claims 4/16, it would be obvious to apply voice activity detection on signals which may contain voice before the effective filing date of the claimed invention. Consider Buck, used to reject these claims under 35 U.S.C. 103 below:
performing voice activity detection (VAD) on the plurality of fourth audio signals respectively to determine the at least one second audio signal ([0021] information about whether this is an actively speaking user in each sound zones 104 may be derived. Additional voice activity detection techniques may additionally be used to determine whether a speaker is present, [0069] The determination of whether there is voice in the microphone signals 112 may be performed using various techniques discussed herein, such as capturing beam-formed signals for each sound zone 104 position to determine a location of a speaker, analysis of the microphone signals 112 to identify changes in energy, spectral, or cepstral distances in the captured microphone signals 112, etc., [Wherein sound zones are taken in view of the multi-tracks/channels/signals of Tan and could be applied to the aforementioned tracks/channels/signals without a change in functionality to the VAD of Buck]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of the instant application to incorporate the teachings of Buck, because of the novel way to reinforce voice signals within noisy environments through cancellation of feedback and echo that is a cause of the environment (as would be determined using VAD) (Buck, [0006]).
Regarding claims 5/17, it would be obvious to determine a voice emission position from a plurality of positions corresponding to audio signals and to determine an audio signal based on the voice emission position before the effective filing date of the claimed invention. Consider Buck further:
determining, according to user feature information ([0021] Additional voice activity detection techniques may additionally be used to determine whether a speaker is present, such as changes in energy, spectral, or cepstral distances in the captured microphone signals 112, [The examiner asserts that voice information, i.e. activity, is user feature information]), at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals ([0017] The reinforcement may include localizing the voice signal within the multiple sound zone environment 102, [Localizing a voice signal indicates a determination of a signal to be voice signal (as would be performed using the previously discussed VAD of Buck) which further indicates a localization corresponding to user information, i.e. voice information, being detected]); and
determining the at least one second audio signal according to the at least one voice emission position ([0069] The determination of whether there is voice in the microphone signals 112 may be performed using various techniques discussed herein, such as capturing beam-formed signals for each sound zone 104 position to determine a location of a speaker, analysis of the microphone signals 112 to identify changes in energy, spectral, or cepstral distances in the captured microphone signals 112, etc., [Determining the location of a speaker indicates determining a second audio signal, i.e. the voice signal, according to where the voice is located, i.e. the emission position, as determined through the beamforming operations]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of the instant application to incorporate the teachings of Buck, because of the novel way to reinforce voice signals within noisy environments through cancellation of feedback and echo that is a cause of the environment (as would be determined using VAD) (Buck, [0006]).
Regarding claims 6/18, it would be obvious to gather user feature information, which is multimodal, use said feature information to determine an emission position, perform recognition on a plurality of positions corresponding to the first audio signals, and determine a voice emission position based on the recognition results before the effective filing date of the claimed invention. Consider Ayrapetian, used to reject claims 6/18 under 35 U.S.C. 103:
wherein the user feature information comprises multimodal information of a user ([Col. 14, Lines 50-56] the device 102 may identify the person and/or a position associated with the person using audio data (e.g., audio beamforming), video data (e.g., facial recognition) and/or other inputs known to one of skill in the art, [Wherein video + audio comprises a multimodal information of a user]); and
the determining, according to user feature information ([In view of the previously cited multimodal data of Ayrapetian]), at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals comprises:
performing recognition on the plurality of positions corresponding to the plurality of first audio signals according to the multimodal information to obtain a recognition result ([Fig. 3A-B, Sections 1-8], [Col. 14, Lines 59-63] the device 102 may receive video data associated with the audio data and may use facial recognition or other techniques to determine a position associated with a face recognized in the video data, [A recognized face is a recognition result. In view of the plurality of sections for detecting people, taken in view of the plurality of sources/channels of Tan/Buck, the examiner asserts that this operation could be extended to multiple sources as disclosed in Tan/Buck without a change in functionality to the recognition result process of Ayrapetian as there are multiple sound sources disclosed in Ayrapetian (Fig. 1, 114a/b)]); and
determining, according to the recognition result ([As previously disclosed]), the at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals ([Col. 17, Lines 10-15] The device 102 may determine the current speech position using voice activity detection, voice recognition, facial recognition (if the device 102 has access to image data), or the like, [Wherein a current speech position tracks to a voice emission position]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of the instant application in view of Ayrapetian, because of the novel way to combine the advantages of single direction acoustic echo cancellation (AEC) with multi-direction AEC, allowing for two output paths which enable automatic speech recognition processing in parallel with voice communication, improving echo cancellation in environments with a changing number of sources (Ayrapetian, [Col. 3, Lines 1-25]).
Regarding claims 7/19, it would have been obvious before the effective filing date of the claimed invention to incorporate multimodal information comprised of image and/or video information. Consider Ayrapetian:
wherein the multimodal information comprises image information or video information ([Col. 14, Lines 50-56] the device 102 may identify the person and/or a position associated with the person using audio data (e.g., audio beamforming), video data (e.g., facial recognition) and/or other inputs known to one of skill in the art).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of the instant application in view of Ayrapetian, because of the novel way to combine the advantages of single direction acoustic echo cancellation (AEC) with multi-direction AEC, allowing for two output paths which enable automatic speech recognition processing in parallel with voice communication, improving echo cancellation in environments with a changing number of sources (Ayrapetian, [Col. 3, Lines 1-25]).
Claim Objections
Claims 7 and 19 are objected to because of the following informalities: the claims read “wherein the multimodal information comprises image information or video information”. It is unclear to the examiner how information represented in one mode, i.e. image or video, can be multimodal when only one mode is required to satisfy the claim as currently constructed. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-7, 13-19 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Independent claim(s) 1 recite:
acquiring a plurality of first audio signals in a space of a mobile terminal;
determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal; and
performing audio mixing based on at least one second audio signal to obtain a third audio signal.
These limitations, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. For example, the claim(s) read(s) on listening to audio signals, separating audio signals based on position of signals, and mixing audio signals based on a position-determined signal. There is nothing in the claim elements which preclude the steps from practically being performed in the mind.
For example, considering being in a car where every seat is occupied (considering Applicant’s disclosure of the mobile terminal to be containing vehicles, [0043]). With regard to the “acquiring” step, the examiner asserts that listening to audio is acquiring said audio through the ears to be processed by the brain. With regard to the “determining” step, the examiner asserts that, generally, one who hears sounds has the ability to look and determine where said sound sources are located. This is visually determining audio signals corresponding to positions in the space of the mobile terminal, wherein a full vehicle will have multiple speakers and/or sources of audio. Any of these can be a second audio signal. With regard to the “performing audio mixing”, the examiner asserts that one can mentally choose whether or not to perceive what is produced from surrounding sound sources, i.e. ignoring annoying sounds. Ignoring an annoying signal to hear other signals is a mental process. Further, the examiner asserts that in the environment of a vehicle, one could speak to other passengers and ask a sound/speaker level to be adjusted, affecting the mix of audio as received by the asking person. This is an audio mixing based on previously, mentally determined separated sound source signals, as obtained by the ears to be processed by the brain.
All of these steps can be performed in the mind and/or using pen and paper. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea (Step 2A, Prong one, Yes).
This judicial exception is not integrated into a practical application because the addition of generically recited computer elements does not add a meaningful limitation to the abstract idea because they amount to simply implementing the abstract idea on a computer. The claims are directed to an abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception ( Step 2A, Prong two, No). As discussed above, with respect to integration of an abstract idea into a practical application, the additional element of “acquiring”, “determining”, “performing”, are merely for the purpose of data gathering, storing, processing, and/or insignificant extra-solution activity that amount to no more than mere instructions to apply the exception using a generic computer component. Fig. 11 of the instant application disclose(s) applying the method to a generic computing device such as a PC (the figure discloses a generic computing device with generic components). Mere instructions to apply the exception using a generic computer component cannot provide an inventive concept. Therefore, the claims are not patent eligible (Step 2B, No).
Claim 13 recites “A non-transitory computer readable storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, causes the processor to implement the audio processing method according to claim 1.” The examiner asserts that using generic computing components, i.e. a non-transitory computer readable storage medium and processor, to perform a mental process (as previously discussed for claim 1) does not provide an inventive concept. See MPEP 2106.04(2).III.C, “Performing a mental process on a generic computer” and/or “Using a computer as a tool to perform a mental process”; therefore, claim 13 is also rejected under 35 U.S.C. 101 as being directed to a mental process.
Similarly, dependent claim(s) 2-7, 14-19 include additional steps that are considered “insignificant extra-solution activity to judicial exception” because they fail to provide meaningful significance that goes beyond generally linking the use of an abstract idea to a particular technological environment.
For example, claim 2 reads on performing separation processing on first audio signals to obtain fourth audio signals and determining a second audio signal based on the separated audio signals. As previously discussed, separation of sound sources is a mental process. Choosing one identified source is also a mental process.
Claim 3 reads on using a neural network to perform the audio source separation. Incorporating a generic software computing component to perform a mental process does not provide an inventive concept. See MPEP 2106.04(2).III.C, “Performing a mental process on a generic computer” and/or “Using a computer as a tool to perform a mental process”.
Claim 4 reads on performing voice activity detection on the separated sound signals to select one as a second audio signal. Determining if a sound is comprised of speech is a mental process associated with listening to the sounds. Determining a signal based on the result is the mental process of determining which have speaking and which don’t. Any signal selected is selected based on the voice activity detection.
Claim 5 reads on the plurality of audio signals corresponding to a plurality of positions of in the space of a mobile terminal, wherein the determining a second audio signal is based determining voice emission position according to user feature information and selecting the second audio signal based on the voice emission positions. With regard to a plurality of audio signals in a plurality of positions in space in the mobile terminal, the examiner asserts that each individual and/or audio source with a vehicle tracks to a plurality of positions within the mobile terminal of a vehicle. With regard to the determining voice emission position based on user feature information, the examiner asserts that this can also be performed mentally based on heard signals. If the listener wants to ask someone to quiet down, they will associate the speech of the speaker (user feature information) with where the speech is coming from (voice emission position). This is determined mentally. As previously disclosed, selecting a second signal is a mental process. Selecting said signal based on the mental process of determining voice emission position is also a mental process.
Claim 6 reads on the user feature information comprising multimodal information, performing recognition on the audio signals according to multimodal information to obtain a recognition result, and determining a voice emission position based on the recognition result. Multimodal information can be information which is seen and heard to be processed by the brain. Performing recognition on audio signals according to multimodal information is also a mental process. Consider listening to an unknown speaker to determine age, gender, etc., wherein the recognition result is the classification of the speaker determined through visual and sound, i.e. multimodal, analyses. Further, determining a voice emission position based on the recognition result is looking to find the source which you have previously mentally recognized to be producing a perceived sound.
Claim 7 reads on the multimodal information being comprised of image or video information. The examiner asserts that a user can mentally look at people to determine who is speaking when they are speaking.
Therefore, these claims are also not patent eligible.
Regarding claims 8 and 20, the examiner asserts that the aforementioned claims contain patent eligible subject matter due to the inclusion of steps which incorporate signal superposition and audio mixing of a signal with a preset signal. Superposition of two signals cannot be reasonably performed in the mind as the signals in a multi-signal environment will already by superposed as they are received by the ears. One can mentally generate and imagine what two signals would sound like when superimposed, but this does not result in a meaningful signal and is the input itself. Further, mixing a preset signal with a previously superimposed signal is another process which cannot be reasonably be asserted to be mental as there is no meaningful superimposed signal to add music to (Step, 2A, Prong 1, NO). Claims 9-11 are eligible as they further define the eligible subject matter of claim 8.
Regarding claim 12, the examiner asserts that playing a mixed audio inside the space of a mobile terminal and/or outside the space of the mobile terminal is not something which can be performed in the mind (Step 2A, Prong 1, NO). A user will be able to mentally “mix” audio based on what they would prefer to listen to, but this does not result in a meaningful signal as output to be played inside and/or outside of a mobile terminal. Further, should this process be interpreted to be generic post-solution activity, the examiner asserts that, given the context of the application to be an improvement of karaoke in mobile terminals, namely, vehicles, the examiner asserts that playing a mixed output based on localized sound sources within the mobile terminal directs the mental process of claim 1 to a practical application in view of the specification as the perceived karaoke quality being output has an improved intelligibility and reduced noise, see [0009], [0025], [0033] of the instant application (Step 2A, Prong 2, YES).
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-3, 12-15 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Tan et al. (US-20250324218-A1), hereinafter Tan.
Regarding claim 1, Tan discloses:
an audio processing method ([0071] method for converting an in-vehicle audio into a multi-channel surround sound, [Wherein converting audio is a form of processing]), comprising:
acquiring a plurality of first audio signals ([Fig. 2, Mobile Phone, Vehicle-mounted App music into Audio decoding], [Sending a signal from both a mobile phone and a mounted music app into audio decoding indicates each input to be its own audio signal to be decoded which must be acquired by said decoder in order to be decoded]) in a space of a mobile terminal ([0096] acquiring an in-vehicle audio signal, [The examiner asserts that a vehicle tracks to a mobile terminal in view of [0042] of the instant application which claims the mobile terminal to be a vehicle]);
determining, based on the plurality of first audio signals ([As previously disclosed in claim 1]), a second audio signal corresponding to at least one position in the space of the mobile terminal ([Fig. 2, Audio separation model], [0125] performing classification according to the extracted features by using an audio separation model, so as to separate the target audio data into a plurality of target single audio signals, [Wherein any of the separated single audio signals represents a second audio signal]);
performing audio mixing based on at least one second audio signal to obtain a third audio signal ([0128] mixing the equalized single audio signals by using the optimized sound mixing parameters to obtain a result audio signal, [As currently claimed, audio mixing based on at least one second audio signal does not indicate with what and/or how the second audio signal is mixed; therefore, the examiner asserts that any mixing with a target single audio signal tracks to the claim. The result audio signal is a third audio signal]).
Regarding claim 2, Tan discloses: the method according to claim 1.
Tan further discloses:
wherein the determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal comprises:
performing separation processing on the plurality of first audio signals to obtain a plurality of fourth audio signals ([Fig. 2, Audio separation model], [0125] performing classification according to the extracted features by using an audio separation model, so as to separate the target audio data into a plurality of target single audio signals, [0100] the audio separation model is an AI model in which structures of a convolutional neural network and a long-short term memory neural network are used to construct an encoder-decoder model , [Wherein any of the separated single audio signals represents a fourth audio signal. Further, wherein the audio separation model tracks to separation processing as [0036] of the instant app discloses separation processing to be performed by an artificial intelligence sound source separation]); and
determining the at least one second audio signal based on the plurality of fourth audio signals ([0118] manually checking the plurality of single audio signals for debugging separated by the audio separation model; if the check is passed, keeping the current single audio signal for debugging, [Selecting a checked signal indicates the selected signal to be a second audio signal, wherein the second audio signal is chosen from the separated, i.e. fourth, signals]).
Regarding claim 3, Tan discloses: the method according to claim 2.
Tan further discloses:
wherein the performing separation processing on the plurality of first audio signals to obtain a plurality of fourth audio signals comprises:
inputting the plurality of first audio signals into a first neural network model ([Fig. 2, decoded audio going into Audio separation model], [0100] the audio separation model is an AI model in which structures of a convolutional neural network and a long-short term memory neural network are used to construct an encoder-decoder model, [Wherein the decoded audio represents audio before speech separation is performed, tracking to first audio signals, wherein the separation model is containing a CNN and/or LSTM]), and
outputting the plurality of fourth audio signals respectively through a plurality of output channels of the first neural network model ([Fig. 2, Multi-channel sound element], [0125] wherein as shown in FIG. 2, the audio separation model separates the sound into multi-track sound elements, [Wherein tracks and channels appear to be synonymous terms with regard to Tan as seen between the figure and description of said figure. Further, using a neural network to separate input into tracks/channel output indicates each output to be from a respective output channel of the neural network]).
Regarding claim 12, Tan discloses: the method according to claim 1.
Tan further discloses:
playing the third audio signal inside the space of the mobile terminal ([0129] outputting the mixed result audio signal, wherein specifically, as shown in FIG. 2, the outputting the mixed result audio signal by an in-vehicle loudspeaker, [Wherein the third audio signal tracks to that which has been mixed. In-vehicle in inside the space of the vehicle, i.e. mobile terminal]) and/or outside the space of the mobile terminal ([The examiner asserts that, due to the disjunctive nature with which the claim is constructed, this element does not require a mapping]).
Regarding claim 13, Tan discloses: a non-transitory computer readable storage medium, on which a computer program is stored ([Abstract, non-transitory computer readable medium stores a computer-program product includes instructions]), wherein the computer program, when executed by a processor ([0051] wherein the processor is configured to execute the instructions), causes the processor to implement the audio processing method according to claim 1 ([See above rejection of claim 1. The same rationale would be extended to this claim as applied by the non-transitory computer readable storage medium]).
Regarding claim 14, Tan discloses: the non-transitory computer readable storage medium according to claim 13.
Tan further discloses:
wherein the determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal comprises:
performing separation processing on the plurality of first audio signals to obtain a plurality of fourth audio signals ([Fig. 2, Audio separation model], [0125] performing classification according to the extracted features by using an audio separation model, so as to separate the target audio data into a plurality of target single audio signals, [0100] the audio separation model is an AI model in which structures of a convolutional neural network and a long-short term memory neural network are used to construct an encoder-decoder model , [Wherein any of the separated single audio signals represents a fourth audio signal. Further, wherein the audio separation model tracks to separation processing as [0036] of the instant app discloses separation processing to be performed by an artificial intelligence sound source separation]); and
determining the at least one second audio signal based on the plurality of fourth audio signals ([0118] manually checking the plurality of single audio signals for debugging separated by the audio separation model; if the check is passed, keeping the current single audio signal for debugging, [Selecting a checked signal indicates the selected signal to be a second audio signal, wherein the second audio signal is chosen from the separated, i.e. fourth, signals]).
Regarding claim 15, Tan discloses: the non-transitory computer readable storage medium according to claim 14.
Tan further discloses:
wherein the performing separation processing on the plurality of first audio signals to obtain a plurality of fourth audio signals comprises:
inputting the plurality of first audio signals into a first neural network model ([Fig. 2, decoded audio going into Audio separation model], [0100] the audio separation model is an AI model in which structures of a convolutional neural network and a long-short term memory neural network are used to construct an encoder-decoder model, [Wherein the decoded audio represents audio before speech separation is performed, tracking to first audio signals, wherein the separation model is containing a CNN and/or LSTM]), and
outputting the plurality of fourth audio signals respectively through a plurality of output channels of the first neural network model ([Fig. 2, Multi-channel sound element], [0125] wherein as shown in FIG. 2, the audio separation model separates the sound into multi-track sound elements, [Wherein tracks and channels appear to be synonymous terms with regard to Tan as seen between the figure and description of said figure. Further, using a neural network to separate input into tracks/channel output indicates each output to be from a respective output channel of the neural network]).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 4-5, 16-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tan in view of Buck et al. (US-20230215449-A1), hereinafter Buck.
Regarding claim 4, Tan discloses: the method according to claim 2.
Tan does not disclose:
wherein the determining the at least one second audio signal based on the plurality of fourth audio signals comprises:
performing voice activity detection (VAD) on the plurality of fourth audio signals respectively to determine the at least one second audio signal.
Buck discloses:
wherein the determining the at least one second audio signal based on the plurality of fourth audio signals comprises:
performing voice activity detection (VAD) on the plurality of fourth audio signals respectively to determine the at least one second audio signal ([0021] information about whether this is an actively speaking user in each sound zones 104 may be derived. Additional voice activity detection techniques may additionally be used to determine whether a speaker is present, [0069] The determination of whether there is voice in the microphone signals 112 may be performed using various techniques discussed herein, such as capturing beam-formed signals for each sound zone 104 position to determine a location of a speaker, analysis of the microphone signals 112 to identify changes in energy, spectral, or cepstral distances in the captured microphone signals 112, etc., [Wherein sound zones are taken in view of the multi-tracks/channels/signals of Tan and could be applied to the aforementioned tracks/channels/signals without a change in functionality to the VAD of Buck]).
Tan and Buck are considered analogous art within signal mixing within vehicles. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan to incorporate the teachings of Buck, because of the novel way to reinforce voice signals within noisy environments through cancellation of feedback and echo that is a cause of the environment (as would be determined using VAD) (Buck, [0006]).
Regarding claim 5, Tan discloses: the method according to claim 1.
Tan further discloses:
wherein the plurality of first audio signals correspond to a plurality of positions in the space of the mobile terminal ([Fig. 2, Mobile Phone, Vehicle-mounted App music into Audio decoding], [Sending a signal from both a mobile phone and a mounted music app into audio decoding indicates each input to be its own position in space of the mobile terminal unless the mobile phone were within the vehicle-mounted music app (which the examiner is interpreting to be a car radio or the like). The examiner asserts that it is not reasonable to put a mobile device within a car radio loudspeaker]).
Tan does not disclose:
determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal comprises:
determining, according to user feature information, at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals; and
determining the at least one second audio signal according to the at least one voice emission position.
Buck discloses:
determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal comprises:
determining, according to user feature information ([0021] Additional voice activity detection techniques may additionally be used to determine whether a speaker is present, such as changes in energy, spectral, or cepstral distances in the captured microphone signals 112, [The examiner asserts that voice information, i.e. activity, is user feature information]), at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals ([0017] The reinforcement may include localizing the voice signal within the multiple sound zone environment 102, [Localizing a voice signal indicates a determination of a signal to be voice signal (as would be performed using the previously discussed VAD of Buck) which further indicates a localization corresponding to user information, i.e. voice information, being detected]); and
determining the at least one second audio signal according to the at least one voice emission position ([0069] The determination of whether there is voice in the microphone signals 112 may be performed using various techniques discussed herein, such as capturing beam-formed signals for each sound zone 104 position to determine a location of a speaker, analysis of the microphone signals 112 to identify changes in energy, spectral, or cepstral distances in the captured microphone signals 112, etc., [Determining the location of a speaker indicates determining a second audio signal, i.e. the voice signal, according to where the voice is located, i.e. the emission position, as determined through the beamforming operations]).
Tan and Buck are considered analogous art within signal mixing within vehicles. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan to incorporate the teachings of Buck, because of the novel way to reinforce voice signals within noisy environments through cancellation of feedback and echo that is a cause of the environment (as would be determined using VAD) (Buck, [0006]).
Regarding claim 16, Tan discloses: the non-transitory computer readable medium according to claim 14.
Tan does not disclose:
wherein the determining the at least one second audio signal based on the plurality of fourth audio signals comprises:
performing voice activity detection (VAD) on the plurality of fourth audio signals respectively to determine the at least one second audio signal.
Buck discloses:
wherein the determining the at least one second audio signal based on the plurality of fourth audio signals comprises:
performing voice activity detection (VAD) on the plurality of fourth audio signals respectively to determine the at least one second audio signal ([0021] information about whether this is an actively speaking user in each sound zones 104 may be derived. Additional voice activity detection techniques may additionally be used to determine whether a speaker is present, [0069] The determination of whether there is voice in the microphone signals 112 may be performed using various techniques discussed herein, such as capturing beam-formed signals for each sound zone 104 position to determine a location of a speaker, analysis of the microphone signals 112 to identify changes in energy, spectral, or cepstral distances in the captured microphone signals 112, etc., [Wherein sound zones are taken in view of the multi-tracks/channels/signals of Tan and could be applied to the aforementioned tracks/channels/signals without a change in functionality to the VAD of Buck]).
Tan and Buck are considered analogous art within signal mixing within vehicles. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan to incorporate the teachings of Buck, because of the novel way to reinforce voice signals within noisy environments through cancellation of feedback and echo that is a cause of the environment (as would be determined using VAD) (Buck, [0006]).
Regarding claim 17, Tan discloses: the non-transitory computer readable medium according to claim 13.
Tan further discloses:
wherein the plurality of first audio signals correspond to a plurality of positions in the space of the mobile terminal ([Fig. 2, Mobile Phone, Vehicle-mounted App music into Audio decoding], [Sending a signal from both a mobile phone and a mounted music app into audio decoding indicates each input to be its own position in space of the mobile terminal unless the mobile phone were within the vehicle-mounted music app (which the examiner is interpreting to be a car radio or the like). The examiner asserts that it is not reasonable to put a mobile device within a car radio loudspeaker]).
Tan does not disclose:
determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal comprises:
determining, according to user feature information, at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals; and
determining the at least one second audio signal according to the at least one voice emission position.
Buck discloses:
determining, based on the plurality of first audio signals, a second audio signal corresponding to at least one position in the space of the mobile terminal comprises:
determining, according to user feature information ([0021] Additional voice activity detection techniques may additionally be used to determine whether a speaker is present, such as changes in energy, spectral, or cepstral distances in the captured microphone signals 112, [The examiner asserts that voice information, i.e. activity, is user feature information]), at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals ([0017] The reinforcement may include localizing the voice signal within the multiple sound zone environment 102, [Localizing a voice signal indicates a determination of a signal to be voice signal (as would be performed using the previously discussed VAD of Buck) which further indicates a localization corresponding to user information, i.e. voice information, being detected]); and
determining the at least one second audio signal according to the at least one voice emission position ([0069] The determination of whether there is voice in the microphone signals 112 may be performed using various techniques discussed herein, such as capturing beam-formed signals for each sound zone 104 position to determine a location of a speaker, analysis of the microphone signals 112 to identify changes in energy, spectral, or cepstral distances in the captured microphone signals 112, etc., [Determining the location of a speaker indicates determining a second audio signal, i.e. the voice signal, according to where the voice is located, i.e. the emission position, as determined through the beamforming operations]).
Tan and Buck are considered analogous art within signal mixing within vehicles. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan to incorporate the teachings of Buck, because of the novel way to reinforce voice signals within noisy environments through cancellation of feedback and echo that is a cause of the environment (as would be determined using VAD) (Buck, [0006]).
Claim(s) 6-7, 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tan in view of Buck, further in view of Ayrapetian et al. (US-9818425-B1), hereinafter Ayrapetian.
Regarding claim 6, Tan in view of Buck discloses: the method according to claim 5.
Tan in view of Buck does not disclose:
wherein the user feature information comprises multimodal information of a user; and
the determining, according to user feature information, at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals comprises:
performing recognition on the plurality of positions corresponding to the plurality of first audio signals according to the multimodal information to obtain a recognition result; and
determining, according to the recognition result, the at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals.
Ayrapetian discloses:
wherein the user feature information comprises multimodal information of a user ([Col. 14, Lines 50-56] the device 102 may identify the person and/or a position associated with the person using audio data (e.g., audio beamforming), video data (e.g., facial recognition) and/or other inputs known to one of skill in the art, [Wherein video + audio comprises a multimodal information of a user]); and
the determining, according to user feature information ([In view of the previously cited multimodal data of Ayrapetian]), at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals comprises:
performing recognition on the plurality of positions corresponding to the plurality of first audio signals according to the multimodal information to obtain a recognition result ([Fig. 3A-B, Sections 1-8], [Col. 14, Lines 59-63] the device 102 may receive video data associated with the audio data and may use facial recognition or other techniques to determine a position associated with a face recognized in the video data, [A recognized face is a recognition result. In view of the plurality of sections for detecting people, taken in view of the plurality of sources/channels of Tan/Buck, the examiner asserts that this operation could be extended to multiple sources as disclosed in Tan/Buck without a change in functionality to the recognition result process of Ayrapetian as there are multiple sound sources disclosed in Ayrapetian (Fig. 1, 114a/b)]); and
determining, according to the recognition result ([As previously disclosed]), the at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals ([Col. 17, Lines 10-15] The device 102 may determine the current speech position using voice activity detection, voice recognition, facial recognition (if the device 102 has access to image data), or the like, [Wherein a current speech position tracks to a voice emission position]).
Tan, Buck, and Ayrapetian are considered analogous art within voice signal mixing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan in view of Buck to incorporate the teachings of Ayrapetian, because of the novel way to combine the advantages of single direction acoustic echo cancellation (AEC) with multi-direction AEC, allowing for two output paths which enable automatic speech recognition processing in parallel with voice communication, improving echo cancellation in environments with a changing number of sources (Ayrapetian, [Col. 3, Lines 1-25]).
Regarding claim 7, Tan in view of Buck, further in view of Ayrapetian discloses: the method according to claim 6.
Ayrapetian further discloses:
wherein the multimodal information comprises image information or video information ([Col. 14, Lines 50-56] the device 102 may identify the person and/or a position associated with the person using audio data (e.g., audio beamforming), video data (e.g., facial recognition) and/or other inputs known to one of skill in the art).
Regarding claim 18, Tan in view of Buck discloses: the non-transitory computer readable medium according to claim 17.
Tan in view of Buck does not disclose:
wherein the user feature information comprises multimodal information of a user; and
the determining, according to user feature information, at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals comprises:
performing recognition on the plurality of positions corresponding to the plurality of first audio signals according to the multimodal information to obtain a recognition result; and
determining, according to the recognition result, the at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals.
Ayrapetian discloses:
wherein the user feature information comprises multimodal information of a user ([Col. 14, Lines 50-56] the device 102 may identify the person and/or a position associated with the person using audio data (e.g., audio beamforming), video data (e.g., facial recognition) and/or other inputs known to one of skill in the art, [Wherein video + audio comprises a multimodal information of a user]); and
the determining, according to user feature information ([In view of the previously cited multimodal data of Ayrapetian]), at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals comprises:
performing recognition on the plurality of positions corresponding to the plurality of first audio signals according to the multimodal information to obtain a recognition result ([Fig. 3A-B, Sections 1-8], [Col. 14, Lines 59-63] the device 102 may receive video data associated with the audio data and may use facial recognition or other techniques to determine a position associated with a face recognized in the video data, [A recognized face is a recognition result. In view of the plurality of sections for detecting people, taken in view of the plurality of sources/channels of Tan/Buck, the examiner asserts that this operation could be extended to multiple sources as disclosed in Tan/Buck without a change in functionality to the recognition result process of Ayrapetian as there are multiple sound sources disclosed in Ayrapetian (Fig. 1, 114a/b)]); and
determining, according to the recognition result ([As previously disclosed]), the at least one voice emission position from the plurality of positions corresponding to the plurality of first audio signals ([Col. 17, Lines 10-15] The device 102 may determine the current speech position using voice activity detection, voice recognition, facial recognition (if the device 102 has access to image data), or the like, [Wherein a current speech position tracks to a voice emission position]).
Tan, Buck, and Ayrapetian are considered analogous art within voice signal mixing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan in view of Buck to incorporate the teachings of Ayrapetian, because of the novel way to combine the advantages of single direction acoustic echo cancellation (AEC) with multi-direction AEC, allowing for two output paths which enable automatic speech recognition processing in parallel with voice communication, improving echo cancellation in environments with a changing number of sources (Ayrapetian, [Col. 3, Lines 1-25]).
Regarding claim 19, Tan in view of Buck, further in view of Ayrapetian discloses: the non-transitory computer readable medium according to claim 18.
Ayrapetian further discloses:
wherein the multimodal information comprises image information or video information ([Col. 14, Lines 50-56] the device 102 may identify the person and/or a position associated with the person using audio data (e.g., audio beamforming), video data (e.g., facial recognition) and/or other inputs known to one of skill in the art).
Claim(s) 8-11, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Tan in view of Kotulla et al. (US-20240331719-A1), hereinafter Kotulla.
Regarding claim 8, Tan discloses: the method according to claim 1.
Tan does not disclose:
wherein the performing audio mixing based on the at least one second audio signal to obtain a third audio signal comprises:
performing signal superposition on the at least one second audio signal to obtain a fifth audio signal; and
performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal.
Kotulla discloses:
wherein the performing audio mixing based on the at least one second audio signal to obtain a third audio signal comprises:
performing signal superposition on the at least one second audio signal to obtain a fifth audio signal ([0050] Each audio outputting element 7 may be assigned to a specific location or space, i.e. particularly to a specific seat (not shown), in the vehicle cabin 4, [0064] The audio receiving device 8 may be configured to receive a human voice signal 9 of at least one first person P located in the vehicle cabin 4 and a human voice signal 9′ of at least one further person P′ located in the vehicle cabin 4. Thereby, the processing device 11 may be further configured to modify the received human voice signal 9 of the at least one first person P located in the vehicle cabin 4 with at least one first acoustic modification parameter and to modify the received human voice signal 9′ of the at least one further person P’ located in the vehicle cabin 4 with at least one further acoustic modification parameter. It may be required to separate the received human voice signal 9 of at least one first person P, which embraces also a group of first persons, from the received human voice signal of at least one further person P’, which embraces also a group of further persons, for processing the respective received human voice signals 9, 9′ differently, [0065] the processing device 11 or a suppressing device 13 assignable or assigned to the processing device 11 may be further configured to suppress the received human voice signal 9′ of the at least one further person P′ located in the vehicle cabin 4 and to generate a resulting received human voice signal which contains (only) the human voice signal 9 of the at least one first person P and a suppressed human voice signal 9′ of the at least one further person P′…Suppressing the received human voice signal 9′ of the at least one further person P′ may require separating the human voice signals 9′ of the at least one further person P′ from the human voice signals 9 of the at least one first person P or vice versa, [Wherein second audio signals are those which have been spatially separated into sound sources. In view of the multiple, separated human voice signals of Kotulla, the examiner asserts that suppressing an undesired voice signal (with respect to a desired voice signal) to generate a resulting human voice signal indicates the combined human voice signal to be a fifth audio signal, wherein the multiple speakers are superposed before separation/suppression. Further, it is unclear what is superposed with one second audio signal. There appears to be nothing to superpose. The examiner asserts that there must be something to superpose the second audio onto; therefore, signal superposing can also just be gathering a second, separated audio signal as currently claimed, superposed with nothing]); and
performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal ([0054] the combined audio signal allows for simultaneously outputting the audio signal 3, e.g. a musical piece, and the received human voice signal 9, e.g. the voice of a person P singing along the audio signal 3, [Wherein the combined audio signal (containing a combination of singing and music) tracks to a third audio signal representing mixing of noise suppressed voice signals (fifth audio signal) with a music track (preset signal)]).
Tan and Kotulla are considered analogous art within adaptive audio mixing within vehicles. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan to incorporate the teachings of Kotulla, because of the novel way to adaptively combine human voice signals with other human voice signals and/or non-human music overlay tracks, improving implementation of special operational modes such as karaoke in vehicles without requiring external hardware which needs to be set up in the cabin of the vehicle (Kotulla, [0004]).
Regarding claim 9, Tan in view of Kotulla discloses: the method according to claim 8.
Kotulla further discloses:
wherein before the performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal, the method further comprises:
performing audio effect processing on the fifth audio signal to obtain a sixth audio signal ([0064] The audio receiving device 8 may be configured to receive a human voice signal 9 of at least one first person P located in the vehicle cabin 4 and a human voice signal 9′ of at least one further person P′ located in the vehicle cabin 4. Thereby, the processing device 11 may be further configured to modify the received human voice signal 9 of the at least one first person P located in the vehicle cabin 4 with at least one first acoustic modification parameter and to modify the received human voice signal 9′ of the at least one further person P’ located in the vehicle cabin 4 with at least one further acoustic modification parameter. It may be required to separate the received human voice signal 9 of at least one first person P, which embraces also a group of first persons, from the received human voice signal of at least one further person P’, which embraces also a group of further persons, for processing the respective received human voice signals 9, 9′ differently, [Modifying voice signals based on acoustic modification parameters indicates the signals output from said modification is a sixth voice signal, i.e. based on the noise cancellation performed after determining the signals to be suppressed from a combined, e.g. superposed/fifth, signal. Further, considering the superposition of second audio is with respect to nothing, the claim essentially performs the audio effect processing on second audio signals, i.e. those which have been spatially separated (clearly indicated through locations of P and P’)]); and
the performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal comprises:
performing audio mixing processing on the sixth audio signal and the preset signal to obtain the third audio signal ([0054] the combined audio signal allows for simultaneously outputting the audio signal 3, e.g. a musical piece, and the received human voice signal 9, e.g. the voice of a person P singing along the audio signal 3, [Wherein the combined audio signal (containing a combination of singing and music) tracks to a third audio signal representing mixing of noise suppressed voice signals (sixth audio signal) with a music track (preset signal)]).
Regarding claim 10, Tan in view of Kotulla discloses: the method according to claim 8.
Tan further discloses:
wherein the acquiring a plurality of first audio signals in a space of a mobile terminal comprises:
acquiring sound signals at a plurality of positions in the space of the mobile terminal through a plurality of transducers to obtain the plurality of first audio signals ([Fig. 2, Mobile phone and Vehicle-mounted App music into Audio decoding], [Passing audio from two different sources into audio decoding indicates at least two first audio signals in the space of a mobile terminal, necessarily acquired by the decoder]).
Regarding claim 11, Tan in view of Kotulla discloses: the method according to claim 8.
Tan further discloses:
wherein the acquiring a plurality of first audio signals in a space of a mobile terminal comprises:
acquiring a plurality of seventh audio signals in the space of the mobile terminal ([Fig. 2, Mobile phone and Vehicle-mounted App music into Audio decoding], [Passing audio from two different sources into audio decoding indicates at least two seventh audio signals in the space of a mobile terminal, necessarily acquired by the decoder]).
Kotulla further discloses:
eliminating interference signals in the plurality of seventh audio signals, respectively, to obtain the plurality of first audio signals ([0066] the processing device 11 or a suppressing device 14 assignable or assigned to the processing device 11 may be configured to suppress the noise acoustically perceivable in the vehicle cabin 4 and to generate a resulting received human voice signal which contains the human voice signal 9 of the person P and suppressed noise, [The examiner asserts that suppressing noise is eliminating interference, in view of the plurality of input signals of Tan]).
Regarding claim 20, Tan discloses: the non-transitory computer readable storage medium according to claim 13.
Tan does not disclose:
wherein the performing audio mixing based on the at least one second audio signal to obtain a third audio signal comprises:
performing signal superposition on the at least one second audio signal to obtain a fifth audio signal; and
performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal.
Kotulla discloses:
wherein the performing audio mixing based on the at least one second audio signal to obtain a third audio signal comprises:
performing signal superposition on the at least one second audio signal to obtain a fifth audio signal ([0050] Each audio outputting element 7 may be assigned to a specific location or space, i.e. particularly to a specific seat (not shown), in the vehicle cabin 4, [0064] The audio receiving device 8 may be configured to receive a human voice signal 9 of at least one first person P located in the vehicle cabin 4 and a human voice signal 9′ of at least one further person P′ located in the vehicle cabin 4. Thereby, the processing device 11 may be further configured to modify the received human voice signal 9 of the at least one first person P located in the vehicle cabin 4 with at least one first acoustic modification parameter and to modify the received human voice signal 9′ of the at least one further person P’ located in the vehicle cabin 4 with at least one further acoustic modification parameter. It may be required to separate the received human voice signal 9 of at least one first person P, which embraces also a group of first persons, from the received human voice signal of at least one further person P’, which embraces also a group of further persons, for processing the respective received human voice signals 9, 9′ differently, [0065] the processing device 11 or a suppressing device 13 assignable or assigned to the processing device 11 may be further configured to suppress the received human voice signal 9′ of the at least one further person P′ located in the vehicle cabin 4 and to generate a resulting received human voice signal which contains (only) the human voice signal 9 of the at least one first person P and a suppressed human voice signal 9′ of the at least one further person P′…Suppressing the received human voice signal 9′ of the at least one further person P′ may require separating the human voice signals 9′ of the at least one further person P′ from the human voice signals 9 of the at least one first person P or vice versa, [Wherein second audio signals are those which have been spatially separated into sound sources. In view of the multiple, separated human voice signals of Kotulla, the examiner asserts that suppressing an undesired voice signal (with respect to a desired voice signal) to generate a resulting human voice signal indicates the combined human voice signal to be a fifth audio signal, wherein the multiple speakers are superposed before separation/suppression. Further, it is unclear what is superposed with one second audio signal. There appears to be nothing to superpose. The examiner asserts that there must be something to superpose the second audio onto; therefore, signal superposing can also just be gathering a second, separated audio signal as currently claimed, superposed with nothing]); and
performing audio mixing processing on the fifth audio signal and a preset signal to obtain the third audio signal ([0054] the combined audio signal allows for simultaneously outputting the audio signal 3, e.g. a musical piece, and the received human voice signal 9, e.g. the voice of a person P singing along the audio signal 3, [Wherein the combined audio signal (containing a combination of singing and music) tracks to a third audio signal representing mixing of noise suppressed voice signals (fifth audio signal) with a music track (preset signal)]).
Tan and Kotulla are considered analogous art within adaptive audio mixing within vehicles. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Tan to incorporate the teachings of Kotulla, because of the novel way to adaptively combine human voice signals with other human voice signals and/or non-human music overlay tracks, improving implementation of special operational modes such as karaoke in vehicles without requiring external hardware which needs to be set up in the cabin of the vehicle (Kotulla, [0004]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Willis et al. (US-20220210593-A1) discloses “In some embodiments, methods and systems for combining a prerecorded performance with a live performance within a space (such as a vehicle) are disclosed. A first performance (which may be a prerecorded performance) may be started within the space, and a second performance (such as a live performance) may subsequently be initiated within the space as well. The second performance may be detected, and based upon the detection, one or more volume parameters of the first performance may be modified. One or more properties of the second performance may also be modified, and the modified first performance may then be combined with the modified second performance to create a combined performance” (abstract). See entire document.
Huang et al. (US-20240137721-A1) discloses “A sound-making apparatus control method includes that a first device obtains position information of a plurality of areas in which a plurality of users is located. The first device controls, based on the position information of the plurality of areas and position information of a plurality of sound-making apparatuses, the plurality of sound-making apparatuses to work” (abstract). See entire document.
Sporer et al. (US-20220159403-A1) discloses “A system and a corresponding method for assisting selective hearing are provided. The system includes a detector for detecting an audio source signal portion of one or more audio sources by using at least two received microphone signals of a hearing environment. In addition, the system includes a position determiner for allocating position information to each of the one or more audio sources. In addition, the system includes an audio type classifier for assigning an audio source signal type to the audio source signal portion of each of the one or more audio sources. In addition, the system includes a signal portion modifier for varying the audio source signal portion of at least one audio source of the one or more audio sources depending on the audio signal type of the audio source signal portion of the at least one audio source so as to obtain a modified audio signal portion of the at least one audio source. In addition, the system includes a signal generator” (abstract). See entire document.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THEODORE JOHN WITHEY whose telephone number is (703)756-1754. The examiner can normally be reached Monday - Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571) 272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THEODORE WITHEY/Examiner, Art Unit 2655
/ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655