Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This is in response to Applicants Request for Reconsideration filed 11/17/25 which has been entered. Claims 1-12 and 14-18 have been amended. Claim 13 has been cancelled. Claims 31-33 have been added. Claims 1-12, 14-18, and 31-33 are still pending in this application, with Claims 1, 7, and 31 being independent.
Claim Objections
Claim 2 is objected to because of the following informalities: Claim 2 recites the limitation “identify the individual sound source using where the mute function users one or more…” The language is confusing and not right. Examiner interprets as identify the individual sound source using one or more. Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-12, 14-18, and 31-33 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Referring to claims 1-12, 14-18, and 31-33, claim 1, 7, and 31 recite the limitation “the plurality of separated audio signals”. There is insufficient antecedent basis for this limitation in the claims. Examiner interprets as delay an input audio signal from the plurality of input audio signals to time-align with a separated audio signal corresponding to the individual sound source. Claims 2-6 and 23-33 depend from claim 1, claims 8-12 depend from claim 7, and claims 14-18 depend from claim 31, therefore, they are rejected for the same reasons.
Referring to claim 15, claim 15 recites the limitation “the individual talkers”. There is insufficient antecedent basis for this limitation in the claim. The claim has also had new language added but just tacked onto the beginning of the old language and it does not really make sense. Examiner interprets as if all the old language “claim according…the individual talkers” is deleted.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3-4, 6-7, 9-10, 12, 15-16, 18, and 31-33 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mauchly et al. US Publication No. 20130044893 (from IDS) in view of Visser et al. US Publication No. 20080208538.
Referring to claim 1, Mauchly et al. teaches an apparatus comprising:
a processor that executes instructions stored in a memory to configure the processor (Fig. 3: processor 32, memory 34; para 0031: “The processor 32 includes an audio processor 37 configured to process audio to remove sound from a sound source. The audio processor 37 is operable, for example, to process audio signals, determine the direction of a sound, separate out different signals (voices), and subtract (cancel, filter) signals of a muted sound source, as described in detail below. In one embodiment, the audio processor 37 first digitizes the sound received from all of the microphones”):
receive, via a microphone array comprising a plurality of microphones, a plurality of input audio signals (Fig. 2: microphones 24; para 0036: “the processing system 31 receives audio from a plurality of microphones 24”) from a plurality of sound sources (para 0031: “different signals (voices)”);
using a mute function, identify in real time an individual sound source among the plurality of sound sources (abstract: “identifying a sound source to be muted, processing the audio to remove sound received from the sound source at each of the microphones”; para 0038: “Real-Time Face Detection”; para 0031: “The audio processor 37 is operable, for example, to process audio signals, determine the direction of a sound, separate out different signals (voices)”);
mute the individual sound source (para 0036: “One or more participants may select a mute option to remove their voice from an audio output”);
in response to the individual sound source being muted, subtract the separated audio signal from the input audio signal to generate an output audio signal (para 0036: “The processing system outputs an audio signal or signals which contain the summed sound of the individual speakers in the room, minus sound from the muted sound source.”).
However, Mauchly et al. does not teach delaying a signal, but Visser et al. teaches delay an input audio signal from the plurality of input audio signals to time-align with a separated audio signal, from the plurality of separated audio signals, corresponding to the individual sound source (Fig. 19B: input channel I2a is delayed to time align with source separated signal input channel I1f). It would have been obvious to one having ordinary skill in the art before the effective filing date to delay an input audio signal, as taught in Visser et al., in the apparatus of Mauchly et al. because the delay is “provided to compensate for processing delay of the corresponding source separator (e.g., to synchronize the input channels of the subsequent stage)”.
Referring to claim 3, Mauchly et al. teaches receive a video image from a video camera; and identify the individual sound source based on the video image (para 0022: “the camera 25 may be used to track participants in the room for use in identifying a sound to be muted.”; para 0037: “The camera 25 is used to track the person in the room (step 52). The system associates the person (visually detected face or face and body) with a voice (audio detected sound source)”).
Referring to claim 4, Mauchly et al. teaches receive, via a user interface, an input from a user to mute or unmute the individual sound source (para 0023: “a user interface 26 for use in selecting a mute option”; para 0024: “The user interface 26 may be associated with one of the microphones 24, a zone 28, or a user 20”).
Referring to claim 6, Mauchly et al. teaches execute a speaker separation function on the plurality of input audio signals to separate the plurality of sound sources and generate a plurality of separated audio signals, each corresponding to a respective individual sound source (para 0031: “separate out different signals (voices)”).
Referring to claim 7, Mauchly et al. teaches a method comprising:
by a processor (Fig. 3: processor 32; para 0031: “The processor 32 includes an audio processor 37 configured to process audio to remove sound from a sound source. The audio processor 37 is operable, for example, to process audio signals, determine the direction of a sound, separate out different signals (voices), and subtract (cancel, filter) signals of a muted sound source, as described in detail below. In one embodiment, the audio processor 37 first digitizes the sound received from all of the microphones”), receiving, via a microphone array comprising a plurality of microphones, a plurality of input audio signals (Fig. 2: microphones 24; para 0036: “the processing system 31 receives audio from a plurality of microphones 24”) from a plurality of sound sources (para 0031: “different signals (voices)”);
by the processor, identifying an individual sound source among the plurality of sound sources in real time using a mute function (abstract: “identifying a sound source to be muted, processing the audio to remove sound received from the sound source at each of the microphones”; para 0038: “Real-Time Face Detection”; para 0031: “The audio processor 37 is operable, for example, to process audio signals, determine the direction of a sound, separate out different signals (voices)”);
by the processor, muting the individual sound source (para 0036: “One or more participants may select a mute option to remove their voice from an audio output”);
in response to the individual sound source being muted, subtracting, by the processor, the separated audio signal from the input audio signal to generate an output audio signal (para 0036: “The processing system outputs an audio signal or signals which contain the summed sound of the individual speakers in the room, minus sound from the muted sound source.”).
However, Mauchly et al. does not teach delaying a signal, but Visser et al. teaches by the processor, delaying an input audio signal from the plurality of input audio signals to time-align with a separated audio signal, from the plurality of separated audio signals, corresponding to the individual sound source (Fig. 19B: input channel I2a is delayed to time align with source separated signal input channel I1f). It would have been obvious to one having ordinary skill in the art before the effective filing date to delay an input audio signal, as taught in Visser et al., in the method of Mauchly et al. because the delay is “provided to compensate for processing delay of the corresponding source separator (e.g., to synchronize the input channels of the subsequent stage)”.
Referring to claim 9, Mauchly et al. teaches receiving a video image from a video camera; and identifying the individual sound source based on the video image (para 0022: “the camera 25 may be used to track participants in the room for use in identifying a sound to be muted.”; para 0037: “The camera 25 is used to track the person in the room (step 52). The system associates the person (visually detected face or face and body) with a voice (audio detected sound source)”).
Referring to claim 10, Mauchly et al. teaches receiving, via a user interface, an input from a user to selectively mute or unmute the individual sound source (para 0023: “a user interface 26 for use in selecting a mute option”; para 0024: “The user interface 26 may be associated with one of the microphones 24, a zone 28, or a user 20”).
Referring to claim 12, Mauchly et al. teaches executing a speaker separation function on the plurality of input audio signals to separate the plurality of sound sources and generate a plurality of separated audio signals, each corresponding to a respective individual sound source (para 0031: “separate out different signals (voices)”).
Referring to claim 15, Mauchly et al. teaches receiving a video image from a video camera; and identifying the individual sound source based on the video image claim according to [Claim 13] that further includes at least one video camera that uses facial recognition and/or mouth-movement detection to assist in the learning and identifying the individual talkers (para 0022: “the camera 25 may be used to track participants in the room for use in identifying a sound to be muted.”; para 0037: “The camera 25 is used to track the person in the room (step 52). The system associates the person (visually detected face or face and body) with a voice (audio detected sound source)”).
Referring to claim 16, Mauchly et al. teaches receiving, via a user interface, an input from a user to selectively mute or unmute the individual sound source (para 0023: “a user interface 26 for use in selecting a mute option”; para 0024: “The user interface 26 may be associated with one of the microphones 24, a zone 28, or a user 20”).
Referring to claim 18, Mauchly et al. teaches executing a speaker separation function on the plurality of input audio signals to separate the plurality of sound sources and generate a plurality of separated audio signals, each corresponding to a respective individual sound source (para 0031: “separate out different signals (voices)”).
Referring to claim 31, Mauchly et al. teaches a non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform (Fig. 3: processor 32; para 0031: “The processor 32 includes an audio processor 37 configured to process audio to remove sound from a sound source. The audio processor 37 is operable, for example, to process audio signals, determine the direction of a sound, separate out different signals (voices), and subtract (cancel, filter) signals of a muted sound source, as described in detail below. In one embodiment, the audio processor 37 first digitizes the sound received from all of the microphones”):
receiving, via a microphone array comprising a plurality of microphones, a plurality of input audio signals (Fig. 2: microphones 24; para 0036: “the processing system 31 receives audio from a plurality of microphones 24”) from a plurality of sound sources (para 0031: “different signals (voices)”);
identifying an individual sound source among the plurality of sound sources in real time using a mute function (abstract: “identifying a sound source to be muted, processing the audio to remove sound received from the sound source at each of the microphones”; para 0038: “Real-Time Face Detection”; para 0031: “The audio processor 37 is operable, for example, to process audio signals, determine the direction of a sound, separate out different signals (voices)”);
muting the individual sound source (para 0036: “One or more participants may select a mute option to remove their voice from an audio output”);
in response to the individual sound source being muted, subtracting the separated audio signal from the input audio signal to generate an output audio signal (para 0036: “The processing system outputs an audio signal or signals which contain the summed sound of the individual speakers in the room, minus sound from the muted sound source.”).
However, Mauchly et al. does not teach delaying a signal., but Visser et al. teaches delaying an input audio signal from the plurality of input audio signals to time-align with a separated audio signal, from the plurality of separated audio signals, corresponding to the individual sound source (Fig. 19B: input channel I2a is delayed to time align with source separated signal input channel I1f). It would have been obvious to one having ordinary skill in the art before the effective filing date to delay an input audio signal, as taught in Visser et al., in the medium of Mauchly et al. because the delay is “provided to compensate for processing delay of the corresponding source separator (e.g., to synchronize the input channels of the subsequent stage)”.
Referring to claim 32, Mauchly et al. teaches the processor is configured to: remove contributions of the individual sound source from others of the plurality of input audio signals when the individual sound source is muted (para 0040: “Sound received at the microphone 24 is input for use in cancelling sound received from the sound source at other microphones”).
Referring to claim 33, Mauchly et al. teaches the processor is configured to: identify an exclusion zone based on a range of directions estimated for the plurality of sound sources; and identify that the individual sound source is in the exclusion zone, and when the processor selectively mutes the individual sound source, the processor is configured to: mute the individual sound source based on the individual sound source being in the exclusion zone (para 0025: “Muting of the microphone 24 prevents sound from a sound source (e.g., participant 20 or participants within zone 28) from being transmitted”; para 0026: “the processing system can use a two-dimensional overhead mapping of a location of the sound source relative to the microphones for use in identifying a sound received from the sound source. As described below, the processing system is coupled to the microphones 24 and may be configured to generate audio data and direction information indicative of the direction of sound received at the microphones”; para 0037: “the audio is processed to generate audio data and direction information indicative of the direction of sound received at the microphones 24. The location of a person 20 may be mapped relative to the microphones 24 and the approximate distance and angle from the microphones used to identify sound received from the person (step 54). The sound that is identified as coming from the person is separated from the other audio received at the microphones and rejected”).
Claim(s) 2, 5, 8, 11, 14, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mauchly et al. and Visser et al., as applied to claims 1, 7, and 31 above, and further in view of Sommers et al. US Publication No. 20180047395.
Referring to claim 2, Mauchly et al. teaches the mute function (para 0036), however, Mauchly et al. and Visser et al. do not teach machine learning per se, but Sommers et al. teaches identify the individual sound source using where the function uses one or more of: artificial intelligence, machine learning, or deep learning (para 0093). It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to use machine learning, as taught in Sommers et al., in the apparatus of Mauchly et al. and Visser et al. because it helps to more accurately identify the speakers.
Referring to claim 5, Mauchly et al. and Visser et al. do not teach diarization per se, but Sommers et al. teaches execute a diarization function on the plurality of input audio signals to identify the individual sound source (para 0093). It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to use a diarization function, as taught in Sommers et al., in the apparatus of Mauchly et al. and Visser et al. because it helps to more accurately identify the speakers.
Referring to claim 8, Mauchly et al. teaches the mute function (para 0036), however, Mauchly et al. and Visser et al. do not teach machine learning per se, but Sommers et al. teaches identifying the individual sound source using one or more of: artificial intelligence, machine learning, or deep learning (para 0093). It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to use machine learning, as taught in Sommers et al., in the method of Mauchly et al. and Visser et al. because it helps to more accurately identify the speakers.
Referring to claim 11, Mauchly et al. and Visser et al. do not teach diarization per se, but Sommers et al. teaches executing a diarization function on the plurality of input audio signals to identify the individual sound source (para 0093). It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to use a diarization function, as taught in Sommers et al., in the method of Mauchly et al. and Visser et al. because it helps to more accurately identify the speakers.
Referring to claim 14, Mauchly et al. teaches the mute function (para 0036), however, Mauchly et al. and Visser et al. do not teach machine learning per se, but Sommers et al. teaches identifying the individual sound source using one or more of: artificial intelligence, machine learning, or deep learning (para 0093). It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to use machine learning, as taught in Sommers et al., in the method of Mauchly et al. and Visser et al. because it helps to more accurately identify the speakers.
Referring to claim 17, Mauchly et al. and Visser et al. do not teach diarization per se, but Sommers et al. teaches executing a diarization function on the plurality of input audio signals to identify the individual sound source (para 0093). It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to use a diarization function, as taught in Sommers et al., in the method of Mauchly et al. and Visser et al. because it helps to more accurately identify the speakers.
Response to Arguments
Applicant’s arguments with respect to the claims have been considered but are moot because the new ground of rejection does not rely on the combination of references applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Examiner respectfully requests, in response to this Office Action, support be shown for language added to any original claims on amendment and any new claims. That is, indicate support for newly added claim language by specifically pointing to page(s) and line number(s) in the specification and/or drawing figure(s). This will assist Examiner in prosecuting the application.
When responding to this Office Action, Applicant is advised to clearly point out the patentable novelty which he or she thinks the claims present, in view of the state of the art disclosed by the references cited or the objections made. He or she must also show how the amendments avoid such references or objections. See 37 CFR 1.111(c).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KATHERINE A FALEY whose telephone number is (571)272-3453. The examiner can normally be reached on Monday to Wednesday, 9am-5pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ahmad Matar can be reached on (571)272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Any response to this action should be mailed to:
Commissioner of Patents and Trademarks
P.O. Box 1450
Alexandria, Va. 22313-1450
Or faxed to:
(571) 273-8300, for formal communications intended for entry and for
informal or draft communications, please label “PROPOSED” or “DRAFT”.
Hand-delivered responses should be brought to:
Customer Service Window
Randolph Building
401 Dulany Street
Arlington, VA 22314
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KATHERINE A FALEY/Primary Examiner, Art Unit 2693