Prosecution Insights
Last updated: October 02, 2026
Application No. 19/059,802

DEVICE FOR RECOGNIZING MULTI-CHANNEL INPUT VOICE INDEPENDENT ON MICROPHONE ARRAY FORM AND LEARNING METHOD THEREOF

Non-Final OA §103
Filed
Feb 21, 2025
Priority
Mar 13, 2024 — RE 10-2024-0035056
Examiner
LEE, EUNICE SOMIN
Art Unit
Tech Center
Assignee
Electronics and Telecommunications Research Institute
OA Round
1 (Non-Final)
89%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
40 granted / 45 resolved
+28.9% vs TC avg
Strong +26% interview lift
Without
With
+25.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
16 currently pending
Career history
57
Total Applications
across all art units

Statute-Specific Performance

§101
20.7%
-19.3% vs TC avg
§103
67.7%
+27.7% vs TC avg
§102
6.5%
-33.5% vs TC avg
§112
1.9%
-38.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 45 resolved cases

Office Action

§103
DETAILED ACTION This communication is in response to the Application filed on February 21, 2025. Claims 1 - 20 are pending and have been examined. Claims 1 and 11 are independent. Foreign priority: March 13, 2024. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on February 21, 2025 and December 9, 2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements were considered by the examiner. Drawings The drawings filed on February 21, 2025 have been accepted and considered by the Examiner. Specification Applicant is reminded of the proper language and format for an abstract of the disclosure. The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words. The form and legal phraseology often used in patent claims, such as "means" and "said," should be avoided. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details. The abstract of the disclosure is objected to because the abstract of 151 words exceeds 150 words. Correction is required. See MPEP § 608.01(b). 35 U.S.C. 112(f) Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “noise mask estimator”, “beamformer estimator”, and “learning machine” in Claim 1, “microphone array selector” in Claim 2, “end-to-end voice recognition model” in Claim 6 - 8. Note the varied definition of this phrase in the supporting Specification which indicates that the “estimator”, “machine”, “selector”. “model” was intended as a generic placeholder. These limitations are generic in the context of the art and don’t refer to any specific structure and only serve as placeholders for the structure that performs the associated function(s) without providing any information about what that structure is. MPEP 2181 I A says: For a term to be considered a substitute for "means," and lack sufficient structure for performing the function, it must serve as a generic placeholder and thus not limit the scope of the claim to any specific manner or structure for performing the claimed function. It is important to remember that there are no absolutes in the determination of terms used as a substitute for "means" that serve as generic placeholders. The examiner must carefully consider the term in light of the specification and the commonly accepted meaning in the technological art. Every application will turn on its own facts. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. Claims 1, 6, 11 - 12 and 17 are rejected under 35 U.S.C. 103(a) as being unpatentable over Eskimez et al., (U.S. Patent Application Publication 2023/0116052), hereinafter referred to as Eskimez, in view of Yoshioka et al., (“VarArray: Array-geometry-agnostic continuous speech separation," 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022), hereinafter referred to as Yoshioka. Regarding Claims 1 and 11, Eskimez teaches: 1. A device for recognizing a multi-channel input voice independent on a microphone array form, the device, and 11. A learning method for multi-channel input voice recognition, the method performed by a device for recognizing a multi-channel input voice independent on a microphone array form comprising: a time-frequency transformer configured to receive a plurality of channel audio signals extracted from voice data recorded through a plurality of microphones having unspecified microphone array forms and to transform the plurality of channel audio signals into a plurality of time-frequency domain signals; [Eskimez, “Aspects of the disclosure describe both (i) array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE),” Par. 0024; “Input audio 112 is provided to an STFT block 302 (i.e., the claimed “time-frequency transformer”), Par. 0042; Figure 3 clearly shows “Short Time Fourier Transform (i.e., the claimed “time-frequency transformer”), 302.”; “In some examples, producing output data 114 using trained PSE model 110 comprises isolating speech data (i.e., the claimed “voice data”) of target speaker 102 in a manner that is agnostic of a configuration of microphone array 200 (i.e., the claimed “plurality of microphones having unspecified microphone array forms”).” Par. 0100] a speaker and noise mask estimator configured to receive the plurality of time- frequency domain signals and to estimate a time-frequency domain mask for voices and noise for a plurality of speakers; [Eskimez, “Aspects of the disclosure describe both (i) array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE),” Par. 0024; “The disclosure below extends PSE to utilize the microphone arrays for environments where strong noise, reverberation, and an interfering speaker are present.” Par. 0056; “The estimated mask (i.e., the claimed “speaker and noise mask estimator”) is applied to the first microphone in both approaches.” Par. 0061; “The model outputs a complex ratio mask which is multiplied with the input mixture to estimate the clean speech (i.e., the claimed “speaker and noise mask estimator”).” Par. 0045] a beamformer estimator configured to estimate a time-frequency domain signal for voice signals of the plurality of speakers from which the noise has been removed from the plurality of time-frequency domain signals by using the time-frequency domain mask; [Eskimez, “Multi-channel PSE uses inputs from multiple microphones, and geometry agnostic operation indicates that the PSE solution does not need to know the positions of the microphones in order to reduce or eliminate non-speech noise and speech data of an interfering speaker.” Par. 0027; “Aspects of the disclosure are also operable with MVDR beamforming followed by a single-channel PSE.” Par. 0084] a time-frequency inverse transformer configured to inversely transform the time- frequency domain signal from which the noise has been removed into a time domain signal; and [Eskimez, “fed into an inverse STFT (i.e., the claimed “time-frequency inverse transformer”) block 618 to produce output data 114,” Par. 0067; Figure 3 clearly shows “Inverse Short Time Fourier Transform (i.e., the claimed “time-frequency inverse transformer”), 322.”] a learning machine configured to train the speaker and noise mask estimator based on a loss function obtained based on results of a comparison between the inversely-transformed time domain signal and a pre-defined answer signal. [Eskimez, “PSE model (i.e., the claimed “learning machine”) is able to implicitly learn the spectral and spatial information from the microphone array.” Par. 0059; “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array, producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045] Eskimez fails to explicitly teach beamformer estimator. a time-frequency transformer configured to receive a plurality of channel audio signals extracted from voice data recorded through a plurality of microphones having unspecified microphone array forms and to transform the plurality of channel audio signals into a plurality of time-frequency domain signals; [Yoshioka, “The proposed model is applicable to any number of microphones (i.e., the claimed “plurality of microphones having unspecified microphone array forms”) without retraining,” Pg 6027; “Let Xmft denote the short Fourier transform (STFT) coefficients of the audio signal observed by the mth microphone (i.e., the claimed “plurality of microphones having unspecified microphone array forms”), where f and t represent the frequency and time indices (i.e., the claimed “time-frequency domain signals”), respectively.” Pg. 6028; “expose the model to a variety of array geometries (i.e., the claimed “plurality of microphones having unspecified microphone array forms”) and inter-microphone spacing patterns.” Pg. 6029] a speaker and noise mask estimator configured to receive the plurality of time- frequency domain signals and to estimate a time-frequency domain mask for voices and noise for a plurality of speakers; [Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “plurality of time-frequency domain signals”) for accurate TF (i.e., the claimed “time-frequency”) mask estimation.” Pg. 6028; “The objective is to construct a speech separation model that receives the feature sequency (zmt) from all the input channels and generates TF mask Msft (i.e., the claimed “time-frequency domain mask”),” Pg. 6028] a beamformer estimator configured to estimate a time-frequency domain signal for voice signals of the plurality of speakers from which the noise has been removed from the plurality of time-frequency domain signals by using the time-frequency domain mask; [Yoshioka, “The masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028; “The first two sources correspond to the two speakers (i.e., the claimed “voice signals of the plurality of speakers”).” Pg. 6028] a learning machine configured to train the speaker and noise mask estimator based on a loss function obtained based on results of a comparison between the inversely-transformed time domain signal and a pre-defined answer signal. [Yoshioka, “Results of fine-tuning the speech separation model to AMI with an ASR based loss function are also presented,” Pg. 6027; “During training, we minimized the uPIT style cross entropy loss between two predicted hypotheses and reference transcriptions (i.e., the claimed “comparison between the inversely transformed time domain signal and a pre-defined answer signal”), Pg. 6029] Eskimez and Yoshioka pertain to microphone geometry agnostic systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the microphone geometry agnostic systems art to modify Eskimez’s teachings of “array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE)” (Eskimez, Par. 0024) with the explicit teachings of “masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”)” (Yoshioka, Pg. 6028) taught by Yoshioka in order to “improve ASR accuracy” (Yoshioka, Pg. 6028). Regarding Claims 6 and 17, Eskimez in view of Yoshioka has been discussed above. The combination further teaches: further comprising an end-to-end voice recognition model configured to receive a time-frequency domain signal from which noise has been removed, which is obtained through the time-frequency transformer, the speaker and noise mask estimator, and the beamformer estimator and to output results of voice recognition, based on voice data recorded through a plurality of microphones having specific microphone array forms. [Yoshioka, “3.3. End-to-end optimization (i.e., the claimed “end-to-end voice recognition model”), Pg. 6029; “E2E model (i.e., the claimed “end-to-end voice recognition model”),” Pg. 6029; Eskimez, “Aspects of the disclosure describe both (i) array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE),” Par. 0024; “Input audio 112 is provided to an STFT block 302 (i.e., the claimed “time-frequency transformer”), Par. 0042; Figure 3 clearly shows “Short Time Fourier Transform (i.e., the claimed “time-frequency transformer”), 302.”; “In some examples, producing output data 114 using trained PSE model 110 comprises isolating speech data (i.e., the claimed “voice data”) of target speaker 102 in a manner that is agnostic of a configuration of microphone array 200 (i.e., the claimed “plurality of microphones having unspecified microphone array forms”).” Par. 0100; “Aspects of the disclosure describe both (i) array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE),” Par. 0024; “The disclosure below extends PSE to utilize the microphone arrays for environments where strong noise, reverberation, and an interfering speaker are present.” Par. 0056; “The estimated mask (i.e., the claimed “speaker and noise mask estimator”) is applied to the first microphone in both approaches.” Par. 0061; “fed into an inverse STFT (i.e., the claimed “time-frequency inverse transformer”) block 618 to produce output data 114,” Par. 0067; Figure 3 clearly shows “Inverse Short Time Fourier Transform (i.e., the claimed “time-frequency inverse transformer”), 322.”; “PSE model (i.e., the claimed “learning machine”) is able to implicitly learn the spectral and spatial information from the microphone array.” Par. 0059; “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array, producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045; Yoshioka, “The proposed model is applicable to any number of microphones (i.e., the claimed “plurality of microphones having unspecified microphone array forms”) without retraining,” Pg 6027; “Let Xmft denote the short Fourier transform (STFT) coefficients of the audio signal observed by the mth microphone (i.e., the claimed “plurality of microphones having unspecified microphone array forms”), where f and t represent the frequency and time indices (i.e., the claimed “time-frequency domain signals”), respectively.” Pg. 6028; “expose the model to a variety of array geometries (i.e., the claimed “plurality of microphones having unspecified microphone array forms”) and inter-microphone spacing patterns.” Pg. 6029; Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “plurality of time-frequency domain signals”) for accurate TF (i.e., the claimed “time-frequency”) mask estimation.” Pg. 6028; “The objective is to construct a speech separation model that receives the feature sequency (zmt) from all the input channels and generates TF mask Msft (i.e., the claimed “time-frequency domain mask”),” Pg. 6028; “The masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028; “The first two sources correspond to the two speakers (i.e., the claimed “voice signals of the plurality of speakers”).” Pg. 6028; “Results of fine-tuning the speech separation model to AMI with an ASR based loss function are also presented,” Pg. 6027; “During training, we minimized the uPIT style cross entropy loss between two predicted hypotheses and reference transcriptions (i.e., the claimed “comparison between the inversely transformed time domain signal and a pre-defined answer signal”), Pg. 6029] Regarding Claim 12, Eskimez in view of Yoshioka has been discussed above. The combination further teaches: wherein the extracting of the plurality of channel audio signals from the voice data recorded through the plurality of microphones having the unspecified microphone array forms and transforming the plurality of channel audio signals into the plurality of time-frequency domain signals comprises: [Eskimez, see mapping applied to claim 11; Yoshioka, see mapping applied to claim 11] selecting and receiving the voice data recorded through the plurality of microphones having the unspecified microphone array forms; [Eskimez, see mapping applied to claim 11; Yoshioka, see mapping applied to claim 11] extracting a plurality of channel audio signals corresponding to the plurality of microphones from the voice data; and [Eskimez, see mapping applied to claim 11; Yoshioka, see mapping applied to claim 11] transforming the plurality of channel audio signals into a plurality of time- frequency domain signals. [Eskimez, see mapping applied to claim 11; Yoshioka, see mapping applied to claim 11] Claims 2 and 13 are rejected under 35 U.S.C. 103(a) as being unpatentable over Eskimez in view of Yoshioka as applied in claim 1 above, and in further view of Shumard et al., (U.S. Patent Application Publication 2021/0058702), hereinafter referred to as Shumard. Regarding Claims 2 and 13, Eskimez in view of Yoshioka has been discussed above. The combination further teaches: among voice data recorded through microphones having a plurality of microphone array forms, which are mounted on various types of devices. [Eskimez, In some examples, producing output data 114 using trained PSE model 110 comprises isolating speech data (i.e., the claimed “voice data”) of target speaker 102 in a manner that is agnostic of a configuration of microphone array 200 (i.e., the claimed “microphones having a plurality of microphone array forms”).” Par. 0100; “multiple microphone array devices (i.e., the claimed “plurality of microphone array forms”) without needing to train the model for each user device (i.e., the claimed “various types of devices”) model (e.g., different models of notebook computers (i.e., the claimed “various types of devices”), with different numbers and/or placement of microphones).” Par. 0057] The combination fails to teach selector. However, Shumard teaches: further comprising a microphone array selector configured to select and receive the voice data recorded through a microphone having any one microphone array form, [Shumard, “Array microphones can have different configurations (i.e., the claimed microphone having any one microphone array form”) and frequency responses depending on the placement of the microphones relative to each other and the direction of arrival for sound waves.” Par. 0007; “That is, the array microphone 100 may be agnostic (i.e., the claimed microphone having any one microphone array form”) to the direction of arrival within the x-y plane.” Par. 0053; “plurality of microphones and based thereon, generate an array output with a directional polar pattern that is selected (i.e., the claimed “microphone array selector”) based on a direction of arrival of the audio signals,” Par. 0016] among voice data recorded through microphones having a plurality of microphone array forms, which are mounted on various types of devices. [Shumard, “Thus, microphones are available in a variety of sizes, form factors (i.e., the claimed “plurality of microphone array forms”), mounting options (i.e., the claimed “mounted on various types of devices”), and wiring options to suit the needs of a given application.” Par. 0005] Eskimez, Yoshioka and Shumard pertain to microphone geometry agnostic systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the microphone geometry agnostic systems art to modify Eskimez’s teachings of “array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE)” (Eskimez, Par. 0024) with the explicit teachings of “masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”)” (Yoshioka, Pg. 6028) taught by Yoshioka and “array output with a directional polar pattern that is selected (i.e., the claimed “microphone array selector”)” (Shumard, Par. 0016) taught by Shumard in order to “improve ASR accuracy” (Yoshioka, Pg. 6028) and “improve frequency-dependent directivity, particularly in the audio frequencies that are important for intelligibility, and the ability to reject unwanted sounds and reflections within a given environment, so as to provide full, natural-sounding speech pickup suitable for conferencing applications” (Shumard, Par. 0011). Claims 3 - 5 and 14 - 16 are rejected under 35 U.S.C. 103(a) as being unpatentable over Eskimez in view of Yoshioka as applied in claim 1 above, and in further view of Kurtz et al., (U.S. Patent 11,734,570), hereinafter referred to as Kurtz. Regarding Claims 3 and 14, Eskimez in view of Yoshioka has been discussed above. The combination further teaches: wherein the learning machine performs the training by tuning a model parameter that constitutes the speaker and noise mask estimator based on the loss function obtained by calculating a sum of reconstruction loss functions each defined as a distance between the inversely-transformed time domain signal and the answer signal for each speaker. [Eskimez, “The disclosure below extends PSE to utilize the microphone arrays for environments where strong noise, reverberation, and an interfering speaker are present.” Par. 0056; “The estimated mask (i.e., the claimed “speaker and noise mask estimator”) is applied to the first microphone in both approaches.” Par. 0061; “PSE model (i.e., the claimed “learning machine”) is able to implicitly learn the spectral and spatial information from the microphone array.” Par. 0059; “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array, producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045; Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “plurality of time-frequency domain signals”) for accurate TF (i.e., the claimed “time-frequency”) mask estimation.” Pg. 6028; “The objective is to construct a speech separation model that receives the feature sequency (zmt) from all the input channels and generates TF mask Msft (i.e., the claimed “time-frequency domain mask”),” Pg. 6028; Yoshioka, “Results of fine-tuning the speech separation model (i.e., the claimed “tuning model parameter”) to AMI with an ASR based loss function are also presented,” Pg. 6027; “During training, we minimized the uPIT style cross entropy loss between two predicted hypotheses and reference transcriptions (i.e., the claimed “comparison between the inversely transformed time domain signal and a pre-defined answer signal”), Pg. 6029] The combination fails to teach sum of reconstruction loss functions each defined as a distance. However, Kurtz teaches: wherein the learning machine performs the training by tuning a model parameter that constitutes the speaker and noise mask estimator based on the loss function obtained by calculating a sum of reconstruction loss functions each defined as a distance between the inversely-transformed time domain signal and the answer signal for each speaker. [Kurtz, “array of microphones (including a plurality of microphones),” Col. 6:9; “Overall loss 270 is then the weighted sum of both utility loss 240 and contractive loss 260 (i.e., the claimed “calculating a sum of reconstruction loss functions”),” Col. 9:7-9; “loss encouraging the distance between the two features (i.e., the claimed “distance between the inversely-transformed time domain signal and the answer signal for each speaker”) in a feature space,” Col. 8: 67-68] Eskimez, Yoshioka and Kurtz pertain to microphone array systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the microphone array systems art to modify Eskimez’s teachings of “array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE)” (Eskimez, Par. 0024) with the explicit teachings of “masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”)” (Yoshioka, Pg. 6028) taught by Yoshioka and “overall loss is then the weighted sum of both utility loss and contractive loss (i.e., the claimed “calculating a sum of reconstruction loss functions”)” (Kurtz, Col. 9:7-9) taught by Kurtz in order to “improve ASR accuracy” (Yoshioka, Pg. 6028) and “locate the source of sound in space of the real environment” (Kurtz, Col. 6:11-12. Regarding Claims 4 and 15, Eskimez in view of Yoshioka and Kurtz has been discussed above. The combination further teaches: wherein the learning machine obtains the loss functions having a number corresponding to various types of devices having different microphone array forms, and [Eskimez, “PSE model (i.e., the claimed “learning machine”) is able to implicitly learn the spectral and spatial information from the microphone array.” Par. 0059; “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array (i.e., the claimed “various types of devices having different microphone array forms”), producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045; Kurtz, “array of microphones (including a plurality of microphones),” Col. 6:9; “Overall loss 270 is then the weighted sum of both utility loss 240 and contractive loss 260 (i.e., the claimed “loss functions”),” Col. 9:7-9] updates the model parameter with temporary model parameters of the microphone array forms having the number corresponding to the various types of devices through a gradient decent algorithm. [Eskimez, “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array (i.e., the claimed “various types of devices having different microphone array forms”), producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045; “update the parameters (i.e., the claimed “model parameters”) of the PSE network.” Par. 0053; “Gradient Descent, with an additional loss applied to the optimization objective (e.g., minimizing the loss).” Col. 8: 55-57] Regarding Claims 5 and 16, Eskimez in view of Yoshioka and Kurtz has been discussed above. The combination further teaches: wherein the learning machine updates a model parameter of the speaker and noise mask estimator with an optimal model parameter derived by performing training in a way to minimize the loss function by applying the gradient decent algorithm to a loss function that is obtained after the temporary model parameter is updated. [Eskimez, “PSE model (i.e., the claimed “learning machine”) is able to implicitly learn the spectral and spatial information from the microphone array.” Par. 0059; “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array (i.e., the claimed “various types of devices having different microphone array forms”), producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045; Kurtz, “array of microphones (including a plurality of microphones),” Col. 6:9; “Overall loss 270 is then the weighted sum of both utility loss 240 and contractive loss 260 (i.e., the claimed “loss functions”),” Col. 9:7-9; Eskimez, “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array (i.e., the claimed “various types of devices having different microphone array forms”), producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045; “update the parameters (i.e., the claimed “model parameters”) of the PSE network.” Par. 0053; “Gradient Descent, with an additional loss applied to the optimization objective (e.g., minimizing the loss (i.e., the claimed “minimize the loss function”)).” Col. 8: 55-57] Claims 7 and 18 are rejected under 35 U.S.C. 103(a) as being unpatentable over Eskimez in view of Yoshioka as applied in claim 1 above, and in further view of Li et al., (U.S. Patent Application Publication 2022/0335947), hereinafter referred to as Li. Regarding Claims 7 and 18, Eskimez in view of Yoshioka has been discussed above. The combination further teaches: the end-to-end voice recognition model defines a loss function for fine-tuning by comparing the results of the voice recognition and pre-prepared answer information, and [Yoshioka, “3.3. End-to-end optimization (i.e., the claimed “end-to-end voice recognition model”), Pg. 6029; “E2E model (i.e., the claimed “end-to-end voice recognition model”),” Pg. 6029; Eskimez, “Aspects of the disclosure describe both (i) array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE),” Par. 0024; “Input audio 112 is provided to an STFT block 302 (i.e., the claimed “time-frequency transformer”), Par. 0042; Figure 3 clearly shows “Short Time Fourier Transform (i.e., the claimed “time-frequency transformer”), 302.”; “In some examples, producing output data 114 using trained PSE model 110 comprises isolating speech data (i.e., the claimed “voice data”) of target speaker 102 in a manner that is agnostic of a configuration of microphone array 200 (i.e., the claimed “plurality of microphones having unspecified microphone array forms”).” Par. 0100; “Aspects of the disclosure describe both (i) array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE),” Par. 0024; “The disclosure below extends PSE to utilize the microphone arrays for environments where strong noise, reverberation, and an interfering speaker are present.” Par. 0056; “The estimated mask (i.e., the claimed “speaker and noise mask estimator”) is applied to the first microphone in both approaches.” Par. 0061; “fed into an inverse STFT (i.e., the claimed “time-frequency inverse transformer”) block 618 to produce output data 114,” Par. 0067; Figure 3 clearly shows “Inverse Short Time Fourier Transform (i.e., the claimed “time-frequency inverse transformer”), 322.”; “PSE model (i.e., the claimed “learning machine”) is able to implicitly learn the spectral and spatial information from the microphone array.” Par. 0059; “using the trained geometry-agnostic PSE model (i.e., the claimed “learning machine”) without geometry information for the microphone array, producing output data, the output data comprising estimated clean speech data of the first target speaker with a reduction of speech data of the interfering speaker (i.e., the claimed “noise”).” Par. 0005; “The model (i.e., the claimed “learning machine”) is trained with a power-law compressed phase-aware mean-squared error (MSE) loss function.” Par. 0045; Yoshioka, “The proposed model is applicable to any number of microphones (i.e., the claimed “plurality of microphones having unspecified microphone array forms”) without retraining,” Pg 6027; “Let Xmft denote the short Fourier transform (STFT) coefficients of the audio signal observed by the mth microphone (i.e., the claimed “plurality of microphones having unspecified microphone array forms”), where f and t represent the frequency and time indices (i.e., the claimed “time-frequency domain signals”), respectively.” Pg. 6028; “expose the model to a variety of array geometries (i.e., the claimed “plurality of microphones having unspecified microphone array forms”) and inter-microphone spacing patterns.” Pg. 6029; Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “plurality of time-frequency domain signals”) for accurate TF (i.e., the claimed “time-frequency”) mask estimation.” Pg. 6028; “The objective is to construct a speech separation model that receives the feature sequency (zmt) from all the input channels and generates TF mask Msft (i.e., the claimed “time-frequency domain mask”),” Pg. 6028; “The masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028; “The first two sources correspond to the two speakers (i.e., the claimed “voice signals of the plurality of speakers”).” Pg. 6028; “Results of fine-tuning the speech separation model to AMI with an ASR based loss function are also presented,” Pg. 6027; “During training, we minimized the uPIT style cross entropy loss between two predicted hypotheses and reference transcriptions (i.e., the claimed “comparison between the inversely transformed time domain signal and a pre-defined answer signal”), Pg. 6029] the loss function for the fine-tuning is constructed by adding a connectionist temporal classification loss function and a cross-entropy loss function for each speaker. [Yoshioka, “Results of fine-tuning the speech separation model to AMI with an ASR based loss function are also presented,” Pg. 6027; “During training, we minimized the uPIT style cross entropy loss (i.e., the claimed “cross entropy loss function”) between two predicted hypotheses and reference transcriptions, Pg. 6029; “The first two sources correspond to the two speakers.” Pg. 6028] The combination fails to teach connectionist temporal classification. However, Li teaches: the loss function for the fine-tuning is constructed by adding a connectionist temporal classification loss function and a cross-entropy loss function for each speaker. [Li, “connectionist temporal classification (CTC)”, Par. 0018] Eskimez, Yoshioka and Li pertain to speech recognition systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the speech recognition systems art to modify Eskimez’s teachings of “array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE)” (Eskimez, Par. 0024) with the explicit teachings of “masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”)” (Yoshioka, Pg. 6028) taught by Yoshioka and “connectionist temporal classification (CTC)”, (Li, Par. 0018) taught by Li in order to “improve ASR accuracy” (Yoshioka, Pg. 6028) and improve “end-to-end framework” so that “various pieces of information concerning speech audio are routed through multiple processing operations in which data is analyzed and transformed in multiple ways to derive a transcript of the contents of the speech audio, and then to derive insights concerning those contents” (Li, Par. 0254). Claim 8 is rejected under 35 U.S.C. 103(a) as being unpatentable over Eskimez in view of Yoshioka and Li as applied in claim 7 above, and in further view of Wang et al., (EP 4006901), hereinafter referred to as Wang. Regarding Claim 8, Eskimez in view of Yoshioka and Li has been discussed above. The combination further teaches: wherein the end-to-end voice recognition model comprises: [Yoshioka, “3.3. End-to-end optimization (i.e., the claimed “end-to-end voice recognition model”), Pg. 6029; “E2E model (i.e., the claimed “end-to-end voice recognition model”),” Pg. 6029; Li, “end-to-end framework,” Par. 0254] a voice recognition encoder configured to receive a time-frequency domain signal of a specific speaker from which noise has been removed and to embed the time- frequency domain signal as a vector that constitutes a hidden vector space; and [Eskimez, “FIG. 4 illustrates an example encoder (i.e., the claimed “voice recognition encoder”)/decoder that may be used in various architectures described herein, for example as encoder blocks 4001a and 4001f, and decoder blocks 4002a and 4002f.” Par. 0046;“The disclosure below extends PSE to utilize the microphone arrays for environments where strong noise, reverberation, and an interfering speaker (i.e., the claimed “specific speaker from which noise has been removed”) are present.” Par. 0056; Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “plurality of time-frequency domain signals”) for accurate TF (i.e., the claimed “time-frequency”) mask estimation.” Pg. 6028; “speaker embeddings 606 are introduced into each complex encoder (i.e., the claimed “voice recognition encoder”),” Par. 0070; “The speaker (i.e., the claimed “specific speaker”) embeddings 308 (d-vector) (i.e., the claimed “embed as a vector”),” Par. 0042] a voice recognition decoder configured to receive the vector that constitutes the hidden vector space and to output results of voice recognition of the specific speaker. [Eskimez, “FIG. 4 illustrates an example encoder /decoder (i.e., the claimed “voice recognition decoder”) that may be used in various architectures described herein, for example as encoder blocks 4001a and 4001f, and decoder blocks 4002a and 4002f.” Par. 0046;“The disclosure below extends PSE to utilize the microphone arrays for environments where strong noise, reverberation, and an interfering speaker (i.e., the claimed “specific speaker from which noise has been removed”) are present.” Par. 0056; Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “plurality of time-frequency domain signals”) for accurate TF (i.e., the claimed “time-frequency”) mask estimation.” Pg. 6028; “output of decoder block 4002a) to produce output data (i.e., the claimed “output results of voice recognition of the specific speaker.”) 114,” Par. 0051] The combination fails to teach hidden vector space. However, Wang teaches: a voice recognition encoder configured to receive a time-frequency domain signal of a specific speaker from which noise has been removed and to embed the time- frequency domain signal as a vector that constitutes a hidden vector space; and [Wang, “encoder network is a four-layer BLSTM network, each hidden layer has 600 nodes, and the output layer is a fully connected layer, which can map a 600-dimensional hidden vector (output feature) outputted by the last hidden layer to a 275*40-dimensional high-dimensional embedding space v (i.e., the claimed “hidden vector space”),” Par. 0172] Eskimez, Yoshioka, Li and Wang pertain to speech recognition systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the speech recognition systems art to modify Eskimez’s teachings of “array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE)” (Eskimez, Par. 0024) with the explicit teachings of “masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”)” (Yoshioka, Pg. 6028) taught by Yoshioka, “connectionist temporal classification (CTC)”, (Li, Par. 0018) taught by Li, and “encoder network is a four-layer BLSTM network, each hidden layer has 600 nodes, and the output layer is a fully connected layer, which can map a 600-dimensional hidden vector (output feature) outputted by the last hidden layer to a 275*40-dimensional high-dimensional embedding space v (i.e., the claimed “hidden vector space”) (Wang, Par. 0172) taught by Wang in order to “improve ASR accuracy” (Yoshioka, Pg. 6028), improve “end-to-end framework” so that “various pieces of information concerning speech audio are routed through multiple processing operations in which data is analyzed and transformed in multiple ways to derive a transcript of the contents of the speech audio, and then to derive insights concerning those contents” (Li, Par. 0254), and “separate an independent audio signal of each speaker talking at the same time” (Wang, Par. 0003). Claims 9 - 10 and 19 - 20 are rejected under 35 U.S.C. 103(a) as being unpatentable over Eskimez in view of Yoshioka as applied in claim 1 above, and in further view of Li Y et al., (CN116453533A), hereinafter referred to as Li Y. Regarding Claims 9 and 19, Eskimez in view of Yoshioka has been discussed above. The combination further teaches: wherein the beamformer estimator receives the time-frequency domain mask and the plurality of time-frequency domain signals, [Eskimez, “Multi-channel PSE uses inputs from multiple microphones, and geometry agnostic operation indicates that the PSE solution does not need to know the positions of the microphones in order to reduce or eliminate non-speech noise and speech data of an interfering speaker.” Par. 0027; “Aspects of the disclosure are also operable with MVDR beamforming followed by a single-channel PSE.” Par. 0084; Yoshioka, “The masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028; “The first two sources correspond to the two speakers (i.e., the claimed “voice signals of the plurality of speakers”).” Pg. 6028] calculates a power spectrum density matrix for the voices and noise for the plurality of speakers, and [Eskimez, “The PSE described herein is capable of removing both the interfering speakers (i.e., the claimed voices and noise for the plurality of speakers”) and environmental noises (i.e., the claimed “noise for the plurality of speakers”),” Par. 0028] calculates a filter coefficient of the beamformer estimator based on the power spectrum density matrix, and [Eskimez, “Multi-channel PSE uses inputs from multiple microphones, and geometry agnostic operation indicates that the PSE solution does not need to know the positions of the microphones in order to reduce or eliminate non-speech noise and speech data of an interfering speaker.” Par. 0027; “Aspects of the disclosure are also operable with MVDR beamforming followed by a single-channel PSE.” Par. 0084; Yoshioka, “The masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028; “The first two sources correspond to the two speakers (i.e., the claimed “voice signals of the plurality of speakers”).” Pg. 6028] estimates the time-frequency domain signals for the voice signals of the plurality of speakers from which noise has been removed by applying the filter coefficient to time bin information and frequency bin information in the time-frequency domain signal for the voice and noise of the speaker. [Eskimez,, frequency bins (i.e., the claimed “frequency bin information”), and time frames (i.e., the claimed “time bin information”),” Par. 0044; Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “time-frequency domain signals for the voice signals of the plurality of speakers”) for accurate TF (i.e., the claimed “time-frequency”) mask estimation.” Pg. 6028; Yoshioka, “The masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028] The combination fails to teach power spectrum density matrix and filter coefficient. However, Li Y teaches: calculates a power spectrum density matrix for the voices and noise for the plurality of speakers, and [Li Y, “power spectrum density matrix of the noise according to said embodiments to obtain the optimal filter coefficient (i.e., the claimed “filter coefficient”)” Par. n0011] calculates a filter coefficient of the beamformer estimator based on the power spectrum density matrix, and estimates the time-frequency domain signals for the voice signals of the plurality of speakers from which noise has been removed by applying the filter coefficient to time bin information and frequency bin information in the time-frequency domain signal for the voice and noise of the speaker. [Li Y, “power spectrum density matrix of the noise according to said embodiments to obtain the optimal filter coefficient (i.e., the claimed “filter coefficient”)” Par. n0011] Eskimez, Yoshioka and Li Y pertain to speech recognition systems and are analogous to the instant application. Accordingly, it would have been obvious to one of ordinary skill in the speech recognition systems art to modify Eskimez’s teachings of “array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE)” (Eskimez, Par. 0024) with the explicit teachings of “masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”)” (Yoshioka, Pg. 6028) taught by Yoshioka and “power spectrum density matrix”, (Li Y, Par. n0011) taught by Li Y in order to “improve ASR accuracy” (Yoshioka, Pg. 6028) and “signal of the microphone to be noise-reduced is recovered using the optimal filter” (Li Y, Par. n0011). Regarding Claims 10 and 20, Eskimez in view of Yoshioka and Li Y has been discussed above. The combination further teaches: wherein the beamformer estimator calculates the power spectrum density matrix, [Eskimez, “Multi-channel PSE uses inputs from multiple microphones, and geometry agnostic operation indicates that the PSE solution does not need to know the positions of the microphones in order to reduce or eliminate non-speech noise and speech data of an interfering speaker.” Par. 0027; “Aspects of the disclosure are also operable with MVDR beamforming followed by a single-channel PSE.” Par. 0084; Yoshioka, “The masks (i.e., the claimed “time-frequency domain mask”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028; “The first two sources correspond to the two speakers (i.e., the claimed “voice signals of the plurality of speakers”).” Pg. 6028; Li Y, “power spectrum density matrix of the noise according to said embodiments to obtain the optimal filter coefficient (i.e., the claimed “filter coefficient”)” Par. n0011] based on a mask coefficient for the time bin information and frequency bin information in the time-frequency domain signal for the voice and noise of the speaker and a column vector comprising time-frequency signals of the time bin information and frequency bin information for all of channels. [Eskimez,, frequency bins (i.e., the claimed “frequency bin information”), and time frames (i.e., the claimed “time bin information”),” Par. 0044; Yoshioka, “VarArray, our proposed speech separation model, can efficiently utilize the spatial and temporal information from any number of microphone inputs (i.e., the claimed “time-frequency domain signals for the voice signals of the plurality of speakers”) for accurate TF (i.e., the claimed “time-frequency”) mask (i.e., the claimed “mask coefficient”) estimation.” Pg. 6028; Eskimez, “Aspects of the disclosure describe both (i) array geometry agnostic (i.e., the claimed “unspecificied microphone array forms”) multi-channel (i.e., the claimed “plurality of channel audio signals”) personalized speech enhancement (PSE),” Par. 0024; “The disclosure below extends PSE to utilize the microphone arrays for environments where strong noise, reverberation, and an interfering speaker are present.” Par. 0056; “The estimated mask (i.e., the claimed “mask coeffcient”) is applied to the first microphone in both approaches.” Par. 0061; “The model outputs a complex ratio mask (i.e., the claimed “mask coefficient”) which is multiplied with the input mixture to estimate the clean speech.” Par. 0045; Yoshioka, “The masks (i.e., mask coefficient”) are used to perform minimum variance distortionless response (MVDR) beamforming to estimate (i.e., the claimed “beamformer estimator”) the clean signals (i.e., the claimed “noise has been removed”).” Pg. 6028; “The first two sources correspond to the two speakers (i.e., the claimed “voice signals of the plurality of speakers”).” Pg. 6028] Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Taherian et al., (“One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement,” arXiv:2110.10330v1, 2021) teaches array agnostic multi-channel personalized speech enhancement. Any inquiry concerning this communication or earlier communications from the examiner should be directed to EUNICE LEE whose telephone number is 571-272-1886. The examiner can normally be reached M-F 8:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /EUNICE LEE/Examiner, Art Unit 2656 /BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Feb 21, 2025
Application Filed
Sep 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749482
METHOD, DEVICE, COMPUTER PROGRAM AND COMPUTER READABLE STORAGE MEDIUM FOR DETERMINING A COMMAND
3y 1m to grant Granted Sep 29, 2026
Patent 12725629
METHOD AND APPARATUS FOR MEASURING SPEECH-IMAGE SYNCHRONICITY, AND METHOD AND APPARATUS FOR TRAINING MODEL
2y 8m to grant Granted Sep 01, 2026
Patent 12725623
VOICE PROCESSING SYSTEM AND VOICE PROCESSING METHOD
2y 7m to grant Granted Sep 01, 2026
Patent 12718027
SPEECH SIGNAL PROCESSING USING ARTIFICIAL INTELLIGENCE
3y 1m to grant Granted Aug 25, 2026
Patent 12706111
RECEIVE-SIDE AUDIO PROCESSING FOR CALLS IN A WEB CONFERENCING CLIENT
2y 6m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
89%
Grant Probability
99%
With Interview (+25.7%)
2y 7m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 45 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month