DETAILED ACTION
This communication is in response to the Application filed on 03/31/2023. Claims 1-18, 22 and 25 are pending and have been examined. Claims 1, 12 and 22 are independent. Claims 19-21 and 23-24 are canceled. This Application was published as U.S. Pub. No. 2024/0331705A1.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/31/2023 was filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claim 1 is objected to because of the following informalities: the limitation, “instructions” in line 3 of claim 1 does not adequately defines a feature of an apparatus claim. Features of an apparatus may be recited either structurally or functionally (MPEP 2114). “instructions,” by itself, recites neither structure nor functionality (or functional connection with other components of claim). Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-12, 14-18, 22 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Jang et al., (US Pub No. 2020/0321022, hereinafter, Jang) in view of Khoury et al., (US Pub No. 2021/0326421, hereinafter, Khoury).
Regarding claim 1,
Jang discloses an apparatus comprising: interface circuitry; instructions; and programmable circuitry to at least one of execute or instantiate the instructions to (Jang, Fig.5, par[055], "…the device 500, the memory 586, the processor 506, the processors 510, a component of the system-on-chip device 522, such as an interface or a controller"; par [050], "…a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC)..."):
calculate a sample embedding vector that characterizes a speaker based on a first audio signal (Jang, Fig.1, par [023], "…The speaker vector extractor 130 is configured to receive a frame 116 of an audio signal and to generate a speaker vector 132 that corresponds to the frame 116...speaker vector 132 that are indicative of a particular speaker..."; par [026], "…The automatic speech recognition engine 150 is configured to process one or more frames of the audio signal 114 that include the utterance 106…");
perform a first update of a personal embedding vector based on the sample embedding vector, the updated personal embedding vector to characterize the speaker based on a second audio signal and the first audio signal (Jang, Figs. 2-3, paras [024, 031, 034], "…the speaker vector comparator 140 is configured to compare the speaker vector 132 to at least one previously processed speaker vector that includes the utterance 106…"); and
Jang does not explicitly discloses the limitation, "perform a second update of the personal embedding vector based on the first update and a universal embedding vector." However, Khoury, in the analogous field of endeavor, discloses the analytical server (Fig.1, paras [110-113]) executing embedding extraction (paras [114-115]) and performing pre-processing (e.g., voice activity detection to differentiate between background noise, silence, and speakers in the audio signal, paras [116]) and continuous and passive enrollment (paras [044-049, 128-130]).
Khoury discloses perform a second update of the personal embedding vector based on the first update and a universal embedding vector (Khoury, paras [133-135], "…the analytics server 102 will evaluate the similarity of the embeddings with a set of voiceprints... voiceprints associated particular speaker characteristics, and voiceprints associated with particular speaker-independent characteristics…", "…The analytics server 102 iterates the clustering process until a stopping criteria is met...").
Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified an end-of-utterance detection device of Jang with passive and continuous enrollment of speaker embedding and clustering by the similarity score of Khoury with a reasonable expectation of success to employ the voice biometrics instead of time-consuming, potentially inaccurate, and stale active/static enrollment process (Khoury, paras [008-011]).
Regarding claim 3,
Jang in view of Khoury discloses the apparatus of claim 1, wherein the apparatus corresponds to a first user and a second user, and the apparatus further includes: identifier circuitry to identify, based on the personal embedding vector after the second update, the speaker as the first user or the second user (Jang, Fig.4, paras [042-050], "…at the end-of-utterance detector and based on the speaker vector, an indicator that indicates whether the frame corresponds to an end of an utterance of a particular speaker, at 406...").
Regarding claim 4,
Jang in view of Khoury discloses the apparatus of claim 1, wherein the first audio signal is not obtained as part of an enrollment process ( Khoury, continuous and passive enrollment (paras [044-049, 128-130]).
Regarding claim 5,
Jang in view of Khoury discloses the apparatus of claim 1, wherein to perform the first update of the personal embedding vector, the programmable circuitry is to: determine a ratio; and combine a first vector and a second vector, the first vector based on a previous version of the personal embedding vector and the ratio, the second vector based on the sample embedding vector and the ratio (Jang, par [024], "…the speaker vector comparator 140 is configured to compare the speaker vector 132 to a moving average of speaker vectors that include the at least one previously
processed speaker vector..."; par [034], "…The moving average unit 320 is configured to determine a moving average of speaker vector values associated with an utterance and to detect when a particular speaker vector differs from the moving average by more than a threshold amount. Such a detected difference in the speaker vector...").
Regarding claim 6,
Jang in view of Khoury discloses the apparatus of claim 1, wherein: the speaker is a first speaker; the sample embedding vector is a second sample embedding vector corresponding to the first audio signal; and the programmable circuitry is to (Jang, Fig.4, par [050], "…The method 400 of FIG. 4 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC)..."):
identify an unknown speaker in a third audio signal (Khoury, Fig.6, steps 602-608);
calculate a third sample embedding vector based on the third audio signal; and determine whether to perform an additional update of the personal embedding vector with the third sample embedding vector, the determination based on a distance calculation between the personal embedding vector and the third sample embedding vector (step 606, similarity score check and step 614, optionally re-clustering operations to update clusters and update speaker profile database).
Regarding claim 7,
Jang in view of Khoury discloses the apparatus of claim 1, wherein to perform the second update of the personal embedding vector, the programmable circuitry is to: determine a ratio; and combine a first vector and a second vector, the first vector based on the personal embedding vector after the first update and the ratio, the second vector based on the universal embedding vector and the ratio (Jang, par [024], "…the speaker vector comparator 140 is configured to compare the speaker vector 132 to a moving average of speaker vectors that include the at least one previously
processed speaker vector..."; par [034], "…The moving average unit 320 is configured to determine a moving average of speaker vector values associated with an utterance and to detect when a particular speaker vector differs from the moving average by more than a threshold amount. Such a detected difference in the speaker vector...").
Regarding claim 8,
Jang in view of Khoury discloses the apparatus of claim 7, wherein the programmable circuitry is to change the ratio over subsequent iterations so that a magnitude of the first vector increases and a magnitude of the second vector decreases (Khoury, par [135], "…the analytics server 102 measures the distances between the embeddings and the centroids based on the correlation of the features in the embeddings. The distance between the embeddings and the centroids are indicated using a similarity score. The more similar the embeddings are to the centroid, the higher the similarity score…").
Regarding claim 9,
Jang in view of Khoury discloses the apparatus of claim 1, wherein the second audio signal is obtained before the first audio signal (Jang, Figs. 2-3, paras [024, 031, 034], "…the speaker vector comparator 140 is configured to compare the speaker vector 132 to at least one previously processed speaker vector that includes the utterance 106…").
Regarding claim 10,
Jang in view of Khoury discloses the apparatus of claim 1, wherein the universal embedding vector characterizes human voice (Khoury, paras [133-135], "…voiceprints associated particular speaker characteristics, and voiceprints associated with particular speaker-independent characteristics (i.e., universal embedding vector of human voice)...");
Regarding claim 11,
Jang in view of Khoury discloses the apparatus of claim 1, wherein the programmable circuitry includes one or more of: at least one of a central processor unit, a graphics processor unit, or a digital signal processor, the at least one of the central processor unit, the graphics processor unit, or the digital signal processor having control circuitry to control data movement within the programmable circuitry, arithmetic and logic circuitry to perform one or more first operations corresponding to machine-readable data, and one or more registers to store a result of the one or more first operations, the machine-readable data in the apparatus; a Field Programmable Gate Array (FPGA), the FPGA including logic gate circuitry, a plurality of configurable interconnections, and storage circuitry, the logic gate circuitry and the plurality of the configurable interconnections to perform one or more second operations, the storage circuitry to store a result of the one or more second operations; or Application Specific Integrated Circuitry (ASIC) including logic gate circuitry to perform one or more third operations (Jang, Fig.4, par [050], "…The method 400 of FIG. 4 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller,
another hardware device, firmware device, or any combination thereof...performed by a processor that executes instructions...").
Claim 12 is a non-transitory machine readable storage medium claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale.
Rationale for combination is similar to that provided for Claim 1.
Claim 14 is a non-transitory machine readable storage medium claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale.
Claim 15 is a non-transitory machine readable storage medium claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale.
Claim 16 is a non-transitory machine readable storage medium claim with limitations similar to the limitations of Claim 5 and is rejected under similar rationale.
Claim 17 is a non-transitory machine readable storage medium claim with limitations similar to the limitations of Claim 6 and is rejected under similar rationale.
Claim 18 is a non-transitory machine readable storage medium claim with limitations similar to the limitations of Claim 7 and is rejected under similar rationale.
Claim 22 is a method claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 1.
Claim 25 is a method claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale.
Claims 2 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Jang in view of Khoury further in view of Eskimez et al., (US Pub No. 2024/0135949, hereinafter, Eskimez).
Regarding claim 2,
Jang in view of Khoury discloses the apparatus of claim 1, wherein the first audio signal includes audio from the speaker and parasitic noise (Khoury, par [116], "…The analytics server 102 performs audio event detection or other voice activity detection to differentiate between background noise, silence, and speakers in the audio signal..."), but does not explicitly discloses the dynamic noise suppression circuitry based on the personal embedding vector.
However, Eskimez, in the analogous field of endeavor, discloses the apparatus further includes: dynamic noise suppression circuitry to output, based on the personal embedding vector after the second update, main speaker audio that includes the speaker but not the parasitic noise (Eskimez, Abstract, par [004], "…perform personalized noise suppression (PNS) to remove speech from one or more interfering speakers and acoustic echo cancellation (AEC) to remove echoes..."; Figs.3,4,6C paras [061-068], "…an operation 680 of generating human speech feature information by removing noise and echoes identified in the near-end signal based on the first feature information, the second feature information, and the alignment information…."; "…an operation 682 of generating target speaker feature information by analyzing the human speech feature information to exclude speech from the first interfering speaker..."); and
transceiver circuitry to transmit the main speaker audio (Fig.6C, par [068], "…The process 670 includes an operation 684 of decoding the target speaker future information to obtain an output audio signal comprising the speech of the target speaker...").
Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified the speaker diarization/separation method/system of Jang in view of Khoury with the personalized noise suppression (PNS), automatic echo cancellation (AEC), and mask prediction layers of Eskimez with a reasonable expectation of success to output an audio signal comprising speech of the target speaker (Eskimez, Abstract).
Claim 13 is a non-transitory machine readable storage medium claim with limitations similar to the limitations of Claim 2 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 2.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Wang et al., ("Similarity measurement of segment-level speaker embeddings in speaker diarization." IEEE/ACM Transactions on Audio, Speech, and Language Processing 30 (2022): 2645-2658, hereinafter, Wang) discloses a neural-network-based similarity measurement method to learn the similarity between any two speaker embeddings, where both previous and future contexts are considered and the segmental pooling strategy and jointly train the speaker embedding network along with the similarity measurement model (Wang, Abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JANGWOEN LEE whose telephone number is (703)756-5597. The examiner can normally be reached Monday-Friday 8:00 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BHAVESH MEHTA can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JANGWOEN LEE/Examiner, Art Unit 2656
/BHAVESH M MEHTA/Supervisory Patent Examiner, Art Unit 2656