Prosecution Insights
Last updated: October 01, 2026
Application No. 18/898,006

SPEAKER IDENTIFICATION METHOD, SPEAKER IDENTIFICATION DEVICE, AND NON-TRANSITORY COMPUTER READABLE RECORDING MEDIUM STORING SPEAKER IDENTIFICATION PROGRAM

Non-Final OA §103
Filed
Sep 26, 2024
Priority
Mar 29, 2022 — JP 2022-053033 +1 more
Examiner
MANOHARAN, SHASHIDHAR SHANKAR
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Panasonic Holdings Corporation
OA Round
2 (Non-Final)
80%
Grant Probability
Favorable
2-3
OA Rounds
2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
4 granted / 5 resolved
+18.0% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
26 currently pending
Career history
33
Total Applications
across all art units

Statute-Specific Performance

§101
17.2%
-22.8% vs TC avg
§103
64.8%
+24.8% vs TC avg
§102
4.7%
-35.3% vs TC avg
§112
9.4%
-30.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 5 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendments filed have been accepted and considered in this office action. Claims 2-4 have been cancelled. Claims 1, 10 and 11 have been amended. Claims 1, 5-11 are pending. Response to Arguments Applicant’s arguments with respects to the 35 U.S.C. 103 claims 1, 5-11 have been considered and are persuasive. The previous 35 U.S.C. 103 rejection has been withdrawn. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1 and 5-11 are rejected under 35 U.S.C. 103 as being unpatentable over Sharifi et al. (hereinafter Sharifi) (US 20160275953 A1) in view of Park et al. (hereinafter Park) (US 20210125617 A1), in further view of Zhang et al. (hereinafter Zhang) (US 6735562 B1) (see attached copy for page numbers). Regarding claim 1, Sharifi teaches: A speaker identification method in a computer (Sharifi, Abstract), the speaker identification method comprising: acquiring voice data to be identified (Sharifi, P[0062]: “the speaker identification system 110 receives data that identifies an utterance 160 of a speaker to be identified.”); selecting a registered speaker of registered voice data corresponding to a highest similarity from among a plurality of calculated similarities (Sharifi, P[0069]: “the speaker vector having the highest degree of similarity to the utterance vector 162 may be selected. ”); Sharifi does not explicitly teach: acquiring a plurality of pieces of registered voice data that are registered in advance; calculating a similarity between the voice data to be identified and each of the plurality of pieces of registered voice data; determining, based on the plurality of calculated similarities, whether or not the voice data to be identified is suitable for speaker identification; determining, based on the highest similarity, whether or not to identify the selected registered speaker as a speaker to be identified of the voice data to be identified in a case where the voice data to be identified is determined to be suitable for the speaker identification; and outputting an identification result wherein in determination as to whether or not the voice data to be identified is suitable for the speaker identification, a variance value of the plurality of calculated similarities is calculated, whether or not the calculated variance value is higher than a first threshold is determined, and in a case where the variance value is determined to be higher than the first threshold, the voice data to be identified is determined to be suitable for the speaker identification However, Park teaches: acquiring a plurality of pieces of registered voice data that are registered in advance (Park, P[0077]: “The speaker recognition apparatus stores generated feature vectors in a registration database… The registered feature model may include an average and a variance of a plurality of registered feature vectors”; calculating a similarity between the voice data to be identified and each of the plurality of pieces of registered voice data (Park, P[0078]: “The speaker recognition apparatus compares the registered data… and the input feature vector”, “The similarity between the registered data and the input feature vector 260 may be calculated using, for example, a distance between two vectors and a cosine similarity); determining, based on the plurality of calculated similarities, whether or not the voice data to be identified is suitable for speaker identification (Park, P[0084]-P[0085]: “The speaker recognition apparatus selectively performs a speaker recognition operation, a registration database updating operation, and a candidate list constructing operation depending on the similarity determination.”, “ the speaker recognition apparatus determines that the received speaker's voice signal is sufficiently the same as the registered speaker's voice signal (for example, determines that a similarity between at least one registered data, e.g., at least one registered speaker's voice signal, vector, or model of the same, included in the registration database and at least one input feature vector corresponding to the speaker's voice signal is greater than or equal to a first threshold), but the similarity between the received speaker's voice signal and the registered speaker's voice signal does not meet, e.g., does not exceed, a predetermined threshold” (threshold based suitability reads on determining suitability)); determining, based on the highest similarity, whether or not to identify the selected registered speaker as a speaker to be identified of the voice data to be identified in a case where the voice data to be identified is determined to be suitable for the speaker identification (Park, P[0092]: “The speaker recognition apparatus may further perform the speaker recognition operation based on an input feature vector, among the one or more input feature vectors corresponding to the voice signal of the speaker, of which a similarity with the registered data meets, e.g., is greater than or equal to, a third threshold.”); and outputting an identification result (Park, P[0101], reads on verification result if authentication criteria is met and when it is not met). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Sharifi in view of Park. Doing so would have provided the threshold-based similarity determination of Park (Park P[0021) with the similarity based speaker identification techniques of Sharifi (Sharifi, P[0068]) in order to improve the reliability of speaker identification in view of known variability in voice signals (Park, P[0004]). The combination of Park and Sharifi do not explicitly teach: wherein in determination as to whether or not the voice data to be identified is suitable for the speaker identification, a variance value of the plurality of calculated similarities is calculated, whether or not the calculated variance value is higher than a first threshold is determined, and in a case where the variance value is determined to be higher than the first threshold, the voice data to be identified is determined to be suitable for the speaker identification However, Zhang teaches: wherein in determination as to whether or not the voice data to be identified is suitable for the speaker identification, a variance value of the plurality of calculated similarities is calculated (Zhang, P(12): "comparing the input utterance with a plurality of predetermined models of possible utterances to provide a plurality of scores indicating a degree of similarity between the input utterance and the plurality of predetermined models, determining a variance of a predetermined number of the plurality of scores", Zhang's plurality of scores indicate the respective degrees of similarity between the same input utterance and the plurality of predetermined models and Zhang determines a variance from a plurality of those calculated similarity scores, thus, reading on calculating a variance value of the plurality of calculated similarities), whether or not the calculated variance value is higher than a first threshold is determined, and in a case where the variance value is determined to be higher than the first threshold, the voice data to be identified is determined to be suitable for the speaker identification (Zhang, P(4): "The confidence measure (CM) is equivalent to the normalized variance of the N-Best scores", P(16): "the CM is compared to the weighted threshold value Tw (step 10). If the CM greater than Tw, it is assumed that the pre-trained model which had the best score is a correct identification of the input utterance and that pre-trained model is accepted", Zhang's confidence measure is the normalized variance of the plurality of best similarity scores, Zhang compares that variance-based value to a threshold, and when the variance-based value is greater than the threshold, Zhang accepts the identification of the input utterance, thus, determining the input voice data to be sufficiently suitable/reliable for identification). It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Sharifi in view of Park and in further view of Zhang. Doing so would have provided the variance based confidence determination of Zhang (Zhang, Abstract, P(4), P(12)) threshold-based similarity determination of Park (Park P[0021) with the similarity based speaker identification techniques of Sharifi (Sharifi, P[0068]). This would improved the reliability of speaker identification by rejecting or avoiding identification when the similarity results do not provide sufficient confidence and accepting identification when the similarity results provide sufficient discrimination among candidate models. Regarding claim 5, Sharifi in view of Park and in further view of Zhang discloses the speaker identification method disclosed in claim 1. Park further teaches: wherein the plurality of pieces of registered voice data include a plurality of pieces of first registered voice data in which voice uttered by a plurality of registered speakers to be identified is registered in advance, and a plurality of pieces of second registered voice data in which voice uttered by a plurality of other registered speakers other than the plurality of registered speakers to be identified is registered in advance (Park, P[0015] and P[0113], shows utilizing a target database and a ‘rejection list’ of non-target voices to distinguish speakers) in calculation of the similarity, a first similarity between the voice data to be identified and each of the plurality of pieces of first registered voice data is calculated, and a second similarity between the voice data to be identified and each of the plurality of pieces of second registered voice data is calculated (Park, P[0089] and P[0112], Park describes performing distance/similarity checks against both target and rejection models), in selection of the registered speaker, a registered speaker of first registered voice data corresponding to a highest first similarity among a plurality of calculated first similarities is selected (Park, P[0089] and Equation 1, system mathematically selects closes match (highest similarity) among registered set) and in determination as to whether or not the voice data to be identified is suitable for the speaker identification, whether or not a highest first similarity or a highest second similarity among the plurality of calculated first similarities and the plurality of calculated second similarities is higher than a first threshold is determined (Park, P[0006]: “determining whether an input feature vector… meets a candidate similarity criterion”, (Park uses the first threshold to determine if the voice is ‘candidate’ material or suitable for processing)), and in a case where the highest first similarity or the highest second similarity is determined to be higher than the first threshold, the voice data to be identified is determined to be suitable for the speaker identification (Park, P[0038]: “select to perform the construction of the candidate list based on the input feature when the input feature vector meets the candidate similarity criterion and the input feature vector meets the registered user similarity criterion.”, (shows meeting first threshold validates the voice as suitable for the identification workflow)). Regarding claim 6, Sharifi in view of Park and in further view of Zhang the speaker identification method disclosed in claim 5. Sharifi further teaches: wherein the plurality of pieces of second registered voice data do not include noise and include only the voice uttered by the other registered speakers (Sharafi, P[0049]: "A speaker vector 130 may be data that indicates characteristics of a speaker's voice." (speaker vectors represent characteristics of speaker's voices derived from utterances meaning stored registered voice data corresponds to voice data of speakers not environmental noise)). Regarding claim 7, Sharifi in view of Park and in further view of Zhang the speaker identification method disclosed in claim 5. Park further teaches: wherein in determination as to whether or not to identify the selected registered speaker as a speaker to be identified of the voice data to be identified, whether or not a highest first similarity among the plurality of calculated first similarities is higher than a second threshold higher than the first threshold is determined (Park, P[0092]: “The third threshold, which is an example similarity constraint condition or registered user similarity criterion for the speaker recognition operation, may be greater than the first threshold” (confirms specific use of the higher threshold for final target identification)), and in a case where the highest first similarity is determined to be higher than the second threshold, the selected registered speaker is identified as the speaker to be identified of the voice data to be identified (Park, P[0101]: “verifies the speaker having uttered the input voice 510 as a registered user” (identification is finalized once the target-specific higher threshold is met)). Regarding claim 8, Sharifi in view of Park and in further view of Zhang discloses the speaker identification method disclosed in claim 1. Park further teaches: further comprising outputting an error message prompting the speaker to be identified to reinput the voice data to be identified in a case where the voice data to be identified is determined not to be suitable for the speaker identification (Park, P[0070]: “For example, the speaker recognition apparatus 120 outputs a message indicating “unregistered user” and rejects or fails to perform the additional operation corresponding to the voice signal uttered by the speaker.” (rejection prompt reads on error message)). Regarding claim 9, Sharifi in view of Park and in further view of Zhang discloses the speaker identification method disclosed in claim 1. Park further teaches: wherein in acquisition of the voice data to be identified, the voice data to be identified obtained by cutting out a predetermined section from voice data uttered by the speaker to be identified is acquired (Park, “one or more input feature vectors corresponding to a voice signal of a speaker” (input feature vectors are mathematically segmented sections of audio signal)), and the speaker identification method further comprises acquiring another piece of voice data to be identified obtained by cutting out a section different from the predetermined section from the voice data in a case where the voice data to be identified is determined not to be suitable for the speaker identification (Wang, P[0095]: “rescue (previous) input feature vectors that were dropped” and P[0097]: “the input feature vector 430 may be additionally registered at a later time, e.g., subsequent to a time the second input feature vector 420 is input and found to have a similarity with the registered data 410 that meets the second threshold and is added to the registered data 410, based on a renewed similarity determination between the input feature vector”, (rescuing other sections/vectors from the stream to attempt identification if initial section is unsuitable)) Regarding claim 10, claim 10 recites the speaker identification device corresponding to the speaker identification method presented in claim 1 and is rejected under the same grounds stated above. The combination further teaches: A speaker identification device comprising: a processor (Sharifi, P[0084); and a memory including a program that, when executed by the processor, causes the processor to (Sharifi, P[0084]): Regarding claim 11, claim 11 recites the non-transitory computer readable recording medium storing a speaker identification program corresponding to the speaker identification method presented in claim 1 and is rejected under the same grounds stated above. The combination further teaches A non-transitory computer readable recording medium storing a speaker identification program that causes a computer to function to (Sharifi, claim 16): Conclusion A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHASHIDHAR S MANOHARAN whose telephone number is (571)272-6772. The examiner can normally be reached M-F 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHASHIDHAR SHANKAR MANOHARAN/Examiner, Art Unit 2655 /ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Sep 26, 2024
Application Filed
Mar 20, 2026
Non-Final Rejection mailed — §103
Jun 22, 2026
Response Filed
Sep 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737537
GOVERNANCE AND CONFIDENCE ASSESSMENT OF LLM
2y 6m to grant Granted Sep 15, 2026
Patent 12682890
MASK-CONFORMER AUGMENTING CONFORMER WITH MASK-PREDICT DECODER UNIFYING SPEECH RECOGNITION AND RESCORING
2y 4m to grant Granted Jul 14, 2026
Patent 12682173
MODULAR FRAMEWORK FOR EVALUATING LANGUAGE MODELS
2y 4m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+33.3%)
2y 2m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 5 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month