Prosecution Insights
Last updated: October 04, 2026
Application No. 19/012,768

SPEAKER VERIFICATION DEVICE, METHOD OF CONTROLLING SPEAKER VERIFICATION DEVICE, AND SPEAKER VERIFICATION SYSTEM

Non-Final OA §101§103
Filed
Jan 07, 2025
Priority
Jan 17, 2024 — RE 10-2024-0007538
Examiner
SCHMIEDER, NICOLE A K
Art Unit
Tech Center
Assignee
Iucf-Hyu(Industry-University Cooperation Foundation Hanyang University)
OA Round
1 (Non-Final)
68%
Grant Probability
Favorable
1-2
OA Rounds
12m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
120 granted / 176 resolved
+8.2% vs TC avg
Strong +34% interview lift
Without
With
+34.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
24 currently pending
Career history
199
Total Applications
across all art units

Statute-Specific Performance

§101
21.8%
-18.2% vs TC avg
§103
48.3%
+8.3% vs TC avg
§102
13.6%
-26.4% vs TC avg
§112
11.5%
-28.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 176 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim(s) 1-16 is/are pending and has/have been examined. Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statement (IDS) submitted on 05/13/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-16 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim(s) 1, 9, and 16, the limitation(s) of determine a first loss function, determine a second loss function, update parameters, update parameters, and (claim 16) receive, input, perform, and send, as drafted, are processes that, under broadest reasonable interpretation, covers performance of the limitation in the mind and/or with pen and paper but for the recitation of generic computer components, as well as mathematical calculations in prose. More specifically, the mental process of a human calculating the difference between two results of working through two algorithms, calculating the difference between two results of working through one of the algorithms a second time, updating both algorithms using specific input information, (claim 16) reading input data regarding speech, using an algorithm to calculate a result using the data, and writing down a conclusion that can be drawn from the result. . If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind and/or with pen and paper but for the recitation of generic computer components, then it falls within the --Mental Processes-- grouping of abstract ideas. Accordingly, the claim(s) recite(s) an abstract idea. This judicial exception is not integrated into a practical application because the recitation of a device, memory, and processor in claims 1 and 9, and a system, terminal, device, memory, and processor in claim 16, reads to generalized computer components, based upon the claim interpretation wherein the structure is interpreted using [0036-40],[0051-62], and [0122-3] in the specification. The recitations of neural networks, student networks, and teacher networks, read to algorithms that can be adjusted and used to calculate specific output. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim(s) is/are directed to an abstract idea. The claim(s) do(es) not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using generalized computer components to determine, determine, update, update, receive, input, perform, and send, amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim(s) is/are not patent eligible. With respect to claim(s) 2, 3, and 10, the claim(s) recite(s) characteristics of the student and teacher networks, and (claims 3 and 10) determining a loss function and updating the parameters, which reads on the algorithms having specific structure/equations, and a human performing a specific calculation and using the result to adjust the equations in one of the algorithms. No additional limitations are present. With respect to claim(s) 4, 5, 11, and 12, the claim(s) recite(s) determine a component, which reads on a human performing specific calculations based on a particular set of data. No additional limitations are present. With respect to claim(s) 6, 7, 8, 13, 14, and 15, the claim(s) recite(s) performing pre-training and fine-tuning using specific information, which reads on a human iteratively adjusting the different sets of algorithms using specific information and variables for each iteration. No additional limitations are present. These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 9, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hu et al. (“DOMAIN ROBUST DEEP EMBEDDING LEARNING FOR SPEAKER RECOGNITION”, ICASSP 2022), hereinafter Hu, in view of Roblek et al. (U.S. PG Pub No. 2015/0294670), hereinafter Roblek. Regarding claims 1, 9, and 16, Hu teaches (claim 1) A speaker verification device comprising (speaker verification system (Intro)): (claim 9) A method for controlling a speaker verification device comprising speaker verification system (Intro)) (claim 16) A speaker verification system comprising (speaker verification system (Intro)): determine a first loss function based on a difference between outputs of the student network and the teacher network for unlabeled utterances (the student and teacher models are both fed unlabeled target domain data of utterances, i.e. the student network and the teacher network for unlabeled utterances, where the loss between the teacher and student outputs is determined, i.e. determine a first loss function based on a difference between outputs Fig. 1a,(Sec. 2)); determine a second loss function based on a difference between an output of the student network for labeled utterances and a one-hot encoding result for the labeled utterances (the student model is sent source labeled data, i.e. output of the student network for labeled utterances, to determine a cross entropy loss, i.e. determine a second loss function based on a difference, with one-hot ground truth labels, i.e. and a one-hot encoding result for the labeled utterances Fig. 1a,(Sec. 2)); update parameters of the student network based on the first and second loss functions (the student model is updated by stochastic gradient descent, where the total loss is a weighted sum of the losses from both domains Fig. 1a,(Sec. 2)); and update parameters of the teacher network based on an exponential moving average of the updated parameters of the student network (the parameters of the teacher model are updated by accumulating the parameters of the student model every iteration through an exponential moving average strategy Fig. 1a,(Sec. 2)). While Hu provides training the models for a speaker verification system, Hu does not specifically teach the memory, processor, or user terminal, and thus does not teach (claim 1) a memory configured to store a neural network model including a student network and a teacher network; and (claim 1) at least one processor configured to train the neural network model, wherein the at least one processor is configured to: (claim 9)…a memory configured to store a neural network model including a student network and a teacher network, and at least one processor configured to train the neural network model, the method comprising: (claim 16) a user terminal; and (claim 16) a speaker verification device which receives utterances and a speaker verification request from the user terminal, inputs the received utterances into a neural network model, performs speaker verification based on an output of the neural network model, and sends a speaker verification result to the user terminal, the speaker verification device including: (claim 16) a memory configured to store the neural network model; and (claim 16) at least one processor configured to train the neural network model, wherein the at least one processor is configured to: Roblek, however, teaches (claim 1) a memory configured to store a neural network model including a student network and a teacher network (speaker verification model is a trained neural network stored in a memory [0031-2]); and (claim 1) at least one processor configured to train the neural network model, wherein the at least one processor is configured to (processors execute the functions, which includes training the neural network [0027],[0119]): (claim 9)…a memory configured to store a neural network model including a student network and a teacher network (speaker verification model is a trained neural network stored in a memory [0031-2]), and at least one processor configured to train the neural network model, the method comprising (processors execute the functions, which includes training the neural network [0027],[0119]): (claim 16) a user terminal (a client device [0031-2]); and (claim 16) a speaker verification device which receives utterances and a speaker verification request from the user terminal, inputs the received utterances into a neural network model, performs speaker verification based on an output of the neural network model, and sends a speaker verification result to the user terminal, the speaker verification device including (the client device, such as a phone, i.e. speaker verification device…user terminal, receives a verification utterance when the user attempts to gain access, i.e. receives utterances and a speaker verification request from the user terminal, and the speaker verification model outputs an evaluation vector that corresponds to the verification utterance, i.e. inputs the received utterances into a neural network model, where the evaluation vector is compared with the reference vector to determine whether the verification utterance was spoken by the user, i.e. performs speaker verification based on an output of the neural network model, and the client device provides an indication that represents a verification result to the user, such as an output message on the display, i.e. sends a speaker verification result to the user terminal [0031-2],[0045-7]): (claim 16) a memory configured to store the neural network model (speaker verification model is a trained neural network stored in a memory [0031-2]); and (claim 16) at least one processor configured to train the neural network model, wherein the at least one processor is configured to (processors execute the functions, which includes training the neural network [0027],[0119]): Where Hu teaches that the neural network includes a student model and a teacher model. Fig. 1a,(Sec. 2) Hu and Roblek are analogous art because they are from a similar field of endeavor in training a neural network model for speaker verification. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the training the models for a speaker verification system teachings of Hu with the use of a client device as taught by Roblek. It would have been obvious to combine the references to enable providing client device access to a user and performing requests without requiring further user inputs (Roblek [0047]). Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hu, in view of Roblek, and further in view of Han et al. (“Self-Supervised Learning With Cluster-Aware-DINO for High-Performance Robust Speaker Verification”, IEEE/ACM, 2023), hereinafter Han. Regarding claim 2, Hu in view of Roblek teaches claim 1. While Hu in view of Roblek provides the use of a student model and a teacher model, Hu in view of Roblek does not specifically teach the structure of the models including encoders or projection heads, and thus does not teach each of the student network and the teacher network includes: an encoder configured to convert segments of utterances into speaker embeddings; and a projection head configured to map the speaker embeddings into a high-dimensional space, and the projection head includes: a multilayer perceptron; and a projection layer configured to receive a representation output from the multilayer perceptron and map the representation into a high-dimensional space. Han, however, teaches each of the student network and the teacher network includes: an encoder configured to convert segments of utterances into speaker embeddings (the student and teacher models include encoders that extract speaker embeddings from utterances Fig. 1,(Sec. III.A.)); and a projection head configured to map the speaker embeddings into a high-dimensional space (the speaker embeddings are passed through a projection head that provides two normalizations, i.e. map the speaker embeddings into a high-dimensional space Fig. 1,(Sec. III.A.)), and the projection head includes: a multilayer perceptron (the projection head includes a 3-layer perceptron Fig. 1,(Sec. III.A.)); and a projection layer configured to receive a representation output from the multilayer perceptron and map the representation into a high-dimensional space (the projection head includes a 3-layer perceptron, followed by two normalization layers to provide a final output, i.e. projection layer…map the representation into a high-dimensional space Fig. 1,(Sec. III.A.)). Hu, Roblek, and Han are analogous art because they are from a similar field of endeavor in training a neural network model for speaker verification. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the use of a student model and a teacher model teachings of Hu, as modified by Roblek, with the use of an encoder and projection head as taught by Han. It would have been obvious to combine the references to enable an improvement to performance when compared to the best-known self-supervised speaker verification systems (Han Abstract)). Claim(s) 6 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hu, in view of Roblek, and further in view of Sang et al. (“Open-set Short Utterance Forensic Speaker Verification using Teacher-Student Network with Explicit Inductive Bias”, arXiv:2009.09556v1, 21 Sep 2020), hereinafter Sang. Regarding claims 6 and 13, Hu in view of Roblek teaches claims 1 and 9. While Hu in view of Roblek provides training student and teacher models, Hu in view of Roblek does not specifically teach the use of penalties, and thus does not teach sequentially perform pre-training without a margin penalty and fine-tuning training with a margin penalty when updating the parameters of the student network and the teacher network. Sang, however, teaches sequentially perform pre-training without a margin penalty and fine-tuning training with a margin penalty when updating the parameters of the student network and the teacher network (the speaker classification loss is calculated after the initial training of the student model, i.e. sequentially perform pre-training, where the loss may be an angular softmax loss that maximizes the angular margin between speaker embeddings, i.e. margin penalty, where the pretrained student model is then fine-tuned by a set of loss values including the classification loss, i.e. pre-training without a margin penalty…fine-tuning training with a margin penalty, where the teacher and student models are trained independently, i.e. when updating the parameters of the student network and the teacher network Fig. 1,(Sec. 2, Table 1)). Hu, Roblek, and Sang are analogous art because they are from a similar field of endeavor in training a neural network for speaker verification. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the training student and teacher models teachings of Hu, as modified by Roblek, with the use of an angular softmax loss as taught by Sang. It would have been obvious to combine the references to achieve the best performances and highest reductions of EER when A-softmax is included (Sang (Sec. 4)). Claim(s) 7, 8, 14, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hu, in view of Roblek, in view of Sang, and further in view of Sadhukhan et al. (“KNOWLEDGE DISTILLATION INSPIRED FINE-TUNING OF TUCKER DECOMPOSED CNNs AND ADVERSARIAL ROBUSTNESS ANALYSIS”, IEEE, 2020), hereinafter Sadhukhan. Regarding claims 7 and 14, Hu in view of Roblek and Sang teaches claims 6 and 13, and Hu further teaches control a temperature of a … function, which controls sharpness of network outputs to be higher in the student network than in the teacher network during the pre-training (only the teacher model output prediction is smoothed, i.e. controls sharpness of network outputs to be higher in the student network than in the teacher network, using a temperature, i.e. control a temperature Fig. 1,(Sec. 3.1)). Where Sang teaches that the teacher model is part of the pre-training of the student model. Fig. 1. While Hu in view of Roblek and Sang provides softmax losses, Hu in view of Roblek and Sang does not specifically teach the temperature of a softmax function, and thus does not teach control a temperature of a softmax function…. Sadhukhan, however, teaches control a temperature of a softmax function… (the temperature parameter is an argument of the softmax (Sec. 3 Intro para)). Hu, Roblek, Sang, and Sadhukhan are analogous art because they are from a similar field of endeavor in neural networks to perform classification tasks. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the use of softmax losses teachings of Hu, as modified by Roblek and Sang, with the control of a temperature parameter that is an argument of the softmax as taught by Sadhukhan. It would have been obvious to combine the references to improve the accuracy and adversarial robustness of student networks (Sadhukhan Abstract). Regarding claim 8, Hu in view of Roblek, Sang, and Sadhukhan teaches claim 7, and Sadhukhan further teaches control the temperature to be the same for both the student network and the teacher network during the fine-tuning training (during fine-tuning of the student networks, i.e. during the fine-tuning training, the temperature of the softmax layer is kept as 1, and the temperature used for the KL-divergence loss between the softmax probabilities of the teacher and student is the value tau, i.e. control the temperature to be the same for both the student network and the teacher network (Sec. 3 Intro para)). Where the motivation to combine is the same as previously presented. Regarding claim 15, Hu in view of Roblek and Sang teaches claim 13. While Hu in view of Roblek and Sang provides controlling a temperature, Hu in view of Roblek and Sang does not specifically teach the temperature is the same during fine-tuning, and thus does not teach controlling the temperature to be the same for both the student network and the teacher network during the fine-tuning training (during fine-tuning of the student networks, i.e. during the fine-tuning training, the temperature of the softmax layer is kept as 1, and the temperature used for the KL-divergence loss between the softmax probabilities of the teacher and student is the value tau, i.e. controlling the temperature to be the same for both the student network and the teacher network (Sec. 3 Intro para)). Hu, Roblek, Sang, and Sadhukhan are analogous art because they are from a similar field of endeavor in neural networks to perform classification tasks. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the use of controlling a temperature teachings of Hu, as modified by Roblek and Sang, with the control of a temperature parameter specifically during fine-tuning as taught by Sadhukhan. It would have been obvious to combine the references to improve the accuracy and adversarial robustness of student networks (Sadhukhan Abstract). Allowable Subject Matter Claims 3-5 and 10-12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims, as well as being rewritten or amended to overcome the rejection(s) under 35 U.S.C. 101, set forth in this Office action. The following is a statement of reasons for the indication of allowable subject matter: The closest prior art of Hu teaches the use of losses in training the models. However, Hu does not teach the determination of a contrastive loss based on a cosine similarity specifically between the output of the student network multilayer perceptron and the speaker embedding output of the encoder of the teacher network, and using the loss function to update the parameters of the student network. Roblek teaches training a neural network, but does not teach the determination of a contrastive loss based on a cosine similarity specifically between the output of the student network multilayer perceptron and the speaker embedding output of the encoder of the teacher network, and using the loss function to update the parameters of the student network. Han teaches the determination of a cosine similarity between the encoders of the student and teacher models. However, Han does not teach the determination of a contrastive loss based on a cosine similarity specifically between the output of the student network multilayer perceptron and the speaker embedding output of the encoder of the teacher network, and using the loss function to update the parameters of the student network. Sang teaches the determination of a cosine distance between the embedding layers of the student and teacher models. However, Sang does not teach the determination of a contrastive loss based on a cosine similarity specifically between the output of the student network multilayer perceptron and the speaker embedding output of the encoder of the teacher network, and using the loss function to update the parameters of the student network. Sadhukhan teaches the determination of different losses for the training of teacher and student networks. However, Sadhukhan does not teach the determination of a contrastive loss based on a cosine similarity specifically between the output of the student network multilayer perceptron and the speaker embedding output of the encoder of the teacher network, and using the loss function to update the parameters of the student network. None of Hu, Roblek, Han, Sang, and Sadhukhan, either alone or in combination, teaches or makes obvious the determination of a contrastive loss based on a cosine similarity specifically between the output of the student network multilayer perceptron and the speaker embedding output of the encoder of the teacher network, and using the loss function to update the parameters of the student network. Therefore, none of the cited prior art either alone or in combination, teaches or makes obvious the combination of limitations as recited in the dependent claims including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICOLE A K SCHMIEDER whose telephone number is (571)270-1474. The examiner can normally be reached 8:00 - 5:00 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NICOLE A K SCHMIEDER/Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jan 07, 2025
Application Filed
Sep 04, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12730984
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND STORAGE MEDIUM
2y 9m to grant Granted Sep 08, 2026
Patent 12730965
HIGH-PERFORMANCE MICROCODED TEXT PARSER
2y 5m to grant Granted Sep 08, 2026
Patent 12718018
TEXT SPAN PREDICTION BASED ON ENTITY TYPE
2y 9m to grant Granted Aug 25, 2026
Patent 12706196
INFORMATION TERMINAL, RECORDING MEDIUM, INFORMATION PROCESSING SYSTEM, AND INFORMATION PROCESSING METHOD
3y 7m to grant Granted Aug 11, 2026
Patent 12705435
TEXTUAL INPUT ANALYSIS METHODS AND SYSTEMS FOR DETERMINING DEGREE OF CORRECTNESS
2y 5m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
68%
Grant Probability
99%
With Interview (+34.0%)
2y 8m (~12m remaining)
Median Time to Grant
Low
PTA Risk
Based on 176 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month