DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to the amendment filed July 7, 2026. Claims 1, 10 and 16 have been amended. Claims 1, 2, 4, 10-11, 13, and 16-18 remain pending.
Claim Rejections - 35 USC § 101
The rejections under 35 USC 101 have been withdrawn. Paragraphs [0015-0016] of the specification describes a technological improvement achieved by the invention. Thus, the abstract idea of the claims provide an improvement in the field of speech intelligibility analysis, and the rejection is withdrawn.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 4, 10-11, 13, 16-18 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Thomson (US Patent Application Publication No. 2021/0375288) in view of Stonehocker et al (US Patent Application Publication No. 2023/0082955), hereinafter Stonehocker, in view of Proenca et al (US Patent Application Publication No. 2021/0225389), hereinafter Proenca.
Thomson discloses a method for transcription generation technique selection. Regarding claim 1, Thomson teaches providing the speech recording [para 0020 – audio]; generating a first transcript for the speech recording using a first automatic speech recognition (ASR) model [Fig 1 (134); para00015; 0025-0028; 0036-0037 – first transcription generation technique]; providing a second transcript for the speech recording [Fig 1 (136); para00015; 0025-0028; 0036-0037 – second transcription generation technique]; wherein the second ASR model has a different transcription accuracy than the first ASR model [Thomson para 0025-0028; 0036-0037; 0040-0041]; comparing the first transcript and the second transcript [para 0036 -- monitor comparisons between the performance of the first transcription generation technique 134 and the performance of the second transcription generation technique 136]. Thomson fails to teach having a the first and second ASR models having first and second number of layers and a first and second number of parameters, where the first and second ASR models are selected based on the first ASR model being more accurate than the second ASR models as indicated by the first number of layers being greater than the second number of layers and the first number of parameters being greater than the second number of parameters. In a similar field of endeavor, Stonehocker teaches automatic speech recognition wherein one or more acoustic models may include a first acoustic model that uses a first neural network (NN) having a first number of hidden layers and a second acoustic model may include a second NN having a second number of hidden layers that is less than the first number [para 0025; 0058; 0062]; measures model performance for accuracy rate and latency [para 0063]; and provides for selecting models based on service levels and accuracy [para 0070] and specifically teaches the system is advantageous in providing different service levels for automatic speech recognition or for other services [para 0001]. One having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the differing layers/parameters for different models and selecting the different models based on service/accuracy as suggested by Stonehocker, in the system of Thomson, and the results would have been predictable and provided an improved and efficient recognition system that provides different servicing levels based on system needs, and thereby improve system performance and the user’s experience. Thomson fails to teach determining an intelligibility score indicating intelligibility of the speech based on determining a similarity metric between the first transcript and the second transcript. In a similar field of endeavor, Proenca generates an intelligibility score representing the intelligibility of the utterance based on the output and the sample text, where generating the intelligibility score may involve (1) calculating conditional intelligibility value(s) for the N recognition result(s), and (2) determining the intelligibility score based on the conditional intelligibility value of the most intelligible recognition result [para 0017-0018; 0047-0061] and teaches the intelligibility score represents the intelligibility of the user’s utterance [para 0073]. One having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the intelligibility scoring processing suggested by Proenca, in the system of Thomson, for the purpose of ensuring the best transcription mode is selected for the provided audio, thereby providing the most accurate transcriptions to the user.
Regarding claim 2, the combination of Thomson, Stonehocker and Proenca teaches determining a similarity metric between the first transcript and the second transcript is calculated based on assigning one transcript selected from the group consisting of the first transcript and the second transcript as a ground truth [Proenca at para 0047-0061—scoring based conditional intelligibility].
Regarding claim 4, the combination of Thomson, Stonehocker and Proenca teaches converting the similarity metric is based on threshold values of the similarity metric [Proenca at para 0047-0061].
Claims 10, 11, 13, and 16-18 are rejected under similar rationale as claims 1, 2, and 4.
Response to Arguments
Applicant's arguments filed July 7, 2028 with respect to the rejections under 35 USC 101 have been fully considered but they are not persuasive.
Applicant argues the combination of Thomson, Stonehocker, and Proenca fails to teach or suggest the claimed invention as a whole. The Examiner respectfully disagrees. As indicated in the rejections above, Thomson teaches providing the speech recording; generating a first transcript for the speech recording using a first automatic speech recognition (ASR) model; providing a second transcript for the speech recording; wherein the second ASR model has a different transcription accuracy than the first ASR model and comparing the first transcript and the second transcript. Thomson fails to teach having a the first and second ASR models having first and second number of layers and a first and second number of parameters, where the first and second ASR models are selected based on the first ASR model being more accurate than the second ASR models as indicated by the first number of layers being greater than the second number of layers and the first number of parameters being greater than the second number of parameters. Stonehocker was cited as teaching automatic speech recognition wherein one or more acoustic models may include a first acoustic model that uses a first neural network (NN) having a first number of hidden layers and a second acoustic model may include a second NN having a second number of hidden layers that is less than the first number [para 0025; 0058; 0062]; measures model performance for accuracy rate and latency [para 0063]; and provides for selecting models based on service levels and accuracy [para 0070] and specifically teaches the system is advantageous in providing different service levels for automatic speech recognition or for other services [para 0001]. One having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the differing layers/parameters for different models and selecting the different models based on service/accuracy as suggested by Stonehocker, in the system of Thomson, and the results would have been predictable and provided an improved and efficient recognition system that provides different servicing levels based on system needs, and thereby improve system performance and the user’s experience. Thomson fails to teach determining an intelligibility score indicating intelligibility of the speech based on determining a similarity metric between the first transcript and the second transcript. Proenca generates an intelligibility score representing the intelligibility of the utterance based on the output and the sample text, where generating the intelligibility score may involve (1) calculating conditional intelligibility value(s) for the N recognition result(s), and (2) determining the intelligibility score based on the conditional intelligibility value of the most intelligible recognition result [para 0017-0018; 0047-0061] and teaches the intelligibility score represents the intelligibility of the user’s utterance [para 0073]. One having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the intelligibility scoring processing suggested by Proenca, in the system of Thomson, for the purpose of ensuring the best transcription mode is selected for the provided audio, thereby providing the most accurate transcriptions to the user.
Applicant appears to argue the combination of the applied prior art fails to teach “when speech is highly intelligible, both models produce similar transcripts, whereas when speech is less intelligible, the transcripts diverge, and this divergence is used to determine an intelligibility score.” In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., when speech is highly intelligible, both models produce similar transcripts, whereas when speech is less intelligible, the transcripts diverge, and this divergence is used to determine an intelligibility score) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). In this instance, the claims broadly recite a similarity metric between the transcripts and using the metric to determine intelligibility. As indicated in the rejection and argued above, the combination of Thomson, Stonehocker, and Proenca provide adequate support for the invention as claimed.
Applicant argues “The Examiner has not articulated a sufficient reason why one of ordinary skill in the art would have been motivated to combine the teachings of Thomson, Stonehocker, and Proenca in the manner claimed. The proposed combination appears to be based on impermissible hindsight reconstruction of the claimed invention. In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANGELA A ARMSTRONG whose telephone number is (571)272-7598. The examiner can normally be reached M,T,TH,F 11:30-8:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
ANGELA A. ARMSTRONG
Primary Examiner
Art Unit 2659
/ANGELA A ARMSTRONG/Primary Examiner, Art Unit 2659