DETAILED ACTION
Response to Arguments
Claims 1-20 are currently pending. Claims 1, 10, and 19 were amended.
Re: Claim Rejections Under 35 U.S.C. § 103
Applicant argues on pg. 8 of the response (“REMARKS”) filed on June 5, 2026 that MUNTASIR (US 2025/0118298) applied in the 35 U.S.C. § 103 rejection discloses performing authentication before audio processing (e.g., sound cancelling techniques to isolate audio data) is performed. However, the Examiner respectfully disagrees. Voice biometrics are used to authenticate a user according to [MUNTASIR, ¶0031]:
Referring back to FIG. 1, the voice signal received from the user's 101 speech input is further analyzed using the voice biometrics module 104e. The voice biometrics module 104e corresponds to a voice biometrics system adapted to authenticate a user based on speech diagnostics corresponding to the received speech audio input. The voice biometrics module 104e further connects to and performs user authentication for, the user identification module 111. The user identification module authenticates the user to a user interaction session using the user's caller number or a unique identification number assigned to the user.
Therefore, speech/audio data must be recorded and processed in order to perform the voice biometrics. Also see [MUNTASIR, ¶0031]: “…authenticating a user to the user interaction session using the user's caller number or a unique identification number assigned to the user further comprises the step of utilizing voice biometrics to identify and register the user participating in the user interaction session.” It is also noted that MUNTASIR was not previously used to teach features for isolating the audio/speech prior to authentication, which were disclosed in RATHAUR (US 2022/036615). These features were further amended to describe generating a sound bubble around the user and the mobile device of the user in part of the sound cancelling techniques. Applicant argued that RATHAUR does not disclose the amended features. Examiner agreed that RATHAUR does not disclose the amended features and has withdrawn the rejection of MUNTASIR in view of RATHAUR. However, a new ground of rejection is asserted in the current Office Action. See Claim Rejections - 35 USC § 103 below for details.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 2, 5-7, 10-11, 14-16, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over US 2025/0118298 to Muntasir et al. (hereinafter, “MUNTASIR”) in view of US 2016/0019026 to Gupta et al. (hereinafter, “GUPTA”).
As per claim 1: MUNTASIR discloses: A computing platform, comprising: at least one processor; a communication interface communicatively coupled to the at least one processor; and a memory storing computer-readable instructions that, when executed by the at least one processor, cause the computing platform to (an interaction voice response (IVR) communication system 100 comprising various components [MUNTASIR, ¶0020; Fig. 1]): receive audio data from a user, wherein the audio data is captured by a mobile device of the user and the audio data is received via the mobile device of the user (a user using a client device that can be a smartphone, feature phone, or any electronic device [MUNTASIR, ¶0022]; receiving conversation data [MUNTASIR, ¶0024]); initiate, based on the received audio data, a transaction session (initiating a call by the user to communicate with the IVR communication system [MUNTASIR, ¶0024]; receiving speech input of the user [MUNTASIR, ¶0028, 0032]); (user identification module 111 of the IVR distinguishes the user’s human voice in the received conversation data from the user’s stored voice biometrics [MUNTASIR, ¶0024, 0031]); responsive to determining that the user is not authenticated to the transaction session, terminate the transaction session; responsive to determining that the user is authenticated to the transaction session (initiating the conversation controller module 109 once the user is successfully authenticated into the IVR communication system [MUNTASIR, ¶0025]; the opposite would have been inherently deny initiating with the IVR communication system when the user fails to authenticate): extract, from the audio data, features, wherein extracting the features results in an audio signal formatted for further processing (the conversation controller module receives audio features [MUNTASIR, ¶0026]); execute one or more speech recognition techniques on the audio signal to generate a plurality of phonetic units (choosing an automated speech recognition (ASR) model with determining speech segments, non-speech segments, turn-taking speech segments and barge-in speech segments to accomplish a targeted service [MUNTASIR, ¶0026]); execute one or more machine learning models, wherein executing the one or more machine learning models includes inputting, to the one or more machine learning models, the plurality of phonetic units to generate an output (a natural language unit (NLU) is further included that is capable of determining and generating transcribed context, intent, use-cases, entities, and metadata of the conversation with the user [MUNTASIR, ¶0041]; furthermore, a large language model (LLM) is also included to determine use-case from the user's speech input as conversational data and generate a reply text based on the conversational data information received [MUNTASIR, ¶0063]); and transmit, to the mobile device of the user, the generated output for confirmation (a dialogue engine dispatcher 106 is configured to receive the response action input and deliver a corresponding response in the form of one or more text message to a TTS module for outputting to the user [MUNTASIR, ¶0059-0060]).
MUNTASIR does not explicitly disclose, but GUPTA discloses: execute sound cancelling techniques to isolate the audio data, wherein executing the sound cancelling techniques includes generating a sound bubble around the user and the mobile device of the user (blind source separation (BSS) technique is used to provide a speech interface to any device for separating (i.e., isolating) speech/voice data for recognition [GUPTA, ¶0016]; separation bubbles are generated around a recognition device for a user to separate one source from other sources and noise [GUPTA, ¶0022, 0039]).
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to filter the noise of the user’s speech input in MUNTASIR by utilizing separation bubbles as suggested in GUPTA. Noise cancellation would have reduced and/or removed unwanted sound from the speech input to enhance the clarity of said input for recognizing a particular user’s speech in voice biometrics.
As per claim 2: MUNTASIR in view of GUPTA disclose all limitations of claim 1. Furthermore, MUNTASIR discloses: further including instructions that, when executed, cause the computing platform to: receive, in response to transmitting the generated output, feedback data (a feedback classifier model included using the user's reaction compared to his experience in a happiness-index feedback block, wherein parameters are updated in a corresponding user profile to be used for future optimization [MUNTASIR, ¶0057]); and update the one or more machine learning models based on the feedback data (updating and training the ASR and NLU models associated with the user profile [MUNTASIR, ¶0098]).
As per claim 5: MUNTASIR in view of GUPTA disclose all limitations of claim 1. Furthermore, MUNTASIR discloses: further including instructions that, when executed, cause the computing platform to: post-process the output to improve accuracy of the output (after receiving a generated text response from the LLM module, the action server performs post-processing to refine the generated text [MUNTASIR, ¶0064]).
As per claim 6: MUNTASIR in view of GUPTA disclose all limitations of claim 5. Furthermore, MUNTASIR discloses: wherein the post-processing includes error correction (post-processing includes removing redundant information [MUNTASIR, ¶0064]).
As per claim 7: MUNTASIR in view of GUPTA disclose all limitations of claim 5. Furthermore, MUNTASIR discloses: wherein the post-processing is performed prior to transmitting the generated output for confirmation (the action server performs the post-processing within a dialogue engine and prior to sending the response to the dispatcher, which generates and delivers a corresponding response [MUNTASIR, ¶0089, 0092; Fig. 1a]).
As per claim 10: Claim 10 is different from overall scope from claim 1. Claim 10 recites a method comprising the same steps as the functions of the computing platform recited in claim 1. Therefore, the response provided for claim 1 is equally applicable to claim 10.
As per claim 11: Claim 11 incorporates all limitations of claim 10. Claim 11 recites a method comprising the same steps as the functions of the computing platform recited in claim 2. Therefore, the responses provided for claims 2 and 10 are applicable to claim 11.
As per claim 14: Claim 14 incorporates all limitations of claim 10. Claim 14 recites a method comprising the same steps as the functions of the computing platform recited in claim 5. Therefore, the responses provided for claims 5 and 10 are applicable to claim 14.
As per claim 15: Claim 15 incorporates all limitations of claim 14. Claim 15 recites a method comprising the same steps as the functions of the computing platform recited in claim 6. Therefore, the responses provided for claims 6 and 14 are applicable to claim 15.
As per claim 16: Claim 16 incorporates all limitations of claim 14. Claim 16 recites a method comprising the same steps as the functions of the computing platform recited in claim 7. Therefore, the responses provided for claims 7 and 14 are applicable to claim 16.
As per claim 19: Claim 19 is different from overall scope from claim 1. Claim 19 recites one or more non-transitory computer-readable media storing executable instructions to cause a computing platform to perform the same steps as the functions of the computing platform recited in claim 1. Therefore, the response provided for claim 1 is equally applicable to claim 19.
As per claim 20: Claim 20 incorporates all limitations of claim 19. Claim 18 recites a one or more non-transitory computer-readable media storing executable instructions to cause a computing platform to perform the same steps as the functions of the computing platform recited in claim 2. Therefore, the responses provided for claims 2 and 19 are applicable to claim 20.
Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over MUNTASIR in view of GUPTA and in further view of US 2022/0139388 to Sharifi et al. (hereinafter, “SHARIFI”).
As per claim 3: MUNTASIR in view of GUPTA disclose all limitations of claim 1, MUNTASIR and GUPTA do not explicitly disclose, but SHARIFI discloses: wherein extracting the features includes executing one or more of: mel-frequency cepstral coefficients (MFCCs) or spectrograms (detecting acoustic features in audio data that may be mel-frequency cepstral coefficients that are representations of short-term power spectrums of utterances [SHARIFI, ¶0016]; producing filtered spectrograms through a generative adversarial network in the voice filter engine [SHARIFI, ¶0035]).
Thus, it would have obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to incorporate MFCCs and/or spectrograms in MUNTASIR into the ASR. MFCCs and spectrograms were commonly used features in conventional speech recognition processes. It would have been obvious to try these techniques with reasonable success as they were some of the well-established forms of audio representations known in the field.
As per claim 12: Claim 12 incorporates all limitations of claim 10. Claim 12 recites a method comprising the same steps as the functions of the computing platform recited in claim 3. Therefore, the responses provided for claims 3 and 10 are applicable to claim 12.
Claims 4, 8, 13, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over MUNTASIR in view of GUPTA and in further view of US 12,531,056 to Liu et al. (hereinafter, “LIU”).
As per claim 4: MUNTASIR in view of GUPTA disclose all limitations of claim 1. MUNTASIR and GUPTA do not explicitly disclose, but LIU discloses: wherein executing the one or more speech recognition techniques includes executing one or more of deep learning models, acoustic models or language models (automated speech recognition can use a language model, such as a large language model built using deep learning, to processing spoken inputs [LIU, col. 2, lines 11-34]; acoustic models may include models corresponding to speech, noise, or silence to determine whether speech is present in audio data [LIU, col. 19, lines 1-20]).
Thus, it would have obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to incorporate any known type of machine learning models for the ASR in MUNTASIR. Machine learning models were (and currently still is) a fast-growing field due to the increase of available computing resources over the years. The accuracy of these models is consistently improving, and it would have advantageous to incorporate them into speech recognition systems.
As per claim 8: MUNTASIR in view of GUPTA disclose all limitations of claim 1. MUNTASIR and GUPTA do not explicitly disclose, but SHARIFI discloses: wherein the mobile device is a wearable device (an assistant-enabled device that a user may interact with through speech can include a smart watch [SHARIFI, ¶0015]).
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to modify the client device in MUNTASIR to be a smart watch. According to [MUNTASIR, ¶0022], the client device may “correspond to a wide variety of electronic devices” and include smartphones in one example embodiment. Wearable smart watches offer similar functions of a smartphone in addition to tracking and monitoring the user’s health.
As per claim 13: Claim 13 incorporates all limitations of claim 10. Claim 13 recites a method comprising the same steps as the functions of the computing platform recited in claim 4. Therefore, the responses provided for claims 4 and 10 are applicable to claim 13.
As per claim 17: Claim 17 incorporates all limitations of claim 10. Claim 17 recites a method comprising the same steps as the functions of the computing platform recited in claim 8. Therefore, the responses provided for claims 8 and 10 are applicable to claim 17.
Claims 9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over MUNTASIR in view of GUPTA and in further view of US 2026/0010713 to Zhou (hereinafter, “ZHOU”).
As per claim 9: MUNTASIR in view of GUPTA disclose all limitations of claim 1. MUNTASIR and GUPTA do not explicitly disclose, but ZHOU discloses: wherein the executing the one or more speech recognition techniques and the executing the one or more machine learning models is performed within an artificial intelligence trust, risk and security management (AITRiSM) framework (a generative machine learning framework helps mitigate the risk of system failures, transaction errors, and security breaches [ZHOU, ¶0139]; the generative ML framework may enable robust security and compliance by prioritizing security and compliance measures to safeguard sensitive data [ZHOU, ¶0145]).
Thus, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to incorporate a secure framework into the system of MUNTASIR for machine learning processing of the ASR, NLU, and LLM models. As personal data (i.e., a user’s speech) are collected and processed, it would have been advantageous to implement security protocols to preserve the privacy of the user for the speech recognition and processing system in MUNTASIR.
As per claim 18: Claim 18 incorporates all limitations of claim 10. Claim 18 recites a method comprising the same steps as the functions of the computing platform recited in claim 9. Therefore, the responses provided for claims 9 and 10 are applicable to claim 18.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 2024/0212689: Zones separate speech by location of each user, wherein the zones apply speaker-specific speech enhancement for isolating a user’s speech. See ¶0082.
US 2020/0245065: Virtual microphones are established on a per configured spatial region basis for signal processing both desired and undesired sound sources in 3D space. See ¶0016.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ROBERT B LEUNG whose telephone number is (571)270-1453. The examiner can normally be reached Mon - Thurs: 10am-7pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, JUNG KIM can be reached at 571-272-3804. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ROBERT B LEUNG/Primary Examiner, Art Unit 2494