Prosecution Insights
Last updated: August 17, 2026
Application No. 18/930,230

VOICE APPLICATION PROTECTION

Non-Final OA §101§103
Filed
Oct 29, 2024
Examiner
ANKRUM, ALEC CHRISTOPHER
Art Unit
2434
Tech Center
2400 — Computer Networks
Assignee
Intuit Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-58.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
12 currently pending
Career history
13
Total Applications
across all art units

Statute-Specific Performance

§101
9.4%
-30.6% vs TC avg
§103
59.4%
+19.4% vs TC avg
§112
28.1%
-11.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Status Claims 1-20 are under examination. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-3, 5-8, 14, and 18-20 are rejected under 35 U.S.C. § 101 under the 2019 PEG framework. The claims recite a judicial exception (an abstract idea falling within the “data analysis” groupings) that is not integrated into a practical application. Step 1: Statutory Category. Claims 1-20 satisfy the statutory category requirement because it is directed to a method and system under 35 U.S.C. § 101(a). The claims recite a series of steps involving receiving data (audio), sending the data to an LLM, receiving a response from the LLM and performing a preemptive action. Claim 20’s steps being preformed on a computer system’s processor based on instructions in memory. Step 2A, Prong 1: Identification of Judicial Exception. Claim 1 and 20 recite a judicial exception within the abstract idea category. First, the claims recite data analysis and information processing: “receiving an audio transmission”, “providing a request” receiving a response”, and “performing … actions” constitute mental processes and data manipulation that could be performed in the human mind. Step 2A, Prong 2: Integration into a Practical Application. Claim 1 and 20 fail integration analysis. The claims do not recite any specific technological improvement to computer functionality. The claims merely state “performing one or more preemptive actions” without describing the model’s innovation or technological contribution, which is insufficient under Alice and Mayo. The preambles recite “processors”, “memory”, and “instructions” but these are generic computer implementation language; there is no claim to specialized hardware, FPGA, ASIC, or particular machine architecture. Per Alice, 573 U.S. at 221, merely implementing an abstract idea on a generic computer does not confer eligibility. The claim does not transform a tangible article into a different state or thing; data manipulation alone—receiving audio, providing a request, receiving a response, and preemptive action—is not transformation. Per Bilski, 561 U.S. at 618, and PEG p. 56, transformation requires a change in physical properties or state of a tangible article. The additional elements (processor, memory, computer system, LLM, computing device) are routine, conventional steps in cybersecurity and constitute insignificant extra-solution activity. Per PEG p. 56, field-of-use limitations do not suffice. The claim does not recite a specific technological problem solved; the specification describes a business problem (LLM systems needing their inputs and outputs properly protected) but not how the claimed method and system improve upon existing cybersecurity tools for LLMs in a non-conventional manner. Step 2B: Significantly More / WURC Analysis. The additional elements— processor, memory, computer system, LLM, and computing device—are all well-understood, routine, and conventional (WURC) in the field of cybersecurity and LLM protection as of the priority date of October 29, 2024. Processors, memory, computer systems, LLMs, and computing device are standard, commercially available components. The specification provides no factual evidence demonstrating that these elements, individually or in combination, represent a non-conventional or inventive approach. Under Berkheimer v. HP, Inc., 881 F.3d 1360 (Fed. Cir. 2018), the examiner would need to establish a factual record that these elements are not WURC, and no such record exists. Therefore, the claim fails Step 2B as well. A similar analysis can be applied to dependent claims 2-3, 5-8, 14, and 18-19. Claims 2-3, 5-8, 14, and 18-19 further recite and describe data manipulation using adversarial/genuine data, unauthorized/authorized requests/responses, and AI based application interfaces that are routine and conventional, therefore they are directed to a judicial exception. The judicial exception is not integrated into a practical application and the claims do not recite additional elements that amount to significantly more than the judicial exception. Claims 4, 9-13, and 15-17 are excluded from this rejection. Claims 4, 9, 13, and 15-17 integrate the abstract ideas into a practical application as claim 4 has a firewall between the application and the LLM which is significantly more and claims 9, 13, and 15-17 clearly recite what the preemptive action is and integrate it into a practical application. Claims 10-12 are excluded based on their dependency of claim 9. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-3, 5-7, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Le Roux et al. (United States Patent Publication No. 2021/0319784), hereinafter Le Roux. Regarding claim 1, a method for protecting a voice application (Le Roux ¶5: “Embodiments of the present disclosure relate to systems and methods for detecting adversarial attacks on a linguistic system and providing a defense against such adversarial attacks.”) communicably coupled to a large language model (LLM) (Le Roux ¶17: “the linguistic system produces each transcription subject to one or more language models. The language model can be used directly or indirectly with a transcribing neural network to produce accurate transcription of the input.”), the method performed by one or more processors of a computing system and comprising (Le Roux ¶123: “The system 300 includes a processor 319 configured to execute stored instructions 321, as well as a memory 323 that stores instructions that are executable by the processor 319.”): receiving an audio transmission over a communications network (Le Roux ¶118: “The input may be obtained wirelessly through a communication channel.”) from a computing device associated with a user of the voice application (Le Roux ¶141: “The smart home system may be implemented with the proposed linguistic system 100. Therefore, when the voice command from the user or generally some audio signal that triggers the smart home system is received by the smart home system it is first provided to the proposed linguistic system…”); providing a request to the LLM based on the audio transmission (Le Roux ¶11: “the linguistic system on reception of the input signal, executes the neural network multiple times to produce multiple transcriptions corresponding to the input signal.”); receiving a response to the request from the LLM (Le Roux ¶69: “When the classifier module 109 classifies the audio input as the legitimate input, the linguistic system outputs a transcription of the input.”); and performing one or more preemptive actions based on anticipating an anomaly in at least one of the request or the response (Le Roux ¶69: “Otherwise, when the input is classified as the illegitimate input, the linguistic system executes a counter-measure routine.”). Le Roux does not disclose all the above in the same embodiment, however it would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to combine the embodiments of Le Roux since it is suggested by Le Roux ¶143: “The description provides exemplary embodiments only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the following description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing one or more exemplary embodiments. Contemplated are various changes that may be made in the function and arrangement of elements without departing from the spirit and scope of the subject matter disclosed as set forth in the appended claims.” This motivation to combine the embodiments of Le Roux applies to the dependent claims as well. Claim 20 is substantially similar to claim 1 and is rejected under the same rationale. In addition, Le Roux teaches claim 20’s a system for protecting a voice application communicably coupled to a large language model (LLM), the system comprising: one or more processors; and at least one memory coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the system to perform operations (Le Roux ¶123: “The system 300 includes a processor 319 configured to execute stored instructions 321, as well as a memory 323 that stores instructions that are executable by the processor 319.”). Regarding claim 2, the method of claim 1, wherein the audio transmission includes a genuine portion and an adversarial portion, wherein the adversarial portion is a perturbation that modifies, combines with, or replaces the genuine portion (Le Roux 41: “This adversarial input can be used to manipulate the linguistic system to generate malicious or illegitimate transcriptions. The adversarial input comprises small imperceptible perturbations added to an original input waveform (or legitimate input).”). Regarding claim 3, the method of claim 2, wherein the perturbation is background noise injected into the audio transmission (Le Roux ¶140: “It is possible that the adversaries may use an adversarial attack such that the smart home system may add noise to the received voice command which manipulates the smart home system to execute malicious activity … Adversaries may broadcast via a radio or a TV program a speech signal that is perceived by a user as innocuous, but is in fact an adversarial attack targeted at the smart home system to execute malicious activity.”). Regarding claim 5, the method of claim 1, wherein the voice application is an artificial intelligence (AI)-based application that provides an interface for the user to submit requests to the LLM (Le Roux ¶58: “The linguistic system 100 further includes a processor in communication with a metric calculation module 107, a classifier module 109, and a decision module 111. Some or each of these modules is implemented using neural networks. On reception of the input 101 at an input interface of the linguistic system 100, the neural network 103 is executed multiple times…”). Regarding claim 6, the method of claim 1, wherein the audio transmission includes a genuine portion and an adversarial portion (Le Roux ¶41: “The adversarial input comprises small imperceptible perturbations added to an original input waveform (or legitimate input).”), wherein the genuine portion includes an authorized request from the user (Le Roux ¶140: “a user may be living in a smart home where a smart home system enables the user to use voice commands in order to control different elements of the smart home.”), wherein the adversarial portion includes an unauthorized request from a third party (Le Roux ¶140: “It is possible that the adversaries may use an adversarial attack such that the smart home system may add noise to the received voice command which manipulates the smart home system to execute malicious activity … Adversaries may broadcast via a radio or a TV program a speech signal that is perceived by a user as innocuous, but is in fact an adversarial attack targeted at the smart home system to execute malicious activity.”), and wherein the request provided to the LLM includes a combination of the authorized request and the unauthorized request (Le Roux ¶140-141: “when the voice command from the user or generally some audio signal that triggers the smart home system is received by the smart home system it is first provided to the proposed linguistic system 100 to classify the voice command as legitimate or illegitimate. On receiving the voice command, the linguistic system 100 produces multiple internal transcriptions of the voice command by executing the neural network 103 multiple times.”). Regarding claim 7, the method of claim 6, wherein the response includes at least an unauthorized response to the unauthorized request (Le Roux ¶53: “Adversaries typically exploit loopholes within the neural network by crafting an input perturbation such that small finely-tuned differences accumulate within the neural network to eventually result in a malicious output.”). Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Venkata (United States Patent Publication No. 2026/0141061). Regarding claim 8, Le Roux teaches the method of claim 1, but fails to explicitly teach further comprising: retrieving user data associated with the user from a user database; and including the user data with the request provided to the LLM. However, Venkata teaches retrieving user data associated with the user from a user database (Venkata ¶27: “The electronic device includes a data integration engine configured to obtain a user-specific data from a plurality of sources for a plurality of users. Further, the electronic device includes an LLM engine configured to utilize an LLM technique and a natural language processing framework to process and analyse textual interactions included in the obtained user-specific data.”); and including the user data with the request provided to the LLM (Venkata ¶27: “the electronic device includes an anomaly detection engine configured to trigger anomaly detection using the LLM technique by comparing the user behaviour with the collect user-specific data.”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Venkata to utilize user stored behavior for greater adaptability of anomaly detection (Venkata ¶02: “Conventional fraud prevention systems often rely on static rules or predefined patterns, which may not effectively adapt to evolving fraud techniques. Although machine learning-based approaches offer greater adaptability, they may face challenges in capturing the subtleties of user behavior”). Claims 4, 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Kawasaki et al. (United States Patent Publication No. 2025/0337775), hereinafter Kawasaki. Regarding claim 4, Le Roux teaches the method of claim 1, but fails to explicitly teach wherein the computing system is at least one of an artificial intelligence (AI) firewall communicably coupled between the voice application and the LLM or integrated with the voice application. However, Kawasaki teaches wherein the computing system is at least one of an artificial intelligence (AI) firewall communicably coupled between the voice application and the LLM or integrated with the voice application (Kawasaki ¶45: “one or more of the analysis engine 152 and the remediation engine 180 can be encapsulated or otherwise within the proxy 150. In this arrangement, the local analysis engine 152 can analyze inputs and/or outputs of the MLA 130 in order to determine, for example, whether to pass on such inputs and/or outputs to the monitoring environment 160 for further analysis.”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Kawasaki to prevent attacks on artificial intelligence models before they happen (Kawasaki ¶16: “The subject matter described herein provides many technical advantages. For example, the current subject matter can be used to identify and stop adversarial prompt injection attacks on artificial intelligence models including large language models. Further, the current subject matter can provide enhanced visibility into the health and security of an enterprise's machine learning assets.”). Regarding claim 9, Le Roux teaches the method of claim 1, but fails to teach wherein: the LLM is a multimodal LLM (MLLM); the request provided to the MLLM is a voice request; anticipating the anomaly includes, prior to providing the request to the MLLM: processing the request using an audio analysis model; and detecting the anomaly in the audio transmission based on results from the audio analysis model; and performing one or more preemptive actions includes, prior to providing the request to the MLLM, removing the detected anomaly from the request. However, Kawasaki teaches the LLM is a multimodal LLM (MLLM); the request provided to the MLLM is a voice request (Kawasaki ¶37: “the MLA 130 can be a multimodal model that accepts multiple mode (i.e., two or more of audio, video, text, images, etc.) input/prompts.”); anticipating the anomaly includes, prior to providing the request to the MLLM: processing the request using an audio analysis model; and detecting the anomaly in the audio transmission based on results from the audio analysis model (Kawasaki ¶40: “The analysis engine 170 can analyze the relayed queries and/or information in order to make an assessment or other determination as to whether the queries are indicative of being malicious. In some cases, a remediation engine 180 which can form part of the monitoring environment 160 (or be external such as illustrated in FIG. 2) can take one or more remediation actions in response to a determination of a query as being malicious. These remediation actions can take various forms including transmitting data to the proxy 150 which causes the query to be blocked before ingestion by the MLA 130.”); and performing one or more preemptive actions includes, prior to providing the request to the MLLM, removing the detected anomaly from the request (Kawasaki ¶40: “the remediation engine 180 can cause data to be transmitted to the proxy 150 which causes the query to be modified in order to be non-malicious, to remove sensitive information, and the like. Such queries, after modification, can be ingested by the MLA 130 and the output provided to the requesting client device 110.”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Kawasaki to prevent attacks on artificial intelligence models before they happen (Kawasaki ¶16: “The subject matter described herein provides many technical advantages. For example, the current subject matter can be used to identify and stop adversarial prompt injection attacks on artificial intelligence models including large language models. Further, the current subject matter can provide enhanced visibility into the health and security of an enterprise's machine learning assets.”). Regarding claim 10, Le Roux and Kawasaki teach the method of claim 9, but Le Roux fails to teach wherein the audio analysis model includes at least one of an artificial intelligence (AI) firewall, an audio filter, a feature extraction operation, or a signal processing application. However, Kawasaki teaches wherein the audio analysis model includes at least one of an artificial intelligence (AI) firewall (Kawasaki ¶45: “one or more of the analysis engine 152 and the remediation engine 180 can be encapsulated or otherwise within the proxy 150. In this arrangement, the local analysis engine 152 can analyze inputs and/or outputs of the MLA 130 in order to determine, for example, whether to pass on such inputs and/or outputs to the monitoring environment 160 for further analysis.”), an audio filter, a feature extraction operation (Kawasaki ¶39: “The proxy 150 can also or alternatively relay information which characterizes the received queries (e.g., excerpts, extracted features, metadata, etc.) to the monitoring environment 160 prior to ingestion by the MLA 130.”), or a signal processing application. It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Kawasaki to prevent attacks on artificial intelligence models before they happen (Kawasaki ¶16: “The subject matter described herein provides many technical advantages. For example, the current subject matter can be used to identify and stop adversarial prompt injection attacks on artificial intelligence models including large language models. Further, the current subject matter can provide enhanced visibility into the health and security of an enterprise's machine learning assets.”). Claims 11 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Kawasaki in view of Venkata. Regarding claim 11, Le Roux and Kawasaki teach the method of claim 9, but fail to explicitly teach wherein the anomaly includes at least one of a rate of speech above a first threshold, a pitch of speech above a second threshold, a pitch of speech below a third threshold, or a volume of speech below a fourth threshold. However, Venkata teaches wherein the anomaly includes at least one of a rate of speech above a first threshold, a pitch of speech above a second threshold, a pitch of speech below a third threshold, or a volume of speech below a fourth threshold (Venkata ¶37: “The LLM builds a vocal profile for each customer based on historical call data, capturing their typical tone, speech patterns, and typical topics of inquiry or concern. In situations where a customer's account might be compromised, and the fraudster tries to gain information or perform transactions over the phone, the LLM detects anomalies in the voice pattern, stress levels, or conversation topics. Variations from the established vocal profile, such as differences in pitch or unusual hesitations”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux and Kawasaki in view of Venkata to identify anomalies in a vocal request that a human cannot hear (Venkata ¶33: “Subtle deviations in tone, speed, or accent that might not be immediately apparent to human listeners are flagged by the LLM.”). Regarding claim 12, Le Roux, Kawasaki, and Venkata teach the method of claim 11, but Le Roux and Kawasaki fail to explicitly teach wherein at least one of the first, second, third, or fourth threshold is defined based on an expected rate, pitch, or volume predetermined for the user based on one or more previous requests received from the user. However, Venkata teaches wherein at least one of the first, second, third, or fourth threshold is defined based on an expected rate, pitch, or volume predetermined for the user based on one or more previous requests received from the user (Venkata ¶37: “The LLM builds a vocal profile for each customer based on historical call data, capturing their typical tone, speech patterns, and typical topics of inquiry or concern. In situations where a customer's account might be compromised, and the fraudster tries to gain information or perform transactions over the phone, the LLM detects anomalies in the voice pattern, stress levels, or conversation topics. Variations from the established vocal profile, such as differences in pitch or unusual hesitations”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux and Kawasaki in view of Venkata to identify anomalies in a vocal request that a human cannot hear (Venkata ¶33: “Subtle deviations in tone, speed, or accent that might not be immediately apparent to human listeners are flagged by the LLM.”). Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Kawasaki in view of Rossi et al. (United States Patent Publication No. 2024/0347058), hereinafter Rossi. Regarding claim 13, Le Roux teaches the method of claim 1, but fails to explicitly teach wherein: the LLM is a multimodal LLM (MLLM); the response received from the MLLM is a voice response; anticipating the anomaly includes, after receiving the response from the MLLM: processing the response using an audio analysis model; and detecting the anomaly in the response based on results from the audio analysis model; and performing one or more preemptive actions includes, prior to providing the response to the user, removing the detected anomaly from the response. However, Kawasaki teaches wherein: the LLM is a multimodal LLM (MLLM) (Kawasaki ¶37: “the MLA 130 can be a multimodal model that accepts multiple mode (i.e., two or more of audio, video, text, images, etc.) input/prompts.”); [the response received from the MLLM is a voice response]; anticipating the anomaly includes, after receiving the response from the MLLM: processing the response using an [audio] analysis model; and detecting the anomaly in the response based on results from the [audio] analysis model; and performing one or more preemptive actions includes, prior to providing the response to the user, removing the detected anomaly from the response (Kawasaki ¶42: “These remediation actions can take various forms including transmitting data to the proxy 150 which causes the output of the MLA 130 to be blocked prior to transmission to the requesting client device 110 … the remediation engine 180 can cause data to be transmitted to the proxy 150 which causes the output for transmission to the requesting client device 110 to be modified in order to be non-malicious, to remove sensitive information, and the like.”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Kawasaki to prevent attacks on artificial intelligence models before they happen (Kawasaki ¶16: “The subject matter described herein provides many technical advantages. For example, the current subject matter can be used to identify and stop adversarial prompt injection attacks on artificial intelligence models including large language models. Further, the current subject matter can provide enhanced visibility into the health and security of an enterprise's machine learning assets.”). Le Roux and Kawasaki fail to explicitly teach the response received from the MLLM is a voice response and an audio analysis model. However, Rossi teaches the response received from the MLLM is a voice response (Rossi ¶63: “The transceiver module 330 is configured to transmit the recorded audio file to the conversation system 110, which in turn performs voice recognition, obtains a text response from the LLM system 120, and generates a speech response based on the text response. The speech response is then sent back to the transceiver module 330 of the client device 130, which causes the playing module 320 to play the speech response to the user.”) and an audio analysis model (Rossi ¶44: “the STT module 210 is configured to analyze the audio signal to identify distinct features like pitch, tone, and speed. These features are used to help distinguish between different sounds and understand speech patterns. In some embodiments, the STT module 210 may be part of the LLM system 120.”) It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux and Kawasaki in view of Rossi to enhance the user’s experience with a more dynamic, auditory conversation with the MLLM (Rossi Abstract: “This process ensures that user interruptions are effectively managed, allowing for a more dynamic and interactive conversation with the LLM and enhancing the user's experience by adapting the conversation flow to real-time inputs.”). Claims 14 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Ohayon et al. (United States Patent Publication No. 2026/0017386), hereinafter Ohayon based on its priority date of 4/19/2021. Regarding claim 14, Le Roux teaches the method of claim 1, but fails to explicitly teach wherein: anticipating the anomaly includes generating defensive instructions for the LLM; and performing one or more preemptive actions includes providing the defensive instructions to the LLM with the request. However, Ohayon teaches Regarding claim 14, the method of claim 1, wherein: anticipating the anomaly includes generating defensive instructions for the LLM (Ohayon ¶114: “an Online AI-Based Defense Unit 158, which may utilize one or more (e.g., remote, cloud-based) AI/ML/DL engines (e.g., using DNN, Random Forest, Evolutionary/Genetic algorithms, or other techniques) to select which protection scheme or protection operations to apply (or not to apply) to the protected ML/DL/AI Engine 101 and/or with regard to a particular input (or set of inputs) that is incoming to the protected ML/DL/AI Engine 101”); and performing one or more preemptive actions includes providing the defensive instructions to the LLM with the request (Ohayon ¶112: “in which the protected ML/DL/AI Engine 101 firstly sends an input item (e.g., an input image) to the Online Defense and Protection Unit 157, and then receives back from it ... an indication that the input item is malicious or adversarial and should be discarded or should not be processed by the protected ML/DL/AI Engine 101”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Ohayon to reduce the vulnerability of the system (Ohayon ¶6: “An Online Protection Unit is configured to perform analysis of at least one of: (i) inputs that are directed to be inputs of the Protected Engine, (ii) outputs that are generated by the Protected Engine; and based on the analysis, to dynamically perform online fortification of the Protected Engine against attacks by dynamically changing operational properties or operational parameters of the Protected Engine to reduce its vulnerability to attacks.”). Regarding claim 18, Le Roux and Ohayon teach the method of claim 14, but Le Roux fails to explicitly teach wherein providing the defensive instructions to the LLM includes at least one of combining the defensive instructions and the request into a single prompt or embedding the defensive instructions in a system prompt for the LLM separate from the request. However, Ohayon teaches wherein providing the defensive instructions to the LLM includes at least one of combining the defensive instructions and the request into a single prompt or embedding the defensive instructions in a system prompt for the LLM separate from the request (Ohayon ¶112: “in which the protected ML/DL/AI Engine 101 firstly sends an input item (e.g., an input image) to the Online Defense and Protection Unit 157, and then receives back from it ... an indication that the input item is malicious or adversarial and should be discarded or should not be processed by the protected ML/DL/AI Engine 101”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Ohayon to reduce the vulnerability of the system (Ohayon ¶6: “An Online Protection Unit is configured to perform analysis of at least one of: (i) inputs that are directed to be inputs of the Protected Engine, (ii) outputs that are generated by the Protected Engine; and based on the analysis, to dynamically perform online fortification of the Protected Engine against attacks by dynamically changing operational properties or operational parameters of the Protected Engine to reduce its vulnerability to attacks.”). Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Ohayon in view of Huang et al. (United States Patent 10,565,985), hereinafter Huang. Regarding claim 15, Le Roux and Ohayon teach the method of claim 14, wherein the LLM is a multimodal LLM (MLLM), and wherein the defensive instructions include at least one of an instruction to ignore speech within the request (As taught in claim 14 by Ohayon ¶112) but fail to teach having a rate above a first threshold, an instruction to ignore speech within the request having a pitch above a second threshold, an instruction to ignore speech within the request having a pitch below a third threshold, or an instruction to ignore speech within the request having a volume below a fourth threshold. However, Huang teaches ignore speech within the request having a rate above a first threshold, an instruction to ignore speech within the request having a pitch above a second threshold, an instruction to ignore speech within the request having a pitch below a third threshold, or an instruction to ignore speech within the request having a volume below a fourth threshold (Huang Col. 6 Lines 53-67: “The instance of the digital assistant application 108 on the client device 104 can perform pre-filtering or pre-processing on the input audio signal to remove certain frequencies of audio. The pre-filtering can include filters such as a low-pass filter, high-pass filter, or a bandpass filter. The filters can be applied in the frequency domain. The filters can be applied using digital signal processing techniques. The filter can be configured to keep frequencies that correspond to a human voice or human speech, while eliminating frequencies that fall outside the typical frequencies of human speech. For example, a bandpass filter can be configured to remove frequencies below a first threshold (e.g., 70 Hz, 75 Hz, 80 Hz, 85 Hz, 90 Hz, 95 Hz, 100 Hz, or 105 Hz) and above a second threshold (e.g., 200 Hz, 205 Hz, 210 Hz, 225 Hz, 235 Hz, 245 Hz, or 255 Hz).”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux and Ohayon in view of Huang “To address the potential security vulnerabilities stemming from the interfacing, the present systems and methods can determine when the continuous access of the audio data from the microphone is authorized or unauthorized.” (Huang Col. 3). Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Ohayon in view of Venkata. Regarding claim 17, Le Roux and Ohayon teach the method of claim 14, wherein the LLM is a multimodal LLM (MLLM), and wherein the defensive instructions include an instruction to ignore speech within the request (As taught in claim 14 by Ohayon ¶112) but they fail to teach that deviates from an expected rate, pitch, or volume associated with the user by more than a threshold, wherein the expected rate, pitch, or volume is indicated in user data provided to the MLLM with the request. However, Venkata teaches ignore speech within the request that deviates from an expected rate, pitch, or volume associated with the user by more than a threshold, wherein the expected rate, pitch, or volume is indicated in user data provided to the MLLM with the request (Venkata ¶37: “The LLM builds a vocal profile for each customer based on historical call data, capturing their typical tone, speech patterns, and typical topics of inquiry or concern. In situations where a customer's account might be compromised, and the fraudster tries to gain information or perform transactions over the phone, the LLM detects anomalies in the voice pattern, stress levels, or conversation topics. Variations from the established vocal profile, such as differences in pitch or unusual hesitations, especially in responses to security questions or during high-risk transactions, the LLM alerts the electronic device to potential fraud, leading to immediate security protocols like call escalation or transaction holds”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux and Ohayon in view of Venkata to be able to ignore anomalies in a vocal request that a human cannot hear (Venkata ¶33: “Subtle deviations in tone, speed, or accent that might not be immediately apparent to human listeners are flagged by the LLM.”). Claims 19 is rejected under 35 U.S.C. 103 as being unpatentable over Le Roux in view of Trim et al. (United States Patent Publication No. 2020/0402516), hereinafter Trim. Regarding claim 19, Le Roux teaches the method of claim 1, but fails to teach wherein: anticipating the anomaly includes: processing the audio transmission using an audio analysis model; and predicting, based on an output of the audio analysis model, a likelihood that two or more portions of the audio transmission originated from two or more sources, wherein the two or more sources include at least one of people, devices, protocols, or environments; and the one or more preemptive actions are performed responsive to the predicted likelihood being greater than a threshold. However, Trim teaches wherein: anticipating the anomaly includes: processing the audio transmission using an audio analysis model (Trim ¶36: “FIG. 2 is a flowchart depicting operational steps of prevention program 200, a program for preventing adversarial audio attacks through detecting and isolating inconsistencies utilizing beamforming techniques and IoT devices”); and predicting, based on an output of the audio analysis model, a likelihood that two or more portions of the audio transmission originated from two or more sources (Trim ¶31: “a user is in a room west of listening device 130 listening to music, while wearing client device 120. In this example, an unauthorized person to the south of listening device 130 plays an ultrasound encoded with a voice instruction (e.g., a command), which listening device 130 receives to the south. Additionally, prevention program 200 utilizes beamforming module 134 to determine a source direction of the voice instruction and utilizes data (e.g., PAN signal, GPS, etc.) of client device 120 to determine that the source of the voice instruction is inconsistent with the location of the user.”), wherein the two or more sources include at least one of people, devices, protocols, or environments (Trim ¶52: “prevention program 200 assigns a confidence level to the identified inconsistency. In one embodiment, prevention program 200 identifies sources of information utilized to derive a score of the identified inconsistency and determines a confidence level for the information. For example, prevention program 200 identifies that the location of the user is derived using GPS data of the smart watch and phone (e.g., client device 120) of the user and the location of the source of the verbal instruction derived from data of the beamforming transceiver device.”); and the one or more preemptive actions are performed responsive to the predicted likelihood being greater than a threshold (Trim ¶35: “prevention program 200 compares a score, rank, and/or confidence level to a system-defined threshold level to determine whether to ignore a voice instruction (e.g., command), generate an audible notification to a co-located authorized user, or send a notification to an authorized user requesting permission to complete the action.”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Le Roux in view of Trim to help ensure a voice command is from an authorized user and not an attacker (Trim ¶10: “Various embodiments of the present invention utilize beamforming capabilities of a digital assistants integrated with the sensing capabilities of nearby IoT devices to ensure a voice command received by the digital assistant is consistent with a user issuing them to prevent an adversarial audio attack.”). Allowable Subject Matter Claim 16 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Jackson (United States Patent Publication No. 2025/0103715). Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEC ANKRUM whose telephone number is (571)272-9209. The examiner can normally be reached M-F 7:15am-3:15pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ali Shayanfar can be reached at 571-270-1050. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /A.C.A./Examiner, Art Unit 2434 /NOURA ZOUBAIR/Primary Examiner, Art Unit 2434
Read full office action

Prosecution Timeline

Oct 29, 2024
Application Filed
Aug 06, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month