Prosecution Insights
Last updated: October 02, 2026
Application No. 18/628,410

INTEGRATED MULTIMODAL ARTIFICIAL INTELLIGENCE FRAMEWORK FOR AUTOMATED PROVISIONING SYSTEMS

Non-Final OA §101§103
Filed
Apr 05, 2024
Examiner
RHO, YONG DOO
Art Unit
Tech Center
Assignee
Bank of America Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
12 currently pending
Career history
4
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 8/6/2024 and 8/23/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Status of Claims The present application is being examined under the claims filed on 4/5/2024. Claims 1-20 are rejected. Claims 1-20 are pending. Specification The specification filed on 4/5/2024 is acceptable for examination purposes. Drawings The drawings filed on 4/5/2024 and 7/8/2024 are acceptable for examination purposes. Claim Objections Claim 8 is objected to because of the following informalities: In claim 8, line 11, “training the multimodal” should read “train the multimodal” In claim 8, line 19, “providing a recommendation” should read “provide a recommendation” Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Claim 1, Step 1: Claim 1 is a system claim. Therefore, Claims 1-7 are directed to a machine. Step 2A Prong 1: aggregating raw data from multiple data sources, wherein the data sources comprise logs, text, audio inputs, and visual inputs, resulting in aggregated raw data (mental process - aggregating raw data from multiple data sources, wherein the data sources comprise logs, text, audio inputs, and visual inputs, resulting in aggregated raw data may be performed manually by a user with the aid of pen and paper by observing/analyzing multiple data sources and combining them to create aggregated raw data. See MPEP 2106.04(a)(2)(III)(C).) producing a pre-processed dataset via normalizing and cleansing the aggregated raw data (mental process - producing a pre-processed dataset via normalizing and cleansing the aggregated raw data may be performed manually by a user with the aid of pen and paper by observing/analyzing the aggregated raw data and balancing/cleaning it to create a pre-processed dataset. See MPEP 2106.04(a)(2)(III)(C).) determining extracted features from the pre-processed dataset using a combination of natural language processing for text data and computer vision for visual data (mental process - determining extracted features from the pre-processed dataset using a combination of natural language processing for text data and computer vision for visual data may be performed manually by a user with the aid of pen and paper by observing/analyzing the pre-processed dataset and using natural language processing and computer vision techniques to determine extracted features. See MPEP 2106.04(a)(2)(III)(C).) [integrating the extracted features into a multimodal AI model framework and training the multimodal AI model framework] to recognize patterns and make decisions (mental process – recognizing patterns and making decisions may be performed manually by a user with the aid of pen and paper by observing/analyzing the extracted features and using a judgement to recognize patterns and make decisions. See MPEP 2106.04(a)(2)(III)(C).) validating the trained multimodal AI model framework using a validation dataset to ensure model performance meets predetermined accuracy, precision, and recall benchmarks (mental process – validating the trained multimodal AI model framework using a validation dataset to ensure model performance meets predetermined accuracy, precision, and recall benchmarks may be performed manually by a user with the aid of pen and paper by observing/analyzing the validation dataset, performance metrics and validating/testing the AI model framework. See MPEP 2106.04(a)(2)(III)(C).) [incorporating voice recognition capabilities] to interpret natural language inputs from users and translate the natural language inputs into executable commands (mental process – interpreting natural language inputs from users and translating the natural language inputs into executable commands may be performed manually by a user with the aid of pen and paper by observing/analyzing the natural language inputs from users and translating the natural language inputs into executable commands. See MPEP 2106.04(a)(2)(III)(C).) monitoring application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions (mental process – monitoring application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions may be performed manually by a user with the aid of pen and paper by observing/analyzing the application workflows in real-time with the AI model framework and using a judgement to detect and classify system errors or exceptions. See MPEP 2106.04(a)(2)(III)(C).) executing a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions (mental process – executing a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions may be performed manually by a user with the aid of pen and paper by observing/analyzing the system errors or exceptions and using a judgement to execute an automatic fix or recommend manual intervention. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: a processing device (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) a non-transitory storage device containing instructions when executed by the processing device, causes the processing device to perform the steps of: (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) integrating the extracted features into a multimodal AI model framework and training the multimodal AI model framework (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) incorporating voice recognition capabilities (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as integrating the extracted features into a multimodal AI model framework, training the multimodal AI model framework and incorporating voice recognition capabilities. Additional Elements: a processing device (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) a non-transitory storage device containing instructions when executed by the processing device, causes the processing device to perform the steps of: (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) integrating the extracted features into a multimodal AI model framework and training the multimodal AI model framework (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) incorporating voice recognition capabilities (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-7. The additional limitations of the dependent claims are addressed below. Regarding Claim 2, Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein aggregating raw data further comprises use of application programming interfaces (APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein aggregating raw data further comprises use of application programming interfaces (APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Regarding Claim 3, Step 2A Prong 1: [wherein normalizing and cleansing the aggregated raw data further comprises] use of an outlier detection algorithms to identify and rectify anomalies within the data set (mental process – using of an outlier detection algorithms to identify and rectify anomalies within the data set may be performed manually by a user with the aid of pen and paper by observing/analyzing an outlier detection algorithms and identifying/rectifying anomalies within the data set. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein normalizing and cleansing the aggregated raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein normalizing and cleansing the aggregated raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 4, Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein extracting features using natural language processing and computer vision further comprises applying recurrent neural networks for the text data and convolutional neural networks for the visual data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein extracting features using natural language processing and computer vision further comprises applying recurrent neural networks for the text data and convolutional neural networks for the visual data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 5, Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 5 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein validating the trained multimodal AI model framework is performed continuously as part of an iterative development process, with each iteration refining the model based on feedback from an operational performance metric (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein validating the trained multimodal AI model framework is performed continuously as part of an iterative development process, with each iteration refining the model based on feedback from an operational performance metric (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 6, Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the voice recognition capabilities comprise adapting to user-specific accents, dialects, and languages to improve the accuracy of voice-to-text conversions and system commands (merely reciting the words "apply it" (or an equivalent) with the judicial exception. Adapting to user-specific accents, dialects, and languages relates to additional training. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the voice recognition capabilities comprise adapting to user-specific accents, dialects, and languages to improve the accuracy of voice-to-text conversions and system commands (merely reciting the words "apply it" (or an equivalent) with the judicial exception. Adapting to user-specific accents, dialects, and languages relates to additional training. See MPEP 2106.05(f).) Regarding Claim 7, Step 2A Prong 1: wherein executing corrective actions comprises an escalation protocol notifying a human operator when the error requires intervention other than a predetermined automated corrective measure (mental process – wherein executing corrective actions comprises an escalation protocol notifying a human operator when the error requires intervention other than a predetermined automated corrective measure may be performed manually by a user with the aid of pen and paper by observing/analyzing the error and a predetermined automated corrective measure, and using a judgement/decision to notify a human operator when the error requires intervention. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2 & Step 2B: There are no additional elements. Regarding Claim 8, Step 1: Claim 8 is a non-transitory computer-readable medium claim. Therefore, Claims 8-14 are directed to a manufacture. Step 2A Prong 1: aggregate raw data from multiple data sources, wherein the data sources comprise logs, text, audio inputs, and visual inputs, resulting in aggregated raw data (mental process - aggregating raw data from multiple data sources, wherein the data sources comprise logs, text, audio inputs, and visual inputs, resulting in aggregated raw data may be performed manually by a user with the aid of pen and paper by observing/analyzing multiple data sources and combining them to create aggregated raw data. See MPEP 2106.04(a)(2)(III)(C).) produce a pre-processed dataset via normalizing and cleansing the aggregated raw data (mental process - producing a pre-processed dataset via normalizing and cleansing the aggregated raw data may be performed manually by a user with the aid of pen and paper by observing/analyzing the aggregated raw data and balancing/cleaning it to create a pre-processed dataset. See MPEP 2106.04(a)(2)(III)(C).) determine extracted features from the pre-processed dataset using a combination of natural language processing for text data and computer vision for visual data (mental process - determining extracted features from the pre-processed dataset using a combination of natural language processing for text data and computer vision for visual data may be performed manually by a user with the aid of pen and paper by observing/analyzing the pre-processed dataset and using natural language processing and computer vision techniques to determine extracted features. See MPEP 2106.04(a)(2)(III)(C).) [integrate the extracted features into a multimodal AI model framework and training the multimodal AI model framework] to recognize patterns and make decisions (mental process – recognizing patterns and making decisions may be performed manually by a user with the aid of pen and paper by observing/analyzing the extracted features and using a judgement to recognize patterns and make decisions. See MPEP 2106.04(a)(2)(III)(C).) validate the trained multimodal AI model framework using a validation dataset to ensure model performance meets predetermined accuracy, precision, and recall benchmarks (mental process – validating the trained multimodal AI model framework using a validation dataset to ensure model performance meets predetermined accuracy, precision, and recall benchmarks may be performed manually by a user with the aid of pen and paper by observing/analyzing the validation dataset, performance metrics and validating/testing the AI model framework. See MPEP 2106.04(a)(2)(III)(C).) [incorporate voice recognition capabilities] to interpret natural language inputs from users and translate the natural language inputs into executable commands (mental process – interpreting natural language inputs from users and translating the natural language inputs into executable commands may be performed manually by a user with the aid of pen and paper by observing/analyzing the natural language inputs from users and translating the natural language inputs into executable commands. See MPEP 2106.04(a)(2)(III)(C).) monitor application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions (mental process – monitoring application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions may be performed manually by a user with the aid of pen and paper by observing/analyzing the application workflows in real-time with the AI model framework and using a judgement to detect and classify system errors or exceptions. See MPEP 2106.04(a)(2)(III)(C).) execute a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions (mental process – executing a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions may be performed manually by a user with the aid of pen and paper by observing/analyzing the system errors or exceptions and using a judgement to execute an automatic fix or recommend manual intervention. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: a non-transitory computer-readable medium comprising code causing an apparatus to: (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) integrate the extracted features into a multimodal AI model framework and training the multimodal AI model framework (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) incorporate voice recognition capabilities (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as integrating the extracted features into a multimodal AI model framework, training the multimodal AI model framework and incorporating voice recognition capabilities. Additional Elements: a non-transitory computer-readable medium comprising code causing an apparatus to: (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) integrate the extracted features into a multimodal AI model framework and training the multimodal AI model framework (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) incorporate voice recognition capabilities (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) For the reasons above, Claim 8 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 9-14. The additional limitations of the dependent claims are addressed below. Regarding Claim 9, Step 2A Prong 1: [wherein aggregating raw data further comprises] use of application programming interfaces (APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases (mental process – using of application programming interfaces (APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases may be performed manually by a user with the aid of pen and paper by observing/analyzing the application programming interfaces (APIs) and using them to retrieve data from user interfaces, middleware and backend databases. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein aggregating raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein aggregating raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 10, Step 2A Prong 1: [wherein normalizing and cleansing the aggregated raw data further comprises] use of an outlier detection algorithms to identify and rectify anomalies within the data set (mental process – using of an outlier detection algorithms to identify and rectify anomalies within the data set may be performed manually by a user with the aid of pen and paper by observing/analyzing an outlier detection algorithms and identifying/rectifying anomalies within the data set. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein normalizing and cleansing the aggregated raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein normalizing and cleansing the aggregated raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 11, Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 11 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein extracting features using natural language processing and computer vision further comprises applying recurrent neural networks for the text data and convolutional neural networks for the visual data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein extracting features using natural language processing and computer vision further comprises applying recurrent neural networks for the text data and convolutional neural networks for the visual data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 12, Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 12 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein validating the trained multimodal AI model framework is performed continuously as part of an iterative development process, with each iteration refining the model based on feedback from an operational performance metric (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein validating the trained multimodal AI model framework is performed continuously as part of an iterative development process, with each iteration refining the model based on feedback from an operational performance metric (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 13, Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 13 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the voice recognition capabilities comprise adapting to user-specific accents, dialects, and languages to improve the accuracy of voice-to-text conversions and system commands (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the voice recognition capabilities comprise adapting to user-specific accents, dialects, and languages to improve the accuracy of voice-to-text conversions and system commands (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 14, Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 14 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein executing corrective actions comprises an escalation protocol notifying a human operator when the error requires intervention other than a predetermined automated corrective measure (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein executing corrective actions comprises an escalation protocol notifying a human operator when the error requires intervention other than a predetermined automated corrective measure (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 15, Step 1: Claim 15 is a method claim. Therefore, Claims 15-20 are directed to a process. Step 2A Prong 1: aggregating raw data from multiple data sources, wherein the data sources comprise logs, text, audio inputs, and visual inputs, resulting in aggregated raw data (mental process - aggregating raw data from multiple data sources, wherein the data sources comprise logs, text, audio inputs, and visual inputs, resulting in aggregated raw data may be performed manually by a user with the aid of pen and paper by observing/analyzing multiple data sources and combining them to create aggregated raw data. See MPEP 2106.04(a)(2)(III)(C).) producing a pre-processed dataset via normalizing and cleansing the aggregated raw data (mental process - producing a pre-processed dataset via normalizing and cleansing the aggregated raw data may be performed manually by a user with the aid of pen and paper by observing/analyzing the aggregated raw data and balancing/cleaning it to create a pre-processed dataset. See MPEP 2106.04(a)(2)(III)(C).) determining extracted features from the pre-processed dataset using a combination of natural language processing for text data and computer vision for visual data (mental process - determining extracted features from the pre-processed dataset using a combination of natural language processing for text data and computer vision for visual data may be performed manually by a user with the aid of pen and paper by observing/analyzing the pre-processed dataset and using natural language processing and computer vision techniques to determine extracted features. See MPEP 2106.04(a)(2)(III)(C).) [integrating the extracted features into a multimodal AI model framework and training the multimodal AI model framework] to recognize patterns and make decisions (mental process – recognizing patterns and making decisions may be performed manually by a user with the aid of pen and paper by observing/analyzing the extracted features and using a judgement to recognize patterns and make decisions. See MPEP 2106.04(a)(2)(III)(C).) validating the trained multimodal AI model framework using a validation dataset to ensure model performance meets predetermined accuracy, precision, and recall benchmarks (mental process – validating the trained multimodal AI model framework using a validation dataset to ensure model performance meets predetermined accuracy, precision, and recall benchmarks may be performed manually by a user with the aid of pen and paper by observing/analyzing the validation dataset, performance metrics and validating/testing the AI model framework. See MPEP 2106.04(a)(2)(III)(C).) [incorporating voice recognition capabilities] to interpret natural language inputs from users and translate the natural language inputs into executable commands (mental process – interpreting natural language inputs from users and translating the natural language inputs into executable commands may be performed manually by a user with the aid of pen and paper by observing/analyzing the natural language inputs from users and translating the natural language inputs into executable commands. See MPEP 2106.04(a)(2)(III)(C).) monitoring application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions (mental process – monitoring application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions may be performed manually by a user with the aid of pen and paper by observing/analyzing the application workflows in real-time with the AI model framework and using a judgement to detect and classify system errors or exceptions. See MPEP 2106.04(a)(2)(III)(C).) executing a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions (mental process – executing a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions may be performed manually by a user with the aid of pen and paper by observing/analyzing the system errors or exceptions and using a judgement to execute an automatic fix or recommend manual intervention. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: integrating the extracted features into a multimodal AI model framework and training the multimodal AI model framework (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) incorporating voice recognition capabilities (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as integrating the extracted features into a multimodal AI model framework, training the multimodal AI model framework and incorporating voice recognition capabilities. Additional Elements: integrating the extracted features into a multimodal AI model framework and training the multimodal AI model framework (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) incorporating voice recognition capabilities (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) For the reasons above, Claim 15 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 16-20. The additional limitations of the dependent claims are addressed below. Regarding Claim 16, Step 2A Prong 1: [wherein aggregating raw data further comprises] use of application programming interfaces (APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases (mental process – using of application programming interfaces (APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases may be performed manually by a user with the aid of pen and paper by observing/analyzing the application programming interfaces (APIs) and using them to retrieve data from user interfaces, middleware and backend databases. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein aggregating raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein aggregating raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 17, Step 2A Prong 1: [wherein normalizing and cleansing the aggregated raw data further comprises] use of an outlier detection algorithms to identify and rectify anomalies within the data set (mental process – using of an outlier detection algorithms to identify and rectify anomalies within the data set may be performed manually by a user with the aid of pen and paper by observing/analyzing an outlier detection algorithms and identifying/rectifying anomalies within the data set. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein normalizing and cleansing the aggregated raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein normalizing and cleansing the aggregated raw data further comprises (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 18, Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 18 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein extracting features using natural language processing and computer vision further comprises applying recurrent neural networks for the text data and convolutional neural networks for the visual data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein extracting features using natural language processing and computer vision further comprises applying recurrent neural networks for the text data and convolutional neural networks for the visual data (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 19, Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 19 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the voice recognition capabilities comprise adapting to user-specific accents, dialects, and languages to improve the accuracy of voice-to-text conversions and system commands (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the voice recognition capabilities comprise adapting to user-specific accents, dialects, and languages to improve the accuracy of voice-to-text conversions and system commands (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 20, Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 20 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein executing corrective actions comprises an escalation protocol notifying a human operator when the error requires intervention other than a predetermined automated corrective measure (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein executing corrective actions comprises an escalation protocol notifying a human operator when the error requires intervention other than a predetermined automated corrective measure (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 4-9, 11-16 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (US 20240185602 A1) (hereinafter Liu), in view of Sinha (US 9633674 B2), and further in view of Rakshit et al. (US 11676574 B2) (hereinafter Rakshit). Regarding Claim 1, Liu teaches: “A system for an integrated multimodal artificial intelligence framework for automated provisioning systems, the system comprising:” (preamble) “a processing device” (Liu, Paragraph [0022], “Components of the computing device 100 may include, but is not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.”; Examiner’s note: a processing device (i.e. one or more processors or processing units 110) is taught.) “a non-transitory storage device containing instructions when executed by the processing device, causes the processing device to perform the steps of:” (Liu, Paragraphs [0022] and [0137], “one or more processors or processing units 110 […] the present disclosure provides a computer program product being tangibly stored in a non-transitory computer storage medium and comprising machine-executable instructions which, when executed by a device, causing the device to perform the method of […]”; Examiner’s note: a non-transitory storage device (i.e. a non-transitory computer storage medium) containing instructions (i.e. machine-executable instructions) when executed by the processing device (i.e. one or more processors or processing units 110), causes the processing device to perform the steps of (i.e. performing the method of) is taught.) aggregating raw data from multiple data sources, wherein the data sources comprise text, and visual inputs, resulting in aggregated raw data (Liu, Fig. 1 and Paragraphs [0028] and [0030], “The input device 150 may be one or more various input devices, such as a mouse, a keyboard, a trackball, a voice-input device, and the like […] As shown in FIG. 1, the computing device 100 may receive a training dataset 170 through the input device 150. The training dataset 170 comprises a plurality of image-text pairs, each image-text pair comprising a training image and a training text corresponding to the training image. FIG. 1 shows an example of an image-text pair, i.e., a training image 171 and a training text 172 corresponding to the training image 171.”; Examiner’s note: aggregating raw data (i.e. creating an image-text pair by combining a training image with a training text corresponding to the training image) from multiple data sources (i.e. input devices such as a mouse, a keyboard, a trackball, a voice-input device, and the like), wherein the data sources comprise text (i.e. training texts), and visual inputs (i.e. training images), resulting in aggregated raw data (i.e. the training dataset including a plurality of image-text pairs) is taught. See Fig. 1 below. PNG media_image1.png 330 456 media_image1.png Greyscale ) “producing a pre-processed dataset via normalizing and cleansing the aggregated raw data” (Liu, Paragraph [0033], “The computing device 100 trains a vision-language model 180 by using the training dataset 170. Accordingly, the vision-language model 180 may also be referred to as a “target model.” The trained vision-language model 180 may be used in a vision-language task to determine association information between an image and a text. In some implementations, the training of the vision-language model 180 at the computing device 100 may be pre-training for universal tasks.”; Liu, Paragraph [0088], “[…] the trained vision-language model 180 comprises the trained text feature extraction sub-model 210. The text feature extraction sub-model 210 extracts a set of text features 615 of an input text 602.”; Liu, Paragraph [0099], “[…] the computing device 100 extracts a set of visual features of a training image according to a visual feature extraction sub-model in a target model […] the computing device 100 determines a set of visual semantic features corresponding to the set of visual features based on a visual semantic dictionary.”; Examiner’s note: producing a pre-processed dataset (i.e. the training of the vision-language model referred to as a target model) via normalizing (i.e. extracting text features of an input text) and cleansing the aggregated raw data (i.e. determining a set of visual semantic features based on a visual semantic dictionary) is taught.) “determining extracted features from the pre-processed dataset using a combination of natural language processing for text data and computer vision for visual data” (Liu, Paragraphs [0002], [0038] and [0042], “[…] a set of visual features of a training image is extracted according to a visual feature extraction sub-model in a target model […] A set of text features of a training text corresponding to the training image is extracted according to a text feature extraction sub-model in the target model […] the text feature extraction sub-model 210 may be implemented using a bi-directional encoder representation from transformers (BERT) […] The visual feature extraction sub-model 220 may be implemented as a trainable visual feature encoder.”; Examiner’s note: determining extracted features (i.e. extracting a set of visual features and a set of text features) from the pre-processed dataset (i.e. the training dataset) using a combination of natural language processing for text data (i.e. the text feature extraction sub-model 210 implemented using a bi-directional encoder representation from transformers (BERT)) and computer vision for visual data (i.e. the visual feature extraction sub-model 220 implemented as a trainable visual feature encoder) is taught.) “integrating the extracted features into a multimodal AI model framework and training the multimodal AI model framework to recognize patterns and make decisions” (Liu, Paragraphs [0047] and [0048], “the fusion sub-model 240 in the vision-language model 180 is configured to generate a set of fused features 245 for the training text 172 and the training image 171 based on the visual semantic features 235 and the text features 215 […] a multi-layer transformer may be used to implement the fusion sub-model 240. Such a multi-layer transformer may learn cross-modal representations with the fusion of features in the visual domain and features in the language domain […] In the training stage, an objective function may be determined based on the fused feature 245, and the vision-language model 180 may be trained by minimizing the objective function. In some implementations, the training of the vision-language model 180 may be pre-training. In such implementations, the objective function may be determined with respect to one or more universal tasks for the pre-training. Universal tasks may comprise determining whether the image and text are matched, predicting the masked text feature, predicting the masked visual semantic feature, etc.”; Examiner’s note: integrating the extracted features (i.e. generating a set of fused features 245 for the training text 172 and the training image 171 based on the visual semantic features 235 and the text features 215) into a multimodal AI model framework (i.e. a multi-layer transformer learning cross-modal representations with the fusion of features) and training the multimodal AI model framework (i.e. training of the vision-language model) to recognize patterns and make decisions (i.e. determining whether the image and text are matched, predicting the masked text feature and predicting the masked visual semantic feature, etc.) is taught.) “validating the trained multimodal AI model framework using a validation dataset to ensure model performance meets predetermined accuracy, precision, and recall benchmarks” (Liu, Paragraphs [0080] and [0081], “In some implementations, the training of the vision-language model 180 may be fine-tuning […] In the fine-tuning or training for image-text retrieval, the training dataset comprises both matched image-text pairs and unmatched image-text pairs.”; Examiner’s note: validating the trained multimodal AI model framework (i.e. fine-tuning the vision-language model) using a validation set (i.e. the training dataset including both matched and unmatched image-text pairs) to ensure model performance meets predetermined accuracy, precision, and recall benchmarks (i.e. Accuracy, precision and recall are well-known and conventional performance metrics.1) is taught.) Liu does not explicitly teach: aggregating raw data from multiple data sources, wherein the data sources comprise logs, and audio inputs, resulting in aggregated raw data “incorporating voice recognition capabilities to interpret natural language inputs from users and translate the natural language inputs into executable commands” “monitoring application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions” “executing a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions” Sinha teaches: aggregating raw data from multiple data sources, wherein the data sources comprise logs, and audio inputs, resulting in aggregated raw data (Sinha, Col. 1, Lines 46-49, “digital assistants that interact with users via speech inputs and outputs typically employ speech-to-text processing techniques to convert speech inputs to textual forms that can be further processed”; Sinha, Col. 9, Lines 2-3, “error logs, resources usage, etc., of the user device 104 are provided […]”; Examiner’s note: aggregating raw data (i.e. converting to textual forms) from multiple data sources (i.e. speech inputs and error logs), wherein the data sources comprise logs (i.e. error logs), and audio inputs (i.e. speech inputs), resulting in aggregated raw data (i.e. textual forms) is taught.) “incorporating voice recognition capabilities to interpret natural language inputs from users and translate the natural language inputs into executable commands” (Sinha, Col. 1, Lines 33-37, “Such digital assistants can interpret the user's input to infer the user's intent, translate the inferred intent into actionable tasks and parameters, execute operations or deploy services to perform the tasks, and produce outputs that are intelligible to the user.”; Sinha, Col. 7, Lines 25-28, “An audio subsystem 226 is coupled to speakers 228 and a microphone 230 to facilitate voice-enabled functions, such as voice recognition, voice replication, digital recording, and telephony functions.”; Examiner’s note: incorporating voice recognition capabilities (i.e. an audio subsystem 226 coupled to speakers 228 and a microphone 230 to facilitate voice recognition) to interpret natural language inputs from users (i.e. interpreting the user’s input to infer the user’s intent) and translate the natural language inputs into executable commands (i.e. translating the inferred intent into actionable tasks and parameters) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the cross-modal processing for vision and language in Liu, and the system and method for detecting errors in interactions with a voice-based digital assistant as taught in Sinha. Liu teaches aggregating raw data from text and visual inputs, producing a pre-processed dataset, determining extracted features from the pre-processed dataset, integrating the extracted features into a multimodal AI model framework, training the multimodal AI model framework and validating the multimodal AI model framework. Sinha teaches aggregating raw data from logs and audio inputs and incorporating voice recognition capabilities to translate the natural language inputs into executable commands. One of ordinary skill would have motivation to combine Liu and Sinha to “interpret the user's input to infer the user's intent, translate the inferred intent into actionable tasks and parameters, execute operations or deploy services to perform the tasks, and produce outputs that are intelligible to the user” (Sinha, Col. 1, Lines 33-37). Rakshit teaches: “monitoring application workflows in real-time with the trained multimodal artificial intelligent (AI) model framework to detect and classify system errors or exceptions” (Rakshit, Col. 6, Lines 14-15, “The task, and the user's performance of the task, may be monitored in real time.”; Rakshit, Col. 8, Lines 8-9, “The AI voice response system may receive data from the monitoring devices in real time (e.g., a live feed).”; Rakshit, Col. 14, Lines 33-36, “A task monitoring program 110a, 110b provides a way to train an AI voice response system to actively monitor a task being performed by a user and interrupt if a difference is detected from a task sequence.”; Examiner’s note: monitoring application workflows in real-time (i.e. monitoring the task and the user’s performance of the task in real time) with the trained multimodal artificial intelligent (AI) model framework (i.e. the AI voice response system) to detect and classify system errors or exceptions (i.e. interrupting if a difference is detected from a task sequence) is taught.) “executing a corrective action automatically or providing a recommendation for manual intervention to resolve the system errors or exceptions” (Rakshit, Col. 9, Lines 48-56, “The AI voice response system may be trained to determine the acceptable level of deviation from the task sequence automatically (e.g., observing users frequently deviate from a task sequence) and/or manually (e.g., user explaining the deviation from the task sequence was a preference). The task monitoring program 110 may interrupt the user if the deviation (e.g., delta) from the task sequence rises above a predefined threshold (e.g., task monitoring program 110 threshold, threshold set by user).”; Examiner’s note: executing a corrective action automatically (i.e. determining the acceptable level of deviation from the task sequence automatically) or providing a recommendation (i.e. predefined threshold) for manual intervention (i.e. user explaining the deviation from the task sequence) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the cross-modal processing for vision and language in Liu, the system and method for detecting errors in interactions with a voice-based digital assistant in Sinha, and the duration based task monitoring of artificial intelligence voice response systems as taught in Rakshit. Liu teaches aggregating raw data from text and visual inputs, producing a pre-processed dataset, determining extracted features from the pre-processed dataset, integrating the extracted features into a multimodal AI model framework, training the multimodal AI model framework and validating the multimodal AI model framework. Sinha teaches aggregating raw data from logs and audio inputs and incorporating voice recognition capabilities to translate the natural language inputs into executable commands. Rakshit teaches monitoring the workflows in real time to detect the errors and executing a corrective action automatically or manually to resolve the errors. One of ordinary skill would have motivation to combine Liu, Sinha and Rakshit to “improve the technical field of task monitoring by actively monitoring a user's task performance and interrupting when the user deviates from a task sequence” (Rakshit, Col. 4, Lines 10-12). Regarding Claim 2, The combination of Liu, Sinha and Rakshit teaches: “The system of claim 1,” (preamble) “wherein aggregating raw data further comprises use of application programming interfaces (APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases” (Sinha, Col. 5, Lines 29-31, “[…] executing the task flow by invoking programs, methods, services, APIs, or the like; and generating output responses to the user in an audible (e.g. speech) and/or visual form.”; Sinha, Col. 17, Lines 46-50, “The service processing module 338 accesses the appropriate service model for a service and generates requests for the service in accordance with the protocols and APIs required by the service according to the service model.”; Examiner’s note: wherein aggregating raw data further comprises use of application programming interfaces (APIs) (i.e. executing the task flow by invoking APIs) to automatically retrieve data from various application layers including user interfaces, middleware, and backend databases (i.e. service requests in accordance with APIs teach retrieving data with APIs from user interfaces, middleware, and backend databases) is taught.) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 4, The combination of Liu, Sinha and Rakshit teaches: “The system of claim 1,” (preamble) “wherein extracting features using natural language processing and computer vision further comprises applying recurrent neural networks for the text data and convolutional neural networks for the visual data” (Liu, Paragraphs [0002], [0038], [0042] and [0043], “[…] a set of visual features of a training image is extracted according to a visual feature extraction sub-model in a target model […] A set of text features of a training text corresponding to the training image is extracted according to a text feature extraction sub-model in the target model […] the text feature extraction sub-model 210 may be implemented using a bi-directional encoder representation from transformers (BERT) […] The visual feature extraction sub-model 220 may be implemented as a trainable visual feature encoder […] The trainable visual feature extraction sub-model 220, e.g., a Convolutional Neural Network (CNN) encoder, is used in the vision-language model 180.”; Rakshit, Col. 6, Lines 48-49, “The task monitoring program 110 may deploy a recurrent neural network (RNN).”; Examiner’s note: wherein extracting features (i.e. extracting a set of visual features and a set of text features) using natural language processing (i.e. the text feature extraction sub-model 210 implemented using a bi-directional encoder representation from transformers (BERT)) and computer vision (i.e. the visual feature extraction sub-model 220 implemented as a trainable visual feature encoder) further comprises applying recurrent neural networks for the text data (i.e. the task monitoring program deploying a recurrent neural network (RNN)) and convolutional neural networks for the visual data (i.e. visual feature extraction sub-model 220 is a Convolutional Neural Network (CNN)) is taught.) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 5, The combination of Liu, Sinha and Rakshit teaches: “The system of claim 1,” (preamble) “wherein validating the trained multimodal AI model framework is performed continuously as part of an iterative development process, with each iteration refining the model based on feedback from an operational performance metric” (Liu, Paragraphs [0080] and [0081], “In some implementations, the training of the vision-language model 180 may be fine-tuning […] In the fine-tuning or training for image-text retrieval, the training dataset comprises both matched image-text pairs and unmatched image-text pairs.”; Sinha, Col. 17, Lines 64-67, “task flow processing module 336 are used collectively and iteratively to infer and define the user's intent, obtain information to further clarify and refine the user intent […]”; Sinha, Col. 18, Lines 55-58, “the accuracy of the error detection module 339 is increased, as the user can quickly and easily provide definitive feedback to confirm or deny whether an error has actually occurred.”; Examiner’s note: wherein validating the trained multimodal AI model framework (i.e. fine-tuning the vision-language model) is performed continuously as part of an iterative development process (i.e. task flow processing module being used iteratively to infer and define the user’s intent), with each iteration refining the model (i.e. further clarifying and refining the user intent) based on feedback from an operational performance metric (i.e. feedback to confirm or deny whether an error has actually occurred (e.g. accuracy)) is taught.) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 6, The combination of Liu, Sinha and Rakshit teaches: “The system of claim 1,” (preamble) “wherein the voice recognition capabilities comprise adapting to user-specific accents, dialects, and languages to improve the accuracy of voice-to-text conversions and system commands” (Sinha, Col. 7, Lines 25-27, “An audio subsystem 226 is coupled to speakers 228 and a microphone 230 to facilitate voice-enabled functions, such as voice recognition […]”; Sinha, Col. 19, Lines 50-57, “[…] the error analysis repository 340 can be used to identify systemic errors and/or problems, as well as or in addition to errors and/or problems that are specific to individual users (e.g., because of accents or grammatical idiosyncrasies of a particular user). The error analysis module 342 analyzes the information in the error analysis repository 340 to identify individual errors and/or patterns of errors by the digital assistant”; Sinha, Col. 20, Lines 5-10, “In some implementations, the error analysis module 342 also automatically (e.g., without human intervention) adjusts one or more attributes or processes of the digital assistant (e.g., an acoustic or language model of the speech-to-text processing module 330, etc.) in response to detecting the pattern in the error analysis repository 340.”; Examiner’s note: wherein the voice recognition capabilities (i.e. an audio subsystem) comprise adapting to user-specific accents, dialects, and languages (i.e. errors and/or problems that are specific to individual users (e.g., because of accents or grammatical idiosyncrasies of a particular user)) to improve the accuracy of voice-to-text conversions and system commands (i.e. the error analysis module automatically adjusting one or more attributes or processes of the digital assistant (e.g., language model of the speech-to-text processing module)) is taught.) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 7, The combination of Liu, Sinha and Rakshit teaches: “The system of claim 1,” (preamble) “wherein executing corrective actions comprises an escalation protocol notifying a human operator when the error requires intervention other than a predetermined automated corrective measure” (Rakshit, Col. 9, Lines 31-35 and Lines 48-56, “The task monitoring program 110 may utilize a microphone or any other mode of communication the user may specify (e.g., microphone of the AI voice response system, message alert to mobile phone, vibration on a wearable device) […] The AI voice response system may be trained to determine the acceptable level of deviation from the task sequence automatically (e.g., observing users frequently deviate from a task sequence) and/or manually (e.g., user explaining the deviation from the task sequence was a preference). The task monitoring program 110 may interrupt the user if the deviation (e.g., delta) from the task sequence rises above a predefined threshold (e.g., task monitoring program 110 threshold, threshold set by user).”; Examiner’s note: wherein executing corrective actions (i.e. determining the acceptable level of deviation from the task sequence) comprises an escalation protocol (i.e. the task monitoring program interrupting the user) notifying a human operator (i.e. the task monitoring program utilizing any other mode of communication the user may specify (e.g., microphone of the AI voice response system, message alert to mobile phone, vibration on a wearable device)) when the error requires intervention (i.e. user explaining the deviation from the task sequence) other than a predetermined automated corrective measure (i.e. observing users frequently deviate from a task sequence and interrupting the user if the deviation from the task sequence rises above a predefined threshold) is taught.) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 8, Claim 8 recites substantially the same limitations as Claim 1, in the form of a computer program product; therefore, it is rejected under the same rationale. Regarding Claim 9, Claim 9 recites substantially the same limitations as Claim 2, in the form of a computer program product; therefore, it is rejected under the same rationale. Regarding Claim 11, Claim 11 recites substantially the same limitations as Claim 4, in the form of a computer program product; therefore, it is rejected under the same rationale. Regarding Claim 12, Claim 12 recites substantially the same limitations as Claim 5, in the form of a computer program product; therefore, it is rejected under the same rationale. Regarding Claim 13, Claim 13 recites substantially the same limitations as Claim 6, in the form of a computer program product; therefore, it is rejected under the same rationale. Regarding Claim 14, Claim 14 recites substantially the same limitations as Claim 7, in the form of a computer program product; therefore, it is rejected under the same rationale. Regarding Claim 15, Claim 15 recites substantially the same limitations as Claim 1, in the form of a method; therefore, it is rejected under the same rationale. Regarding Claim 16, Claim 16 recites substantially the same limitations as Claim 2, in the form of a method; therefore, it is rejected under the same rationale. Regarding Claim 18, Claim 18 recites substantially the same limitations as Claim 4, in the form of a method; therefore, it is rejected under the same rationale. Regarding Claim 19, Claim 19 recites substantially the same limitations as Claim 6, in the form of a method; therefore, it is rejected under the same rationale. Regarding Claim 20, Claim 20 recites substantially the same limitations as Claim 7, in the form of a method; therefore, it is rejected under the same rationale. Claims 3, 10 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Liu, in view of Sinha, and further in view of Rakshit as applied in claim 1, and further in view of Vu et al. (US 20240428124 A1) (hereinafter Vu). Regarding Claim 3, The combination of Liu, Sinha and Rakshit teaches: “The system of claim 1,” (preamble) wherein normalizing and cleansing the aggregated raw data further comprises identifying and rectifying anomalies within the data set (Liu, Paragraph [0081], “To enable the vision-language model 180 to predict correct classification for the matched image-text pairs and the unmatched image-text pairs, the fine-tuning or training for the image-text retrieval may be considered as a binary classification problem.”; Liu, Paragraph [0088], “[…] the trained vision-language model 180 comprises the trained text feature extraction sub-model 210. The text feature extraction sub-model 210 extracts a set of text features 615 of an input text 602.”; Liu, Paragraph [0099], “[…] the computing device 100 extracts a set of visual features of a training image according to a visual feature extraction sub-model in a target model […] the computing device 100 determines a set of visual semantic features corresponding to the set of visual features based on a visual semantic dictionary.”; Examiner’s note: wherein normalizing (i.e. extracting text features of an input text) and cleansing the aggregated raw data (i.e. determining a set of visual semantic features based on a visual semantic dictionary) further comprises identifying and rectifying anomalies within the data set (i.e. fine-tuning or training for the image-text retrieval) is taught.) The combination of Liu, Sinha and Rakshit does not explicitly teach: wherein normalizing and cleansing the aggregated raw data further comprises use of an outlier detection algorithms Vu teaches: wherein normalizing and cleansing the aggregated raw data further comprises use of an outlier detection algorithms (Vu, Paragraph [0050], “Given this derived data set D. “Isolation Forest” and “Average KNN” outlier detection methods are performed over D. The outlier, inlier labels are used to calculate received operator characteristic (ROC) scores. If the ROC scores generated by “Isolation Forest” and “Average KNN” are both greater than one half (0.5), D is considered an outlier data set in the collection.”; Examiner’s note: wherein normalizing and cleansing the aggregated raw data (i.e. this derived data set D) further comprises use of an outlier detection algorithms (i.e. Isolation Forest and Average KNN) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the cross-modal processing for vision and language in Liu, the system and method for detecting errors in interactions with a voice-based digital assistant in Sinha, the duration based task monitoring of artificial intelligence voice response systems in Rakshit, and the outlier detection with transfer learning as taught in Vu. Liu teaches aggregating raw data from text and visual inputs, producing a pre-processed dataset, determining extracted features from the pre-processed dataset, integrating the extracted features into a multimodal AI model framework, training the multimodal AI model framework and validating the multimodal AI model framework. Sinha teaches aggregating raw data from logs and audio inputs and incorporating voice recognition capabilities to translate the natural language inputs into executable commands. Rakshit teaches monitoring the workflows in real time to detect the errors and executing a corrective action automatically or manually to resolve the errors. Vu teaches the outlier detection algorithm. One of ordinary skill would have motivation to combine Liu, Sinha, Rakshit and Vu to “ensure that data analysis is performed on good, reliable data” (Vu, Paragraph [0004]). Regarding Claim 10, Claim 10 recites substantially the same limitations as Claim 3, in the form of a computer program product; therefore, it is rejected under the same rationale. Regarding Claim 17, Claim 17 recites substantially the same limitations as Claim 3, in the form of a method; therefore, it is rejected under the same rationale. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to YONG D RHO whose telephone number is (571)270-0194. The examiner can normally be reached 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at 5712705871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YONG DOO RHO/Examiner, Art Unit 2147 /MARC S SOMERS/Primary Examiner, Art Unit 2159 1 Accuracy, precision, recall, and F1 score are the four most commonly used metrics. <What Is Accuracy, Precision, Recall, and F1 Score? — labelf.ai>
Read full office action

Prosecution Timeline

Apr 05, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month