Prosecution Insights
Last updated: August 17, 2026
Application No. 17/548,182

COMPUTER OPTIMIZATION OF TASK PERFORMANCE THROUGH DYNAMIC SENSING

Non-Final OA §103
Filed
Dec 10, 2021
Examiner
TRAN, DAVID HOANG
Art Unit
2147
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
3 (Non-Final)
21%
Grant Probability
At Risk
3-4
OA Rounds
0m
Est. Remaining
35%
With Interview

Examiner Intelligence

Grants only 21% of cases
21%
Career Allowance Rate
4 granted / 19 resolved
-33.9% vs TC avg
Moderate +14% lift
Without
With
+14.1%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
20 currently pending
Career history
57
Total Applications
across all art units

Statute-Specific Performance

§101
29.4%
-10.6% vs TC avg
§103
48.3%
+8.3% vs TC avg
§102
8.5%
-31.5% vs TC avg
§112
12.7%
-27.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 19 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 05/20/2026 has been entered. Response to Arguments Applicant’s arguments on pages 10-17 of Remarks dated 05/20/2026 regarding the rejection under 35 U.S.C. 103 with respect to claims 1-6 and 9-21 have been fully considered but are not persuasive. Beginning on page 10 of Remarks, Applicant asserts that equating the inference with both a task and a classification is inaccurate and they cannot both logically read on the “inference”. However, under the broadest reasonable interpretation Examiner is interpreting inference as a result being produced for the task or classification as shown in paragraph [0009] of Yadav, “wherein the classification output is a result generated by performing the task;” Beginning on page 11, Applicant’s arguments regarding Yadav not teaching the “main modality” have been fully considered but are moot. New reference Wood has been incorporated to teach the main modality. Beginning on page 13, Applicant’s arguments regarding the limitation “engaging, by one or more processors, based on a request for an inference, from a group of sensors of multiple modalities at a physical location in a physical setting comprising environmental detriments to utilizing sensors comprising Internet of Things devices, wherein the group of sensors of the multiple modalities are integrated into a roaming edge device,” have been fully considered but are moot in view of the new combination of references Yadav, Wood, Fayyad and Wood. See updated rejection below. Beginning on page 17, Applicant’s arguments regarding the rejection under 35 U.S.C. 103 with respect to claim 17 has been fully considered but are moot. New reference Lee has been incorporated below to teach the newly presented limitations. CRM Examiner is interpreting computer readable medium as non-transitory in view of paragraph [0098] of the Specification. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 3, 4, 6, 10, 11, 12, 13, 14, 16 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Yadav et al. (US 20210201091 A1); hereinafter Yadav in view of Wood et al. (US 10665251 B1); hereinafter Wood in view of Fayyad et al. (Deep Learning Sensor Fusion for Autonomous Vehicle Perception and Localization: A Review); hereinafter Fayyad and in further view of Yamato et al. (US 20200412807 A1); hereinafter Yamato Claim 1 is rejected over Yadav, Wood, Fayyad and Yamato. Regarding claim 1, Yadav teaches a computer-implemented method, comprising: engaging, by one or more processors, based on a request for an inference, from a group of sensors of multiple modalities at a physical location, at least one sensor [of a main modality] to provide data to a pipeline to generate the inference, (Yadav [0044]: “As illustrated in FIG. 2, sensor data 18 a, 18 b, 18 c generated by each sensor 14 a, 14 b, 14 c of the plurality of sensors 14, including first sensor data 18 a from the first sensor 14 a, second sensor data 18 b from the second sensor 14 b, and third sensor data 18 c from the third sensor 14 c, is received by the data quality module 26. The data quality module 26 is configured to assess the quality of the sensor data 18 a, 18 b, 18 c generated from each sensor 14 a, 14 b, 14 c.”; [0038]: “The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor, a thermopile sensor, a microwave sensor, an image sensor, a sound sensor, and/or any other sensor modality.”; and [0009]: “the method further comprises: receiving a classification output from each selected machine learning algorithm for each selected sensor, wherein the classification output is a result generated by performing the task;”; Note: The classification is the inference.) utilizing sensors comprising Internet of Things devices, (Yadav [0007]: “The context-switching algorithm for adaptive fusion of sensor data may be applied to Internet of Things (IoT) lighting systems. The IoT lighting system may comprise luminaires arranged to illuminate a room. Multiple sensors (for example, image sensors, sound sensors, infrared sensors, etc.) may be arranged in the room to detect what is happening in the room and to provide data to accomplish a task.”) wherein the pipeline comprises one or more machine learning models, and (See Figure 2 of Yadav to see that there is a pipeline of an Artificial Neural Net (ANNN) for each Sensor N (SN).) wherein the one or more machine learning models generate the inference for a downstream task; (Yadav [0009]: “the method further comprises: receiving a classification output from each selected machine learning algorithm for each selected sensor, wherein the classification output is a result generated by performing the task;”; Note: The classification is the inference.) determining, by the one or more processors, based on the one or more machine learning models and the main modality, one or more modalities which provide data to generate the inference for a downstream task in addition to the main modality data, (Yadav [0047]: “The ability of a particular machine learning algorithm to use sensor data 18 a, 18 b, 18 c to accurately accomplish a task depends on many factors, including the type of sensor data 18 a, 18 b, 18 c, the quality of sensor data 46 a, 46 b, 46 c, and the specific task. A non-exhaustive list of exemplary machine learning algorithms which can be used to accomplish a task includes support vector machines, nearest neighbors, decision trees, Bayesian algorithms, neural networks, deep learning based algorithms, or any other known machine learning algorithms or combinations thereof. Machine learning algorithms selected to be included set of predetermine machine learning algorithms 44 may be selected based on the known performance of those algorithms in accomplishing the task. As another example, the machine learning algorithms included in the set of predetermine machine learning algorithms 44 may be selected by evaluating their performance accomplishing a given task using a benchmark data set meeting certain quality thresholds.”; [0038]: “The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor, a thermopile sensor, a microwave sensor, an image sensor, a sound sensor, and/or any other sensor modality.”; and [0006]: “the system selects a specific sensor or a combination of sensors along with a corresponding machine learning algorithm to achieve enhanced performance in performing a task.”;) based on determining the one or more modalities which provide data to generate the inference for a downstream task in addition to the main modality data, (Yadav [0038]: “The plurality of sensors 14 may comprise sensors 14 a, 14 b, 14 c of one type or multiple types and are arranged to provide data to perform a set of preselected tasks. The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor, a thermopile sensor, a microwave sensor, an image sensor, a sound sensor, and/or any other sensor modality.”) automatically engaging, by the one or more processors, at least one sensor of at least one different modality than the [main] modality from the group of sensors of multiple modalities; (Yadav [0038]: “The plurality of sensors 14 may comprise sensors 14 a, 14 b, 14 c of one type or multiple types and are arranged to provide data to perform a set of preselected tasks. The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor, a thermopile sensor, a microwave sensor, an image sensor, a sound sensor, and/or any other sensor modality.”) based on the automatically engaging of the at least one sensor of the at least one different modality, obtaining, by the one or more processors, new raw data from the at least one sensor of the at least one different modality; and (Yadav [0030]: “Multiple sensors (for example, image sensors, sound sensors, infrared sensors, etc.) may be arranged in the room to detect what is happening in the room. The context-switching algorithm can be utilized to determine which of the multiple sensors to use given the quality of the data that each sensor is generating.”; Note: See Figure 2 of Yadav to see that each Sensor N (SN) processes its own Sensor Data (D¬N), therefore each Sensor Data (D¬N) after D1 is new raw data.) applying, by the one or more processors, the one or more machine learning models to the new raw data to derive the inference. (Yadav [0049]: “For every sensor (e.g., 14 a, 14 b, 14 c) of the plurality of sensors 14 which the assessment artificial intelligence program 62 determines to use, a machine learning algorithm (from the predefined set of machine learning algorithms 44 (shown in FIG. 3)) is used to provide a classification output (e.g., 50 a, 50 b). A classification output 50 a, 50 b is an output from the machine learning algorithm (e.g., 44 a, 44 b, 44 c, 44 d) for a particular sensor (e.g., 14 a, 14 b, 14 c) and is a result generated by performing the task 16.”; Note: A classification output is the derived inference.) Yadav does not appear to explicitly teach engaging, by one or more processors, based on a request for an inference, from a group of sensors of multiple modalities at a physical location, at least one sensor of a main modality to provide data to a pipeline to generate the inference, based on the engaging of the at least one sensor of the main modality, obtaining, by the one or more processors, raw data from the at least one sensor of the main modality; wherein the at least one sensor of the at least one different modality than the main modality comprises the one or more modalities. automatically engaging, by the one or more processors, at least one sensor of at least one different modality than the main modality from the group of sensors of multiple modalities; However, Wood teaches engaging, by one or more processors, based on a request for an inference, from a group of sensors of multiple modalities at a physical location, at least one sensor of a main modality to provide data to a pipeline to generate the inference, (Wood [col. 3, lines 5-9]: “combining 5 audio with multiple other supporting modalities, using the results from the supporting modalities to eliminate the large number of anomalous devices of this nature.”; [col. 1, lines 25-29]: “receive primary sensor data from a primary 25 sensor for a first device and determining a baseline from the primary sensor data for the first device.”; and [col. 5, lines 38-44]: “The multi-modal anomaly detection program 138 may generate and use the anomaly dependency graph 136 to detect an anomaly in the device 140. The multi-modal anomaly detection program 138 may receive the primary sensor data 114 and the secondary sensor data 124 a-124 c, which may be received and/or collected by the server 130 and stored as the program data 134 in the program database 132.”; Note: The primary sensor is the main modality.) based on the engaging of the at least one sensor of the main modality, obtaining, by the one or more processors, raw data from the at least one sensor of the main modality; (Wood [col. 3, lines 5-9]: “combining 5 audio with multiple other supporting modalities, using the results from the supporting modalities to eliminate the large number of anomalous devices of this nature.”; [col. 1, lines 25-29]: “receive primary sensor data from a primary 25 sensor for a first device and determining a baseline from the primary sensor data for the first device.”; and [col. 5, lines 38-44]: “The multi-modal anomaly detection program 138 may generate and use the anomaly dependency graph 136 to detect an anomaly in the device 140. The multi-modal anomaly detection program 138 may receive the primary sensor data 114 and the secondary sensor data 124 a-124 c, which may be received and/or collected by the server 130 and stored as the program data 134 in the program database 132.”; Note: The primary sensor is the main modality.) wherein the at least one sensor of the at least one different modality than the main modality comprises the one or more modalities. (Wood [col. 3, lines 5-9]: “combining 5 audio with multiple other supporting modalities, using the results from the supporting modalities to eliminate the large number of anomalous devices of this nature.”; [col. 1, lines 25-29]: “receive primary sensor data from a primary 25 sensor for a first device and determining a baseline from the primary sensor data for the first device.”; and [col. 5, lines 38-44]: “The multi-modal anomaly detection program 138 may generate and use the anomaly dependency graph 136 to detect an anomaly in the device 140. The multi-modal anomaly detection program 138 may receive the primary sensor data 114 and the secondary sensor data 124 a-124 c, which may be received and/or collected by the server 130 and stored as the program data 134 in the program database 132.”; Note: The primary sensor is the main modality.) automatically engaging, by the one or more processors, at least one sensor of at least one different modality than the main modality from the group of sensors of multiple modalities; (Wood [col. 4, lines 43-49]: “While three secondary sensors 120 a-120 c are depicted it can be appreciated that the multi-modal anomaly detection system 100 may include any number of secondary sensors including fewer than three secondary sensors 120 or more than three secondary sensors. The secondary sensors 120 a-120 c are described in more detail with reference to FIG. 4.”) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the anomaly detection in the primary sensor of Wood for improved maintenance (Wood, col. 1, Background). Yadav and Wood are analogous art because they both concern sensor selection based on quality of sensor data. Yadav does not appear to explicitly teach in a physical setting comprising environmental detriments to utilizing sensors However, Fayyad teaches in a physical setting comprising environmental detriments to utilizing sensors (Fayyad [page 4]: “thermal cameras have been fused with either RGB-D [24] or LiDAR sensors [25,26] to add depth, and hence improve the system performance; however, this advantage can be dramatically compromised in extreme weather conditions, such as high temperatures”;) It would have been obvious before the effective filing date to combine the infrared sensor, image sensor and sound sensor of Yadav with the combinations of different sensors of Fayyad to effectively compensate for losses due to outages in certain environmental conditions (Fayyad, page 5). Yadav and Fayyad are analogous art because they both concern multimodal sensors. Yadav does not appear to explicitly teach wherein the group of sensors of the multiple modalities are integrated into a roaming edge device, applying, by the one or more processors, an outlier detector to the raw data to determine if there is an outlier in the raw data; based on determining that there is an outlier in the raw data, However, Yamato teaches wherein the group of sensors of the multiple modalities are integrated into a roaming edge device, (Yamato [0045]: “Each real sensor 12 may be a stationary sensor, or a mobile sensor, such as a mobile phone, a smartphone, or a tablet. Each real sensor 12 may be a single sensing device, or may include multiple sensing devices”) applying, by the one or more processors, an outlier detector to the raw data to determine if there is an outlier in the raw data; (Yamato [0007]: “This session control apparatus switches the first device that outputs input data to the processing module to the second device when the input data fails to satisfy the condition (outlier). Thus, the session control apparatus discontinues input of data failing to satisfy the condition into the processing module and can maintain the quality of input data.”; and [0064]: “The quality conditions relate to the quality of input data (sensing data). Examples of the conditions include an outlier condition, a data dropout frequency condition, and a manufacturer condition. Besides these, the quality conditions may also include a sensor state condition, a sensor installation condition, a sensor maintenance history condition, a data specification condition, and a data resolution condition.”) based on determining that there is an outlier in the raw data, (Yamato: [0045]: “Each real sensor 12 observes a target to obtain sensing data. Each real sensor 12 may be an image sensor (camera), a temperature sensor, a humidity sensor, an illumination sensor, a force sensor, a sound sensor, a radio frequency identification (RFID) sensor, an infrared sensor, a posture sensor, a rain sensor, a radiation sensor, or a gas sensor. Each real sensor 12 may be any other sensor.”; and [0040]: “The session control apparatus 130 according to the present embodiment switches the input sensor to another real sensor 12 when the quality of input data output to the processing module 150 fails to satisfy conditions regarding the quality of input data output to the processing module 150. Thus, the session control apparatus 130 discontinues input of input data failing to satisfy the conditions into the processing module 150 and can maintain the quality of input data. “; Note: These real sensors are of different modalities and conditions include an outlier condition.) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the sensor switching unit of Yamato effectively use sensors based on data quality (Yamato, [0007]). Yadav and Yamato are analogous art because they both concern sensor selection based on quality of sensor data. Claim 2 is rejected over Yadav, Wood, Fayyad and Yamato with the incorporation of claim 1. Regarding claim 2, Yadav teaches based on determining that there is no outlier in the raw data, applying, by the one or more processors, the one or more machine learning models to the raw data to derive the inference. (Yadav [0050]: “selecting, via an assessment artificial intelligence program 62, one or more sensors 14 a, 14 b, 14 c from the plurality of sensors 14 based on the determined quality metric 46 a, 46 b, 46 c of the sensor data 18 a, 18 b, 18 c that yields a desired accuracy for performing the given task (step 230);”; Note: Yielding a desired accuracy shows that there is no outlier.) Claim 3 is rejected over Yadav, Wood, Fayyad and Yamato with the incorporation of claim 1. Regarding claim 3, Yadav teaches determining, by the one or more processors, the [main] modality of multiple modalities for sensor data provided to the pipeline to generate the inference; (Yadav [0006]: “the quality of the data generated from the different sensors is estimated using sensor-relevant data quality estimation techniques. Then, the system selects a specific sensor or a combination of sensors along with a corresponding machine learning algorithm to achieve enhanced performance in performing a task.”) obtaining, by the one or more processors, data from the group of sensors of the multiple modalities; (Yadav [0038]: “The plurality of sensors 14 may comprise sensors 14 a, 14 b, 14 c of one type or multiple types and are arranged to provide data to perform a set of preselected tasks. The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor, a thermopile sensor, a microwave sensor, an image sensor, a sound sensor, and/or any other sensor modality.”) utilizing, by the one or more processors, the data from the group of sensors to train the one or more machine learning models, based on the physical location; and (Yadav [0044]: “the sensor data 18 a, 18 b, 18 c and the sensor data quality metric 46 a, 46 b, 46 c for each sensor 14 a, 14 b, 14 c are inputted into an assessment artificial intelligence program 62 (e.g., artificial neural network 40 a, 40 b, 40 c) operated by processor 22. The assessment artificial intelligence program 62 is a trained, supervised machine learning algorithm whose objective it is to learn if sensor data 18 a, 18 b, 18 c having a determined quality metric 46 a, 46 b, 46 c can lead to an accurate and/or satisfactory execution of a task.”;) Yadav does not appear to explicitly teach determining, by the one or more processors, the main modality of multiple modalities for sensor data provided to the pipeline to generate the inference; However, Wood teaches determining, by the one or more processors, the main modality of multiple modalities for sensor data provided to the pipeline to generate the inference; (Wood [col. 3, lines 5-9]: “combining 5 audio with multiple other supporting modalities, using the results from the supporting modalities to eliminate the large number of anomalous devices of this nature.”; [col. 1, lines 25-29]: “receive primary sensor data from a primary 25 sensor for a first device and determining a baseline from the primary sensor data for the first device.”; and [col. 5, lines 38-44]: “The multi-modal anomaly detection program 138 may generate and use the anomaly dependency graph 136 to detect an anomaly in the device 140. The multi-modal anomaly detection program 138 may receive the primary sensor data 114 and the secondary sensor data 124 a-124 c, which may be received and/or collected by the server 130 and stored as the program data 134 in the program database 132.”; Note: The primary sensor is the main modality.) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the anomaly detection in the primary sensor of Wood for improved maintenance (Wood, col. 1, Background). Yadav and Wood are analogous art because they both concern sensor selection based on quality of sensor data. Yadav does not appear to explicitly teach generating, by the one or more processors, an outlier detector for each of the one or more machine learning models, based on the data from the group of sensors. However, Yamato teaches generating, by the one or more processors, an outlier detector for each of the one or more machine learning models, based on the data from the group of sensors. (Yamato [0077]: “When multiple pieces of real sensor information are received, the prioritizing unit 113 prioritizes the real sensors 12. The prioritizing unit 113 may prioritize each of the real sensors 12 with any criterion. When, for example, the outlier condition is prioritized among the conditions contained in the user data catalogue, the prioritizing unit 113 may assign a higher priority to the real sensor 12 having a lower outlier frequency.”;) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the sensor switching unit based on outlier condition in sensor data of Yamato effectively use sensors based on data quality (Yamato, [0007]). Yadav and Yamato are analogous art because they both concern sensor selection based on quality of sensor data. Claim 4 is rejected over Yadav, Wood, Fayyad and Yamato with the incorporation of claim 1. Regarding claim 4, Yadav teaches generating, by one or more processors, the pipeline. (Yadav [0043]: “Referring to FIG. 2, a data quality module 26 is configured to receive sensor data 18 a, 18 b, 18 c generated by a plurality of sensors 14 in the environment 2 (shown in FIG. 1). Processor 22 (shown in FIG. 1) is arranged to operate an assessment artificial intelligence program 62, such as an artificial neural network 40, which can make determinations about which sensors 14 a, 14 b, 14 c to use, based on the quality of the sensor data 18 a, 18 b, 18 c, to perform a task 16 (shown in FIG. 3). Additionally, the assessment artificial intelligence program 62 can determine which machine learning algorithm from a predetermined set of machine learning algorithms 44 (shown in FIG. 3) should be used to perform the task.”; Note: See Figure 2 of Yadav to see that each sensor has its own artificial intelligence pipeline.) Claim 6 is rejected over Yadav, Wood, Fayyad and Yamato with the incorporation of claim 1. Regarding claim 6, Yadav teaches wherein the pipeline is an artificial intelligence pipeline and the task is an artificial intelligence task. (Yadav [0043]: “Referring to FIG. 2, a data quality module 26 is configured to receive sensor data 18 a, 18 b, 18 c generated by a plurality of sensors 14 in the environment 2 (shown in FIG. 1). Processor 22 (shown in FIG. 1) is arranged to operate an assessment artificial intelligence program 62, such as an artificial neural network 40, which can make determinations about which sensors 14 a, 14 b, 14 c to use, based on the quality of the sensor data 18 a, 18 b, 18 c, to perform a task 16 (shown in FIG. 3). Additionally, the assessment artificial intelligence program 62 can determine which machine learning algorithm from a predetermined set of machine learning algorithms 44 (shown in FIG. 3) should be used to perform the task.”; Note: See Figure 2 of Yadav to see that each sensor has its own artificial intelligence pipeline.) Claim 10 is rejected over Yadav, Wood, Fayyad and Yamato with the incorporation of claim 1. Regarding claim 10, Yadav teaches wherein the at least one sensor of at least one different modality [than the main modality comprises all available sensors at the location.] (Yadav [0038]: “The plurality of sensors 14 may comprise sensors 14 a, 14 b, 14 c of one type or multiple types and are arranged to provide data to perform a set of preselected tasks. The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor, a thermopile sensor, a microwave sensor, an image sensor, a sound sensor, and/or any other sensor modality.”) Yadav does not appear to explicitly teach wherein the at least one sensor of at least one different modality than the main modality comprises all available sensors at the location. However, Wood teaches wherein the at least one sensor of at least one different modality than the main modality comprises all available sensors at the location. (Wood [col. 4, lines 43-49]: “While three secondary sensors 120 a-120 c are depicted it can be appreciated that the multi-modal anomaly detection system 100 may include any number of secondary sensors including fewer than three secondary sensors 120 or more than three secondary sensors. The secondary sensors 120 a-120 c are described in more detail with reference to FIG. 4.”) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the anomaly detection in the primary sensor of Wood for improved maintenance (Wood, col. 1, Background). Yadav and Wood are analogous art because they both concern sensor selection based on quality of sensor data. Claim 11 is rejected over Yadav, Wood, Fayyad and Yamato. Regarding claim 11, Yadav teaches a computer program product comprising: a computer readable storage medium readable by one or more processors of a shared computing environment comprising a computing system and storing instructions for execution by the one or more processors for performing a method comprising: (Yadav [0053]: “The present disclosure may be implemented as a system, a method, and/or a computer program product at any possible technical detail level of integration. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure”) The remainder of claim 11 is claim 1 in the form of a computer readable storage medium and is rejected for the same reasons as claim 1 stated above. Dependent claim 12 is claim 2 in the form of a computer program product and is rejected for the same reasons as claim 2 stated above. For the rejection of the limitations specifically pertaining to the computer program product of claim 11, see the rejection of claim 11 above. Dependent claim 13 is claim 3 in the form of a computer program product and is rejected for the same reasons as claim 3 stated above. For the rejection of the limitations specifically pertaining to the computer program product of claim 11, see the rejection of claim 11 above. Dependent claim 14 is claim 4 in the form of a computer program product and is rejected for the same reasons as claim 4 stated above. For the rejection of the limitations specifically pertaining to the computer program product of claim 11, see the rejection of claim 11 above. Dependent claim 16 is claim 6 in the form of a computer program product and is rejected for the same reasons as claim 6 stated above. For the rejection of the limitations specifically pertaining to the computer program product of claim 11, see the rejection of claim 11 above. Claim 21 is rejected over Yadav, Wood, Fayyad and Yamato with the incorporation of claim 1. Regarding claim 21, Yadav does not appear to explicitly teach wherein the environmental detriments comprise elevated temperatures. However, Fayyad teaches wherein the environmental detriments comprise elevated temperatures. (Fayyad [page 4]: “thermal cameras have been fused with either RGB-D [24] or LiDAR sensors [25,26] to add depth, and hence improve the system performance; however, this advantage can be dramatically compromised in extreme weather conditions, such as high temperatures”;) It would have been obvious before the effective filing date to combine the infrared sensor, image sensor and sound sensor of Yadav with the combinations of different sensors of Fayyad to effectively compensate for losses due to outages in certain environmental conditions (Fayyad, page 5). Yadav and Fayyad are analogous art because they both concern multimodal sensors. Claims 5, 9 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Yadav, Wood, Fayyad and Yamato in view of Gonzalez Aguirre et al. (US 20190135300 A1); hereinafter Gonzalez Aguirre Claim 5 is rejected over Yadav, Wood, Fayyad, Yamato and Gonzalez Aguirre with the incorporation of claim 1. Regarding claim 5, Yadav does not appear to explicitly teach wherein the raw data comprises unlabeled data. However, Gonzalez Aguirre teaches wherein the raw data comprises unlabeled data. (Gonzalez Aguirre [0069]: “FIGS. 10A and 10B depict an example end-to-end system training data flow 1000 of the anomaly detection apparatus 306 of FIGS. 3 and 4 to perform unsupervised multimodal anomaly detection for autonomous vehicles using the example feature fusion and deviation data flow 700 of FIG. 7. The example end-to-end system training data flow 1000 is shown as including six phases. At an example first phase (1) 1002 (FIG. 10A), the sensor data interface 402 (FIG. 4) obtains multimodal raw sensor data samples (Ii(x,y,t)) (unlabeled data) from a database to train auto-encoders for corresponding ones of the sensors 202, 204, 206, 304 (FIGS. 2 and 3).”) It would have been obvious before the effective filing date to combine the multimodal sensors of Yadav with the unsupervised multimodal anomaly detection of Gonzalez Aguirre to effectively perform anomaly detection for autonomous systems (Gonzalez Aguirre, [0017]). Yadav and Gonzalez Aguirre are analogous art because they both concern processing data using multimodal sensors. Claim 9 is rejected over Yadav, Wood, Fayyad, Yamato and Gonzalez Aguirre with the incorporation of claim 1. Regarding claim 9, Yadav teaches wherein the [main] modality is selected from the group consisting of: optical, audio, infrared, and (Yadav [0038]: “The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor (infrared), a thermopile sensor, a microwave sensor, an image sensor (optical), a sound sensor (audio), and/or any other sensor modality.”;) Yadav does not appear to explicitly teach light detecting and ranging. However, Gonzalez Aguirre teaches light detecting and ranging. (Gonzalez Aguirre [0016]: “Autonomous robotic systems such as autonomous vehicles use multiple cameras as well as range sensors to perceive characteristics of their environments. The different sensor types (e.g., infrared (IR) sensors, red-green-blue (RGB) color cameras, Light Detection and Ranging (LIDAR) sensors, Radio Detection and Ranging (RADAR) sensors, SOund Navigation And Ranging (SONAR) sensors, etc.) can be used together in heterogeneous sensor configurations useful for performing various tasks of autonomous vehicles”;) It would have been obvious before the effective filing date to combine the infrared sensor, image sensor and sound sensor of Yadav with the LIDAR sensor of Gonzalez Aguirre to effectively perform various autonomous vehicle tasks (Gonzalez Aguirre, [0016]). Yadav and Gonzalez Aguirre are analogous art because they both concern processing data using multimodal sensors. Yadav does not appear to explicitly teach wherein the main modality is selected However, Wood teaches wherein the main modality is selected (Wood [col. 3, lines 33-37]: “The primary sensor device 110 may include the primary sensor database 112. The primary sensor 110 may be any device capable of capturing the primary sensor data 114. The primary sensor data 114 may include, but is not limited to, visual, audio, textual data, and/or physical data”) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the anomaly detection in the primary sensor of Wood for improved maintenance (Wood, col. 1, Background). Yadav and Wood are analogous art because they both concern sensor selection based on quality of sensor data. Dependent claim 15 is claim 5 in the form of a computer program product and is rejected for the same reasons as claim 5 stated above. For the rejection of the limitations specifically pertaining to the computer program product of claim 11, see the rejection of claim 11 above. Claims 17, 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Yadav, Wood, Fayyad and Yamato in view of Lee et al. (Making Sense of Vision and Touch: Self-Supervised Learning of Multimodal Representations for Contact-Rich Tasks); hereinafter Lee Claim 17 is rejected over Yadav, Wood, Fayyad, Yamato and Lee. Regarding claim 17, Yadav teaches a computer system comprising: a group of sensors of multiple modalities communicatively coupled to one or more processors; (Yadav [0011]: “In an aspect, the plurality of sensors include at least one of a PIR sensor, a thermopile sensor, a microwave sensor, an image sensor, and a sound sensor.”) a memory; (Yadav [0013]: “The plurality of non-transitory computer readable instructions arranged to be stored and executed on a memory and a processor.”) the one or more processors in communication with the memory; program instructions executable by the one or more processors to perform a method, the method comprising: (Yadav [0013]: “The plurality of non-transitory computer readable instructions arranged to be stored and executed on a memory and a processor.”) engaging, by the one or more processors, based on a request for an inference, at least one sensor of [a main modality] of the multiple modalities to provide data to a pipeline to generate the inference, (Yadav [0044]: “As illustrated in FIG. 2, sensor data 18 a, 18 b, 18 c generated by each sensor 14 a, 14 b, 14 c of the plurality of sensors 14, including first sensor data 18 a from the first sensor 14 a, second sensor data 18 b from the second sensor 14 b, and third sensor data 18 c from the third sensor 14 c, is received by the data quality module 26. The data quality module 26 is configured to assess the quality of the sensor data 18 a, 18 b, 18 c generated from each sensor 14 a, 14 b, 14 c.”; [0038]: “The plurality of sensors 14 may be selected from a passive infrared (“PIR”) sensor, a thermopile sensor, a microwave sensor, an image sensor, a sound sensor, and/or any other sensor modality.”; and [0009]: “the method further comprises: receiving a classification output from each selected machine learning algorithm for each selected sensor, wherein the classification output is a result generated by performing the task;”; [0009]; Note: The classification is the inference.) wherein the pipeline comprises one or more machine learning models, and wherein the one or more machine learning models generate the inference for a downstream task, (Yadav [0009]: “the method further comprises: receiving a classification output from each selected machine learning algorithm for each selected sensor, wherein the classification output is a result generated by performing the task;”; Note: The classification is the inference. See Figure 2 of Yadav to see that there is a pipeline of an Artificial Neural Net (ANNN) for each Sensor N (SN).) Yadav does not appear to explicitly teach engaging, by the one or more processors, based on a request for an inference, at least one sensor of a main modality of the multiple modalities to provide data to a pipeline to generate the inference, However, Wood teaches engaging, by the one or more processors, based on a request for an inference, at least one sensor of a main modality of the multiple modalities to provide data to a pipeline to generate the inference, (Wood [col. 3, lines 5-9]: “combining 5 audio with multiple other supporting modalities, using the results from the supporting modalities to eliminate the large number of anomalous devices of this nature.”; [col. 1, lines 25-29]: “receive primary sensor data from a primary 25 sensor for a first device and determining a baseline from the primary sensor data for the first device.”; and [col. 5, lines 38-44]: “The multi-modal anomaly detection program 138 may generate and use the anomaly dependency graph 136 to detect an anomaly in the device 140. The multi-modal anomaly detection program 138 may receive the primary sensor data 114 and the secondary sensor data 124 a-124 c, which may be received and/or collected by the server 130 and stored as the program data 134 in the program database 132.”; Note: The primary sensor is the main modality.) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the anomaly detection in the primary sensor of Wood for improved maintenance (Wood, col. 1, Background). Yadav and Wood are analogous art because they both concern sensor selection based on quality of sensor data. Yadav does not appear to explicitly teach wherein the inference comprises a self-supervised joint representation generated by the one or more machine learning models, and wherein the self- supervised joint representation is utilized to infer a representation associated with the main modality based on samples associated with the one or more modalities in addition to the main modality; However, Lee teaches wherein the inference comprises a self-supervised joint representation generated by the one or more machine learning models, and wherein the self- supervised joint representation is utilized to infer a representation associated with the main modality based on samples associated with the one or more modalities in addition to the main modality; (Lee [page 3, A. Modality Encoders]: “The resulting three feature vectors are concatenated into one vector and passed through the multimodal fusion module (2-layer MLP) to produce the final 128-d multimodal representation.”; [page 1, Abstract]: “We use self-supervision to learn a compact and multimodal representation of our sensory inputs,”; and [page 5, Figure 4]: “Ablative study of representations trained on different combinations of sensory modalities. We compare our full model, trained with a combination of visual and haptic feedback and proprioception, with baselines that are trained without vision, or haptics, or either.”) It would have been obvious before the effective filing date to combine the adaptive multimodal sensors of Yadav with the self-supervised joint representation of Lee for improved sample efficiency (Lee, Abstract). Yadav and Lee are analogous art because they both concern sensor selection based on quality of sensor data. The remainder of claim 17 is claim 1 in the form of a system and is rejected for the same reasons as claim 1 stated above. Dependent claim 18 is claim 2 in the form of a computer system and is rejected for the same reasons as claim 2 stated above. For the rejection of the limitations specifically pertaining to the computer program product of claim 17, see the rejection of claim 17 above. Dependent claim 19 is claim 3 in the form of a computer system and is rejected for the same reasons as claim 3 stated above. For the rejection of the limitations specifically pertaining to the computer program product of claim 17, see the rejection of claim 17 above. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Yadav, Wood, Fayyad, Yamato and Lee in view of Gonzalez Aguirre Claim 20 is rejected over Yadav, Wood, Fayyad, Yamato, Lee and Gonzalez Aguirre with the incorporation of claim 17. Regarding claim 20, Yadav does not appear to explicitly teach wherein a roaming edge device comprises the group of sensors of the multiple modalities communicatively and the one or more processors. However, Gonzalez Aguirre teaches wherein a roaming edge device comprises the group of sensors of the multiple modalities communicatively and the one or more processors. (Gonzalez Aguirre [0019]: “In the heterogeneous sensor configuration of FIG. 1, the autonomous vehicle 100 (roaming edge device) is provided with camera sensors, RADAR sensors, and LIDAR sensors”;) It would have been obvious before the effective filing date to combine the infrared sensor, image sensor and sound sensor of Yadav with the LIDAR sensor of Gonzalez Aguirre to effectively perform various autonomous vehicle tasks (Gonzalez Aguirre, [0016]). Yadav and Gonzalez Aguirre are analogous art because they both concern processing data using multimodal sensors. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVID H TRAN whose telephone number is (703)756-1525. The examiner can normally be reached M-F 9:30 am - 5:30 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DAVID H TRAN/Examiner, Art Unit 2147 /VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147
Read full office action

Prosecution Timeline

Dec 10, 2021
Application Filed
Oct 25, 2023
Response after Non-Final Action
Aug 15, 2025
Non-Final Rejection mailed — §103
Nov 17, 2025
Response Filed
Feb 20, 2026
Final Rejection mailed — §103
May 20, 2026
Request for Continued Examination
May 23, 2026
Response after Non-Final Action
Jul 06, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12632724
CANONICALIZATION OF DATA WITHIN OPEN KNOWLEDGE GRAPHS
4y 8m to grant Granted May 19, 2026
Patent 12579404
PROCESSOR FOR NEURAL NETWORK, PROCESSING METHOD FOR NEURAL NETWORK, AND NON-TRANSITORY COMPUTER READABLE STORAGE MEDIUM
4y 2m to grant Granted Mar 17, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
21%
Grant Probability
35%
With Interview (+14.1%)
4y 5m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 19 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month