Prosecution Insights
Last updated: October 01, 2026
Application No. 18/941,269

INFERENCING USING NEURAL NETWORKS

Non-Final OA §101§103
Filed
Nov 08, 2024
Priority
Jul 30, 2021 — continuation of 12/183,050
Examiner
YANG, WEI WEN
Art Unit
Tech Center
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
560 granted / 684 resolved
+21.9% vs TC avg
Moderate +12% lift
Without
With
+11.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
31 currently pending
Career history
705
Total Applications
across all art units

Statute-Specific Performance

§101
7.8%
-32.2% vs TC avg
§103
75.0%
+35.0% vs TC avg
§102
9.3%
-30.7% vs TC avg
§112
7.8%
-32.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 684 resolved cases

Office Action

§101 §103
DETAILED ACTION Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-4, 7-8, 13-14, and 16-18 are subject to the claim rejection under 35 U.S.C. 101 because the each claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more: According to the 2024 Guidance Update on Patent Subject Matter Eligibility, Step 1 analysis: claims 1-4, and 7 are identified as directed to a process; 8, 13-14, and 16-18 are identified as directed to a machine; In Step 2A, Prong One Analysis evaluates whether the claim recites a judicial exception: independent claim 1 recites “identifying one or more objects in one or more images”, “comparing a size of the one or more objects in a first image of the one or more images to a size of the one or more objects in a second image of the one or more images”; and “re-inferencing the one or more objects based, at least in part, on the comparison”; which are as abstract idea of Mental processes, that can be mentally carried out solely by human, and without any particular algorithms or improvements, a person would be able to identifying one or more objects in images, and comparing a size of the one or more objects in images, and re-inferencing the one or more objects based, at least in part, on the comparison, because these are human observation and perception, as the mental processes based on the person’s observation and assessments. Step (2) Analysis, each of independent claim 1 recites abstract idea of Mental processes, those steps are human observation, perception and mental processes. There is no limitation, or combination of another element, and there is no additional element. The claim does not limit any other particular device, or algorithms, or any modification or improvement of the device. In Step 2B, Prong Two Analysis to evaluate whether the claimed additional elements amount to significantly more than the recited judicial exception itself: There are no any additional elements recited in claim beyond the Mental processes. As no particular computer function, algorithm or configuration is practically improved, and no any particular application, or improvement in any meaningful practice is defined. Therefore, independent claim 1 do not have additional elements amount to significantly more than the recited judicial exception itself. Overall, Claim 1 does not limit what’s specialty, algorithms and characteristics of a method, or a device with any specific algorithms or improvements. Therefore, Claim 1 is considered as human mental processing, and a conventional data and image mental processes, and would be conventional practice of human observation and mental processing. Therefore above mentioned steps are considered as an abstract idea of Mental processes. Dependent claims 2-4 merely limit comparing size by determining a width/or a height changes/increases, and claim 4 further limit comparing size determining whether a result of the comparison exceeds a predetermined threshold; claim 7 limits “classifying the one or more objects in the first image according to one or more classifications; and applying at least one classification of the one or more classifications to the one or more objects in the second image…”, however, since the contents, characteristics, attributions, or types of the classifying and classifications are not specifically defined, “classifying the one or more objects…in the first image”, and “applying at least one classification in the second image” are further regarded as an abstract idea of Mental processes; so that overall claims 1-4, and 7 are human observation, perception and mental processes. Similarly, claims 8, 13-14, 16-18 limit a system, comprising: one or more processors comprising one or more circuits applied to steps, which otherwise are being determined as the abstract idea of mental processing as discussed above for claims 1-4, and 7; and “one or more processors comprising one or more circuits” is considered as a generic device, or a generic computer; since there is no particular computer function, algorithm or configuration is practically improved, and no any particular application, or improvement in any meaningful practice is defined. Therefore, independent claim 8, 13-14, 16-18 do not have additional elements other than the “generic device, processor, or computer” amount to significantly more than the recited judicial exception itself. On the other hands, claims 5-6, 9-12, 15, and 19-20, recites additional limitation with specific implementations respectively, which are considered as an additional element require a neural network to process the obtained data, features from the images and information. Thus, claims 5-6, 9-12, 15, and 19-20 are eligible in Patent Subject Matter Eligibility analysis. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kumar (US 20230042004 A1, DATE FILED: 2021-02-05, and claims priority of us-provisional-application US 62971862 20200207, as provided in IDS), in view of ROZNER (US 20200097778 A1, as provided in IDS). Re Claim 1, Kumar discloses a method, comprising: identifying one or more objects in one or more images (see Kumar: e.g., --generating one or more views materializing one or more tensors generated by a convolutional neural network operating on a frame of a video; determining a first output of the convolutional neural network operating on a first subsequent frame of the video to identify an object in the video, the output being determined by at least generating a first query on the one or more views, the first query performing, based at least on a first change in the first subsequent frame, a first partial re-computation of the one or more views; and determining, based at least on the output of the convolutional neural network, a label identifying the object in the video.--, in [0030]; and, -- convolutional neural networks (CNNs), a type of deep machine learning model, may be especially adept at performing analytical tasks on images, videos, and time series data....convolutional neural network inference may consume significant time and computational resource. For instance, inferences performed using a VGG16 (OxfordNet) convolutional neural network requires 15 gigaflops of computations. Moreover, many analytical tasks involving convolutional neural networks require repeated convolutional neural network inferences, amplifying the computational cost and latency. For this reason, the adoption of convolutional neural networks may be especially unwieldy for interactive and/or resource-constrained settings such as mobile, browser, and edge devices. The adoption of convolutional neural networks may also be excessively costly for regular server and cloud settings. [0077] In some example embodiments, a machine learning engine may apply one or more query optimization techniques to reduce the computational cost associated with convolutional neural network inferences. For example, the machine learning engine may perform convolutional neural network inferences as queries that reuse, for each convolutional neural network inference, as much of the results of previous computations as possible. As such, query-based convolutional neural network inferences may be performed for tasks that require repeated convolutional neural network inferences on slightly modified inputs. For instance, query-based convolutional neural network inferences may be performed for the analytical task of occlusion-based explanation (OBE) in which an explanation of a convolutional neural network, such as a heatmap, is generated to identify regions of an input (e.g., an image) responsible for an output (e.g., a label classifying the image) of the convolutional neural network. Alternatively and/or additionally, query-based convolutional neural network inferences may be performed for the analytical task of object recognition in video (ORV) in which a convolutional neural network operates to identify an object over multiple frames of a video. [0078] To perform the task of occlusion-based explanation (OBE), the machine learning engine may perform convolutional neural network inferences on an image with a patch occluding different portions (e.g., pixels) of the image. For example, the machine learning engine may perform a first convolutional neural network inference on the image with a patch occluding a first portion of the image and a second convolutional neural network inference on the image with the patch occluding a second portion of the image. The machine learning engine may further track changes in the output of the convolutional neural network when different portions of the image are occluded with the patch. For instance, occluding the first portion of the image may have more effect on the output of the convolutional neural network than occluding the second portion of the image. Accordingly, the machine learning engine may generate a heatmap having different representations (e.g., colors, symbols, and/or the like) for the first portion of the image and the second portion of the image in order to indicate that the first portion of the image is more (or less) responsible for the output of the convolutional neural network than the second portion of the image. Because the task of occlusion based explanation (OBE) requires multiple convolutional neural network inferences on slightly modified versions of the image, the machine learning model may perform query-based convolutional neural network inferences to minimize the overhead associated with the repeated inferences. [0079] The machine learning engine may, as noted, also perform convolutional neural network inferences to accomplish the task of object recognition in videos (ORV). Machine learning enabled object recognition in videos may be especially popular due to the mass deployment of video cameras in applications such as security surveillance, traffic monitoring, wild animal and livestock tracking, and/or the like. In object recognition in videos, each frame of a video may be treated as an individual image. A convolutional neural network may be trained to process each frame of the video to identify an object appearing in the video. For example, the trained convolutional neural network may perform single object recognition over a fixed-angle camera video feed. As such, the convolutional neural network may perform multiple inferences over largely identical video frames, which is why the machine learning engine may apply query-based convolutional neural network inferences to the task of object recognition in videos (ORV) to minimize the computational overhead associated with the task.--, in [0076]-[0080], and [0150]-[0152]); comparing a size of the one or more objects in a first image of the one or more images to a size of the one or more objects in a second image of the one or more images (see Kumar: e.g., Fig.1, --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022]; and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “comparing a size of the one or more objects in a first image of the one or more images to a size of the one or more objects in a second image of the one or more images”, because the “identify one or more portions” corresponding to “a size of the one or more objects“, as see Kumar’s Fig. 1} in [0033]-[0034]; also see: , -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134]); re-inferencing the one or more objects (see Kumar: e.g., Fig.1, --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022]; and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086]; and, {re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134], {so that “given threshold fraction τ (0<τ≤1)” read on “adjusting, a set of conditions” by setting the threshold fraction}; also see: -- this optimization also builds upon the incremental inference optimizations described earlier, but may be orthogonal to projective field thresholding and can therefore be used in unison. [0146] The notion of theoretical speedup for the adaptive drill-down optimization may be independent of the theoretical speedup associated with incremental inference. At the outset, setting the parameters r.sub.drill-down and S.sub.1 may be an application-specific balancing act. For example, if r.sub.drill-down is low, only a small region will require re-inferencing at the original resolution, which will save a lot of FLOPs. However, this may miss some regions of interest and thus compromise important explanation details. Similarly, a large S.sub.1 may also save a lot of FLOPs by reducing the number of re-inference queries in stage one but doing so may run the risk of misidentifying interesting regions, especially when the size of those regions are smaller than the occlusion patch.--, in [0145]-[0146]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]), Kumar however does not explicitly disclose re-inferencing the one or more objects based, at least in part, on the comparison, ROZNER discloses re-inferencing the one or more objects based, at least in part, on the comparison (see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068]); Kumar and ROZNER are combinable as they are in the same field of endeavor: using machine learning models in object inferences and optimizations of inferences. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify Kumar’s method using ROZNER’s teachings by including re-inferencing the one or more objects based, at least in part, on the comparison {in adjusting a set of conditions such as adjusting weights in re-models for re-inferences} to Kumar’s adjusting/setting regions in order to use optimized neural network to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize (see ROZNER: e.g. in [0058]-[0059], and [0065]-[0068]). Re Claim 2, Kumar as modified by ROZNER further disclose wherein comparing the size of the one or more objects in the first image to the size of the one or more objects in the second image comprises determining a change in a width and/or a height of the one or more objects between the first image to the second image (see Kumar: e.g., --A convolutional neural network, such as the convolutional neural network 115, may be organized as layers of various types, each of which transforming one tensor (e.g., a multidimensional array that is typically 3-D) into another tensor. The convolution layer may use image filters from graphics to extract features, but with parametric filter weights (learned during training). The pooling layer may subsample features in a spatial-aware manner, the batch-normalization layer may normalize the output tensor, the non-linearity layer may apply an element-wise non-linear function (e.g., ReLU), and the fully-connected layer may include an ordered collection of perceptrons. [0094] The output tensor of a layer can have a different width, height, and/or depth than the input tensor of the layer. An image may be viewed as a tensor, e.g., a 224×224 RGB image is a 3-D tensor with a width and height of 224 and a depth of 3.--, in [0091]-[0094], and, --8] Local spatial context granularity layers may perform weighted aggregations of slices of the input tensor, called local spatial contexts, by multiplying them with a filter kernel (a tensor of weights). Thus, input tensors and output tensors can differ in width, height, and depth. If the input is incrementally updated, the region of the output that will be affected is not straightforward to ascertain but requires non-trivial and careful calculations due to the overlapping nature of how filters get applied to local spatial contexts. Both convolution layers and pooling layers fall under this category. As shown in FIG. 5, such layers typically account for the bulk of the computational cost of deep convolutional neural network inferences (over 90%). Thus, enabling incremental inference for such layers in the context of occlusion-based explanation may be key to increasing the computational efficiency and reducing the latency of the task. [0099] A convolution layer l of the convolutional neural network f may have C.sub.O:l 3-D filter kernels arranged as a 4-D array K.sub.conv:l, with each having a smaller spatial width W.sub.K:l and height H.sub.K:l than the width W.sub.I:l and height H.sub.I:l of the input tensor I.sub.:l but the same depth C.sub.I:l. During inference, c.sup.th filter kernel may be “strided” along the width and height dimensions of the input to produce a 2-D “activation map” A.sub.:c=(a.sub.y,x:c)∈custom-character.sup.H.sup.O:l.sup.×W.sup.O:l by computing element-wise products between the kernel and the local spatial context and adding a bias value as per Equation (7) below. The computational cost of each of these small matrix products may be proportional to the volume of the filter kernel.--, in [0098]-[0099]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]). Re Claim 3, Kumar as modified by ROZNER further disclose re-inferencing the one or more objects based, at least in part, on an increase in the width and/or the height of the one or more objects from the first image to the second image (see Kumar: e.g., --A convolutional neural network, such as the convolutional neural network 115, may be organized as layers of various types, each of which transforming one tensor (e.g., a multidimensional array that is typically 3-D) into another tensor. The convolution layer may use image filters from graphics to extract features, but with parametric filter weights (learned during training). The pooling layer may subsample features in a spatial-aware manner, the batch-normalization layer may normalize the output tensor, the non-linearity layer may apply an element-wise non-linear function (e.g., ReLU), and the fully-connected layer may include an ordered collection of perceptrons. [0094] The output tensor of a layer can have a different width, height, and/or depth than the input tensor of the layer. An image may be viewed as a tensor, e.g., a 224×224 RGB image is a 3-D tensor with a width and height of 224 and a depth of 3.--, in [0091]-[0094], and, --8] Local spatial context granularity layers may perform weighted aggregations of slices of the input tensor, called local spatial contexts, by multiplying them with a filter kernel (a tensor of weights). Thus, input tensors and output tensors can differ in width, height, and depth. If the input is incrementally updated, the region of the output that will be affected is not straightforward to ascertain but requires non-trivial and careful calculations due to the overlapping nature of how filters get applied to local spatial contexts. Both convolution layers and pooling layers fall under this category. As shown in FIG. 5, such layers typically account for the bulk of the computational cost of deep convolutional neural network inferences (over 90%). Thus, enabling incremental inference for such layers in the context of occlusion-based explanation may be key to increasing the computational efficiency and reducing the latency of the task. [0099] A convolution layer l of the convolutional neural network f may have C.sub.O:l 3-D filter kernels arranged as a 4-D array K.sub.conv:l, with each having a smaller spatial width W.sub.K:l and height H.sub.K:l than the width W.sub.I:l and height H.sub.I:l of the input tensor I.sub.:l but the same depth C.sub.I:l. During inference, c.sup.th filter kernel may be “strided” along the width and height dimensions of the input to produce a 2-D “activation map” A.sub.:c=(a.sub.y,x:c)∈custom-character.sup.H.sup.O:l.sup.×W.sup.O:l by computing element-wise products between the kernel and the local spatial context and adding a bias value as per Equation (7) below. The computational cost of each of these small matrix products may be proportional to the volume of the filter kernel.--, in [0098]-[0099]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]; also see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068]). Re Claim 4, Kumar as modified by ROZNER further disclose determining whether a result of the comparison exceeds a predetermined threshold (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086], {re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134], {so that “given threshold fraction τ (0<τ≤1)” read on “adjusting, a set of conditions” by setting the threshold fraction}; also see: -- this optimization also builds upon the incremental inference optimizations described earlier, but may be orthogonal to projective field thresholding and can therefore be used in unison. [0146] The notion of theoretical speedup for the adaptive drill-down optimization may be independent of the theoretical speedup associated with incremental inference. At the outset, setting the parameters r.sub.drill-down and S.sub.1 may be an application-specific balancing act. For example, if r.sub.drill-down is low, only a small region will require re-inferencing at the original resolution, which will save a lot of FLOPs. However, this may miss some regions of interest and thus compromise important explanation details. Similarly, a large S.sub.1 may also save a lot of FLOPs by reducing the number of re-inference queries in stage one but doing so may run the risk of misidentifying interesting regions, especially when the size of those regions are smaller than the occlusion patch.--, in [0145]-[0146]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]); and re-inferencing the one or more objects based, at least in part, on the result of the comparison exceeding the predetermined threshold (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086], {re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134], {so that “given threshold fraction τ (0<τ≤1)” read on “adjusting, a set of conditions” by setting the threshold fraction}; also see: -- this optimization also builds upon the incremental inference optimizations described earlier, but may be orthogonal to projective field thresholding and can therefore be used in unison. [0146] The notion of theoretical speedup for the adaptive drill-down optimization may be independent of the theoretical speedup associated with incremental inference. At the outset, setting the parameters r.sub.drill-down and S.sub.1 may be an application-specific balancing act. For example, if r.sub.drill-down is low, only a small region will require re-inferencing at the original resolution, which will save a lot of FLOPs. However, this may miss some regions of interest and thus compromise important explanation details. Similarly, a large S.sub.1 may also save a lot of FLOPs by reducing the number of re-inference queries in stage one but doing so may run the risk of misidentifying interesting regions, especially when the size of those regions are smaller than the occlusion patch.--, in [0145]-[0146]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]). Re Claim 5, Kumar as modified by ROZNER further disclose wherein the predetermined threshold is dynamically adjusted based, at least in part, on an availability of computing resources to be used to perform re-inferencing on the one or more objects (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086], { re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134], {so that “given threshold fraction τ (0<τ≤1)” read on “adjusting, a set of conditions” by setting the threshold fraction}; also see: -- this optimization also builds upon the incremental inference optimizations described earlier, but may be orthogonal to projective field thresholding and can therefore be used in unison. [0146] The notion of theoretical speedup for the adaptive drill-down optimization may be independent of the theoretical speedup associated with incremental inference. At the outset, setting the parameters r.sub.drill-down and S.sub.1 may be an application-specific balancing act. For example, if r.sub.drill-down is low, only a small region will require re-inferencing at the original resolution, which will save a lot of FLOPs. However, this may miss some regions of interest and thus compromise important explanation details. Similarly, a large S.sub.1 may also save a lot of FLOPs by reducing the number of re-inference queries in stage one but doing so may run the risk of misidentifying interesting regions, especially when the size of those regions are smaller than the occlusion patch.--, in [0145]-[0146]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]; also see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; -- FIG. 4, a convolution/pooling scheme to analyze image data is presented in CNN convolution process 400. As shown in FIG. 4, pixel data from a photographic image populates an input table 402. Each cell in the input table 402 represents a value of a pixel in the photograph. This value is based on the color and intensity for each pixel. A subset of pixels from the input table 402 is associated with a filter 404. That is, filter 404 is matched to a same-sized subset of pixels (e.g., pixel subset 406) by sliding the filter 404 across the input table 402. The filter 404 slides across the input grid at some predefined stride (i.e., one or more pixels). Thus, if the stride is “1”, then the filter 404 slides over in increments of one (column) of pixels. In the example shown in FIG. 4, this results in the filter 404 sliding over the subset of pixels shown as pixel subset 406 (3,4,3,4,3,1,2,3,5 when read from left to right for each row) followed by filter 404 sliding over the subset of pixels just to the right (4,3,3,3,1,3,2,5,3). If the stride were “2”, then the next subset of pixels that filter 404 would slide to would be (3,3,1,1,3,3,5,3,4).--, in [0061], and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068]). Re Claim 6, Kumar as modified by ROZNER further disclose using one or more neural networks to generate one or more additional features of the one or more objects as a result of re-inferencing the one or more objects (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086], { re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134]; also see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068]; and, -- FIG. 5, inference is the process of using a trained CNN to recognize certain objects from a photograph or other data. In the example in FIG. 5, pixels from photograph 501 are input into a trained CNN (e.g., CNN 500), resulting in the identification and/or labeling (for display on the photograph 501) a particular object, such as the dog. [0071] That is, a CNN is trained to recognize a certain object (e.g., a dog in a photograph). By using a new photograph as an input to the trained CNN, a dog in the new photograph is also identified/labeled using a process known as inferencing. This inferencing occurs in real time, and recognizes specific objects (e.g., a dog) by running the new photograph through the trained CNN.--, in [0070]-[0071], and, -- The object identifier logic 713 is logic (e.g., the CNN logic 717 or an abbreviated version of the CNN logic 717 described herein) used to identify an object within a photograph.--, in [0077]). Re Claim 7, Kumar as modified by ROZNER further disclose classifying the one or more objects in the first image according to one or more classifications (see Kumar: e.g., --generating one or more views materializing one or more tensors generated by a convolutional neural network operating on a frame of a video; determining a first output of the convolutional neural network operating on a first subsequent frame of the video to identify an object in the video, the output being determined by at least generating a first query on the one or more views, the first query performing, based at least on a first change in the first subsequent frame, a first partial re-computation of the one or more views; and determining, based at least on the output of the convolutional neural network, a label identifying the object in the video.--, in [0030]; and, -- convolutional neural networks (CNNs), a type of deep machine learning model, may be especially adept at performing analytical tasks on images, videos, and time series data....convolutional neural network inference may consume significant time and computational resource. For instance, inferences performed using a VGG16 (OxfordNet) convolutional neural network requires 15 gigaflops of computations. Moreover, many analytical tasks involving convolutional neural networks require repeated convolutional neural network inferences, amplifying the computational cost and latency. For this reason, the adoption of convolutional neural networks may be especially unwieldy for interactive and/or resource-constrained settings such as mobile, browser, and edge devices. The adoption of convolutional neural networks may also be excessively costly for regular server and cloud settings. [0077] In some example embodiments, a machine learning engine may apply one or more query optimization techniques to reduce the computational cost associated with convolutional neural network inferences. For example, the machine learning engine may perform convolutional neural network inferences as queries that reuse, for each convolutional neural network inference, as much of the results of previous computations as possible. As such, query-based convolutional neural network inferences may be performed for tasks that require repeated convolutional neural network inferences on slightly modified inputs. For instance, query-based convolutional neural network inferences may be performed for the analytical task of occlusion-based explanation (OBE) in which an explanation of a convolutional neural network, such as a heatmap, is generated to identify regions of an input (e.g., an image) responsible for an output (e.g., a label classifying the image) of the convolutional neural network. Alternatively and/or additionally, query-based convolutional neural network inferences may be performed for the analytical task of object recognition in video (ORV) in which a convolutional neural network operates to identify an object over multiple frames of a video. [0078] To perform the task of occlusion-based explanation (OBE), the machine learning engine may perform convolutional neural network inferences on an image with a patch occluding different portions (e.g., pixels) of the image. For example, the machine learning engine may perform a first convolutional neural network inference on the image with a patch occluding a first portion of the image and a second convolutional neural network inference on the image with the patch occluding a second portion of the image. The machine learning engine may further track changes in the output of the convolutional neural network when different portions of the image are occluded with the patch. For instance, occluding the first portion of the image may have more effect on the output of the convolutional neural network than occluding the second portion of the image. Accordingly, the machine learning engine may generate a heatmap having different representations (e.g., colors, symbols, and/or the like) for the first portion of the image and the second portion of the image in order to indicate that the first portion of the image is more (or less) responsible for the output of the convolutional neural network than the second portion of the image. Because the task of occlusion based explanation (OBE) requires multiple convolutional neural network inferences on slightly modified versions of the image, the machine learning model may perform query-based convolutional neural network inferences to minimize the overhead associated with the repeated inferences. [0079] The machine learning engine may, as noted, also perform convolutional neural network inferences to accomplish the task of object recognition in videos (ORV). Machine learning enabled object recognition in videos may be especially popular due to the mass deployment of video cameras in applications such as security surveillance, traffic monitoring, wild animal and livestock tracking, and/or the like. In object recognition in videos, each frame of a video may be treated as an individual image. A convolutional neural network may be trained to process each frame of the video to identify an object appearing in the video. For example, the trained convolutional neural network may perform single object recognition over a fixed-angle camera video feed. As such, the convolutional neural network may perform multiple inferences over largely identical video frames, which is why the machine learning engine may apply query-based convolutional neural network inferences to the task of object recognition in videos (ORV) to minimize the computational overhead associated with the task.--, in [0076]-[0080], and [0150]-[0152]); and applying at least one classification of the one or more classifications to the one or more objects in the second image as a result of the comparison indicating that the size of the one or more objects in the second image is less than the size of the one or more objects in the first image (see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068]). Re Claims 8-10, claims 8-10 are the corresponding system claim to claims 1, 6-7 respectively. Therefore, claims 8-10 are rejected for the similar reasons for claims 1, 6-7 respectively. Furthermore, Kumar as modified by ROZNER further disclose a system, comprising: one or more processors comprising one or more circuits to perform the method (see Kumar: e.g., -- Systems, methods, and articles of manufacture, including computer program products, are provided for query optimized occlusion-based explanations (OBE). In some example embodiments, there is provided a system that includes at least one processor and at least one memory. The at least one memory may include program code that provides operations when executed by the at least one processor.--, in [0005], and [0024]-[0025]). Re Claim 11, Kumar as modified by ROZNER further disclose further obtain data from inferencing the one or more objects and use the obtained data to perform the comparison (see Kumar: e.g., --generating one or more views materializing one or more tensors generated by a convolutional neural network operating on a frame of a video; determining a first output of the convolutional neural network operating on a first subsequent frame of the video to identify an object in the video, the output being determined by at least generating a first query on the one or more views, the first query performing, based at least on a first change in the first subsequent frame, a first partial re-computation of the one or more views; and determining, based at least on the output of the convolutional neural network, a label identifying the object in the video.--, in [0030]; and, -- convolutional neural networks (CNNs), a type of deep machine learning model, may be especially adept at performing analytical tasks on images, videos, and time series data....convolutional neural network inference may consume significant time and computational resource. For instance, inferences performed using a VGG16 (OxfordNet) convolutional neural network requires 15 gigaflops of computations. Moreover, many analytical tasks involving convolutional neural networks require repeated convolutional neural network inferences, amplifying the computational cost and latency. For this reason, the adoption of convolutional neural networks may be especially unwieldy for interactive and/or resource-constrained settings such as mobile, browser, and edge devices. The adoption of convolutional neural networks may also be excessively costly for regular server and cloud settings. [0077] In some example embodiments, a machine learning engine may apply one or more query optimization techniques to reduce the computational cost associated with convolutional neural network inferences. For example, the machine learning engine may perform convolutional neural network inferences as queries that reuse, for each convolutional neural network inference, as much of the results of previous computations as possible. As such, query-based convolutional neural network inferences may be performed for tasks that require repeated convolutional neural network inferences on slightly modified inputs. For instance, query-based convolutional neural network inferences may be performed for the analytical task of occlusion-based explanation (OBE) in which an explanation of a convolutional neural network, such as a heatmap, is generated to identify regions of an input (e.g., an image) responsible for an output (e.g., a label classifying the image) of the convolutional neural network. Alternatively and/or additionally, query-based convolutional neural network inferences may be performed for the analytical task of object recognition in video (ORV) in which a convolutional neural network operates to identify an object over multiple frames of a video. [0078] To perform the task of occlusion-based explanation (OBE), the machine learning engine may perform convolutional neural network inferences on an image with a patch occluding different portions (e.g., pixels) of the image. For example, the machine learning engine may perform a first convolutional neural network inference on the image with a patch occluding a first portion of the image and a second convolutional neural network inference on the image with the patch occluding a second portion of the image. The machine learning engine may further track changes in the output of the convolutional neural network when different portions of the image are occluded with the patch. For instance, occluding the first portion of the image may have more effect on the output of the convolutional neural network than occluding the second portion of the image. Accordingly, the machine learning engine may generate a heatmap having different representations (e.g., colors, symbols, and/or the like) for the first portion of the image and the second portion of the image in order to indicate that the first portion of the image is more (or less) responsible for the output of the convolutional neural network than the second portion of the image. Because the task of occlusion based explanation (OBE) requires multiple convolutional neural network inferences on slightly modified versions of the image, the machine learning model may perform query-based convolutional neural network inferences to minimize the overhead associated with the repeated inferences. [0079] The machine learning engine may, as noted, also perform convolutional neural network inferences to accomplish the task of object recognition in videos (ORV). Machine learning enabled object recognition in videos may be especially popular due to the mass deployment of video cameras in applications such as security surveillance, traffic monitoring, wild animal and livestock tracking, and/or the like. In object recognition in videos, each frame of a video may be treated as an individual image. A convolutional neural network may be trained to process each frame of the video to identify an object appearing in the video. For example, the trained convolutional neural network may perform single object recognition over a fixed-angle camera video feed. As such, the convolutional neural network may perform multiple inferences over largely identical video frames, which is why the machine learning engine may apply query-based convolutional neural network inferences to the task of object recognition in videos (ORV) to minimize the computational overhead associated with the task.--, in [0076]-[0080], and [0150]-[0152]). Re Claim 12, Kumar as modified by ROZNER further disclose wherein the data obtained from inferencing the one or more objects comprises one or more of: a width, a height, or a position of the one or more objects (zee Kumar: e.g., --A convolutional neural network, such as the convolutional neural network 115, may be organized as layers of various types, each of which transforming one tensor (e.g., a multidimensional array that is typically 3-D) into another tensor. The convolution layer may use image filters from graphics to extract features, but with parametric filter weights (learned during training). The pooling layer may subsample features in a spatial-aware manner, the batch-normalization layer may normalize the output tensor, the non-linearity layer may apply an element-wise non-linear function (e.g., ReLU), and the fully-connected layer may include an ordered collection of perceptrons. [0094] The output tensor of a layer can have a different width, height, and/or depth than the input tensor of the layer. An image may be viewed as a tensor, e.g., a 224×224 RGB image is a 3-D tensor with a width and height of 224 and a depth of 3.--, in [0091]-[0094], and, --8] Local spatial context granularity layers may perform weighted aggregations of slices of the input tensor, called local spatial contexts, by multiplying them with a filter kernel (a tensor of weights). Thus, input tensors and output tensors can differ in width, height, and depth. If the input is incrementally updated, the region of the output that will be affected is not straightforward to ascertain but requires non-trivial and careful calculations due to the overlapping nature of how filters get applied to local spatial contexts. Both convolution layers and pooling layers fall under this category. As shown in FIG. 5, such layers typically account for the bulk of the computational cost of deep convolutional neural network inferences (over 90%). Thus, enabling incremental inference for such layers in the context of occlusion-based explanation may be key to increasing the computational efficiency and reducing the latency of the task. [0099] A convolution layer l of the convolutional neural network f may have C.sub.O:l 3-D filter kernels arranged as a 4-D array K.sub.conv:l, with each having a smaller spatial width W.sub.K:l and height H.sub.K:l than the width W.sub.I:l and height H.sub.I:l of the input tensor I.sub.:l but the same depth C.sub.I:l. During inference, c.sup.th filter kernel may be “strided” along the width and height dimensions of the input to produce a 2-D “activation map” A.sub.:c=(a.sub.y,x:c)∈custom-character.sup.H.sup.O:l.sup.×W.sup.O:l by computing element-wise products between the kernel and the local spatial context and adding a bias value as per Equation (7) below. The computational cost of each of these small matrix products may be proportional to the volume of the filter kernel.--, in [0098]-[0099]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]). Re Claim 13, Kumar as modified by ROZNER further disclose wherein a result of the comparison is evaluated against a set of conditions used to determine whether re-inferencing is to be performed (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086], {re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134], {so that “given threshold fraction τ (0<τ≤1)” read on “adjusting, a set of conditions” by setting the threshold fraction}; also see: -- this optimization also builds upon the incremental inference optimizations described earlier, but may be orthogonal to projective field thresholding and can therefore be used in unison. [0146] The notion of theoretical speedup for the adaptive drill-down optimization may be independent of the theoretical speedup associated with incremental inference. At the outset, setting the parameters r.sub.drill-down and S.sub.1 may be an application-specific balancing act. For example, if r.sub.drill-down is low, only a small region will require re-inferencing at the original resolution, which will save a lot of FLOPs. However, this may miss some regions of interest and thus compromise important explanation details. Similarly, a large S.sub.1 may also save a lot of FLOPs by reducing the number of re-inference queries in stage one but doing so may run the risk of misidentifying interesting regions, especially when the size of those regions are smaller than the occlusion patch.--, in [0145]-[0146]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]; also see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068]). Re Claim 14, Kumar as modified by ROZNER further disclose wherein the set of conditions comprises an object size threshold (see Kumar: e.g., --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086], {re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134], {so that “given threshold fraction τ (0<τ≤1)” read on “adjusting, a set of conditions” by setting the threshold fraction}; also see: -- this optimization also builds upon the incremental inference optimizations described earlier, but may be orthogonal to projective field thresholding and can therefore be used in unison. [0146] The notion of theoretical speedup for the adaptive drill-down optimization may be independent of the theoretical speedup associated with incremental inference. At the outset, setting the parameters r.sub.drill-down and S.sub.1 may be an application-specific balancing act. For example, if r.sub.drill-down is low, only a small region will require re-inferencing at the original resolution, which will save a lot of FLOPs. However, this may miss some regions of interest and thus compromise important explanation details. Similarly, a large S.sub.1 may also save a lot of FLOPs by reducing the number of re-inference queries in stage one but doing so may run the risk of misidentifying interesting regions, especially when the size of those regions are smaller than the occlusion patch.--, in [0145]-[0146]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]). Re Claim 15, Kumar as modified by ROZNER further disclose wherein the one or more processors comprising the one or more circuits are to adjust the set of conditions based, at least in part, on one or more of: graphics processing unit (GPU) utilization, one or more characteristics of the first and second images, a light intensity associated with the one or more objects, or an intensity of the first and second images (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086]; and, -- Algorithm 3 formalizes the object recognition in videos (ORV) workflow at the machine learning engine. For example, Algorithm 3 receives, as input, a video V, a threshold for frame differencing, a max_patch size for scene separation, a KryptonGraph kg for performing the incremental inference, and a batch size for batching multiple incremental inference requests. The base frame is initialized to the first frame of the video followed by an iteration through successive frames in video V calling the procedure FrameDifferencing to find the changed region before appending the result to a batch. Two possible events may occur to trigger an incremental inference on the compiled batch of changed regions. The first event being the changed region size exceeding the max_patch size and upon encountering a new scene, and the second being the current batch size reaching the max batch size. This max batch size may be necessary to avoid the possibility of exhausting hardware resource such as graphic processing unit (GPU) memory. Unlike in occlusion-based explanation where all patches are of a same size, changed regions in object recognition in videos are of arbitrary size. As such, when invoking incremental inference on a batch of changed regions, the machine learning engine 110 may first find the maximum size of the patches as the final patch size.--, in [0154]). Re Claim 16, claim 16 is the corresponding processor claim to claim 1 respectively. Therefore, claim 16 is rejected for the similar reasons for claim 1 respectively. Furthermore, Kumar as modified by ROZNER further disclose one or more processors comprising processing circuits to perform the method (see Kumar: e.g., -- Systems, methods, and articles of manufacture, including computer program products, are provided for query optimized occlusion-based explanations (OBE). In some example embodiments, there is provided a system that includes at least one processor and at least one memory. The at least one memory may include program code that provides operations when executed by the at least one processor.--, in [0005], and [0024]-[0025]). Re Claim 17, Kumar as modified by ROZNER further disclose track the one or more objects using one or more bounding boxes; determine a change in bounding box information of the tracked one or more objects (see Kumar: e.g., -- The operations may include: generating one or more views materializing one or more tensors generated by a convolutional neural network operating on a frame of a video; determining a first output of the convolutional neural network operating on a first subsequent frame of the video to identify an object in the video, the output being determined by at least generating a first query on the one or more views, the first query performing, based at least on a first change in the first subsequent frame, a first partial re-computation of the one or more views; and determining, based at least on the output of the convolutional neural network, a label identifying the object in the video.--, in [0025], [0030]; [0078] To perform the task of occlusion-based explanation (OBE), the machine learning engine may perform convolutional neural network inferences on an image with a patch occluding different portions (e.g., pixels) of the image. For example, the machine learning engine may perform a first convolutional neural network inference on the image with a patch occluding a first portion of the image and a second convolutional neural network inference on the image with the patch occluding a second portion of the image. The machine learning engine may further track changes in the output of the convolutional neural network when different portions of the image are occluded with the patch. For instance, occluding the first portion of the image may have more effect on the output of the convolutional neural network than occluding the second portion of the image. Accordingly, the machine learning engine may generate a heatmap having different representations (e.g., colors, symbols, and/or the like) for the first portion of the image and the second portion of the image in order to indicate that the first portion of the image is more (or less) responsible for the output of the convolutional neural network than the second portion of the image. Because the task of occlusion based explanation (OBE) requires multiple convolutional neural network inferences on slightly modified versions of the image, the machine learning model may perform query-based convolutional neural network inferences to minimize the overhead associated with the repeated inferences. [0079] The machine learning engine may, as noted, also perform convolutional neural network inferences to accomplish the task of object recognition in videos (ORV). Machine learning enabled object recognition in videos may be especially popular due to the mass deployment of video cameras in applications such as security surveillance, traffic monitoring, wild animal and livestock tracking, and/or the like. In object recognition in videos, each frame of a video may be treated as an individual image. A convolutional neural network may be trained to process each frame of the video to identify an object appearing in the video. For example, the trained convolutional neural network may perform single object recognition over a fixed-angle camera video feed. As such, the convolutional neural network may perform multiple inferences over largely identical video frames, which is why the machine learning engine may apply query-based convolutional neural network inferences to the task of object recognition in videos (ORV) to minimize the computational overhead associated with the task.--, in [0076]-[0080], and [0150]-[0152]; also see ROZNER: e.g., -- FIG. 5, inference is the process of using a trained CNN to recognize certain objects from a photograph or other data. In the example in FIG. 5, pixels from photograph 501 are input into a trained CNN (e.g., CNN 500), resulting in the identification and/or labeling (for display on the photograph 501) a particular object, such as the dog. [0071] That is, a CNN is trained to recognize a certain object (e.g., a dog in a photograph). By using a new photograph as an input to the trained CNN, a dog in the new photograph is also identified/labeled using a process known as inferencing. This inferencing occurs in real time, and recognizes specific objects (e.g., a dog) by running the new photograph through the trained CNN.--, in [0070]-[0071], and, -- The object identifier logic 713 is logic (e.g., the CNN logic 717 or an abbreviated version of the CNN logic 717 described herein) used to identify an object within a photograph.--, in [0077]); and re-infer the tracked one or more objects based, at least in part, the change in the bounding box information (see Kumar: e.g., -- The operations may include: generating one or more views materializing one or more tensors generated by a convolutional neural network operating on a frame of a video; determining a first output of the convolutional neural network operating on a first subsequent frame of the video to identify an object in the video, the output being determined by at least generating a first query on the one or more views, the first query performing, based at least on a first change in the first subsequent frame, a first partial re-computation of the one or more views; and determining, based at least on the output of the convolutional neural network, a label identifying the object in the video.--, in [0025], [0030]; [0078] To perform the task of occlusion-based explanation (OBE), the machine learning engine may perform convolutional neural network inferences on an image with a patch occluding different portions (e.g., pixels) of the image. For example, the machine learning engine may perform a first convolutional neural network inference on the image with a patch occluding a first portion of the image and a second convolutional neural network inference on the image with the patch occluding a second portion of the image. The machine learning engine may further track changes in the output of the convolutional neural network when different portions of the image are occluded with the patch. For instance, occluding the first portion of the image may have more effect on the output of the convolutional neural network than occluding the second portion of the image. Accordingly, the machine learning engine may generate a heatmap having different representations (e.g., colors, symbols, and/or the like) for the first portion of the image and the second portion of the image in order to indicate that the first portion of the image is more (or less) responsible for the output of the convolutional neural network than the second portion of the image. Because the task of occlusion based explanation (OBE) requires multiple convolutional neural network inferences on slightly modified versions of the image, the machine learning model may perform query-based convolutional neural network inferences to minimize the overhead associated with the repeated inferences. [0079] The machine learning engine may, as noted, also perform convolutional neural network inferences to accomplish the task of object recognition in videos (ORV). Machine learning enabled object recognition in videos may be especially popular due to the mass deployment of video cameras in applications such as security surveillance, traffic monitoring, wild animal and livestock tracking, and/or the like. In object recognition in videos, each frame of a video may be treated as an individual image. A convolutional neural network may be trained to process each frame of the video to identify an object appearing in the video. For example, the trained convolutional neural network may perform single object recognition over a fixed-angle camera video feed. As such, the convolutional neural network may perform multiple inferences over largely identical video frames, which is why the machine learning engine may apply query-based convolutional neural network inferences to the task of object recognition in videos (ORV) to minimize the computational overhead associated with the task.--, in [0076]-[0080], and [0150]-[0152]; also see ROZNER: e.g., -- FIG. 5, inference is the process of using a trained CNN to recognize certain objects from a photograph or other data. In the example in FIG. 5, pixels from photograph 501 are input into a trained CNN (e.g., CNN 500), resulting in the identification and/or labeling (for display on the photograph 501) a particular object, such as the dog. [0071] That is, a CNN is trained to recognize a certain object (e.g., a dog in a photograph). By using a new photograph as an input to the trained CNN, a dog in the new photograph is also identified/labeled using a process known as inferencing. This inferencing occurs in real time, and recognizes specific objects (e.g., a dog) by running the new photograph through the trained CNN.--, in [0070]-[0071], and, -- The object identifier logic 713 is logic (e.g., the CNN logic 717 or an abbreviated version of the CNN logic 717 described herein) used to identify an object within a photograph.--, in [0077]). Re Claim 18, Kumar as modified by ROZNER further disclose wherein a change in size of the one or more objects in the second image relative to the first image corresponds to the one or more objects moving closer to a camera used to capture the first and second images (see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068]). Re Claim 19, Kumar as modified by ROZNER further disclose wherein the processing circuitry is further to adjust a predetermined threshold, the predetermined threshold used to determine whether to re-infer the one or more objects, based, at least in part, on graphics processing unit (GPU) utilization (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086]; and, -- Algorithm 3 formalizes the object recognition in videos (ORV) workflow at the machine learning engine. For example, Algorithm 3 receives, as input, a video V, a threshold for frame differencing, a max_patch size for scene separation, a KryptonGraph kg for performing the incremental inference, and a batch size for batching multiple incremental inference requests. The base frame is initialized to the first frame of the video followed by an iteration through successive frames in video V calling the procedure FrameDifferencing to find the changed region before appending the result to a batch. Two possible events may occur to trigger an incremental inference on the compiled batch of changed regions. The first event being the changed region size exceeding the max_patch size and upon encountering a new scene, and the second being the current batch size reaching the max batch size. This max batch size may be necessary to avoid the possibility of exhausting hardware resource such as graphic processing unit (GPU) memory. Unlike in occlusion-based explanation where all patches are of a same size, changed regions in object recognition in videos are of arbitrary size. As such, when invoking incremental inference on a batch of changed regions, the machine learning engine 110 may first find the maximum size of the patches as the final patch size.--, in [0154]). Re Claim 20, Kumar as modified by ROZNER further disclose wherein the processing circuitry is further to use one or more neural networks to infer an object class of the one or more objects from the first image and to infer one or more refined classification parameters corresponding to the one or more objects from the second image (see Kumar: e.g., --generating, at a first stride size, a second heatmap; identifying, based at least on the second heatmap, one or more regions of the image exhibiting a largest contribution to the output of the convolutional neural network operating on the image, a quantity of the one or more regions being proportional to a threshold fraction of the image; and determining, at a second stride size, the first output, the second stride size being smaller than the first stride size such that the first heatmap generated based on the first output has a higher resolution than the second heatmap.--, in [0022] {herein thresholding read on “adjusting, a set of conditions”}, and, --comparing the base frame and the first subsequent frame to identify one or more portions of the first subsequent frame that differ from the base frame, the one or more identified portions including at least a threshold quantity of pixels; and [0034] In some variations, the operations may further include: in response to the changed region exceeding a threshold fraction of a size of the base frame, setting the first subsequent frame as the base frame.--, {above underlined disclosures align to “conditions for determining when an object is re-inferenced”, and “thresholding” align to “adjusting these conditions”} in [0033]-[0034]; and, --multiple re-inference requests may be batched to reuse the same materialized views, thus achieving multi-query optimization (MQO) in the form of batched incremental inferences. A graphic processing unit (GPU) optimized kernel may be used for executing these batched incremental inferences. [0086] Where some degradation in the visual quality of the heatmaps generated as a result of occlusion-based explanation (OBE), one or more approximate inference optimizations may be applied to further reduce the runtime associated with the task. These optimizations may build upon the incremental inference optimization to trade off heatmap quality in a user-tunable manner. The first approximate optimization, projective field thresholding, exploits the internal semantics of the convolutional neural network 115.--, in [0085]-[0086], {re-inference is determined based on approximate optimization, projective field thresholding}, and, -- projective field thresholding, which includes truncating the projective field from growing beyond a given threshold fraction τ (0<τ≤1) of the output size. This means that inferences in subsequent layers of the convolutional neural network are approximate. FIG. 12(b) illustrates the concept of projective field thresholding for a filter size of 3 and a stride of 1. One input element may be updated (shown in red/dark) and the change may propagate to 3 elements in the next layer and then to 5 elements in the following layer before being truncated because we set a threshold τ=5/7. This approximation may alter the accuracy of the output values and the visual quality of the resulting heatmap.--, in [0132]-[0134], {so that “given threshold fraction τ (0<τ≤1)” read on “adjusting, a set of conditions” by setting the threshold fraction}; also see: -- this optimization also builds upon the incremental inference optimizations described earlier, but may be orthogonal to projective field thresholding and can therefore be used in unison. [0146] The notion of theoretical speedup for the adaptive drill-down optimization may be independent of the theoretical speedup associated with incremental inference. At the outset, setting the parameters r.sub.drill-down and S.sub.1 may be an application-specific balancing act. For example, if r.sub.drill-down is low, only a small region will require re-inferencing at the original resolution, which will save a lot of FLOPs. However, this may miss some regions of interest and thus compromise important explanation details. Similarly, a large S.sub.1 may also save a lot of FLOPs by reducing the number of re-inference queries in stage one but doing so may run the risk of misidentifying interesting regions, especially when the size of those regions are smaller than the occlusion patch.--, in [0145]-[0146]; and, -- [0152] The machine learning engine 110 may apply an approximate frame differencing technique to identify the single most important changed region for incremental inference in each frame. This technique is formally present in Algorithm 2 and pictorially presented in FIG. 16. Approximate frame differencing may be performed based on inputs including a base_frame which is treated as the background, a new_frame from which to identify the changed region, and a threshold which will be used to identify the changed pixels. By using pixel-subtraction, the machine learning engine 110 may identify all of the changes between the current_frame and the base_frame on a per-pixel basis. Thresholding the resultant data may eliminate noise and restrict the necessary re-inferencing to a more limited scope. The machine learning engine 110 may also calculate bounding boxes for the remaining areas of difference, thereby providing a more regular shape for subsequent inference. These bounding boxes can often overlap, which is why they may be collapsed into larger bounding boxes to eliminate any overlaps. The largest of the resultant bounding boxes may be selected as the most important changed region for incremental inference, with the coordinates and the dimensions of this box being the output of approximate frame differencing. It should be appreciated that smaller threshold values tend to select smaller changed regions and result in higher speedups. However, a smaller threshold may also reduce the accuracy of the generated predictions. The most optimal value for threshold (value between 0 and 255) may be largely dependent on the chosen use case. Empirically, a threshold value of 40 was found to provide a reasonable trade-off between runtime and accuracy--, in [0152]; also see ROZNER: e.g., -- If the output from output layer 307 is a vector that is predetermined to describe a dog (e.g., (1,2,4,10)), then the weights (and alternatively the algorithms) are adjusted until the vector (1,2,4,10), or a vector that is mathematically similar, is output from output layer 307 when pixel data from a photograph of a dog is input into input layer 303. [0059] When automatically adjusted, the weights (and/or algorithms) are adjusted using “back propagation”, in which weight values of the neurons are adjusted by using a “gradient descent” method that determines which direction each weight value should be adjusted to. This gradient descent process moves the weight in each neuron in a certain direction until the output from output layer 307 improves (e.g., gets closer to (1,2,4,10)).--, in [0058]-[0059]; and, -- FIG. 5, the pooling stage and a classification stage (as well as the convolution stage) of a CNN 500 during inference processing is depicted. That is, once the CNN is optimized by adjusting weights and/or algorithms in the neurons (see FIG. 3), by adjusting the stride of movement of the pixel subset 406 (see FIG. 4), and/or by adjusting the filter 404 shown in FIG. 4, then it is trusted to be able to recognize similar objects in similar photographs. This optimized CNN is then used to infer (hence the name inference processing) that the object in a new photograph is the same object that the CNN has been trained to recognize.--, in [0065]-[0068])). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Sajjadi (US 20200293796 A1) disclose object detection and classification including one or more of the pixels within the reduced-size intersection coverage map(s) 126 may have bounding box location and/or dimensionality information encoded thereto. In another example, each pixel or point within the bounding box may have location and/or dimensionality information encoded thereto, but the reduced-size intersection coverage map(s) 126 may be used to determine which pixels or points to leverage (e.g., the pixels or points inside of the reduced-size intersection coverage map(s) 126) when generating the final bounding box proposals {see in Fig. 3, [0045], and, -- [0050] Referring again to FIG. 1, the distance(s) 130 for each intersection in an image representative of sensor data 102 may then be used to determine which intersection is located in which range.--, in [0048]-[0050]). Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIWEN YANG whose telephone number is (571)270-5670. The examiner can normally be reached on Monday-Friday 8:30am-4:30pm east. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WEI WEN YANG/ Primary Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Nov 08, 2024
Application Filed
Aug 24, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743775
BIOMARKERS OF COLLAGEN FIBER ARCHITECTURE IN EPITHELIAL OVERIAN CANCER (EOC) PATIENTS
2y 11m to grant Granted Sep 22, 2026
Patent 12737879
Machine Learning for Detection of Diseases from External Anterior Eye Images
3y 9m to grant Granted Sep 15, 2026
Patent 12737880
SYSTEMS AND METHODS OF ANALYZING MICROBIOMES USING ARTIFICIAL INTELLIGENCE
3y 4m to grant Granted Sep 15, 2026
Patent 12738057
CUT-PASTE TRAINING AUGMENTATION FOR MACHINE LEARNING MODELS
2y 7m to grant Granted Sep 15, 2026
Patent 12737851
ENHANCED QUALITY BOREHOLE IMAGE GENERATION AND METHOD
2y 3m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
93%
With Interview (+11.5%)
2y 5m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 684 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month