Prosecution Insights
Last updated: August 18, 2026
Application No. 18/915,958

HEAT MAP-BASED FACIAL FEATURE KEY POINT OCCLUSION DETECTION FOR OCCUPANT MONITORING SYSTEMS AND APPLICATIONS

Non-Final OA §101§103
Filed
Oct 15, 2024
Examiner
DARDANO, STEFANO ANTHONY
Art Unit
2663
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
69 granted / 88 resolved
+16.4% vs TC avg
Strong +33% interview lift
Without
With
+32.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
14 currently pending
Career history
101
Total Applications
across all art units

Statute-Specific Performance

§101
10.6%
-29.4% vs TC avg
§103
55.3%
+15.3% vs TC avg
§102
18.8%
-21.2% vs TC avg
§112
13.7%
-26.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 88 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Status Claims 1-20 are pending. Information Disclosure Statement The IDS filed 10/15/25 and 01/30/26 have been considered. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: "One or more processors comprising processing circuitry to" in claim 1 The claims do not use “circuitry” in their ordinary meaning (by contrast, see the Abacus v MIT decision cited in MPEP 2181 where circuitry ordinarily means custom hardware); rather, they are understood as a reference to the structures identified in the specification. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 35 U.S.C. 101 requires that a claimed invention must fall within one of the four eligible categories of invention (i.e. process, machine, manufacture, or composition of matter) and must not be directed to subject matter encompassing a judicially recognized exception as interpreted by the courts. MPEP 2106. Three categories of subject matter are found to be judicially recognized exceptions to 35 U.S.C. § 101 (i.e. patent ineligible) (1) laws of nature, (2) physical phenomena, and (3) abstract ideas. MPEP 2106(II). To be patent-eligible, a claim directed to a judicial exception must as whole be integrated into a practical application or directed to significantly more than the exception itself (MPEP 2106). Hence, the claim must describe a process or product that applies the exception in a meaningful way, such that it is more than a drafting effort designed to monopolize the exception. Claims 12-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., an abstract idea) without integration into a practical application or recitation of significantly more. In the analysis below, the system of claim 12 is considered representative of both claims 12 and 20 since all of the independent claims recite very similar steps despite being directed to different statutory matter. Furthermore, each of independent claims 12 and 20 are directed to one of the four statutory categories of eligible subject matter; thus, the claims pass Step 1 of the Subject Matter Eligibility Test (See flowchart in MPEP 2106). Step 2A, Prong 1 Analysis The independent claims are directed to determining a location of individual feature key points of one or more feature key points based at least on a peak value of a heat map representing key point location confidence values; compute an occlusion score for at least one individual feature key point of the one or more feature key points based at least on a statistical distribution of the key point location confidence values on the heat map. An individual can mentally determine determining a location of individual feature key points of one or more feature key points based at least on a peak value of a heat map representing key point location confidence values. Additionally the step of computing an occlusion score is mathematical process. The steps of computing, estimating, and searching are stated as mathematical concepts. “A claim that recites a mathematical calculation, when the claim is given its broadest reasonable interpretation in light of the specification, will be considered as falling within the "mathematical concepts" grouping” (MPEP 2106.04(a)(2) I.C.). The above step is considered to fall under such a grouping. As such, the description in independent claim 12 is an abstract idea – namely, a menta process and mathematical formula. Accordingly, the analysis under prong one of step 2A of the Subject Matter Eligibility Test does not result in a conclusion of eligibility (See flowchart in MPEP 2106). Additional elements The independent claims further recites the additional element of a system comprising one or more processors and outputting key point prediction data based at least on the location of individual feature key points and the occlusion score for the at least one individual feature key point. Step 2A, prong 2 analysis The above-identified additional elements do not integrate the judicial exception into a practical application. Each of the additional elements (system and processors) amounts to merely using a computer as a tool to perform the claimed mental process. Implementing an abstract idea on a computer does not integrate a judicial exception into a practical application (See MPEP 2106.05(f)). The additional elements of outputting key point prediction data based at least on the location of individual feature key points and the occlusion score for the at least one individual feature key point amount to insignificant post-solution activity which does not integrate the abstract idea into a practical application (MPEP 2106.05(g)). Moreover, the additional elements of the claims do not recite an improvement in the functioning of a computer or other technology or technical field, the claimed steps are not performed using a particular machine, the claimed steps do not effect a transformation, and the claims do not apply the judicial exception in any meaningful way beyond generically linking the use of the judicial exception to a particular technological environment (See MPEP 2106.04(d)). Therefore, the analysis under prong two of step 2A of the Subject Matter Eligibility Test does not result in a conclusion of eligibility (See flowchart in MPEP 2106). Step 2B Finally, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Each of the additional elements (system and processors) are generic computer features which perform generic computer functions that are well-understood, routine, and conventional and do not amount to more than implementing the abstract idea with a computerized system. Thus, taken alone, the additional elements do not amount to significantly more than the above-identified judicial exception (the abstract idea). The additional elements of outputting key point prediction data based at least on the location of individual feature key points and the occlusion score for the at least one individual feature key point amount to insignificant post-solution activity which does amount to significantly more (MPEP 2106.05(g)). Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements taken individually. There is no indication that the combination of elements improves the functioning of a computer or improves any other technology. Their collective functions merely provide conventional computer implementation, and mere implementation on a generic computer does not add significantly more to the claims. Accordingly, the analysis under step 2B of the Subject Matter Eligibility Test does not result in a conclusion of eligibility (See flowchart in MPEP 2106). For all of the foregoing reasons, independent claims 1, 10, and 18 do not recite eligible subject matter under 35 USC 101. Examiner recommends amending in the control of the vehicle similar to claim 1 to overcome the 101 rejection and make these claim eligible. Dependent claims 13-19 are dependent on independent claim 12 and therefore include all of the limitations of claim 12. Therefore, claims 13-19 recite the same abstract idea of a mental process and mathematical calculation. Claim 13 recites infer an occlusion classification for the at least one individual feature key point based at least on one or more facial images and the map, wherein the occlusion classification comprises one of: no occlusion, a self-occlusion, an object occlusion, or a truncated occlusion. An individual can mentally infer an occlusion classification, where the occlusion classification comprises one of the listed items. Thus, the feature of claim 13 is directed to the mental process. Accordingly, the claim does not recite any additional limitations that integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claim 14 recites correlate the location of individual feature key points with respect to the map with one or more key point locations with respect to one or more facial images of a subject to generate the key point prediction data. This is related to the post solution activity in claim 12, and does not recite any additional limitations that integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claim 15 recites wherein the heat map representing key point location confidence values is generated using a machine learning model trained based at least on a two-dimensional feature key point data set and a three-dimensional feature key point data set comprising one or more two-dimensional projections of self-occluded feature key points. Further defining the data that is mentally analyzed does not integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Accordingly, the claim does not recite any additional limitations that integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claim 16 recites further comprising a machine learning model trained to generate the heat map based at least on a key point location misalignment loss and an occlusion loss. Further defining the data that is mentally analyzed does not integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Accordingly, the claim does not recite any additional limitations that integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claim 17 recites further apply one or more labels to auto-annotate one or more facial images of a subject based on one or more inferences from the one or more facial images that indicate that the at least one individual facial feature key point is a self-occluded key point with an object occlusion. An individual can mentally (with the assistance of a physical support such as a pen), apply labels to auto annotate keypoints that have been inferred where the keypoints are self-occluded. Thus, the feature of claim 17 is directed to the mental process. Accordingly, the claim does not recite any additional limitations that integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claim 18 recites generate the map as a heat map wherein pixels of the heat map represent one or more data channels, wherein an individual data channel of the one or more data channels represents a respective key point location confidence value for a facial feature key point. Further defining the data that is mentally analyzed does not integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Accordingly, the claim does not recite any additional limitations that integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. Claim 19 recites wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for three-dimensional assets; a system for performing deep learning operations; a system for performing remote operations; a system for performing real-time streaming; a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system implementing one or more language models; a system implementing one or more large language models (LLMs); a system implementing one or more vision language models (VLMs); a system for generating synthetic data; a system for generating synthetic data using AI; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Further defining examples of the hardware of the system does not integrate the judicial exception into a practical application or amount to significantly more than the judicial exception without defining the use in the system. Accordingly, the claim does not recite any additional limitations that integrate the judicial exception into a practical application or amount to significantly more than the judicial exception. By contrast, claim 1 includes additional elements “control one or more operations of a vehicle based at least on the key point prediction” which integrates the abstract idea into a practical application of controlling a vehicle based on processed data. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7, 10-14, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yuen et al. (“An Occluded Stacked Hourglass Approach to Facial Landmark Localization and Occlusion Estimation” Hereinafter “Yuen”) in view of MORIMOTO et al. (US 20210124963 A1 Hereinafter “MORIMOTO”). Regarding claim 1, Yuen teaches one or more processors comprising processing circuitry (Fig. 3 shows the proposed network structure, processors would be required for the structure to operate) to: generate a map of key point location confidence values for one or more feature key points based at least on one or more images of a subject (Page 321, section I: “In an effort to create this system, the Stacked Hourglass, an existing deep neural network implementation designed for body pose joint estimation, is modified to take in a detected face window and output 68 occlusion heatmap images of facial landmarks which can be used to estimate landmark location and occlusion”. The heatmaps output for the keypoints acts as maps that contain confidence values of a keypoint being present “The heatmap image training labels are generally designed such that a small Gaussian blob with amplitude 1 is placed at the location of the landmark relative to the input image with the rest of the heatmap being approximately zero. This results in heatmap image with values roughly in the range of 0 to 1, where high valued areas indicate a high confidence of the landmark being in that location (Page 323, section III)); determine, with respect to the map, a location of individual feature key points of the one or more feature key points based at least on a peak value of the key point location confidence values (Page 325-326, initial landmark estimation: “The location of the landmark is then calculated as the weighted centroid of the group of pixels with the highest value in the magnitude score image (i.e. MATLAB’s regionprops function). This is repeated for the remaining land marks. The locations are then carefully transformed back from the 64 × 64 score image space to the resized 256 × 256 input face image space, and finally back to the original image coordinate space. The detection box is then re-localized based on the minimum and maximum of the x-y coordinates of all landmarks and extending the top of this box by 20% as done during training sample generation of this chapter”. Key points are initially estimated by the maps with the highest magnitude score (which acts as the peak)); compute an occlusion score for at least one individual feature key point of the one or more feature key points based at least on a statistical distribution of the key point location confidence values (Page 326, Refining Landmark Location and Extracting Occlusion Score: “To extract the refined landmark locations, the same step as “Initial Landmark Location” in this subsection are applied using the new 68 landmark score images. The value of the pixel in the corresponding new score image at the refined landmark location is used as the landmark occlusion score”. This landmark occlusion score is determined based on the statistical distribution of the key point location confidence values, “The network is trained by following the original work of the Stacked Hourglass paper with a minor modification of including occlusion information in the heatmap ground truth labels. More specifically, if a landmark is labeled as occluded, then the Gaussian blob placed at the location of the landmark would have an amplitude of −1 instead of +1. The reasoning behind this is so that the location of the landmark can still be derived from looking at the magnitude of the heatmap, while the sign gives an indication of whether it’s occluded or not” (Page 324, Col. 1). The statistical distribution of the key point location values (heat map values) contain signs, these signs dictate whether or not the keypoint is occluded, generating an occlusion score); generate a key point prediction based at least on the location of individual feature key points and the occlusion score for the at least one individual feature key point (Page 324, Fig. 3, occluded stacked hourglass: “The initial bounding box proposal from the face detector is refined by generating a tighter-bounding box around the face using the location of estimated landmarks to provide better localization of the face, while the score is refined using occlusion confidence information”. The initial landmarks are used alongside the occlusion to score to generate an overall prediction of the landmark points “The assumption is that at least one landmark is visible and is detected well by the network, and so the best scoring landmark is given a magnitude score of 1, and the rest of the landmarks are scored relatively to it” (Page 326, col 2)); and While it is presumed Yuen would perform controlling a vehicle based off the result of the landmark detection since they’re creating a detection system for vehicle safety systems (Page 321, section 1: “One way to minimize these events is by introducing an active safety system to assist the driver by monitoring their face in order to give alerts when the driver is not paying attention to a particular surrounding area. By tracking landmarks through the face along with eye information, head pose can be computed to provide the driver’s facing direction. Indicators from eyes and mouth can provide information on eye closure duration, blink rates, or yawning to provide information on the driver’s drowsiness or fatigue [17]. Often when people are drowsy, they may cover their mouths when yawning or rub their eyes which both cause occlusion, which must be detected in order to avoid computing inaccurate measurements on driver’s drowsiness or even the head pose. This allows for the system to warn the driver ahead of time of any unseen danger that they may not have seen or noticed and minimize the probability of an accident. By creating such a system also paves way to other systems which coordinates heads, eyes, and hands [18], systems which looks at humans inside and outside vehicle cabins [19], or distraction [20]. In an effort to create this system, the Stacked Hourglass, an existing deep neural network implementation designed for body pose joint estimation, is modified to take in a detected face window and output 68 occlusion heatmap images of facial landmarks which can be used to estimate landmark location and occlusion”), Yuen does not expressly disclose controlling a vehicle based on the detection of landmarks. However, MORIMOTO teaches controlling a vehicle based on the detection of landmarks ([0078]: “The vehicle control apparatus determines whether the driver is in a drowsiness state or an inoperable state based on the driver's state monitoring information and the driver's face information detected in the captured image by the camera. If it is determined that the driver is in a drowsiness state or an inoperable state, the vehicle control such as deceleration and stop is executed”. The monitoring information of the drivers face comes from detected keypoints “In this embodiment, a human face is recognized based on the relative positions of facial feature points such as eyes and nose specified in the captured image” [0022]). At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify Yuen’s vehicle safety system to include MORIMOTO’s controlling of a vehicle based on drowsiness detection because such a modification is the result of applying a known technique to a known device ready for improvement to yield predictable results. More specifically, MORIMOTO’s controlling of a vehicle based on drowsiness detection permits performing vehicle operations based on detected drowsiness of a driver. This known benefit in MORIMOTO is applicable to Yuen’s vehicle safety system as they both share characteristics and capabilities, namely, they are directed to safety systems for drivers. Yuen would find it beneficial to have a system for operating the vehicle based on the drowsiness detection since the purpose of Yuen’s system is to provide safety for the drive. Therefore, it would have been recognized that modifying Yuen’s vehicle safety system to include MORIMOTO’s controlling of a vehicle based on drowsiness detection would have yielded predictable results because (i) the level of ordinary skill in the art demonstrated by the references applied shows the ability to incorporate MORIMOTO’s controlling of a vehicle based on drowsiness detection in safety systems for drivers and (ii) the benefits of such a combination would have been recognized by those of ordinary skill in the art. Regarding claim 2, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 1, in addition, Yuen further teaches wherein the one or more processors are further to correlate the location of individual feature key points with respect to the map with one or more key point locations with respect to the one or more facial images of the subject to generate the key point prediction (Page 324, Fig. 3, occluded stacked hourglass: “The initial bounding box proposal from the face detector is refined by generating a tighter-bounding box around the face using the location of estimated landmarks to provide better localization of the face, while the score is refined using occlusion confidence information”. The heatmaps are used to obtain the initial keypoints and refined keypoints, those predicted keypoints are correlated to the face of the driver to generate the final predicted keypoints (As seen in the left of Fig. 1)). Regarding claim 3, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 1, in addition, Yuen further teaches wherein the one or more processors are further to generate the map on a per keypoint basis as a heat map wherein pixels represent one or more data channels (Page 321, section I: “In an effort to create this system, the Stacked Hourglass, an existing deep neural network implementation designed for body pose joint estimation, is modified to take in a detected face window and output 68 occlusion heatmap images of facial landmarks which can be used to estimate landmark location and occlusion”. 68 heatmaps are generated for the 68 keypoints, one for each keypoint. These pixels also represent at least one data channel since they are from RGB images), wherein pixel values in an individual data channel of the one or more data channels represents a respective key point location confidence value for a feature key point (Page 323, section III: ““The heatmap image training labels are generally designed such that a small Gaussian blob with amplitude 1 is placed at the location of the landmark relative to the input image with the rest of the heatmap being approximately zero. This results in heatmap image with values roughly in the range of 0 to 1, where high valued areas indicate a high confidence of the landmark being in that location”. The heatmaps output for the keypoints acts as maps that contain confidence values of a keypoint being present). Regarding claim 4, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 1, in addition, Yuen further teaches wherein the one or more processors are further to generate the map based on applying the one or more images to a feature key point estimation model (Page 323-324, Fig. 3, Occluded Stacked Hourglass: “To given an idea of the input-output functionality of the Stacked Hourglass, essentially it is designed to take an input image of fixed dimension and outputs multiple heatmap images, which are specified during training. The heatmap image training labels are generally designed such that a small Gaussian blob with amplitude 1 is placed at the location of the landmark relative to the input image with the rest of the heatmap being approximately zero. This results in heatmap image with values roughly in the range of 0 to 1, where high valued areas indicate a high confidence of the landmark being in that location” This stacked hourglass acts as the feature key point estimation model.) Regarding claim 5, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 4, in addition, Yuen further teaches wherein the feature key point estimation model is further to infer an occlusion classification based on at least one image of the one or more images and the map, wherein the occlusion classification indicates a type occlusion that is obscuring a feature key point of the one or more feature key points (Page 324, Fig. 3, Occluded Stacked Hourglass: “In this paper, the Occluded Stacked Hourglass is presented, in which the training label heat map of the Stacked Hourglass is modified to also estimate occlusion for each facial landmark so that the system can tell which parts of the face is occluded, allowing future face analysis systems to be robust to any occlusions that may occur”. The landmark being occluded is a type of occlusion, so the classification of occlusion is a type of occlusion) Regarding claim 6, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 5, in addition, Yuen further teaches wherein the occlusion classification comprises one of: no occlusion, ((Page 324, Col. 1: “The network is trained by following the original work of the Stacked Hourglass paper with a minor modification of including occlusion information in the heatmap ground truth labels. More specifically, if a landmark is labeled as occluded, then the Gaussian blob placed at the location of the landmark would have an amplitude of −1 instead of +1. The reasoning behind this is so that the location of the landmark can still be derived from looking at the magnitude of the heatmap, while the sign gives an indication of whether it’s occluded or not”. One of the output classifications from the heatmap is non-occluded, which means no occlusion. The “or” language in the claim means the list is alternatives meaning only one condition be met for a prima facia rejection). Regarding claim 7, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 4, in addition, Yuen further teaches wherein the feature key point estimation model is trained using image training data wherein samples of the image training data comprise first annotations based on locations of the one or more feature key points (Page 324, Training Data Generation: “For landmark samples, a combined dataset of about 3,400 faces with 68 landmark annotations which follows the CMU Multi-PIE dataset format”. These sets of training data are used to train the stacked hourglass model) and second annotations based on an indication of occlusion of the one or more feature key points (Page 325, Training Data Generation: “Due to resource constraints, all landmarks (regardless of whether they were occluded in the original image) are marked as non-occluded and only artificially occluded landmarks, using the method described here, are marked as occluded”. These sets of training data are used to train the stacked hourglass model). Regarding claim 10, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 1, in addition, Yuen further teaches one or more processors of claim 1, wherein the one or more images of the subject are captured by one or more image sensors comprising a red, green, blue (RGB) sensor, (Page 324, Fig. 3: “The system takes N detected RGB face images which are preprocessed with CLA Histogram Equalization individually on each channel”. RGB face images require a RGB sensor to capture them). Regarding claim 11, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 1, in addition, Yuen further teaches one or more processors of claim 1, wherein the processing circuitry is comprised in at least one of: a system for performing deep learning operations (Page 324, Fig. 3: The stacked hourglass structure is a system for performing deep learning operations. The “or” language in the claim means the list is alternatives meaning only one condition be met for a prima facia rejection, the crossed out regions do not mean Yuen does not teach the listed items for this claim, they simply mean they are not being mapped to meet the case of rejection); Regarding claim 12, the content of claim 12 is similar to the content of claim 1, therefore it is rejected for the same reasons of obviousness as claim 1. Regarding claim 13, the content of claim 13 is similar to the content of claim 6, therefore it is rejected for the same reasons of obviousness as claim 6. Regarding claim 14, the content of claim 14 is similar to the content of claim 2, therefore it is rejected for the same reasons of obviousness as claim 2. Regarding claim 17, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 12, in addition, Yuen further teaches wherein the one or more processors are further to apply one or more labels to auto-annotate one or more facial images of a subject based on one or more inferences from the one or more facial images that indicate that the at least one individual facial feature key point is a self-occluded key point with an object occlusion (Page 324, Fig. 3, occluded stacked hourglass: “The initial bounding box proposal from the face detector is refined by generating a tighter-bounding box around the face using the location of estimated landmarks to provide better localization of the face, while the score is refined using occlusion confidence information”. The initial landmarks are used alongside the occlusion to score to generate an overall prediction of the landmark points “The assumption is that at least one landmark is visible and is detected well by the network, and so the best scoring landmark is given a magnitude score of 1, and the rest of the landmarks are scored relatively to it” (Page 326, col 2). These points are labeled as occluded or non-occluded on the facial images (using colors) with the score being -1 or +1, and these labels are based on the inferences from the network from the one or more facial images that are input, and the facial image indicate at least one facial keypoint that is a self-occlusion key point with an object detection (drivers own hand occluding their face) “The driver’s mouth are shown closed at (a), then slowly opens from (b) to (e). (f) and (g) shows red occluded landmarks caused by a hand covering the mouth” (Page 329, Fig. 9)). Regarding claim 18, the content of claim 18 is similar to the content of claim 3, therefore it is rejected for the same reasons of obviousness as claim 3. Regarding claim 19, the content of claim 19 is similar to the content of claim 11, therefore it is rejected for the same reasons of obviousness as claim 11. Regarding claim 20, the content of claim 20 is similar to the content of claim 1, therefore it is rejected for the same reasons of obviousness as claim 1. Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Yuen et al. (“An Occluded Stacked Hourglass Approach to Facial Landmark Localization and Occlusion Estimation” Hereinafter “Yuen”) in view of MORIMOTO et al. (US 20210124963 A1 hereinafter “MORIMOTO”) as evidenced by Newell et al. (“Stacked Hourglass Networks for Human Pose Estimation” Hereinafter “Newell”). Regarding claim 8, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 4, in addition, Yuen further teaches one or more processors of claim 4, wherein the feature key point estimation model is trained by optimizing the feature key point estimation model based at least on a key point location misalignment loss and a key point occlusion loss (Page 324, Col. 1, Yuen: “The network is trained by following the original work of the Stacked Hourglass paper with a minor modification of including occlusion information in the heatmap ground truth labels. More specifically, if a landmark is labeled as occluded, then the Gaussian blob placed at the location of the landmark would have an amplitude of −1 instead of +1. The reasoning behind this is so that the location of the landmark can still be derived from looking at the magnitude of the heatmap, while the sign gives an indication of whether it’s occluded or not. During training, techniques from our previous paper is used to apply augmentations to include artificial occlusion data”. Yuen explains that they train their network using Newell’s previous paper “The stacked 2-hourglass network is trained using published code provided by the Newell, Yang, & Deng, 2016, original authors of the paper” (Page 325, section C, Yuen). Newell explains how their model is trained with respect to ground truth images “The same technique as Tompson et al. [15] is used for supervision. A Mean Squared Error (MSE) loss is applied comparing the predicted heatmap to a ground-truth heatmap consisting of a 2D gaussian (with standard deviation of 1 px) centered on the joint location” (Page 490, last paragraph, Newell). Loss is obtained by comparing the ground truth to the predicted heat map. Since Yuen as mapped above has ground truth maps for both the landmark location and the occlusion information for the landmarks, the loss would be associated with the misalignment of keypoints and occlusion). Regarding claim 16, the content of claim 16 is similar to the content of claim 8, therefore it is rejected for the same reasons of obviousness as claim 8. Claims 9 is rejected under 35 U.S.C. 103 as being unpatentable over Yuen et al. (“An Occluded Stacked Hourglass Approach to Facial Landmark Localization and Occlusion Estimation” Hereinafter “Yuen”) in view of MORIMOTO et al. (US 20210124963 A1 hereinafter “MORIMOTO”) in further view of KIM et al. (US 20250128726 A1 Hereinafter “KIM”). Regarding claim 9, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 1, in addition, Yuen further teaches one or more processors of claim 1, wherein the one or more processors are further to compute the statistical distribution (Page 326, Refining Landmark Location and Extracting Occlusion Score: “To extract the refined landmark locations, the same step as “Initial Landmark Location” in this subsection are applied using the new 68 landmark score images. The value of the pixel in the corresponding new score image at the refined landmark location is used as the landmark occlusion score”. This landmark occlusion score is determined based on the statistical distribution of the key point location confidence values, “The network is trained by following the original work of the Stacked Hourglass paper with a minor modification of including occlusion information in the heatmap ground truth labels. More specifically, if a landmark is labeled as occluded, then the Gaussian blob placed at the location of the landmark would have an amplitude of −1 instead of +1. The reasoning behind this is so that the location of the landmark can still be derived from looking at the magnitude of the heatmap, while the sign gives an indication of whether it’s occluded or not” (Page 324, Col. 1). The statistical distribution of the key point location values (heat map values) contain signs, these signs dictate whether or not the keypoint is occluded, generating an occlusion score); The combination of Yuen and MORIMOTO does not expressly disclose determining the statistical distribution using the standard deviation. However, KIM teaches determining the statistical distribution using the standard deviation ([0241]: “As a result, the technique according to an exemplary embodiment of the present disclosure calculates the size of the standard deviation of the values in the heatmap to determine whether an area determined as the eye area by the model is the actual eye area. As an example, the case where the value of the standard deviation of the heatmap is large may indicate that there is a high possibility that occlusion will occur in the image”. Statistical distribution is determined off the standard deviation to determine occlusion in the image). At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify the combination of Yuen and MORIMOTO’s occlusion detection system using heatmaps to include KIM’s use of standard deviation for statistical distribution for determining occlusion because such a modification is the result of applying a known technique to a known device ready for improvement to yield predictable results. More specifically, KIM’s use of standard deviation for statistical distribution for determining occlusion permits an accurate method of occlusion detection in heat maps of keypoints of driver’s facers. This known benefit in KIM is applicable to the combination of Yuen and MORIMOTO’s occlusion detection system using heatmaps as they both share characteristics and capabilities, namely, they are directed to driver face occlusion detection using keypoints from heatmaps. Therefore, it would have been recognized that modifying the combination of Yuen and MORIMOTO’s occlusion detection system using heatmaps to include KIM’s use of standard deviation for statistical distribution for determining occlusion would have yielded predictable results because (i) the level of ordinary skill in the art demonstrated by the references applied shows the ability to incorporate KIM’s use of standard deviation for statistical distribution for determining occlusion in driver face occlusion detection using keypoints from heatmaps and (ii) the benefits of such a combination would have been recognized by those of ordinary skill in the art. Claims 15 is rejected under 35 U.S.C. 103 as being unpatentable over Yuen et al. (“An Occluded Stacked Hourglass Approach to Facial Landmark Localization and Occlusion Estimation” Hereinafter “Yuen”) in view of MORIMOTO et al. (US 20210124963 A1 hereinafter “MORIMOTO”) in further view of Iqbal et al. (US 20190278983 A1 Hereinafter “Iqbal”). Regarding claim 15, the combination of Yuen and MORIMOTO teaches the one or more processors of claim 2, in addition, Yuen further teaches wherein the heat map representing key point location confidence values is generated using a machine learning model trained based at least on a two-dimensional feature key point data set (Page 324, Training Data Generation: “For landmark samples, a combined dataset of about 3,400 faces with 68 landmark annotations which follows the CMU Multi-PIE dataset format”. These sets of training data are used to train the stacked hourglass model) comprising one or more two-dimensional projections of self-occluded feature key points (Page 324, Training Data Generation: “For landmark samples, a combined dataset of about 3,400 faces with 68 landmark annotations which follows the CMU Multi-PIE dataset format”. “Due to resource constraints, all landmarks (regardless of whether they were occluded in the original image) are marked as non-occluded and only artificially occluded landmarks, using the method described here, are marked as occluded”. These sets of training data are used to train the stacked hourglass model. These sets of data include keypoints that are self-occluded by hands “During initial training using only the SUN2012 dataset, there were issues with landmarks being marked as non-occluded even though they were covered by hands and/or sunglasses, and so a small dataset was collected composed of about 400 hands and 140 sunglasses with transparency masks already included from the various sources on the web. The object is chosen by first randomly selecting, from a uniform distribution, one of the three datasets, e.g. SUN2012, hands or sunglasses. An object is then randomly selected from within the chosen dataset, augmented, and placed on top of the positive sample. Due to resource constraints, all landmarks(regardless of whether they were occluded in the original image) are marked as non-occluded and only artificially occluded landmarks, using the method described here, are marked as occluded” (Page 325, Col. 1). This means the training data comprises images and keypoints that are projected over the self-occluded areas). The combination of Yuen and MORIMOTO does not expressly disclose using 3D data for training the keypoint heat map model. However, Iqbal teaches using 3D data for training the keypoint heat map model ([0060]: “In contrast with neural network models that can be trained only given input image and 2D coordinates or 3D coordinates, but not if both data sources are mixed, the neural network models 210 and 212 may be trained using 2D coordinates and 3D coordinates simultaneously, without any constraints”. These networks output heatmaps of keypoints “instead implicitly learned latent 2.5 heatmaps output by the neural network model 212 are converted to the scale normalized 2.5D coordinates in a differentiable manner within the 2.5D keypoint estimation system 250” [0060]) At the time the invention was made, it would have been obvious to one of ordinary skill in the art to modify the combination of Yuen and MORIMOTO’s occlusion detection system using heatmaps to include Iqbal’s neural network’s ability for using 3D data and use of 3D training data for keypoint detection because such a modification is the result of applying a known technique to a known device ready for improvement to yield predictable results. More specifically, Iqbal’s neural network’s ability for using 3D data and use of 3D training data for keypoint detection permits a way to leverage 3D data to train a neural network to generate heat maps of keypoints, allowing for more access to data. This known benefit in Iqbal is applicable to the combination of Yuen and MORIMOTO’s occlusion detection system using heatmaps as they both share characteristics and capabilities, namely, they are directed to detection of keypoints using heatmaps generated by machine learning. Therefore, it would have been recognized that modifying the combination of Yuen and MORIMOTO’s occlusion detection system using heatmaps to include Iqbal’s neural network’s ability for using 3D data and use of 3D training data for keypoint detection would have yielded predictable results because (i) the level of ordinary skill in the art demonstrated by the references applied shows the ability to incorporate Iqbal’s neural network’s ability for using 3D data and use of 3D training data for keypoint detection in detection of keypoints using heatmaps generated by machine learning and (ii) the benefits of such a combination would have been recognized by those of ordinary skill in the art. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Guo et al. (US 20260045091 A1) teaches keypoint detection and probability analysis Park (US 12541979 B2) teaches Driver keypoint heatmaps MANFREDI et al. (US 20250238951 A1) teaches keypoint maps that contain probability measures. Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEFANO A DARDANO whose telephone number is (703)756-4543. The examiner can normally be reached Monday - Friday 11:00 - 7:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Greg Morse can be reached at (571) 272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /STEFANO ANTHONY DARDANO/ Examiner, Art Unit 2663 /GREGORY A MORSE/Supervisory Patent Examiner, Art Unit 2698
Read full office action

Prosecution Timeline

Oct 15, 2024
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705757
Image Processor
2y 10m to grant Granted Aug 11, 2026
Patent 12688567
METHOD FOR DETECTING AND LOCALIZING A FALSIFIED AREA IN JPEG IMAGES
3y 3m to grant Granted Jul 21, 2026
Patent 12688727
DETECTION SYSTEM AND DETECTION METHOD
3y 2m to grant Granted Jul 21, 2026
Patent 12670621
METHOD AND SYSTEM FOR TRACKING A STATE OF A CAMERA
3y 0m to grant Granted Jun 30, 2026
Patent 12670619
EYE TRACKING
2y 8m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+32.6%)
2y 12m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 88 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month