Prosecution Insights
Last updated: August 17, 2026
Application No. 18/243,555

LANDMARK DETECTION WITH AN ITERATIVE NEURAL NETWORK

Non-Final OA §103
Filed
Sep 07, 2023
Priority
Sep 13, 2022 — provisional 63/406,175
Examiner
CZEKAJ, DAVID J
Art Unit
2400
Tech Center
2400 — Computer Networks
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
49%
Grant Probability
Moderate
3-4
OA Rounds
2y 0m
Est. Remaining
40%
With Interview

Examiner Intelligence

Grants 49% of resolved cases
49%
Career Allowance Rate
116 granted / 236 resolved
-8.8% vs TC avg
Minimal -9% lift
Without
With
+-8.9%
Interview Lift
resolved cases with interview
Typical timeline
5y 0m
Avg Prosecution
24 currently pending
Career history
245
Total Applications
across all art units

Statute-Specific Performance

§101
12.2%
-27.8% vs TC avg
§103
68.6%
+28.6% vs TC avg
§102
9.8%
-30.2% vs TC avg
§112
4.4%
-35.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 236 resolved cases

Office Action

§103
DETAILED ACTION This action is in response to the amendment filed on 6/24/2025. Claims 1-12, 14-31 are pending. Response to Arguments Claim Rejections 35 USC §102 & 103 Applicant's amendment filed 6/24/2025 have been fully considered but they are not persuasive. Applicant states: 1. Page 9 “However, the general disclosure in Lin of detecting landmark features from a target image patch using a trained neural network does not teach or suggest any specific process of the trained neural network, let alone teach an iterative neural network that processes an input over a plurality of iterations at inference time to predict one or more landmarks for the input, as applicant claims (see claim language excerpted above)“, any emphasis not shown. Examiner’s response: See the new rejection. Claim Mapping Notation In this office action, following notations are being used to refer to the paragraph numbers or column number and lines of portions of the cited reference. “[0027]…” (Paragraph number [0027]) [4:3-15] ”…” (Column 4 Lines 3-15) Furthermore, unless necessary to distinguish from other references in this action, “et al.” will be omitted when referring to the reference. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under pre-AIA 35 U.S.C. 103(a) are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-5, 7-12, 14-15, 17-19 and 21-24 and 30-31 are rejected under 35 U.S.C. 103 as being unpatentable over Lin et al. (US 20230097869 A1) in view of Golden et al. (US 20180218502 A1). Regarding the claim 1, Lin discloses the invention substantially as claimed. Lin discloses, 1. A method, comprising: at a device: …processing an input, using an iterative neural network, to predict one or more landmarks for the input; and “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art that the training involves iterative process. outputting the one or more landmarks. “[0048]...Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” Lin does not disclose, …at inference time… Golden discloses, …at inference time… “[0185] Inference is the process of utilizing a trained model for prediction on new data… [0186] The inference service is responsible for loading a model, generating contours, and displaying them for the user. After inference is invoked at 902, at 904 images are sent to an inference server. At 906, the production model or network that is used by the inference service is loaded onto the inference server. The network may have been previously selected from the corpus of models trained during hyperparameter search…” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Golden and apply them on the teachings of Lin to incorporate trained model at the time of inference. One would have been motivated as implementing training model during inference is a well-known scheme that would have been expected to be implemented as taught by Golden. Unless stated otherwise, the same explanation for the rationale for the following dependent claims applies as given for the independent claim. 2. The method of claim 1, wherein the input is an image. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 3. The method of claim 1, wherein the input is a video. Lin “[0147]…The display 1050 is able to present, for example, various contents, such as text, images, videos, icons, and symbols.” 4. The method of claim 1, wherein the input depicts an image of an object. Lin “[0130] The first neural network 101 may identify a target object from an original image, using a plurality of convolutional layers. The first neural network 101 may be trained using a plurality of sample images including the same target object, to minimize loss between a predicted heat map and a ground-truth heat map.” 5. The method of claim 1, wherein the one or more landmarks include landmarks located on an object. Lin “[0130] The first neural network 101 may identify a target object from an original image, using a plurality of convolutional layers. The first neural network 101 may be trained using a plurality of sample images including the same target object, to minimize loss between a predicted heat map and a ground-truth heat map.” 7. The method of claim 1, wherein the iterative neural network is trained on labeled still images. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art understands that image includes both still and video. Furthermore, under BRI any syntax related to the images is considered to be “labeled” 8. The method claim 7, wherein the input is a video having a plurality of frames, and wherein the one or more landmarks are predicted for each frame of the plurality of frames. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 9. The method of claim 1, wherein a plurality of videos are labeled, and wherein each of the videos is labeled by: utilizing a plurality of existing trained neural networks to independently predict labels for the video, and Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art understands that image includes both still and video. Furthermore, under BRI any syntax related to the images is considered to be “labeled” Lin “[0067] The contracting path may be formed by a plurality of contracting paths, including a first contracting path to a fourth contracting path. The target image patch may pass through the contracting path, in the order of the first, the second, the third, and the fourth contracting paths. In each contracting path, two or more convolutional layers (e.g., 3×3 convolutional layers), each followed by an activation function layer (e.g., a rectified linear unit (ReLU) activation layer), and a max pooling layer (e.g., a 2×2 max pooling layer) with a stride greater than 1, are provided for down-sampling. The number of feature channels may be doubled in each contracting path.” labeling the video based on statistics of the predicted labels. Lin “[0048]…The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” 10. The method of claim 9, wherein the statistics include a mean of the predicted labels from plurality of existing trained neural networks, Lin “[0048]…The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” wherein the plurality of existing trained neural networks have been trained on still images or on videos. Lin “[0067] The contracting path may be formed by a plurality of contracting paths, including a first contracting path to a fourth contracting path. The target image patch may pass through the contracting path, in the order of the first, the second, the third, and the fourth contracting paths. In each contracting path, two or more convolutional layers (e.g., 3×3 convolutional layers), each followed by an activation function layer (e.g., a rectified linear unit (ReLU) activation layer), and a max pooling layer (e.g., a 2×2 max pooling layer) with a stride greater than 1, are provided for down-sampling. The number of feature channels may be doubled in each contracting path.” “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” Lin 11. The method of claim 1, wherein the iterative neural network predicts each of the one or more landmarks via a corresponding heatmap. “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 12. The method of claim 11, wherein the heatmap is converted into a landmark point representing the corresponding landmark. Lin “[0048]...Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” 14. The method of claim 13, wherein an initial iteration generates at least one initial heatmap for the input, and Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” wherein each subsequent iteration updates each heatmap from a prior iteration. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art that the training involves iterative process which means that the heat map will be updated. 15. The method of claim 13, wherein each initial heatmap corresponds to a landmark to be learned. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 17. The method of claim 1, wherein the iterative neural network includes a recurrent loss term. Lin “[0048] A loss function of the neural network calculates a per-pixel difference between the predicted heat map and the ground-truth heat map. The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE).” 18. The method of claim 1, wherein the input includes a video with a plurality of frames, and Lin “[0147]…The display 1050 is able to present, for example, various contents, such as text, images, videos, icons, and symbols.” wherein the iterative neural network processes the plurality of frames to predict the one or more landmarks for each frame of the plurality of frames, including: for each frame after an initial frame of the video, further processing, by the iterative neural network, a landmark prediction made by the iterative neural network for a prior frame of the video, to predict the one or more landmarks for the frame. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map. The ground-truth heat map is rendered using manually annotated ground-truth landmark locations. A loss function of the neural network calculates a per-pixel difference between the predicted heat map and the ground-truth heat map. The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” 19. The method of claim 18, wherein for the initial frame, the iterative neural network further processes an initialized landmark prediction to predict the one or more landmarks for the initial frame. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 21. The method of claim 18, wherein processing the landmark prediction made by the iterative neural network for the prior frame of the video when processing a current frame of the video provides temporal coherence of predicted landmarks across sequential frames of the video. Lin “[0048] A loss function of the neural network calculates a per-pixel difference between the predicted heat map and the ground-truth heat map. The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE).” Loss function is known to allow the “temporal coherence”. 22. The method of claim 18, wherein for each frame after the initial frame of the video, a number of steps the iterative neural network takes when processing the frame is limited based on a predefined criterion. Lin “[0103] In an example, an average mean square error between the feature vectors of the target image patch and the pre-calculated feature vectors of the template image is computed as a value representing the similarity. When the similarity is greater than a threshold similarity value (or when the average mean square error is lower than a threshold error value), the feature vectors of the target image patch are determined to be aligned with the pre-calculated feature vectors of the template image.” 23. The method of claim 22, wherein the predefined criterion includes a threshold difference between outputs of sequential steps. Lin “[0103] In an example, an average mean square error between the feature vectors of the target image patch and the pre-calculated feature vectors of the template image is computed as a value representing the similarity. When the similarity is greater than a threshold similarity value (or when the average mean square error is lower than a threshold error value), the feature vectors of the target image patch are determined to be aligned with the pre-calculated feature vectors of the template image.” 24. The method of claim 1, wherein the one or more landmarks predicted for the input is output to a downstream task. Lin “[0048]…The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” Regarding the claims 30 and 31, they recite elements that are at least included in the claims1 and 1 above but in a different claim form. Therefore, the same rationale for the rejection of the claims 1 and 1 applies. Regarding the processor, memory and storage medium in the claims, see Lin [177] and [178]. Regarding the claims 30 and 31, they recite elements that are at least included in the claims 1 and 1 above but in a different claim form. Therefore, the same rationale for the rejection of the claims applies. Regarding the processor, memory and storage medium in the claims, see Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under pre-AIA 35 U.S.C. 103(a) are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Lin-Golden further in view of Bai et al. (US 20230306617 A1). Regarding the claim 6, Lin-Golden disclose the invention substantially as claimed as mentioned above for the claim 1. Lin-Golden does not disclose, 6. The method of claim 1, wherein the iterative neural network is a deep equilibrium model (DEQ). Bai discloses, 6. The method of claim 1, wherein the iterative neural network is a deep equilibrium model (DEQ). “[0004] One class of implicit layer models is a deep equilibrium (DEQ) model. DEQ modeling includes specifying a layer that finds the fixed point of some iterative procedure. A multiscale deep equilibrium model (MDEQ) directly solves for and backpropagates through the equilibrium points of multiple feature resolutions simultaneously, using implicit differentiation to avoid storing intermediate states.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Bai and apply them on the teachings of Lin-Golden to incorporate the DEQ in the system of Lin-Golden when predicting landmarks in Lin-Golden as taught by Bai. One would have been motivated as implement the DEQ for the benefit of increasing memory efficiency and simplification of the system. 16. The method of claim 13, wherein the iterative neural network processes the input over the plurality of iterations until an equilibrium is found. Lin “[0004] One class of implicit layer models is a deep equilibrium (DEQ) model. DEQ modeling includes specifying a layer that finds the fixed point of some iterative procedure. A multiscale deep equilibrium model (MDEQ) directly solves for and backpropagates through the equilibrium points of multiple feature resolutions simultaneously, using implicit differentiation to avoid storing intermediate states.” Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Lin-Golden further in view of Tokmakov et al. (US 20220300748 A1). Regarding the claim 20, Lin-Golden discloses the invention substantially as claimed as mentioned above for the claims 1, 18 and 19. Lin-Golden does not disclose, 20. The method of claim 19, wherein the initialized landmark prediction generated for the initial frame of the video is set to zero. Tokmakov discloses, 20. The method of claim 19, wherein the initialized landmark prediction generated for the initial frame of the video is set to zero. “[0039] In such an example, the updated state 310 is determined by a GRU function based on a previous state M.sup.t−1 and the feature map F.sup.t. For an initial frame, the previous state M.sup.t−1 may be initialized to a particular value, such as zero. The updated state 310 M.sup.t may be an example of an output feature map. In the example of FIG. 3, the explicit encoding of the objects in the previous frame H.sup.t−1 (e.g., the heat map of prior tracked objects) is not used because the explicit encoding is captured in the ConvGRU state M.sup.t. Additionally, in the example of FIG. 3, the updated state 310 M.sup.t may be processed by distinct sub-networks 330a f.sub.p, 330b f.sub.s, and 330c f.sub.d to produce predictions for the current frame I.sup.t. The predictions are based on the updated state 310 M.sup.t, which is based on the features of the current frame I.sup.t and the previous frames (I.sup.t-1 to I.sup.i). Each sub-network 330a, 330b, 330c may be a convolutional neural network trained to perform a specific task, such as determine object centers based on features of the updated state 310 M.sup.t, determine bounding box dimensions based on features of the updated state 310 M.sup.t, and determine displacement vectors of the updated state 310 M.sup.t.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Tokmakov and apply them on the teachings of Lin-Golden to incorporate the setting of zero for the initialized landmark prediction generated for the initial frame Lin-Golden when predicting landmarks in Lin-Golden as taught by Tokmakov. One would have been motivated as such setting to zero would have offered the benefits of more simplistic approach and better capture of temporal dynamics. Claims 25-26 are rejected under 35 U.S.C. 103 as being unpatentable over Lin-Golden further in view of Haskin et al. (WO 2024042508 A1). Regarding the claim 25, Lin-Golden discloses the invention substantially as claimed as mentioned above for the claim 24. Lin-Golden does not disclose, 25. The method of claim 24, wherein the downstream task includes a self-driving application. Haskin discloses, 25. The method of claim 24, wherein the downstream task includes a self-driving application. [P186 30-34] “Alternatively or in addition, the vehicle may be an aircraft adapted to fly in air, and the aircraft may be a fixed wing or a rotorcraft aircraft, such as an airplane, a spacecraft, a glider, a drone, or an Unmanned Aerial Vehicle (UAV). Any vehicle herein may be a ground vehicle that may consist of, or may comprise, an autonomous car, which may be according to levels 0, 1, 2, 3, 4, or 5 of the Society of Automotive Engineers (SAE) J3016 standard.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Haskin and apply them on the teachings of Lin-Golden to incorporate the autonomous automobile application as well as facial detection in Lin when predicting landmarks in Lin-Golden as taught by Haskin.. One would have been motivated as such setting to zero would have offered the benefits of more simplistic approach and better capture of temporal dynamics. 26. The method of claim 25, wherein the input depicts a human face of a driver of an automobile, and Haskin [P6 18-25] “A camera with human face detection means is disclosed in U.S. Patent 6,940,545 to Ray et al., entitled: “Face Detecting Camera and Method”, and in U.S. Patent Application Publication No. 2012/0249768 to Binder entitled: ”System and Method for Control Based on Face or Hand Gesture Detection”, which are both incorporated in their entirety for all purposes as if fully set forth herein.” Haskin [P186 30-34] “Alternatively or in addition, the vehicle may be an aircraft adapted to fly in air, and the aircraft may be a fixed wing or a rotorcraft aircraft, such as an airplane, a spacecraft, a glider, a drone, or an Unmanned Aerial Vehicle (UAV). Any vehicle herein may be a ground vehicle that may consist of, or may comprise, an autonomous car, which may be according to levels 0, 1, 2, 3, 4, or 5 of the Society of Automotive Engineers (SAE) J3016 standard.” wherein the self-driving application uses the one or more landmarks predicted for the human face to monitor a state of the driver of the automobile for making autonomous driving policy decisions based thereon. Haskin [P6 18-25] “A camera with human face detection means is disclosed in U.S. Patent 6,940,545 to Ray et al., entitled: “Face Detecting Camera and Method”, and in U.S. Patent Application Publication No. 2012/0249768 to Binder entitled: ”System and Method for Control Based on Face or Hand Gesture Detection”, which are both incorporated in their entirety for all purposes as if fully set forth herein.” Haskin [P186 30-34] “Alternatively or in addition, the vehicle may be an aircraft adapted to fly in air, and the aircraft may be a fixed wing or a rotorcraft aircraft, such as an airplane, a spacecraft, a glider, a drone, or an Unmanned Aerial Vehicle (UAV). Any vehicle herein may be a ground vehicle that may consist of, or may comprise, an autonomous car, which may be according to levels 0, 1, 2, 3, 4, or 5 of the Society of Automotive Engineers (SAE) J3016 standard.” Haskin [P34 6-12] “The computation of sensor motion from sets of displacement vectors obtained from consecutive pairs of images is described in a paper by Wilhelm Burger and Bir Bhanu entitled: “Estimating 3-D Egomotion from Perspective Image Sequences”, published in IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 12, NO. 11, NOVEMBER 1990, which is incorporated in its entirety for all purposes as if fully set forth herein. The problem is investigated with emphasis on its application to autonomous robots and land vehicles.” Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” Claims 27-29 are rejected under 35 U.S.C. 103 as being unpatentable over Lin-Golden further in view of Khakhulin et al. (US 20230154111 A1). Regarding the claim 27, Lin-Golden discloses the invention substantially as claimed as mentioned above for the claim 24. Lin-Golden does not disclose, 27. The method of claim 24, wherein the downstream task includes an avatar-based application. Khakhulin discloses, 27. The method of claim 24, wherein the downstream task includes an avatar-based application. Khakhulin “[0032] Embodiments of the disclosure provide three-dimensional of an object (e.g., a human head) in the form of polygonal mesh using a single image with animation and realistic rendering capabilities for novel head poses. Personalized human avatars are becoming the key technology across several application domains, such as telepresence, virtual worlds, online commerce. In many cases, it is sufficient to personalize only a part of the avatars' body. The remaining body parts may then be either chosen from a certain library of assets or omitted from the interface. Towards this end, many applications require personalization at the head level, e.g., creating person-specific head models. Creating personalized heads is an important and viable intermediate step between personalizing just face (which is often insufficient) and creating personalized full-body models, which is a much harder task that limits quality of the resulting models and/or requires cumbersome data collection.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Khakhulin and apply them on the teachings of Lin-Golden to incorporate the face/body detection and depiction as avatar when predicting landmarks in Lin-Golden as taught by Khakhulin.. One would have been motivated as such avatar related depiction of humans in autonomous vehicles are readily found which would have offered a more convenient and flexible way to display the driver as taught by Khakhulin. 28. The method of claim 27, wherein the input depicts a human face, and wherein the avatar-based application uses the one or more landmarks predicted for the human face to apply a select avatar to the depiction of the human face. Khakhulin “[0029] FIG. 3 illustrates comparison of renders on a VoxCeleb2 dataset. The task is to reenact the source image with the expression and pose of the driver image;” Khakhulin “[0032] Embodiments of the disclosure provide three-dimensional of an object (e.g., a human head) in the form of polygonal mesh using a single image with animation and realistic rendering capabilities for novel head poses. Personalized human avatars are becoming the key technology across several application domains, such as telepresence, virtual worlds, online commerce. In many cases, it is sufficient to personalize only a part of the avatars' body. The remaining body parts may then be either chosen from a certain library of assets or omitted from the interface. Towards this end, many applications require personalization at the head level, e.g., creating person-specific head models. Creating personalized heads is an important and viable intermediate step between personalizing just face (which is often insufficient) and creating personalized full-body models, which is a much harder task that limits quality of the resulting models and/or requires cumbersome data collection.” Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 29. The method of claim 27, wherein the input depicts a human body, and wherein the avatar-based application uses the one or more landmarks predicted for the human body to determine a pose of the human body and to generate the avatar in the pose. Khakhulin “[0032] Embodiments of the disclosure provide three-dimensional of an object (e.g., a human head) in the form of polygonal mesh using a single image with animation and realistic rendering capabilities for novel head poses. Personalized human avatars are becoming the key technology across several application domains, such as telepresence, virtual worlds, online commerce. In many cases, it is sufficient to personalize only a part of the avatars' body. The remaining body parts may then be either chosen from a certain library of assets or omitted from the interface. Towards this end, many applications require personalization at the head level, e.g., creating person-specific head models. Creating personalized heads is an important and viable intermediate step between personalizing just face (which is often insufficient) and creating personalized full-body models, which is a much harder task that limits quality of the resulting models and/or requires cumbersome data collection.” Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAE N. NOH whose telephone number is (571) 270-0686. The examiner can normally be reached on Mon-Fri 8:30AM-5PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, William Vaughn can be reached on (571) 272-3922. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAE N NOH/ Primary Examiner Art Unit 2481
Read full office action

Prosecution Timeline

Show 1 earlier event
Mar 27, 2025
Non-Final Rejection mailed — §103
Jun 24, 2025
Response Filed
Oct 06, 2025
Final Rejection mailed — §103
Dec 08, 2025
Response after Non-Final Action
Dec 30, 2025
Notice of Allowance
Mar 02, 2026
Response after Non-Final Action
Mar 19, 2026
Response after Non-Final Action
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12651328
FULL-SPACE INTELLIGENT DETECTION METHOD AND SYSTEM FOR UNDERGROUND DRAINAGE NETWORKS, AS WELL AS STORAGE MEDIA
2y 0m to grant Granted Jun 09, 2026
Patent 12639952
Egress Obstruction Detection via Computer Vision
2y 0m to grant Granted May 26, 2026
Patent 12634493
METHOD FOR IMAGE COMPRESSION AND APPARATUS FOR IMPLEMENTING THE SAME
3y 3m to grant Granted May 19, 2026
Patent 12608465
BOT DETECTION SYSTEM
3y 0m to grant Granted Apr 21, 2026
Patent 12586173
APPARATUS FOR HAIR INSPECTION ON SUBSTRATE AND METHOD FOR HAIR INSPECTION ON SUBSTRATE
2y 1m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
49%
Grant Probability
40%
With Interview (-8.9%)
5y 0m (~2y 0m remaining)
Median Time to Grant
High
PTA Risk
Based on 236 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month