Prosecution Insights
Last updated: October 02, 2026
Application No. 18/243,555

LANDMARK DETECTION WITH AN ITERATIVE NEURAL NETWORK

Non-Final OA §103
Filed
Sep 07, 2023
Priority
Sep 13, 2022 — provisional 63/406,175
Examiner
CZEKAJ, DAVID J
Art Unit
2400
Tech Center
2400 — Computer Networks
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
50%
Grant Probability
Moderate
3-4
OA Rounds
1y 10m
Est. Remaining
42%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
120 granted / 241 resolved
-8.2% vs TC avg
Minimal -8% lift
Without
With
+-7.5%
Interview Lift
resolved cases with interview
Typical timeline
4y 11m
Avg Prosecution
23 currently pending
Career history
265
Total Applications
across all art units

Statute-Specific Performance

§101
11.3%
-28.7% vs TC avg
§103
67.8%
+27.8% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
4.9%
-35.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 241 resolved cases

Office Action

§103
DETAILED ACTION Response to Arguments Applicant’s arguments with respect to the rejection(s) of the claim(s) have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made as outlined below. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under pre-AIA 35 U.S.C. 103(a) are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-5, 7-12, 14-15, 17-19 and 21-24 and 30-31 are rejected under 35 U.S.C. 103 as being unpatentable over Lin et al. (US 20230097869 A1) in view of Hassani et al. (2023/0260329), (hereinafter referred to as “Hassani”). Regarding the claim 1, Lin discloses the invention substantially as claimed. Lin discloses, 1. A method, comprising: at a device: …processing an input, using an iterative neural network, to predict one or more landmarks for the input; and “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art that the training involves iterative process. outputting the one or more landmarks. “[0048]...Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” Lin does not disclose, …at inference time… Hassani discloses, At an inference time, process an input over a plurality of iterations, using a trained iterative neural network, to predict one or more features for the input Figures 2 and 5; paragraph 0044-0048 It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Hassani and apply them on the teachings of Lin to incorporate trained model at the time of inference. One would have been motivated to so as to provide accurate and timely data regarding objects in an environment (Hassani: paragraph 0002). Unless stated otherwise, the same explanation for the rationale for the following dependent claims applies as given for the independent claim. 2. The method of claim 1, wherein the input is an image. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 3. The method of claim 1, wherein the input is a video. Lin “[0147]…The display 1050 is able to present, for example, various contents, such as text, images, videos, icons, and symbols.” 4. The method of claim 1, wherein the input depicts an image of an object. Lin “[0130] The first neural network 101 may identify a target object from an original image, using a plurality of convolutional layers. The first neural network 101 may be trained using a plurality of sample images including the same target object, to minimize loss between a predicted heat map and a ground-truth heat map.” 5. The method of claim 1, wherein the one or more landmarks include landmarks located on an object. Lin “[0130] The first neural network 101 may identify a target object from an original image, using a plurality of convolutional layers. The first neural network 101 may be trained using a plurality of sample images including the same target object, to minimize loss between a predicted heat map and a ground-truth heat map.” 7. The method of claim 1, wherein the iterative neural network is trained on labeled still images. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art understands that image includes both still and video. Furthermore, under BRI any syntax related to the images is considered to be “labeled” 8. The method claim 7, wherein the input is a video having a plurality of frames, and wherein the one or more landmarks are predicted for each frame of the plurality of frames. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 9. The method of claim 1, wherein a plurality of videos are labeled, and wherein each of the videos is labeled by: utilizing a plurality of existing trained neural networks to independently predict labels for the video, and Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art understands that image includes both still and video. Furthermore, under BRI any syntax related to the images is considered to be “labeled” Lin “[0067] The contracting path may be formed by a plurality of contracting paths, including a first contracting path to a fourth contracting path. The target image patch may pass through the contracting path, in the order of the first, the second, the third, and the fourth contracting paths. In each contracting path, two or more convolutional layers (e.g., 3×3 convolutional layers), each followed by an activation function layer (e.g., a rectified linear unit (ReLU) activation layer), and a max pooling layer (e.g., a 2×2 max pooling layer) with a stride greater than 1, are provided for down-sampling. The number of feature channels may be doubled in each contracting path.” labeling the video based on statistics of the predicted labels. Lin “[0048]…The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” 10. The method of claim 9, wherein the statistics include a mean of the predicted labels from plurality of existing trained neural networks, Lin “[0048]…The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” wherein the plurality of existing trained neural networks have been trained on still images or on videos. Lin “[0067] The contracting path may be formed by a plurality of contracting paths, including a first contracting path to a fourth contracting path. The target image patch may pass through the contracting path, in the order of the first, the second, the third, and the fourth contracting paths. In each contracting path, two or more convolutional layers (e.g., 3×3 convolutional layers), each followed by an activation function layer (e.g., a rectified linear unit (ReLU) activation layer), and a max pooling layer (e.g., a 2×2 max pooling layer) with a stride greater than 1, are provided for down-sampling. The number of feature channels may be doubled in each contracting path.” “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” Lin 11. The method of claim 1, wherein the iterative neural network predicts each of the one or more landmarks via a corresponding heatmap. “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 12. The method of claim 11, wherein the heatmap is converted into a landmark point representing the corresponding landmark. Lin “[0048]...Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” 14. The method of claim 13, wherein an initial iteration generates at least one initial heatmap for the input, and Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” wherein each subsequent iteration updates each heatmap from a prior iteration. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” It is understood by one of ordinary skilled in the art that the training involves iterative process which means that the heat map will be updated. 15. The method of claim 13, wherein each initial heatmap corresponds to a landmark to be learned. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 17. The method of claim 1, wherein the iterative neural network includes a recurrent loss term. Lin “[0048] A loss function of the neural network calculates a per-pixel difference between the predicted heat map and the ground-truth heat map. The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE).” 18. The method of claim 1, wherein the input includes a video with a plurality of frames, and Lin “[0147]…The display 1050 is able to present, for example, various contents, such as text, images, videos, icons, and symbols.” wherein the iterative neural network processes the plurality of frames to predict the one or more landmarks for each frame of the plurality of frames, including: for each frame after an initial frame of the video, further processing, by the iterative neural network, a landmark prediction made by the iterative neural network for a prior frame of the video, to predict the one or more landmarks for the frame. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map. The ground-truth heat map is rendered using manually annotated ground-truth landmark locations. A loss function of the neural network calculates a per-pixel difference between the predicted heat map and the ground-truth heat map. The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” 19. The method of claim 18, wherein for the initial frame, the iterative neural network further processes an initialized landmark prediction to predict the one or more landmarks for the initial frame. Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 21. The method of claim 18, wherein processing the landmark prediction made by the iterative neural network for the prior frame of the video when processing a current frame of the video provides temporal coherence of predicted landmarks across sequential frames of the video. Lin “[0048] A loss function of the neural network calculates a per-pixel difference between the predicted heat map and the ground-truth heat map. The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE).” Loss function is known to allow the “temporal coherence”. 22. The method of claim 18, wherein for each frame after the initial frame of the video, a number of steps the iterative neural network takes when processing the frame is limited based on a predefined criterion. Lin “[0103] In an example, an average mean square error between the feature vectors of the target image patch and the pre-calculated feature vectors of the template image is computed as a value representing the similarity. When the similarity is greater than a threshold similarity value (or when the average mean square error is lower than a threshold error value), the feature vectors of the target image patch are determined to be aligned with the pre-calculated feature vectors of the template image.” 23. The method of claim 22, wherein the predefined criterion includes a threshold difference between outputs of sequential steps. Lin “[0103] In an example, an average mean square error between the feature vectors of the target image patch and the pre-calculated feature vectors of the template image is computed as a value representing the similarity. When the similarity is greater than a threshold similarity value (or when the average mean square error is lower than a threshold error value), the feature vectors of the target image patch are determined to be aligned with the pre-calculated feature vectors of the template image.” 24. The method of claim 1, wherein the one or more landmarks predicted for the input is output to a downstream task. Lin “[0048]…The neural network is trained to minimize the per-pixel difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” difference between the predicted heat map and the ground-truth heat map, for example, using a mean-square error (MSE). Once the predicted heat map of landmarks is obtained from the neural network, a post-processing step may be performed to obtain the locations of the landmarks (e.g., pixel coordinates of the landmarks) from the predicted heat map.” Regarding the claims 30 and 31, they recite elements that are at least included in the claims1 and 1 above but in a different claim form. Therefore, the same rationale for the rejection of the claims 1 and 1 applies. Regarding the processor, memory and storage medium in the claims, see Lin [177] and [178]. Regarding the claims 30 and 31, they recite elements that are at least included in the claims 1 and 1 above but in a different claim form. Therefore, the same rationale for the rejection of the claims applies. Regarding the processor, memory and storage medium in the claims, see Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Lin-Hassani further in view of Bai et al. (US 20230306617 A1). Regarding the claim 6, Lin-Hassani disclose the invention substantially as claimed as mentioned above for the claim 1. Lin-Hassani does not disclose, 6. The method of claim 1, wherein the iterative neural network is a deep equilibrium model (DEQ). Bai discloses, 6. The method of claim 1, wherein the iterative neural network is a deep equilibrium model (DEQ). “[0004] One class of implicit layer models is a deep equilibrium (DEQ) model. DEQ modeling includes specifying a layer that finds the fixed point of some iterative procedure. A multiscale deep equilibrium model (MDEQ) directly solves for and backpropagates through the equilibrium points of multiple feature resolutions simultaneously, using implicit differentiation to avoid storing intermediate states.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Bai and apply them on the teachings of Lin-Hassani to incorporate the DEQ in the system of Lin-Hassani when predicting landmarks in Lin-Hassani as taught by Bai. One would have been motivated as implement the DEQ for the benefit of increasing memory efficiency and simplification of the system. 16. The method of claim 13, wherein the iterative neural network processes the input over the plurality of iterations until an equilibrium is found. Lin “[0004] One class of implicit layer models is a deep equilibrium (DEQ) model. DEQ modeling includes specifying a layer that finds the fixed point of some iterative procedure. A multiscale deep equilibrium model (MDEQ) directly solves for and backpropagates through the equilibrium points of multiple feature resolutions simultaneously, using implicit differentiation to avoid storing intermediate states.” Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Lin-Hassani further in view of Tokmakov et al. (US 20220300748 A1). Regarding the claim 20, Lin-Hassani discloses the invention substantially as claimed as mentioned above for the claims 1, 18 and 19. Lin-Hassani does not disclose, 20. The method of claim 19, wherein the initialized landmark prediction generated for the initial frame of the video is set to zero. Tokmakov discloses, 20. The method of claim 19, wherein the initialized landmark prediction generated for the initial frame of the video is set to zero. “[0039] In such an example, the updated state 310 is determined by a GRU function based on a previous state M.sup.t−1 and the feature map F.sup.t. For an initial frame, the previous state M.sup.t−1 may be initialized to a particular value, such as zero. The updated state 310 M.sup.t may be an example of an output feature map. In the example of FIG. 3, the explicit encoding of the objects in the previous frame H.sup.t−1 (e.g., the heat map of prior tracked objects) is not used because the explicit encoding is captured in the ConvGRU state M.sup.t. Additionally, in the example of FIG. 3, the updated state 310 M.sup.t may be processed by distinct sub-networks 330a f.sub.p, 330b f.sub.s, and 330c f.sub.d to produce predictions for the current frame I.sup.t. The predictions are based on the updated state 310 M.sup.t, which is based on the features of the current frame I.sup.t and the previous frames (I.sup.t-1 to I.sup.i). Each sub-network 330a, 330b, 330c may be a convolutional neural network trained to perform a specific task, such as determine object centers based on features of the updated state 310 M.sup.t, determine bounding box dimensions based on features of the updated state 310 M.sup.t, and determine displacement vectors of the updated state 310 M.sup.t.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Tokmakov and apply them on the teachings of Lin-Hassani to incorporate the setting of zero for the initialized landmark prediction generated for the initial frame Lin-Hassani when predicting landmarks in Lin-Hassani as taught by Tokmakov. One would have been motivated as such setting to zero would have offered the benefits of more simplistic approach and better capture of temporal dynamics. Claims 25-26 are rejected under 35 U.S.C. 103 as being unpatentable over Lin-Hassani further in view of Haskin et al. (WO 2024042508 A1). Regarding the claim 25, Lin-Hassani discloses the invention substantially as claimed as mentioned above for the claim 24. Lin-Hassani does not disclose, 25. The method of claim 24, wherein the downstream task includes a self-driving application. Haskin discloses, 25. The method of claim 24, wherein the downstream task includes a self-driving application. [P186 30-34] “Alternatively or in addition, the vehicle may be an aircraft adapted to fly in air, and the aircraft may be a fixed wing or a rotorcraft aircraft, such as an airplane, a spacecraft, a glider, a drone, or an Unmanned Aerial Vehicle (UAV). Any vehicle herein may be a ground vehicle that may consist of, or may comprise, an autonomous car, which may be according to levels 0, 1, 2, 3, 4, or 5 of the Society of Automotive Engineers (SAE) J3016 standard.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Haskin and apply them on the teachings of Lin-Hassani to incorporate the autonomous automobile application as well as facial detection in Lin when predicting landmarks in Lin-Hassani as taught by Haskin.. One would have been motivated as such setting to zero would have offered the benefits of more simplistic approach and better capture of temporal dynamics. 26. The method of claim 25, wherein the input depicts a human face of a driver of an automobile, and Haskin [P6 18-25] “A camera with human face detection means is disclosed in U.S. Patent 6,940,545 to Ray et al., entitled: “Face Detecting Camera and Method”, and in U.S. Patent Application Publication No. 2012/0249768 to Binder entitled: ”System and Method for Control Based on Face or Hand Gesture Detection”, which are both incorporated in their entirety for all purposes as if fully set forth herein.” Haskin [P186 30-34] “Alternatively or in addition, the vehicle may be an aircraft adapted to fly in air, and the aircraft may be a fixed wing or a rotorcraft aircraft, such as an airplane, a spacecraft, a glider, a drone, or an Unmanned Aerial Vehicle (UAV). Any vehicle herein may be a ground vehicle that may consist of, or may comprise, an autonomous car, which may be according to levels 0, 1, 2, 3, 4, or 5 of the Society of Automotive Engineers (SAE) J3016 standard.” wherein the self-driving application uses the one or more landmarks predicted for the human face to monitor a state of the driver of the automobile for making autonomous driving policy decisions based thereon. Haskin [P6 18-25] “A camera with human face detection means is disclosed in U.S. Patent 6,940,545 to Ray et al., entitled: “Face Detecting Camera and Method”, and in U.S. Patent Application Publication No. 2012/0249768 to Binder entitled: ”System and Method for Control Based on Face or Hand Gesture Detection”, which are both incorporated in their entirety for all purposes as if fully set forth herein.” Haskin [P186 30-34] “Alternatively or in addition, the vehicle may be an aircraft adapted to fly in air, and the aircraft may be a fixed wing or a rotorcraft aircraft, such as an airplane, a spacecraft, a glider, a drone, or an Unmanned Aerial Vehicle (UAV). Any vehicle herein may be a ground vehicle that may consist of, or may comprise, an autonomous car, which may be according to levels 0, 1, 2, 3, 4, or 5 of the Society of Automotive Engineers (SAE) J3016 standard.” Haskin [P34 6-12] “The computation of sensor motion from sets of displacement vectors obtained from consecutive pairs of images is described in a paper by Wilhelm Burger and Bir Bhanu entitled: “Estimating 3-D Egomotion from Perspective Image Sequences”, published in IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL. 12, NO. 11, NOVEMBER 1990, which is incorporated in its entirety for all purposes as if fully set forth herein. The problem is investigated with emphasis on its application to autonomous robots and land vehicles.” Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” Claims 27-29 are rejected under 35 U.S.C. 103 as being unpatentable over Lin-Hassani further in view of Khakhulin et al. (US 20230154111 A1). Regarding the claim 27, Lin-Hassani discloses the invention substantially as claimed as mentioned above for the claim 24. Lin-Hassani does not disclose, 27. The method of claim 24, wherein the downstream task includes an avatar-based application. Khakhulin discloses, 27. The method of claim 24, wherein the downstream task includes an avatar-based application. Khakhulin “[0032] Embodiments of the disclosure provide three-dimensional of an object (e.g., a human head) in the form of polygonal mesh using a single image with animation and realistic rendering capabilities for novel head poses. Personalized human avatars are becoming the key technology across several application domains, such as telepresence, virtual worlds, online commerce. In many cases, it is sufficient to personalize only a part of the avatars' body. The remaining body parts may then be either chosen from a certain library of assets or omitted from the interface. Towards this end, many applications require personalization at the head level, e.g., creating person-specific head models. Creating personalized heads is an important and viable intermediate step between personalizing just face (which is often insufficient) and creating personalized full-body models, which is a much harder task that limits quality of the resulting models and/or requires cumbersome data collection.” It would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to utilize the teachings of Khakhulin and apply them on the teachings of Lin-Hassani to incorporate the face/body detection and depiction as avatar when predicting landmarks in Lin-Hassani as taught by Khakhulin.. One would have been motivated as such avatar related depiction of humans in autonomous vehicles are readily found which would have offered a more convenient and flexible way to display the driver as taught by Khakhulin. 28. The method of claim 27, wherein the input depicts a human face, and wherein the avatar-based application uses the one or more landmarks predicted for the human face to apply a select avatar to the depiction of the human face. Khakhulin “[0029] FIG. 3 illustrates comparison of renders on a VoxCeleb2 dataset. The task is to reenact the source image with the expression and pose of the driver image;” Khakhulin “[0032] Embodiments of the disclosure provide three-dimensional of an object (e.g., a human head) in the form of polygonal mesh using a single image with animation and realistic rendering capabilities for novel head poses. Personalized human avatars are becoming the key technology across several application domains, such as telepresence, virtual worlds, online commerce. In many cases, it is sufficient to personalize only a part of the avatars' body. The remaining body parts may then be either chosen from a certain library of assets or omitted from the interface. Towards this end, many applications require personalization at the head level, e.g., creating person-specific head models. Creating personalized heads is an important and viable intermediate step between personalizing just face (which is often insufficient) and creating personalized full-body models, which is a much harder task that limits quality of the resulting models and/or requires cumbersome data collection.” Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” 29. The method of claim 27, wherein the input depicts a human body, and wherein the avatar-based application uses the one or more landmarks predicted for the human body to determine a pose of the human body and to generate the avatar in the pose. Khakhulin “[0032] Embodiments of the disclosure provide three-dimensional of an object (e.g., a human head) in the form of polygonal mesh using a single image with animation and realistic rendering capabilities for novel head poses. Personalized human avatars are becoming the key technology across several application domains, such as telepresence, virtual worlds, online commerce. In many cases, it is sufficient to personalize only a part of the avatars' body. The remaining body parts may then be either chosen from a certain library of assets or omitted from the interface. Towards this end, many applications require personalization at the head level, e.g., creating person-specific head models. Creating personalized heads is an important and viable intermediate step between personalizing just face (which is often insufficient) and creating personalized full-body models, which is a much harder task that limits quality of the resulting models and/or requires cumbersome data collection.” Lin “[0048] In operation S120, the method S100 includes detecting a plurality of landmark features from the target image patch, using a neural network which has been trained to output a prediction of a heat map of landmarks (a two-dimensional heat map of moon landmarks), using a ground-truth heat map.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAVE J CZEKAJ whose telephone number is (571)272-7327. The examiner can normally be reached 8-6:00 Monday-Thursday and every other Friday. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, please contact Jamie Atala at 571-272-7384. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Dave Czekaj/ Supervisory Patent Examiner, Art Unit 2487
Read full office action

Prosecution Timeline

Show 1 earlier event
Mar 27, 2025
Non-Final Rejection mailed — §103
Jun 24, 2025
Response Filed
Oct 06, 2025
Final Rejection mailed — §103
Dec 08, 2025
Response after Non-Final Action
Dec 30, 2025
Notice of Allowance
Mar 02, 2026
Response after Non-Final Action
Mar 19, 2026
Response after Non-Final Action
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749331
METHOD, DEVICE AND STORAGE MEDIUM FOR RECOGNIZING CHART
2y 10m to grant Granted Sep 29, 2026
Patent 12739317
ADU ASSOCIATION METHOD AND COMPUTER DEVICE
2y 5m to grant Granted Sep 15, 2026
Patent 12651328
FULL-SPACE INTELLIGENT DETECTION METHOD AND SYSTEM FOR UNDERGROUND DRAINAGE NETWORKS, AS WELL AS STORAGE MEDIA
2y 0m to grant Granted Jun 09, 2026
Patent 12639952
Egress Obstruction Detection via Computer Vision
2y 0m to grant Granted May 26, 2026
Patent 12634493
METHOD FOR IMAGE COMPRESSION AND APPARATUS FOR IMPLEMENTING THE SAME
3y 3m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
50%
Grant Probability
42%
With Interview (-7.5%)
4y 11m (~1y 10m remaining)
Median Time to Grant
High
PTA Risk
Based on 241 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month