Prosecution Insights
Last updated: October 02, 2026
Application No. 18/584,191

POSE RELATION TRANSFORMER AND REFINING OCCLUSIONS FOR HUMAN POSE ESTIMATION

Final Rejection §103
Filed
Feb 22, 2024
Priority
Mar 01, 2023 — provisional 63/487,728
Examiner
ABDI, AMARA
Art Unit
2668
Tech Center
2600 — Communications
Assignee
Purdue Research Foundation
OA Round
2 (Final)
83%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
76%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
697 granted / 840 resolved
+21.0% vs TC avg
Minimal -7% lift
Without
With
+-7.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
22 currently pending
Career history
859
Total Applications
across all art units

Statute-Specific Performance

§101
11.0%
-29.0% vs TC avg
§103
64.5%
+24.5% vs TC avg
§102
9.7%
-30.3% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 840 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Applicant's response to the last office action, filed July 27, 2026 has been entered and made of record. Claims 1 and 12 are amended; and claim 19 is cancelled. By this amendment, claims 1-18, and 20 are pending in this application for examination. Response to Arguments Applicant’s arguments with respect to claim(s) 1-18, and 20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 5, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Batra et al, (US-Patent 12,211,307) in view of Ali et al, (US-PGPUB 20230282031) Regarding claim 1, Batra et al discloses a method for human pose estimation, (see at least: Fig. 1, and col. 1, lines 49-50, “devices, systems, and methods for body pose estimation”), comprising: obtaining, with a processor, a plurality of keypoints corresponding to a plurality of joints of a human in an image, (see at least: Fig. 1, first stage TCN, plurality of joints of a human body; and col. 3, lines 39-40, where a network estimates two-dimensional body joints in an image); masking, with the processor, a subset of keypoints in the plurality of keypoints corresponding to occluded joints of the human, (see at least: Fig. 1, first stage TCN, where subset of points corresponding to occluded joints of the human, are masked with rectangular boxes at t=1, t=2, and t=w, on the left side (input) of the first stage TCN 110. Further, col. 6, lines 33-42, “Occluded joints”, where some joints are occluded due to the relative orientation between the human and the camera, and external occlusion caused by other objects in the scene. See also, col. 7, lines 20-21, “masking is performed over the joints”); determining, with the processor, a reconstructed subset of keypoints by reconstructing the masked subset of keypoints using a machine learning model, (see at least: Fig. 1, see first stage TCN 110, “i.e., machine learning model”, which outputs the masked subset of keypoints of the human body. Further, col. 4, lines 40-45, the first stage 110 is a TCN that accepts a window of two-dimensional poses as input and outputs three-dimensional poses, which implicitly comprises the reconstructed masked subset of keypoints of the human body); and forming, with the processor, a refined plurality of keypoints, the refined plurality of keypoints being used by a system to perform a task, (see at least: Fig. 1, where the TCN refiner 120, receives as input the outputs of the first stage 110, and outputs a refined human body, (t=1 …t=W), including the refined keypoints and the masked keypoints reconstructed from first stage TCN 110. See also, col. 4, lines 45-51, the outputs of the first stage 110 are passed to the second stage 120 which is a temporal refiner network that improves the estimated three-dimensional poses, which technically produces reliable three-dimensional poses, (see col. 2, line 29) that enables the human body or robot to perform one or more tasks, “i.e., the refined plurality of keypoints being implicitly used to perform a task”). Batra does not expressly disclose that the refined plurality of keypoints is formed by substituting the reconstructed subset of keypoints in place of the masked subset of keypoints in the plurality of keypoints, each non-masked keypoint in the plurality of keypoints being retained in the refined plurality of keypoints However, Ali discloses substituting the reconstructed subset of keypoints in place of the masked subset of keypoints in the plurality of keypoints, each non-masked keypoint in the plurality of keypoints being retained in the refined plurality of keypoints, (see at least: Par. 0021-0022, the machine learning model is previously trained to receive as input spatial information for n+m joints of the articulated object, where m is greater than or equal to 1. During training, the input data provided to the machine learning model may be progressively masked, e.g., input data corresponding to some of the m joints may be replaced with pre-defined values representing masked joints, or simply omitted, as such the machine learning model may accurately predict the pose of the articulated object at runtime based on sparse inputs—e.g., based at least on spatial information for the n joints, without spatial information for the m joints, [i.e., substituting the reconstructed subset of keypoints in place of the masked subset of keypoints in the plurality of keypoints, “implicit by replacing the m joints with pre-defined values representing masked joints, or simply omitted”, each non-masked keypoint in the plurality of keypoints being retained in the refined plurality of keypoints, “implicit by predicting the pose of the articulated object at runtime based at least on spatial information for the n joints only”]). Batra and Ali are combinable because they are both concerned with pose determination. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify Batra, to incorporate the machine learning model as though by Ali, with the Batra’s TCN refiner 120, in order to predict the pose of the articulated object at runtime based on at least on spatial information for the n joints, without spatial information for the masked m joints, (Ali, Par. 0022), which may provide the technical benefit of reducing consumption of computing resources while accurately recreating a human user's real-world pose, by reducing the amount of data that must be collected as input for the pose prediction process, (Ali, Par. 0023). Regarding claim 5, the combination of Batra and Ali as whole discloses the limitations of claim 1. Batra further discloses that the masking the subset of keypoints further comprising: obtaining, with the processor, a respective confidence value for each keypoint in the plurality of keypoints; and determining, with the processor, the subset of keypoints as those keypoints in the plurality of keypoints having respective confidence values that are less than a predetermined threshold, (see at least: col. 7, lines 12-18, using binary channel as explicit occlusion indicator, where at inference time, the two-dimensional detector confidence can be thresholded and used for the purpose of detecting occluded joints, [i.e., the confidence value for each keypoint is obtained, “implicit by the two-dimensional detector confidence”, and the subset of keypoints corresponding to occluded keypoints are implicitly determined based on thresholding the confidence”]). Regarding claim 20, the combination of Batra and Ali as whole discloses the limitations of claim 1. Batra further discloses wherein the machine learning model has been previously trained by randomly masking keypoints in a training dataset and learning to predict the masked keypoints, (see at least: col. 4, lines 14-15, keypoints are randomly masked during training to provide occlusion data augmentation, which the machine learning model implicitly has been previously trained by randomly masking keypoints in a training dataset and learning to predict the masked keypoints). Claims 2 and 3 are rejected under 35 U.S.C. 103 as being unpatentable over Batra and Ali, as applied to claim 1 above; and further in view of Zheng et al, (US-PGPUB 20230196617) Regarding claim 2, the combination of Batra and Ali as whole discloses the limitations of claim 1. Batra further discloses the obtaining the plurality of keypoints further comprising: receiving, with the processor, the image from an image sensor, the image capturing the human, (see at least: col. 7, lines 58-63, implicit by obtaining a plurality of two-dimensional images of a body); and determining, with the processor, the plurality of keypoints corresponding to the plurality of joints of the human two-dimensional location in the two-dimensional image of one or more joints of the body, [i.e., implicitly determining plurality of keypoints corresponding to joints of the human body, based on determining location of one or more joints of the body). The combination of Batra and Ali as whole does not expressly disclose using a keypoint detection model for determining plurality of keypoints corresponding to the plurality of joints of the human. Zheng discloses using a keypoint detection model for determining plurality of keypoints corresponding to the plurality of joints of the human, (see at least: Fig. 2, Par. 0023, based on an RGB image of the person, a plurality of body keypoints 206, may be extracted from the image, for example, using a first neural network 204 (e.g., a body keypoint extraction network). Batra, Ali, and Zheng are combinable because they are both concerned with body joints detection. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Batra and Ali, to use the body keypoint extraction network 204, as though by Zheng, in order to extract plurality of body keypoints (joints), (Par. 0023) Regarding claim 3, the combination of Batra, Ali, and Zheng as whole discloses the limitations of claim 2. Furthermore, Zheng discloses generating, with the processor, a plurality of heatmaps based on the image; and determining, with the processor, the plurality of keypoints based on the plurality of heatmaps, each respective joint in the plurality of keypoints being determined based on a corresponding respective heatmap in the plurality of heatmaps, where the body keypoints may refer to body parts and/or joints, (Zheng, see at least: Par. 0023, the extracted body keypoints 206 are in the form of one or more heat maps representing the body keypoints 206, [i.e., a plurality of heatmaps are implicitly generated; and the extracted body keypoints are determined based on the heatmaps, and each respective joint, (body keypoints refer to joints), being implicitly determined based on a corresponding respective heatmap of one or more heatmaps). Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Batra, Ali, and Zheng, as applied to claim 3 above; and further in view of Brown et al, (US-Patent 11,875,529) The combination of Batra, Ali, and Zheng as whole discloses the limitations of claim 3. The combination of Batra, Ali, and Zheng as whole does not expressly disclose determining, with the processor, a plurality of confidence values for the plurality of keypoints based on the plurality of heatmaps, each respective confidence value being determined based on a corresponding respective heatmap in the plurality of heatmaps. Brown et al discloses determining, with the processor, a plurality of confidence values for the plurality of keypoints based on the plurality of heatmaps, each respective confidence value being determined based on a corresponding respective heatmap in the plurality of heatmaps, (see at least: col. 5, lines 47-57, Fig. 8 depth estimator generates a separate depth heatmap for each kind of joint, where the value (i.e., darkness) of each point along a given heatmap may represent the confidence of the corresponding joint being at that depth, and the confidence values of all depths may be zero for a not-visible joint, [i.e., generating plurality of confidence values for the plurality of keypoints based on the plurality of heatmaps, each respective confidence value being determined based on a corresponding respective heatmap in the plurality of heatmaps, “implicit by estimating the confidence value for each joint in a given heatmap’]). Batra, Ali, Zheng, and Brown are combinable because they are all concerned with body joints detection. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Batra, Ali, and Zheng, to generate a confidence value for each joint in the heatmap, as though by Brown, in order to estimate a 3D location for each joint, (Brown, col. 5, lines 59-60) Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Batra and Ali, as applied to claim 1 above; and further in view of Yip et al, (US-PGPUB 20240289882) The combination of Batra and Ali as whole discloses the limitations of claim 1. The combination of Batra and Ali as whole does not expressly disclose wherein the machine learning model incorporates a Transformer-based neural network architecture and uses multi-scale graph convolution. However, Yip discloses wherein the machine learning model incorporates a Transformer-based neural network architecture and uses multi-scale graph convolution, (see at least: Par. 0012, adopting a transformer module with a learnable transformer neural network layer, “i.e., Transformer-based neural network architecture”; and in the transformer graph convolved dynamic mode decomposition (TGCDMD), the structural dependence of various assets is captured by a weighted graph, which is fed into a next step that integrates a graph convolution layer and DMD graph convolution, “implicit the multi-scale graph convolution”). Batra, Ali, and Yip et al are combinable because they are both concerned with object recognition. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Batra and Ali, to incorporate the TGCDMD, as though by Yip, with the Batra’s body pose estimator 100, in order to model the poses faster. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Batra and Ali, as applied to claim 1 above; and further in view of Cai et al, (US-PGPUB 20220335654) The combination of Batra and Ali as whole discloses the limitations of claim 1. The combination of Batra and Ali as whole does not expressly disclose that the determining the reconstructed subset of keypoints further comprising determining, with the processor, an initial feature embedding based on the plurality of keypoints. However, Cai et al discloses the determining, with the processor, an initial feature embedding based on the plurality of keypoints, (see at least: Par. 0043, calculate the initial feature of the first point cloud data based on initial weight information of the first encoder, [i.e., the encoded initial feature is calculated based on the point cloud data, “plurality of keypoints”]). Batra, Ali, and Cai et al are combinable because they are both concerned with object recognition. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Batra and Ali, to calculate the initial feature of the first point cloud data, as though by Cai, in order to generating point cloud encoder, (Cai et al, Par. 0036). Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Batra, Ali, and Cai, as applied to claim 7 above; and further in view of Xing et al, (US-PGPUB 20230028046). The combination of Batra, Ali, and Cai as whole discloses the limitations of the claim 7. The combination of Batra, Ali, and Cai as whole does not expressly disclose the determining the initial feature embedding using multi-scale graph convolution. However, Xing discloses determining the initial feature embedding using multi-scale graph convolution, (see at least: Par. 0087, based on the initial feature of the node and the initial feature of each target node corresponding to the node (i.e., each associated feature of the node), the weight of each associated feature is determined through a graph convolution network, “i.e., implicitly using multi-scale graph convolution”, and then the initial feature of the node and the initial feature of each target node may be weighted according to respective corresponding weights of the initial features of the node and the initial feature of each target node corresponding to the node, to obtain each weighted initial feature, “initial feature embedding”, [i.e., [i.e., determining the initial feature embedding, “weighted initial feature”, using the multi-scale graph convolution, “graph convolution network based on multi-weights”]). Batra, Ali, Cai, and Xing are combinable because they are all concerned with feature detection. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combine teaching Batra, Ali, and Cai, to use the weight-based determination graph convolution network, as though by Xing, in order to determine one or more weighted initial feature, (Xing, Par. 0087). Claims 9-15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Batra, Ali, and Cai, as applied to claim 7 above; and further in view of Zhao et al, (US-PGPUB 20210133535) Regarding claim 9, the combination of Batra, Ali, and Cai as whole discloses the limitations of claim 7. The combination of Batra, Ali, and Cai as whole does not expressly disclose the determining, with the processor, based on the initial feature embedding, a plurality of attended feature embeddings using an encoder of the machine learning model, the encoder having a Transformer-based neural network architecture. However, Zhao discloses determining, with the processor, based on the initial feature embedding, a plurality of attended feature embeddings using an encoder of the machine learning model, the encoder having a Transformer-based neural network architecture, (see at least: Par. 0026-0028, where the equivalent number of attention matrices “i.e., the plurality of attended feature embeddings” is determined based on the query, key, and value matrices for each set of weight matrices, “i.e., the initial feature embedding”, using an encoder of the transformer model, which is a deep learning model, “the encoder having a Transformer-based neural network architecture”). Batra, Ali, Cai, and Zhao are combinable because they are all concerned with feature detection. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Batra, Ali, and Cai, to use the transformer model, as though by Zhao, to calculate an equivalent number of attention matrices using the query, key, and value matrices for each set of weight matrices, (Zhao, Par. 0028) Regarding claim 10, the combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 9. Zhao further discloses wherein the encoder has a plurality of encoding layers, the plurality of encoding layers having a sequential order, each respective encoding layer determining a respective attended feature embedding of the plurality of attended feature embeddings, (Zhao, see at least: Par. 0026-0028, the encoder includes a set of encoding layers that processes the input iteratively one layer after another, “plurality of encoding layers having a sequential order”, and the process undertaken by each self-attention layer includes taking in an input including a list of fixed-length vectors, and splitting the input into a set of query, key, and value matrices, “i..e, each respective encoding layer determining a respective attended feature embedding”). Regarding claim 11, the combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 10. Zhao further discloses determining, with the processor, each respective attended feature embedding of the plurality of attended feature embeddings, in a respective encoding layer of the plurality of encoding layers, based on a previous feature embedding, (Zhao, see at least: Par. 0026, the encoder includes a set of encoding layers that processes the input iteratively one layer after another; and Par. 0028, calculating an equivalent number of attention matrices, “each respective attended feature embedding”, using the query, key, and value matrices for each set of weight matrices, “based on a previous feature embedding”); and wherein (i) for a first encoding layer of the plurality of encoding layers, the previous feature embedding is the initial feature embedding and (ii) for each encoding layer of the plurality of encoding layers other than the first encoding layer, the previous feature embedding is that which is output by a previous encoding layer of the plurality of encoding layers, (Zhao, see at least: Par. 0026, the encoder includes a set of encoding layers that processes the input iteratively one layer after another, “which implicit that the first layer encode the initial feature embedding, and for next layers after the first layer, each layer encodes the previous feature output by preceding layer) Regarding claim 12, the combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 11. Furthermore, Zhao discloses determining, with the processor, a respective attention matrix based on the previous feature embedding, (Zhao, see at least: Par. 0029, the output of the attention layer is a matrix, “attention matrix”, including a vector for each entry in the input sequence, “the previous feature embedding”); and determining, with the processor, the respective attended feature embedding based on the respective attention matrix and the previous attended feature embedding, (Zhao, see at least: Par. 0028-0029, the matrix serves as the input of the feed-forward neural network, which the output of feed-forward neural network, implicitly corresponds to the respective attended feature embedding, [i.e., the output of feed-forward neural network, “respective attended feature embedding”, is determined implicitly based on the attention matrix, and the vector entry in the input sequence, “previous attended feature embedding”]). Regarding claim 13, the combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 12. Zhao further discloses that determining the respective attention matrix further comprising: determining, with the processor, a respective multi-head self-attention matrix, (Zhao, see at least: Par. 0027-0028, passing the vectors into a self-attention layer (multi-head attention), to compute the final attention vector for every entry, “implicit the determining respective multi-head self-attention matrix”). Regarding claim 14, the combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 12. Zhao further discloses the determining, with the processor, respective Key, Query, and Value matrices based on the previous feature embedding; and determining, with the processor, the respective attention matrix based on the previous feature embedding and the respective Key, Query, and Value matrices, (Zhao, see at least: Par. 0028, splitting the input into a set of query, key, and value matrices, ….and calculating an equivalent number of attention matrices using the query, key, and value matrices for each set of weight matrices). Regarding claim 15, the combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 14. Zhao further discloses determining the respective attention matrix further comprising: determining, with the processor, the respective Key, Query, and Value matrices using multi-scale graph convolution, (Zhao, see at least: Par. 0026, transformer model is a deep learning model, “implicit the graph convolution”; and from Par. 0028, multiplying the input by a set of weight matrices, where each set includes a query weight matrix, a key weight matrix, and a value weight matrix, to determine the respective Key, Query, and Value matrices, “set of weight matrices implicit the multi-scale of transformer model or graph convolution”). Regarding claim 17, the combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 10. Batra further discloses determining, with the processor, the reconstructed subset of keypoints based on a final attended feature embedding of the plurality of attended feature embeddings, the final attended feature embedding being output by a final encoding layer of the plurality of encoding layers, (Batra, col. 3, lines 57-59, a network includes an attention mechanism added to the TCN, which implicit that the reconstructed subset of keypoints is based on a final attended feature output by a final encoding layer of the plurality of encoding layers, as attention mechanism added to the TCN implicitly comprises plurality of encoding layers) Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Batra, Ali, Cai, and Zhao, as applied to claim 17 above; and further in view of Hwang et al, (US-Patent 12131493) The combination of Batra, Ali, Cai, and Zhao as whole discloses the limitations of claim 17. Batra further discloses determining, with the processor, the reconstructed subset of keypoints based on the final attended feature embedding using sequence, (Batra, col. 3, lines 57-59, a network includes an attention mechanism added to the TCN, which implicit that the reconstructed subset of keypoints is based on a final attended feature output. Further, col. 4, lines 65-67, “the input sequence” implicit the use of sequence). The combination of Batra, Ali, Cai, and Zhao as whole does not expressly disclose that the reconstructed subset of keypoints based on the final attended feature embedding using excitation Hwang discloses the encoder performing convolution by performing a an excitation process for generating first channel information, (step s210 in Fig. 4, col. 8, lines 17-25, and claim 1) Batra, Ali, Cai, Zhao, and Hwang are combinable because they are all concerned with feature detection. Therefore, it would have been obvious to a person of ordinary skill in the art, to modify the combination of Batra, Ali, Cai, and Zhao, to perform a an excitation process, as though by Hwang, in order to generate channel information, (col. 8, lines 17-25). Allowable Subject Matter Claim 16 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. With respect to claim 16, the prior art of record, alone or in reasonable combination, does not teach or suggest the following limitation(s), (in consideration of the claim as a whole): “determining, with the processor, a respective intermediate feature embedding based on the attention matrix and the previous attended feature embedding; and determining, with the processor, the respective attended feature embedding based on the respective intermediate feature embedding using a multi-layer perceptron” The relevant prior art of record, Zhao et al, (US-PGPUB 20210133535) discloses determining, with the processor, a respective attention matrix based on the previous feature embedding, (see at least: Par. 0029, the output of the attention layer is a matrix, “attention matrix”, including a vector for each entry in the input sequence, “the previous feature embedding”); and determining, with the processor, the respective attended feature embedding based on the attention matrix and the previous attended feature embedding, (see at least: Par. 0028-0029, the matrix serves as the input of the feed-forward neural network, which the output of feed-forward neural network, implicitly corresponds to the respective attended feature embedding, [i.e., the output of feed-forward neural network, “respective attended feature embedding”, is determined implicitly based on the attention matrix, and the vector entry in the input sequence, “previous attended feature embedding”]); but fails to teach or suggest, either alone or in combination with the other cited references, determining a respective intermediate feature embedding based on the attention matrix and the previous attended feature embedding; and determining the respective attended feature embedding based on the respective intermediate feature embedding using a multi-layer perceptron. A further prior art of record, Batra et al discloses a method for human pose estimation, (see at least: Fig. 1, and col. 1, lines 49-50, “devices, systems, and methods for body pose estimation”), comprising: obtaining, with a processor, a plurality of keypoints corresponding to a plurality of joints of a human in an image, (see at least: Fig. 1, first stage TCN, plurality of joints of a human body; and col. 3, lines 39-40, where a network estimates two-dimensional body joints in an image); masking, with the processor, a subset of keypoints in the plurality of keypoints corresponding to occluded joints of the human, (see at least: Fig. 1, first stage TCN, where subset of points corresponding to occluded joints of the human, are masked with rectangular boxes at t=1, t=2, and t=w, on the left side (input) of the first stage TCN 110. Further, col. 6, lines 33-42, “Occluded joints”, where some joints are occluded due to the relative orientation between the human and the camera, and external occlusion caused by other objects in the scene. See also, col. 7, lines 20-21, “masking is performed over the joints”); determining, with the processor, a reconstructed subset of keypoints by reconstructing the masked subset of keypoints using a machine learning model, (see at least: Fig. 1, and col. 4, lines 40-45, “see the rejection of claim 1 for more details”); and forming, with the processor, a refined plurality of keypoints based on the plurality of keypoints and the reconstructed subset of keypoints, the refined plurality of keypoints being used by a system to perform a task, (see at least: Fig. 1, col. 2, line 29, and col. 4, lines 45-51, “see the rejection of claim 1 for more details”). However, Batra fails to teach or suggest, either alone or in combination with the other cited references, determining a respective intermediate feature embedding based on the attention matrix and the previous attended feature embedding; and determining the respective attended feature embedding based on the respective intermediate feature embedding using a multi-layer perception. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMARA ABDI whose telephone number is (571)272-0273. The examiner can normally be reached 9:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMARA ABDI/Primary Examiner, Art Unit 2668 09/11/2026
Read full office action

Prosecution Timeline

Feb 22, 2024
Application Filed
Jan 27, 2026
Non-Final Rejection mailed — §103
Jul 27, 2026
Response Filed
Sep 16, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749345
MONITORING AND ANALYZING BODY LANGUAGE WITH MACHINE LEARNING, USING ARTIFICIAL INTELLIGENCE SYSTEMS FOR IMPROVING INTERACTION BETWEEN HUMANS, AND HUMANS AND ROBOTS
2y 7m to grant Granted Sep 29, 2026
Patent 12749346
FINGER ENCODING BASED POSE CLASSIFICATION
2y 6m to grant Granted Sep 29, 2026
Patent 12749307
CAMERA APPARATUS AND METHOD OF ENHANCED FOLIAGE DETECTION
2y 2m to grant Granted Sep 29, 2026
Patent 12743792
SYSTEMS AND METHODS FOR IMAGE PROCESSING
2y 11m to grant Granted Sep 22, 2026
Patent 12728441
ROBOTIC REPAIR CONTROL SYSTEMS AND METHODS
3y 6m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
83%
Grant Probability
76%
With Interview (-7.3%)
2y 6m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 840 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month