Prosecution Insights
Last updated: October 02, 2026
Application No. 18/778,780

METHOD AND APPARATUS FOR POSE ESTIMATION, AND ELECTRONIC DEVICE

Non-Final OA §101§103
Filed
Jul 19, 2024
Priority
Jul 20, 2023 — CN 202310896769.5
Examiner
HAGOS, EYOB
Art Unit
Tech Center
Assignee
Beijing Zitiao Network Technology Co., Ltd.
OA Round
1 (Non-Final)
66%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 66% — above average
66%
Career Allowance Rate
268 granted / 404 resolved
+6.3% vs TC avg
Strong +43% interview lift
Without
With
+43.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
30 currently pending
Career history
431
Total Applications
across all art units

Statute-Specific Performance

§101
24.3%
-15.7% vs TC avg
§103
49.6%
+9.6% vs TC avg
§102
6.3%
-33.7% vs TC avg
§112
17.2%
-22.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 404 resolved cases

Office Action

§101 §103
DETAILED ACTION 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . 2. Claims 1-20 are pending and presented for examination. Claim Rejections - 35 USC § 101 3. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 4. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The representative claim 19 recites: An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when being executed by the one or more processors, the one or more processors implement a method for pose estimation comprising: obtaining an observation information sequence corresponding to a human target part, wherein the human target part comprises at least one of a head or a hand; determining an initial human joint point feature based on the observation information sequence; and performing feature interaction based on the initial human joint point feature, and estimating a human pose with an interaction feature to obtain human pose information. The claim limitations in the abstract idea have been highlighted in bold above; the remaining limitations are “additional elements”. Under step 1 of the eligibility analysis, we determine whether the claims are to a statutory category by considering whether the claimed subject matter falls within the four statutory categories of patentable subject matter identified by 35 U.S.C. 101: process, machine, manufacture, or composition of matter. The above claims are considered to be in a statutory category (process). Under Step 2A, Prong One, we consider whether the claim recites a judicial exception (abstract idea). In the above claim, the highlighted portion constitutes an abstract idea because, under a broadest reasonable interpretation, it recites limitation that fall into/recite abstract idea exceptions. Specifically, under the 2019 Revised Patent Subject Matter Eligibility Guidance, it falls into the grouping of subject matter that, when recited as such in a claim limitation, covers mathematical concepts (mathematical relationships, mathematical formulas or equations, mathematical calculations) and/or mental processes – concepts performed in the human mind including an observation, evaluation, judgement, and/or opinion. Next, under Step 2A, Prong Two, we consider whether the claim that recites a judicial exception is integrated into a practical application. In this step, we evaluate whether the claim recites additional elements that integrate the exception into a practical application of that exception. This judicial exception is not integrated into a practical application because the additional limitations in the claim are only: one or more processors; a storage device having one or more programs stored thereon, when being executed by the one or more processors. The limitations “one or more processors; a storage device having one or more programs stored thereon, when being executed by the one or more processors” are recited at a high level of generality (i.e., as a computer structures performing a generic computer function of processing and storing information) such that they amount no more than mere instructions to apply the exception using generic computer components. Finally, under Step 2B, we consider whether the additional elements are sufficient to amount to significantly more than the abstract idea. Claim 19 does not include additional elements that are sufficient to amount to significantly more than the judicial exception because, as noted above, the additional elements are recited at a high level of generality (i.e., as a computer/computing components performing a generic function of processing and storing information). Further, the additional elements are conventional in the art, as evidenced by the art of record (see, Nakamura et al. US 2022/0108468 (hereinafter, Nakamura), ([0053]), and Knues et al. US 20260237093 (hereinafter, Knues), ([0076]). Therefore, claim 19 is directed to an abstract idea without significantly more. The claim is not patent eligible. Dependent claims 2-18, add further details of the identified abstract idea. The claims are not patent eligible. Independent claims 1 and 20, the claims are rejected with the same rationale as in claim 19 as explained above. Claim Objection 3. Claim 3 is objected to because of the following informalities: Claim 3 recites “…an initial human join point…(line 5)” should read “…[[an]] the initial human join point…” Appropriate correction is required. Claim Rejections - 35 USC § 103 4. In the event the determination of the status of the application as subject to AlA 35 U.S.C. 102 and 103 (or as subject to pre-AlA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 5. Claims 1, 2, 7, 8, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang et al. “AvatarPoser: Articulated Full-Body Pose Tracking from Sparse Motion Sensing”, July 2022 (hereinafter, Jiang), in view of Nakamura et al. US 2022/0108468 (hereinafter, Nakamura). 6. Regarding claim 1, Jiang discloses a method for pose estimation, comprising: obtaining an observation information sequence corresponding to a human target part, wherein the human target part comprises at least one of a head or a hand (Fig. 1: takes as input only the positions and orientations of one headset and two handheld controllers (or hands), and generates a full-body avatar pose over 22 joints….[Further] pages 5-6, sections 3.2 and 3.3: the 6D representation of rotations has proved effective for training neural networks due to its continuity…position and orientation of the headset and hands…and the final input representation is a concatenated vector of position, linear velocity, rotation, and angular velocity from all given sparse inputs), determining [human joint point] feature based on the observation information sequence (page 6, sections 3.3: AvatarPoser is a time series network that takes as input the 6D signals from the sparse trackers over the previous N-1 frames and the current Nth frame and predicts global orientation of the human body as well as the local rotations at each joint with respect to its parent joint… a Transformer model to extract the useful information from time-series data, following its benefits in efficiency, scalability, and long-term modeling capabilities... Next, our Transformer Encoder extracts deep pose features from previous time steps from the headset and hands); and performing feature interaction based on the human joint point feature, and estimating a human pose with an interaction feature to obtain human pose information (page 5, section 3.1, Fig. 1: AvatarPoser reconstructs the position of the articulated joints of the user's full body within the world w… Specifically, we use the SMPL model to represent and animate our human body pose. We use the first 22 joints defined in the kinematic tree of the SMPL human skeleton…. The output of our rotation-based pose estimation network is the local rotation at each joint with respect to the parent joints...As we use 22 joints to represent the full-body motion, the output dimension at each time step is 132…. [Further], pages 6-7, section 3.3, Fig. 2: the Transformer Encoder extracts deep pose features from previous time step signals from the headset and hands, which are split into global and local branches and correspond to global and local pose estimation, respectively. The Stabilizer is responsible for global motion navigation by decoupling global orientation from pose features and estimating global translation from the head position through the body's kinematic chain. The Forward-Kinematics Module calculates joint positions from a human skeleton model and a predicted body pose. The Inverse-Kinematics Module adjusts the estimated rotation angles of joints on the shoulder and elbow to reduce hand position errors). Jiang discloses a time series network that takes as input the 6D signals from the sparse trackers over the previous N-1 frames and the current Nth frame and predicts global orientation of the human body as well as the local rotations at each joint with respect to its parent joint as disclosed above. Jiang does not disclose: initial human joint point feature. However, Nakamura discloses: initial human joint point feature ([0075], [0096], [0101]). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang to use initial human joint point feature as taught by Nakamura. The motivation for doing so would have been in order to obtain predictable results in determining the pose information (Nakamura, [0098]). 7. Regarding claims 19 and 20, the claims are rejected with the same rationale as in claim 1. 8. Regarding claim 2, Jiang in view of Nakamura disclose the method of claim 1, as disclosed above. Jiang further discloses wherein the performing feature interaction based on the initial human joint point feature comprises: performing the feature interaction in a spatial dimension and/or in a temporal dimension based on the initial human joint point feature (pages 5-6, section 3. 2, and 3.3). 8. Regarding claim 7, Jiang in view of Nakamura disclose the method of claim 1, wherein the determining the initial human joint point feature based on the observation information sequence as disclosed above. Jiang further discloses determining human pose information based on the observation information sequence; and correcting the human pose information, and determining the human joint point feature based on the corrected human pose information (pages 6-7, section 3.3, Fig. 2). Jiang does not disclose: initial human joint point feature. However, Nakamura discloses: initial human joint point feature ([0075], [0096], [0101]). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang to use initial human joint point feature as taught by Nakamura. The motivation for doing so would have been in order to obtain predictable results in determining the pose information (Nakamura, [0098]). 8. Regarding claim 8, Jiang in view of Nakamura disclose the method of claim 1, wherein the determining the initial human joint point feature based on the observation information sequence as disclosed above. Jiang further discloses wherein the human pose information comprises a relative rotation angle of a human joint point under a human parameterized grid model and/or a joint point coordinate of the human joint point under the human parameterized grid model, and the observation information sequence comprises a rotation angle and/or a joint point coordinate of a joint point on the human target part (page 7, Fig. 2.); and correcting the human pose information comprises at least one of: replacing a rotation angle of the joint point on the human target part in the human pose information with the rotation angle of the joint point on the human target part in the observation information sequence; replacing a joint point coordinate of the joint point on the human target part in the human pose information with the joint point coordinate of the joint point on the human target part in the observation information sequence (pages 6-7, section 3.3, Fig. 2). Jiang does not disclose: initial human joint point feature. However, Nakamura discloses: initial human joint point feature ([0075], [0096], [0101]). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang to use initial human joint point feature as taught by Nakamura. The motivation for doing so would have been in order to obtain predictable results in determining the pose information (Nakamura, [0098]). 8. Regarding claim 18, Jiang in view of Nakamura disclose the method of claim 1, as disclosed above. Jiang further discloses wherein the observation information comprises a movement velocity and an angular velocity of a joint point (page 5, section 3. 2). 5. Claims 3-6 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang, in view of Nakamura, in further view of Zhang et al. “MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video”, 2022 (hereinafter, Zhang). 8. Regarding claim 3, Jiang in view of Nakamura disclose the method of claim 2, wherein the performing the feature interaction in the spatial dimension and/or in the temporal dimension based on the initial human joint point feature as disclosed above. Jiang further discloses performing interaction on the human joint point feature corresponding to the collection point (pages 6-7, 9, sections 3.3, 4.1, Fig. 2). Further, Nakamura discloses initial human joint point feature ([0075], [0096], [0101]). Jiang in view of Nakamura does not disclose: for each collection point in a collection time period, determining an attention score between an initial human joint point feature corresponding to the collection point and a target input feature, and performing interaction on the initial human joint point feature corresponding to the collection point and the target input feature with the attention score, to obtain an interaction feature in the spatial dimension, wherein the target input feature is obtained by mapping a high-dimensional input feature to a same dimension of the initial human joint point feature corresponding to the collection point, the high-dimensional input feature being determined based on the observation information sequence, and the collection time period is a time period for collecting the observation information sequence. However, Zhang discloses: for each collection point in a collection time period, determining an attention score between an initial human joint point feature corresponding to the collection point and a target input feature, and performing interaction on the initial human joint point feature corresponding to the collection point and the target input feature with the attention score, to obtain an interaction feature in the spatial dimension (pages 3-4, section 3.1.2. Fig. 2: employ the spatial transformer block (STB) to learn spatial correlations among joints in each frame. Given 2D keypoints with N joints, we consider each joint as a token in spatial attention. Firstly, we take 2D keypoints as input and project each keypoint to a high-dimensional feature with the linear embedding layer. The feature is referred to as a spatial token in STB. We then embed the spatial position information with a positional matrix …. After that, spatial tokens Pi ∈ RN×dm of th i-th frame is fed into spatial self-attention mechanism of STB to model dependencies across all joints and output the high-dimensional), wherein the target input feature is obtained by mapping a high-dimensional input feature to a same dimension of the initial human joint point feature corresponding to the collection point, the high-dimensional input feature being determined based on the observation information sequence, and the collection time period is a time period for collecting the observation information sequence (Abstract, pages 3-4, and Fig. 2 ). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura to use for each collection point in a collection time period, determining an attention score between an initial human joint point feature corresponding to the collection point and a target input feature, and performing interaction on the initial human joint point feature corresponding to the collection point and the target input feature with the attention score, to obtain an interaction feature in the spatial dimension, wherein the target input feature is obtained by mapping a high-dimensional input feature to a same dimension of the initial human joint point feature corresponding to the collection point, the high-dimensional input feature being determined based on the observation information sequence, and the collection time period is a time period for collecting the observation information sequence as taught by Zhang. The motivation for doing so would have been in order to determining spatio temporal correlation and learn different motion trajectories of body joints (Zhang, page 1). 8. Regarding claim 4, Jiang in view of Nakamura disclose the method of claim 2, wherein the performing the feature interaction in the spatial dimension and/or in the temporal dimension based on the initial human joint point feature as disclosed above. Jiang further discloses performing interaction on the human joint point feature corresponding to the joint point, to obtain an interaction feature in the temporal dimension, wherein the collection time period is a time period for collecting the observation information sequence (pages 5-6, sections 3.2, 3.3, Fig. 2). Further, Nakamura discloses initial human joint point feature ([0075], [0096], [0101]). Jiang in view of Nakamura does not disclose: for each joint point of human joint points, determining an attention score between a plurality of initial joint point features of the joint point within a collection time period; and performing interaction on the plurality of initial joint point features corresponding to the joint point with the attention score, to obtain an interaction feature in the temporal dimension. However, Zhang discloses: for each joint point of human joint points, determining an attention score between a plurality of initial joint point features of the joint point within a collection time period; and performing interaction on the plurality of initial joint point features corresponding to the joint point with the attention score, to obtain an interaction feature in the temporal dimension (pages 3-4, section 3.1.1, Fig. 2). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura to use for each joint point of human joint points, determining an attention score between a plurality of initial joint point features of the joint point within a collection time period; and performing interaction on the plurality of initial joint point features corresponding to the joint point with the attention score, to obtain an interaction feature in the temporal dimension as taught by Zhang. The motivation for doing so would have been in order to determining spatial temporal correlation and learn different motion trajectories of body joints (Zhang, page 1). 8. Regarding claim 5, Jiang in view of Nakamura disclose the method of claim 2, wherein the performing the feature interaction in the spatial dimension and/or in the temporal dimension based on the initial human joint point feature as disclosed above. Jiang further discloses spatio-temporal representations of multiple pose hypotheses to predict 3D human pose from monocular videos (page 4). Further, Nakamura discloses initial human joint point feature ([0075], [0096], [0101]). Jiang in view of Nakamura does not disclose: inputting the initial human joint point feature and a target input feature into a pre-trained feature interaction network to obtain an interaction feature, wherein the interaction feature comprises an interaction feature of a human joint point in the spatial dimension and an interaction feature of a human joint point in the temporal dimension, and the target input feature is determined based on the observation information sequence. However, Zhang discloses: inputting the initial human joint point feature and a target input feature into a pre-trained feature interaction network to obtain an interaction feature, wherein the interaction feature comprises an interaction feature of a human joint point in the spatial dimension and an interaction feature of a human joint point in the temporal dimension, and the target input feature is determined based on the observation information sequence (Abstract, pages 3-4, Fig. 2). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura to use inputting the initial human joint point feature and a target input feature into a pre-trained feature interaction network to obtain an interaction feature, wherein the interaction feature comprises an interaction feature of a human joint point in the spatial dimension and an interaction feature of a human joint point in the temporal dimension, and the target input feature is determined based on the observation information sequence as taught by Zhang. The motivation for doing so would have been in order to determining spatial temporal correlation and learn different motion trajectories of body joints (Zhang, page 1). 8. Regarding claim 6, Jiang in view of Nakamura disclose the method of claim 5, wherein the performing the feature interaction in the spatial dimension and/or in the temporal dimension based on the initial human joint point feature as disclosed above. Jiang further discloses spatio-temporal representations of multiple pose hypotheses to predict 3D human pose from monocular videos (page 4). Further, Nakamura discloses initial human joint point feature ([0075], [0096], [0101]). Jiang in view of Nakamura does not disclose: wherein the feature interaction network comprises at least two first coding layers and at least two second coding layers, the first coding layer is configured to perform feature interaction in the spatial dimension, and the second coding layer is configured to perform feature interaction in the temporal dimension; and the inputting the initial human joint point feature and the target input feature into the pre-trained feature interaction network to obtain the interaction feature comprises: inputting the initial human joint point feature and the target input feature into alternately arranged first coding layer and second coding layer to obtain the interaction feature. However, Zhang discloses: wherein the feature interaction network comprises at least two first coding layers and at least two second coding layers, the first coding layer is configured to perform feature interaction in the spatial dimension, and the second coding layer is configured to perform feature interaction in the temporal dimension (Abstract, pages 3-4, Fig. 2: MixSTE (Mixed Spatio-Temporal Encoder), which has a temporal transformer block to separately model the temporal motion of each joint and a spatial transformer block to learn inter-joint spatial correlation); and the inputting the initial human joint point feature and the target input feature into the pre-trained feature interaction network to obtain the interaction feature comprises: inputting the initial human joint point feature and the target input feature into alternately arranged first coding layer and second coding layer to obtain the interaction feature (Abstract, pages 3-4, Fig. 2). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura to use wherein the feature interaction network comprises at least two first coding layers and at least two second coding layers, the first coding layer is configured to perform feature interaction in the spatial dimension, and the second coding layer is configured to perform feature interaction in the temporal dimension; and the inputting the initial human joint point feature and the target input feature into the pre-trained feature interaction network to obtain the interaction feature comprises: inputting the initial human joint point feature and the target input feature into alternately arranged first coding layer and second coding layer to obtain the interaction feature as taught by Zhang. The motivation for doing so would have been in order to determining spatial temporal correlation and learn different motion trajectories of body joints (Zhang, page 1). 5. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Jiang, in view of Nakamura, in further view of Nie et al. “Monocular 3D Human Pose Estimation by Predicting Depth on Joints”, 2017 (hereinafter, Nie). 8. Regarding claim 9, Jiang in view of Nakamura disclose the method of claim 1, wherein determining the initial human joint point feature based on the observation information sequence as disclosed above. Jiang further discloses inputting the observation information sequence into a trained joint point prediction sub-model to obtain the human joint point feature; and performing the feature interaction based on the human joint point feature, and estimating the human pose with the interaction feature to obtain the human pose information (pages 6-7, 9, sections 3.3, 4.1, Fig. 2). Further, Nakamura discloses initial human joint point feature ([0075], [0096], [0101]). Jiang in view of Nakamura does not disclose: inputting the observation information sequence into a pre-trained joint point prediction sub-model to obtain the human joint point feature; and performing the feature interaction based on the human joint point feature, and estimating the human pose with the interaction feature to obtain the human pose information comprises: inputting the human joint point feature into a pre-trained pose estimation sub-model to obtain the human pose information (Abstract, pages 3471, section 4). However, Nie discloses: initial human joint point feature ([0075], [0096], [0101]). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura to use inputting the observation information sequence into a pre-trained joint point prediction sub-model to obtain the human joint point feature; and performing the feature interaction based on the human joint point feature, and estimating the human pose with the interaction feature to obtain the human pose information comprises: inputting the human joint point feature into a pre-trained pose estimation sub-model to obtain the human pose information as taught by Nie. The motivation for doing so would have been in order to estimate joint depth and reconstructing of a human pose (Nie, page 3472). 5. Claims 10-14 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang, in view of Nakamura, in view of Nie, in further view of Madadi et al. “SMPLR: Deep SMPL reverse for 3D human pose and shape recover”, 2019 (hereinafter, Madadi). 8. Regarding claim 10, Jiang in view of Nakamura in view of Nie disclose the method of claim 9, as disclosed above. Jiang further discloses determining a loss value with a preset loss function based on the human pose information (page 8: Loss Function). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie does not disclose: adjusting, with the loss value, a model parameter of the joint point prediction sub-model and a model parameter of the pose estimation sub-model, to obtain the adjusted joint point prediction sub-model and the adjusted pose estimation sub-model. However, Madadi discloses: adjusting, with the loss value, a model parameter of the joint point prediction sub-model and a model parameter of the pose estimation sub-model, to obtain the adjusted joint point prediction sub-model and the adjusted pose estimation sub-model (pages 5-6, Table 1: Loss function: fine-tune the network adding L1 loss on the SMPL output. See also pages 7-8, section 4.4.3 and Pose estimation). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie to use adjusting, with the loss value, a model parameter of the joint point prediction sub-model and a model parameter of the pose estimation sub-model, to obtain the adjusted joint point prediction sub-model and the adjusted pose estimation sub-model as taught by Madadi. The motivation for doing so would have been in order to adjust the model and estimate pose accurately (Madadi, pages 7-8). 8. Regarding claim 11, Jiang in view of Nakamura in view of Nie in view of Madadi disclose the method of claim 10, as disclosed above. Jiang further discloses determining the loss value with the preset loss function based on the human pose information (page 8: Loss Function); and determining a difference between the relative rotation angle of the human joint point (page 8 and Fig. 2). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie does not disclose: wherein the human pose information comprises a relative rotation angle of a human joint point under a human parameterized grid model, and the tag pose information comprises a tag relative rotation angle; and determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the relative rotation angle of the human joint point under the human parameterized grid model and the tag relative rotation angle. However, Madadi discloses: wherein the human pose information comprises a relative rotation angle of a human joint point under a human parameterized grid model, and the tag pose information comprises a tag relative rotation angle; and determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the relative rotation angle of the human joint point under the human parameterized grid model and the tag relative rotation angle (page 4: a set of relative rotation matrices are computed for each joint with respect wo their parents…pages 5-6, Table 1: Loss function: fine-tune the network adding L1 loss on the SMPL output. See also pages 7-8, section 4.4.3 and Pose estimation). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie to use wherein the human pose information comprises a relative rotation angle of a human joint point under a human parameterized grid model, and the tag pose information comprises a tag relative rotation angle; and determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the relative rotation angle of the human joint point under the human parameterized grid model and the tag relative rotation angle as taught by Madadi. The motivation for doing so would have been in order to adjust the model and estimate pose accurately (Madadi, pages 7-8). 8. Regarding claim 12, Jiang in view of Nakamura in view of Nie in view of Madadi disclose the method of claim 10, wherein the human pose information comprises a joint point coordinate of a human joint point under a human parameterized grid model, and the tag pose information comprises a tag joint point coordinate as disclosed above. Jiang further discloses determining the loss value with the preset loss function based on the human pose information (page 8: Loss Function); and determining a difference between the joint point coordinate (page 8 and Fig. 2). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie does not disclose: determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the joint point coordinate of the human joint point under the human parameterized grid model and the tag joint point coordinate. However, Madadi discloses: determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the joint point coordinate of the human joint point under the human parameterized grid model and the tag joint point coordinate (page 4: a set of relative rotation matrices are computed for each joint with respect wo their parents…pages 5-6, Table 1: Loss function: fine-tune the network adding L1 loss on the SMPL output. See also pages 7-8, section 4.4.3 and Pose estimation). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie to use determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the joint point coordinate of the human joint point under the human parameterized grid model and the tag joint point coordinate as taught by Madadi. The motivation for doing so would have been in order to adjust the model and estimate pose accurately (Madadi, pages 7-8). 8. Regarding claim 13, Jiang in view of Nakamura in view of Nie in view of Madadi disclose the method of claim 10, as disclosed above. Jiang further discloses wherein the human pose information comprises a joint point coordinate of a hand joint point in a world coordinate system (Abstract), determining the loss value with the preset loss function based on the human pose information (page 8: Loss Function); and determining a difference between the joint point coordinate (page 8 and Fig. 2). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie does not disclose: wherein the human pose information comprises the tag pose information comprises a tag hand joint point coordinate; and determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the joint point coordinate of the hand joint point in the world coordinate system and the tag hand joint point coordinate. However, Madadi discloses: wherein the human pose information comprises the tag pose information comprises a tag hand joint point coordinate; and determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the joint point coordinate of the hand joint point in the coordinate system and the tag hand joint point coordinate (page 4: a set of relative rotation matrices are computed for each joint with respect wo their parents…pages 5-6, Table 1: Loss function: fine-tune the network adding L1 loss on the SMPL output. See also pages 7-8, section 4.4.3 and Pose estimation). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie to use wherein the human pose information comprises the tag pose information comprises a tag hand joint point coordinate; and determining the loss value with the preset loss function based on the human pose information and the tag pose information comprises: determining, as the loss value, a difference between the joint point coordinate of the hand joint point in the coordinate system and the tag hand joint point coordinate as taught by Madadi. The motivation for doing so would have been in order to adjust the model and estimate pose accurately (Madadi, pages 7-8). 8. Regarding claim 14, Jiang in view of Nakamura in view of Nie in view of Madadi disclose the method of claim 10, as disclosed above. Jiang further discloses wherein the human pose information comprises a movement velocity of a human joint point within a preset duration, and the tag pose information comprises a tag movement velocity, determining the loss value with the preset loss function based on the human pose information (page 8: Loss Function), and determining, as the loss value, a difference between the movement velocity of the human joint point within the preset duration and the tag movement velocity (pages 5-8, Fig. 2). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie does not disclose: determining the loss value with the preset loss function based on the human pose information and the tag pose information. However, Madadi discloses: determining the loss value with the preset loss function based on the human pose information and the tag pose information (pages 5-6, Table 1: Loss function: fine-tune the network adding L1 loss on the SMPL output. See also pages 7-8, section 4.4.3 and Pose estimation). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie to use determining the loss value with the preset loss function based on the human pose information and the tag pose information as taught by Madadi. The motivation for doing so would have been in order to adjust the model and estimate pose accurately (Madadi, pages 7-8). 5. Claims 15-17 are rejected under 35 U.S.C. 103 as being unpatentable over Jiang, in view of Nakamura, in view of Nie, in view of Madadi, in further view of Shi et al. “MotioNet: 3D Human Motion Reconstruction from Monocular Video with Skeleton Consistency”, 2020 (hereinafter, Shi). 8. Regarding claim 15, Jiang in view of Nakamura in view of Nie in view of Madadi disclose the method of claim 14, as disclosed above. Jiang further discloses determining a difference between the movement velocity of the human joint point within the preset duration and the tag movement velocity (pages 5-8, Fig. 2). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie in view of Madadi does not disclose: wherein the human joint point comprises a foot joint point; and determining, as the loss value, the difference between the movement velocity of the human joint point within the preset duration and the tag movement velocity comprises: if a foot is placed on the ground within the preset duration, determining, as the loss value, a difference between the movement velocity of the foot joint point within the preset duration and a target velocity. However, Shi discloses: wherein the human joint point comprises a foot joint point; and determining, as the loss value, the difference between the movement velocity of the human joint point within the preset duration and the tag movement velocity comprises: if a foot is placed on the ground within the preset duration, determining, as the loss value, a difference between the movement velocity of the foot joint point within the preset duration and a target velocity (pages 1:5-1:7, 1:12, and Figs. 11, 12). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie in view of Madadi to use wherein the human joint point comprises a foot joint point; and determining, as the loss value, the difference between the movement velocity of the human joint point within the preset duration and the tag movement velocity comprises: if a foot is placed on the ground within the preset duration, determining, as the loss value, a difference between the movement velocity of the foot joint point within the preset duration and a target velocity as taught by Shi. The motivation for doing so would have been in order to account for foot contact during pose estimation (Shi, page 1:7). 8. Regarding claim 16, Jiang in view of Nakamura in view of Nie in view of Madadi disclose the method of claim 10, wherein determining the loss value with the preset loss function based on the human pose information and the tag pose information as disclosed above. Jiang further discloses loss function (page 8). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie in view of Madadi does not disclose: determining, with the human pose information, whether there is a joint point lower than a ground height in human joint points; and in response to that there is a joint point lower than the ground height, determining, as the loss value, a difference between a height of the lowest point in the human joint points lower than the ground and the ground height. However, Shi discloses: determining, with the human pose information, whether there is a joint point lower than a ground height in human joint points; and in response to that there is a joint point lower than the ground height, determining, as the loss value, a difference between a height of the lowest point in the human joint points lower than the ground and the ground height (pages 1:5, 1:7 (1st col), 1:12, and Fig. 12). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie in view of Madadi to use determining, with the human pose information, whether there is a joint point lower than a ground height in human joint points; and in response to that there is a joint point lower than the ground height, determining, as the loss value, a difference between a height of the lowest point in the human joint points lower than the ground and the ground height as taught by Shi. The motivation for doing so would have been in order to account for foot contact during pose estimation (Shi, page 1:7). 8. Regarding claim 17, Jiang in view of Nakamura in view of Nie in view of Madadi disclose the method of claim 10, wherein determining the loss value with the preset loss function based on the human pose information and the tag pose information as disclosed above. Jiang further discloses loss function (page 8). Further, Nie discloses loss function (pages 3471, 3473, sections 3.1 and 4.3). Jiang in view of Nakamura in view of Nie in view of Madadi does not disclose: in response to that there is a joint point lower than the ground height, determining, as the loss value, a difference between a height of the lowest point in the human joint points lower than the ground and the ground height. However, Shi discloses: in response to that there is a joint point lower than the ground height, determining, as the loss value, a difference between a height of the lowest point in the human joint points lower than the ground and the ground height (pages 1:5, 1:7 (1st col), 1:12, and Fig. 12). Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Jiang in view of Nakamura in view of Nie in view of Madadi to use in response to that there is a joint point lower than the ground height, determining, as the loss value, a difference between a height of the lowest point in the human joint points lower than the ground and the ground height as taught by Shi. The motivation for doing so would have been in order to account for foot contact during pose estimation (Shi, page 1:7). Conclusion 28. Examiner has cited particular columns and line numbers, and/or paragraphs, and/or pages in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. In the case of amending the claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. 29. Any inquiry concerning this communication or earlier communications from the examiner should be directed to EYOB HAGOS whose telephone number is (571)272-3508. The examiner can normally be reached on 8:30-5:30PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Shelby Turner can be reached on 571-272-6334. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Eyob Hagos/ Primary Examiner, Art Unit 2857
Read full office action

Prosecution Timeline

Jul 19, 2024
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12741682
TRAIN SPEED ESTIMATION DEVICE AND METHOD BASED ON VIBRATION SIGNALS
4y 12m to grant Granted Sep 22, 2026
Patent 12719855
SYSTEMS AND METHODS FOR GENERATING A SNAPSHOT VIEW OF NETWORK INFRASTRUCTURE
2y 4m to grant Granted Aug 25, 2026
Patent 12708289
Waist swinging estimation device, estimation system, waist swinging estimation method, and recording medium
3y 2m to grant Granted Aug 18, 2026
Patent 12704652
FAULTED SEISMIC HORIZON MAPPING
3y 10m to grant Granted Aug 11, 2026
Patent 12699205
LOCATION-BASED FORECASTING OF WEATHER EVENTS BASED ON USER IMPACT
2y 2m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
66%
Grant Probability
99%
With Interview (+43.1%)
3y 4m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 404 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month