Prosecution Insights
Last updated: October 02, 2026
Application No. 18/640,845

HUMAN BODY MOTION CAPTURE METHOD AND APPARATUS, DEVICE, MEDIUM, AND PROGRAM

Final Rejection §103
Filed
Apr 19, 2024
Priority
Apr 26, 2023 — CN 202310465797.1
Examiner
LANTZ, KARSTEN FOSTER
Art Unit
2664
Tech Center
2600 — Communications
Assignee
Beijing Zitiao Network Technology Co., Ltd.
OA Round
2 (Final)
100%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
8 granted / 8 resolved
+38.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
24 currently pending
Career history
33
Total Applications
across all art units

Statute-Specific Performance

§101
1.8%
-38.2% vs TC avg
§103
83.3%
+43.3% vs TC avg
§102
7.9%
-32.1% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 8 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 6/30/2026 have been fully considered but they are not persuasive. Claims 1,4-6,9-14,16 and 18-26 are pending in this application and have been considered below. Claims 2, 3, 7, 8, 15, and 17 are canceled by the applicant. Argument: The applicant argues Nan does not teach recognizing human body motion or inputting different images to the multiple networks. Response: It is not necessary that Nan expressly disclose the application of human body motion recognition, because it teaches the underlying structural capability. Nan explicitly discloses a shared network backbone that processes input data and distributes it to a plurality of sub-networks "each of the sub-networks 304A-C may perform an individual task." A person of ordinary skill in the art would readily recognize that such a flexible, multi-task architecture utilizing shared image inputs is inherently adaptable to motion recognition tasks without requiring a specific redesign or inventive effort. Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The IDS dated 7/17/2024 that has been previously considered remains placed in the application file. 1st Claim Rejections - 35 USC § 103 Claims 1, 6, 11, 12, 16, 18, 19, 20, 23, 24, and 25 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2024 0215866 A1, (Demaster-Smith et al.) in view of US Patent No. 11,507,203 B1, (Bosworth) and US Patent Publication 2024 0193464 A1, (Nan et al.). Claim 1 Regarding claim 1, Demaster-Smith et al. teach a human body motion capture method, comprising: obtaining a human body image shot ("using image data from a camera, such as a forward-facing camera and/or a downward-facing camera, to infer body posture based on how much of the user's body is in view of the camera," par. 37) by the headset ("the use of cameras onboard the head-mounted device," par. 14); determining motion information of a key node of a human body ("one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) based on the human body image, ("the use of a downward-facing camera to estimate body posture based on image data," par. 28) the motion information of the key node comprising position information or pose information of the key node ("Anatomical points can include estimated points in three-dimensional space associated with a user's anatomy," par. 25); determining motion information of a head based on the pose information of the headset, the motion information of the head comprising position information or pose information of the head; ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information is estimated accordingly. For example, one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) and motion information of the key node and the motion information of the head ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information," par. 27). inputting the human body image to a deep learning network to obtain the motion information of the key node of the human body, wherein inputting the human body image to the deep learning network to obtain the motion information of the key node of the human body comprises: inputting a first human body image in the human body image to the first deep learning subnetwork to obtain motion information and inputting a second human body image in the human body image to the second deep learning subnetwork to obtain motion information ("The body posture estimation process starts with a plurality of sensor data 204 as input to the artificial neural network 202 … the sensor data 204 can include inertial measurement data from at least one IMU sensor and/or image data from at least one camera," par. 22). Demaster-Smith et al. do not explicitly teach all of obtaining pose information of a headset; the key node comprising one or more of a hand key node, a foot key node, or a waist key node; determining, by using an inverse kinematics method, posture information of the human body, wherein determining motion information of the key node of the human body based on the human body image comprises: wherein the deep learning network comprises a first deep learning subnetwork and a second deep learning subnetwork, and wherein the first human body image and the second human body image are different. However, Bosworth teaches obtaining pose information of a headset; ("the system may determine a pose of a headset worn by the user based on sensor data captured by the headset," col. 18, line 46) the key node comprising one or more of a hand key node, a foot key node, or a waist key node; ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc., based on the 3D position of the controller camera, the camera's intrinsic/extrinsic parameters, and the images captured by the camera," col. 2, line 1) and determining, by using an inverse kinematics method, posture information of the human body ("The inverse kinematic optimizer 515 may fit the input keypoints based on the muscular-skeletal model 514 to determine if any key point positions or key point relationships are not complying with the muscular-skeletal model and to make adjustment to accordingly to determine the optimal body pose of the user. The muscular-skeletal model 514 may include a number of constraints limiting the possible body pose of the user and these constraints may be applied by the inverse kinematic optimizer 515. As a result, the refined full body pose 516 may provide more accurate body pose estimation results than the initial full body pose," col. 15, line 63) wherein determining motion information of the key node of the human body based on the human body image comprises: ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc., based on the 3D position of the controller camera, the camera's intrinsic/extrinsic parameters, and the images captured by the camera," col. 2, line 1) a hand key node ("the system may accurately determine at least three keypoints 241, 242A, and 242B associated with the user's head and hands," col. 7, line 24) and a foot key node ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc.," col. 2, line 1), and wherein the first human body image and the second human body image are different ("The system may or may not able to accurately determine the 3D positions of the keypoints based on a single image captured by a single controller camera, but can accurately determine the 3D positions of the keypoints based on the multiple images captured by the multiple controller cameras from different perspectives. In particular embodiments, the system may feed the captured images of the user's body parts to a neural network to determine the corresponding keypoints," col. 11, line 8). Therefore, taking the teachings of Demaster-Smith et al. and Bosworth as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify body posture estimation system as taught by Demaster-Smith et al. to use inverse kinematic optimizer and simultaneous localization and mapping techniques as taught by Bosworth. The suggestion/motivation for doing so would have been that, “the initial full body pose 513 may provide a rough estimation for the user's body pose and may not be perfectly accurate. The system may feed the initial full body pose 513 to an inverse-kinematic optimizer to refine and optimize the results … the headset 210 may include IMUS and cameras (e.g., 211 and 213) which can be used to perform simultaneous localization and mapping (SLAM) for self-localization. Thus, the headset 210 may be used to accurately determine the head position (e.g., as represented by the key point 241) of the user 201 (taking into consideration of the relative position of the headset 210 and the head of the user” as noted by the Bosworth disclosure in paragraphs [43 and 24], which also motivates combination because the combination would predictably have a higher accuracy as there is a reasonable expectation that the inverse kinematic optimizer will accurately convert estimated, rough 3D body pose data into refined joint positions, and that the SLAM techniques will enable precise spatial localization of the headset camera relative to the user's environment, thereby overcoming limitations of the initial rough estimations; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Additionally, Nan et al. teach wherein the deep learning network comprises a first deep learning subnetwork and a second deep learning subnetwork ("one shared backbone network plus multiple subnetworks," par. 16). Therefore, taking the teachings of Demaster-Smith et al., Bosworth, and Nan et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify body posture estimation system as taught by Demaster-Smith et al. and headset capturing and keypoint estimation techniques as taught by Bosworth to use the shared network backbone as taught by Nan et al. The suggestion/motivation for doing so would have been that, “This lightweight network framework can effectively complete multiple deep-learning tasks, such as object detection, image classification and other tasks, in a very small size with faster speed, which can be applied to most mobile use cases” as noted by the Nan et al. disclosure in paragraph [16], which also motivates combination because the combination would predictably improve efficiency as there is a reasonable expectation that applying a shared backbone to both posture estimation and headset keypoint detection would significantly reduce computational redundancy, enhance real-time performance on the headset, and improve the accuracy of keypoint localization by leveraging shared features between the two related tasks; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. The rejection of method claim 1 above applies mutatis mutandis to the corresponding limitations of apparatus claim 18 and device claim 19 while noting that the rejection above cites to both device and method disclosures. Claims 18 and 19 are mapped below for clarity of the record and to specify any new limitations not included in claim 1. Claim 6 Regarding claim 6, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 1 as noted above. Demaster-Smith et al. teach wherein when the motion information of the key node comprises the position information of the key node, ("one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) and the motion information of the head comprises the position information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18), the motion information of the key node and the motion information of the head ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information," par. 27), the position information of the key node and the position information of the head; and the posture information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18). Demaster-Smith et al. do not explicitly teach all of the determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; and determining posture information of another node of the human body based on the posture information of the key node. However, Bosworth teaches the determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; ("The inverse kinematic optimizer 515 may fit the input keypoints based on the muscular-skeletal model 514 to determine if any key point positions or key point relationships are not complying with the muscular-skeletal model and to make adjustment to accordingly to determine the optimal body pose of the user. The muscular-skeletal model 514 may include a number of constraints limiting the possible body pose of the user and these constraints may be applied by the inverse kinematic optimizer 515. As a result, the refined full body pose 516 may provide more accurate body pose estimation results than the initial full body pose," col. 15, line 63) and determining posture information of another node of the human body ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6) based on the posture information of the key node ("the system may use a muscular-skeletal model of human body to (1) infer the positions of the user's body keypoints based on other keypoints," col. 13, line 25). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Claim 11 Regarding claim 11, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 1 as noted above. Demaster-Smith et al. teach computing a loss of the deep learning network based on the label of the human body sample image and estimated motion information output by the deep learning network; ("The artificial neural network then computes a loss value using the target output and the predicted body posture based on a loss function," par. 21) and updating a parameter of the deep learning network based on the loss of the deep learning network ("The computed loss can then be used to adjust parameters and/or weights of the artificial neural network," par. 21). Demaster-Smith et al. do not explicitly teach all of training the deep learning network with a human body sample image, a label of the human body sample image being actual motion information of the key node, and the actual motion information of the key node comprising actual position information or actual pose information of the key node. However, Bosworth teaches training the deep learning network with a human body sample image, ("a ML model that is trained to extract keypoints and determine the 3D positions for these keypoints based in input images," col. 14, line 23) a label of the human body sample image being actual motion information of the key node, and the actual motion information of the key node comprising actual position information or actual pose information of the key node ("the system may first determine all the keypoints of the user's body and use a subset of the known keypoints as the input training samples and another subset of the known keypoints as the ground truth to train the ML model," col. 12, line 9). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Bosworth are combined as per claim 1. Claim 12 Regarding claim 12, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 11 as noted above. Demaster-Smith et al. teach wherein the computing a loss of the deep learning network based on the label of the human body sample image and estimated motion information output by the deep learning network comprises: ("The artificial neural network then computes a loss value using the target output and the predicted body posture based on a loss function," par. 21) computing the loss of the deep learning network based on the label of the human body sample image and the estimated motion information output by the deep learning network ("The artificial neural network then computes a loss value using the target output and the predicted body posture based on a loss function," par. 21). Demaster-Smith et al. do not explicitly teach all of when the key node is visible in the human body sample image. Claim 16 Regarding claim 16, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 1 as noted above. Demaster-Smith et al. do not explicitly teach all of wherein the headset obtains the pose information of the headset by using a visual simultaneous localization and mapping (SLAM) method. However, Bosworth teaches wherein the headset obtains the pose information of the headset by using a visual simultaneous localization and mapping (SLAM) method ("The pose of the headset may include a position and two axis directions of the headset within the three-dimensional space. In particular embodiments, the sensor data captured by the controller may include inertial measurement unit (IMU) data. The pose of the controller may be determined using simultaneous localization and mapping (SLAM) for self-localization," col. 18, line 58). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. However, Bosworth teaches when the key node is visible in the human body sample image ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Claim 18 Regarding claim 18, Demaster-Smith et al. teach a human body motion capture apparatus, comprising: a second obtaining module configured to obtain a human body image shot ("using image data from a camera, such as a forward-facing camera and/or a downward-facing camera, to infer body posture based on how much of the user's body is in view of the camera," par. 37) by the headset ("the use of cameras onboard the head-mounted device," par. 14); a first position determining module configured to determine motion information of a key node of a human body ("one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) based on the human body image, ("the use of a downward-facing camera to estimate body posture based on image data," par. 28) the motion information of the key node comprising position information or pose information of the key node ("Anatomical points can include estimated points in three-dimensional space associated with a user's anatomy," par. 25); a second position determining module configured to determine motion information of a head based on the pose information of the headset, the motion information of the head comprising position information or pose information of the head; ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information is estimated accordingly. For example, one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) and motion information of the key node and the motion information of the head ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information," par. 27) input the human body image to a deep learning network to obtain the motion information of the key node of the human body, wherein when performing the step of inputting the human body image to the deep learning network to obtain the motion information of the key node of the human body comprises: input a first human body image in the human body image to the first deep learning subnetwork to obtain motion information and input a second human body image in the human body image to the second deep learning subnetwork to obtain motion information ("The body posture estimation process starts with a plurality of sensor data 204 as input to the artificial neural network 202 … the sensor data 204 can include inertial measurement data from at least one IMU sensor and/or image data from at least one camera," par. 22). Demaster-Smith et al. do not explicitly teach all of a first obtaining module configured to obtain pose information of a headset; the key node comprising one or more of a hand key node, a foot key node, or a waist key node; a posture estimation module configured to determine, by using an inverse kinematics method, posture information of the human body, wherein when performing the step of determining motion information of the key node of the human body based on the human body image, the first position determining module is configured to: wherein the deep learning network comprises a first deep learning subnetwork and a second deep learning subnetwork, and wherein the first human body image and the second human body image are different. However, Bosworth teaches a first obtaining module configured to obtain pose information of a headset; ("the system may determine a pose of a headset worn by the user based on sensor data captured by the headset," col. 18, line 46) the key node comprising one or more of a hand key node, a foot key node, or a waist key node; ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc., based on the 3D position of the controller camera, the camera's intrinsic/extrinsic parameters, and the images captured by the camera," col. 2, line 1) and a posture estimation module configured to determine, by using an inverse kinematics method, posture information of the human body ("The inverse kinematic optimizer 515 may fit the input keypoints based on the muscular-skeletal model 514 to determine if any key point positions or key point relationships are not complying with the muscular-skeletal model and to make adjustment to accordingly to determine the optimal body pose of the user. The muscular-skeletal model 514 may include a number of constraints limiting the possible body pose of the user and these constraints may be applied by the inverse kinematic optimizer 515. As a result, the refined full body pose 516 may provide more accurate body pose estimation results than the initial full body pose," col. 15, line 63) wherein when performing the step of determining motion information of the key node of the human body based on the human body image, the first position determining module is configured to: ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc., based on the 3D position of the controller camera, the camera's intrinsic/extrinsic parameters, and the images captured by the camera," col. 2, line 1) a hand key node ("the system may accurately determine at least three keypoints 241, 242A, and 242B associated with the user's head and hands," col. 7, line 24) and a foot key node ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc.," col. 2, line 1), and wherein the first human body image and the second human body image are different ("The system may or may not able to accurately determine the 3D positions of the keypoints based on a single image captured by a single controller camera, but can accurately determine the 3D positions of the keypoints based on the multiple images captured by the multiple controller cameras from different perspectives. In particular embodiments, the system may feed the captured images of the user's body parts to a neural network to determine the corresponding keypoints," col. 11, line 8). Additionally, Nan et al. teach wherein the deep learning network comprises a first deep learning subnetwork and a second deep learning subnetwork ("one shared backbone network plus multiple subnetworks," par. 16). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Claim 19 Regarding claim 19, Demaster-Smith et al. teach an electronic device, comprising: a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to invoke and run the computer program stored in the memory, to perform: ("the computing system 100 includes a computing device 102 that further includes a processor 104 (e.g., central processing units, or “CPUs”), an input/output (I/O) module 106, volatile memory 108, and non-volatile memory 110. The different components are operatively coupled to one another. The non-volatile memory 110 stores a body posture estimation program 112, which contains instructions for the various software modules described herein for execution by the processor," par. 16) obtaining a human body image shot ("using image data from a camera, such as a forward-facing camera and/or a downward-facing camera, to infer body posture based on how much of the user's body is in view of the camera," par. 37) by the headset ("the use of cameras onboard the head-mounted device," par. 14); determining motion information of a key node of a human body ("one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) based on the human body image, ("the use of a downward-facing camera to estimate body posture based on image data," par. 28) the motion information of the key node comprising position information or pose information of the key node ("Anatomical points can include estimated points in three-dimensional space associated with a user's anatomy," par. 25); determining motion information of a head based on the pose information of the headset, the motion information of the head comprising position information or pose information of the head; ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information is estimated accordingly. For example, one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) and motion information of the key node and the motion information of the head ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information," par. 27). inputting the human body image to a deep learning network to obtain the motion information of the key node of the human body, wherein inputting the human body image to the deep learning network to obtain the motion information of the key node of the human body comprises: inputting a first human body image in the human body image to the first deep learning subnetwork to obtain motion information and inputting a second human body image in the human body image to the second deep learning subnetwork to obtain motion information ("The body posture estimation process starts with a plurality of sensor data 204 as input to the artificial neural network 202 … the sensor data 204 can include inertial measurement data from at least one IMU sensor and/or image data from at least one camera," par. 22). Demaster-Smith et al. do not explicitly teach all of obtaining pose information of a headset; the key node comprising one or more of a hand key node, a foot key node, or a waist key node; determining, by using an inverse kinematics method, posture information of the human body, wherein determining motion information of the key node of the human body based on the human body image comprises: wherein the deep learning network comprises a first deep learning subnetwork and a second deep learning subnetwork, and wherein the first human body image and the second human body image are different. However, Bosworth teaches obtaining pose information of a headset; ("the system may determine a pose of a headset worn by the user based on sensor data captured by the headset," col. 18, line 46) the key node comprising one or more of a hand key node, a foot key node, or a waist key node; ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc., based on the 3D position of the controller camera, the camera's intrinsic/extrinsic parameters, and the images captured by the camera," col. 2, line 1) and determining, by using an inverse kinematics method, posture information of the human body ("The inverse kinematic optimizer 515 may fit the input keypoints based on the muscular-skeletal model 514 to determine if any key point positions or key point relationships are not complying with the muscular-skeletal model and to make adjustment to accordingly to determine the optimal body pose of the user. The muscular-skeletal model 514 may include a number of constraints limiting the possible body pose of the user and these constraints may be applied by the inverse kinematic optimizer 515. As a result, the refined full body pose 516 may provide more accurate body pose estimation results than the initial full body pose," col. 15, line 63) wherein determining motion information of the key node of the human body based on the human body image comprises: ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc., based on the 3D position of the controller camera, the camera's intrinsic/extrinsic parameters, and the images captured by the camera," col. 2, line 1) a hand key node ("the system may accurately determine at least three keypoints 241, 242A, and 242B associated with the user's head and hands," col. 7, line 24) and a foot key node ("The controller may determine the 3D locations of the keypoints related to knees, legs, feet, etc.," col. 2, line 1), and wherein the first human body image and the second human body image are different ("The system may or may not able to accurately determine the 3D positions of the keypoints based on a single image captured by a single controller camera, but can accurately determine the 3D positions of the keypoints based on the multiple images captured by the multiple controller cameras from different perspectives. In particular embodiments, the system may feed the captured images of the user's body parts to a neural network to determine the corresponding keypoints," col. 11, line 8). Additionally, Nan et al. teach wherein the deep learning network comprises a first deep learning subnetwork and a second deep learning subnetwork ("one shared backbone network plus multiple subnetworks," par. 16). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Claim 20 Regarding claim 20, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 1 as noted above. Demaster-Smith et al. teach a computer-readable storage medium, configured to store a computer program, wherein the computer program causes a computer to perform ("such methods and processes may be implemented as a computer-application program or service," par. 42). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Claim 23 Regarding claim 23, Demaster-Smith et al., Bosworth, and Nan et al. teach the electronic device according to claim 19 as noted above. Demaster-Smith et al. teach wherein when the motion information of the key node comprises the position information of the key node, ("one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) and the motion information of the head comprises the position information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18), the motion information of the key node and the motion information of the head, the processor is configured to perform: ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information," par. 27), the position information of the key node and the position information of the head; and the posture information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18). Demaster-Smith et al. do not explicitly teach all of when performing the step of determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; and determining posture information of another node of the human body based on the posture information of the key node. However, Bosworth teaches the when performing the step of determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; ("The inverse kinematic optimizer 515 may fit the input keypoints based on the muscular-skeletal model 514 to determine if any key point positions or key point relationships are not complying with the muscular-skeletal model and to make adjustment to accordingly to determine the optimal body pose of the user. The muscular-skeletal model 514 may include a number of constraints limiting the possible body pose of the user and these constraints may be applied by the inverse kinematic optimizer 515. As a result, the refined full body pose 516 may provide more accurate body pose estimation results than the initial full body pose," col. 15, line 63) and determining posture information of another node of the human body ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6) based on the posture information of the key node ("the system may use a muscular-skeletal model of human body to (1) infer the positions of the user's body keypoints based on other keypoints," col. 13, line 25). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Claim 24 Regarding claim 24, Demaster-Smith et al., Bosworth, and Nan et al. teach the electronic device according to claim 19 as noted above. Demaster-Smith et al. teach computing a loss of the deep learning network based on the label of the human body sample image and estimated motion information output by the deep learning network; ("The artificial neural network then computes a loss value using the target output and the predicted body posture based on a loss function," par. 21) and updating a parameter of the deep learning network based on the loss of the deep learning network ("The computed loss can then be used to adjust parameters and/or weights of the artificial neural network," par. 21). Demaster-Smith et al. do not explicitly teach all of training the deep learning network with a human body sample image, a label of the human body sample image being actual motion information of the key node, and the actual motion information of the key node comprising actual position information or actual pose information of the key node. However, Bosworth teaches training the deep learning network with a human body sample image, ("a ML model that is trained to extract keypoints and determine the 3D positions for these keypoints based in input images," col. 14, line 23) a label of the human body sample image being actual motion information of the key node, and the actual motion information of the key node comprising actual position information or actual pose information of the key node ("the system may first determine all the keypoints of the user's body and use a subset of the known keypoints as the input training samples and another subset of the known keypoints as the ground truth to train the ML model," col. 12, line 9). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. Claim 25 Regarding claim 25, Demaster-Smith et al., Bosworth, and Nan et al. teach the electronic device according to claim 24 as noted above. Demaster-Smith et al. teach wherein the computing a loss of the deep learning network based on the label of the human body sample image and estimated position information output by the deep learning network comprises: ("The artificial neural network then computes a loss value using the target output and the predicted body posture based on a loss function," par. 21) computing the loss of the deep learning network based on the label of the human body sample image and the estimated motion information output by the deep learning network ("The artificial neural network then computes a loss value using the target output and the predicted body posture based on a loss function," par. 21). Demaster-Smith et al. do not explicitly teach all of when the key node is visible in the human body sample image. However, Bosworth teaches when the key node is visible in the human body sample image ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6). Demaster-Smith et al., Bosworth, and Nan et al. are combined as per claim 1. 2nd Claim Rejections - 35 USC § 103 Claims 4, 5, 9, 10, 14, 21, and 22 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2024 0215866 A1, (Demaster-Smith et al.), US Patent No. 11,507,203 B1, (Bosworth), and US Patent Publication 2024 0193464 A1, (Nan et al.) in view of US Patent Publication 2022 0409991 A1, (Ohashi et al.). Claim 4 Regarding claim 4, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 1 as noted above. Demaster-Smith et al. teach wherein the second deep learning subnetwork further outputs motion information ("The body posture estimation process starts with a plurality of sensor data 204 as input to the artificial neural network 202 … the sensor data 204 can include inertial measurement data from at least one IMU sensor and/or image data from at least one camera," par. 22). Demaster-Smith et al. do not explicitly teach all of the waist key node. However, Ohashi et al. teach the waist key node ("the human body model used here includes, as illustrated in FIG. 5, a head node P1, a neck node P2, a chest node P3, a waist node," par. 60). Therefore, taking the teachings of Demaster-Smith et al., Bosworth, Nan et al., and Ohashi et al. as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify body posture estimation system as taught by Demaster-Smith et al., headset capturing and keypoint estimation techniques as taught by Bosworth, and the shared network backbone as taught by Nan et al. to use the waist node as taught by Ohashi et al. The suggestion/motivation for doing so would have been that, “ in a case where the sensor device 20 is not attached to the body part corresponding to the node, the body parts position velocity estimation section 33 estimates the information regarding the position and posture of the node by using the information regarding the position and posture of the sensor device 20 attached to another body part” as noted by the Ohashi et al. disclosure in paragraph [62], which also motivates combination because the combination would predictably have additional utility as there is a reasonable expectation that the system will gain an additional region for tracking movement; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Claim 5 Regarding claim 5, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 1 as noted above. Demaster-Smith et al. teach inputting the human body image to the third deep learning subnetwork to obtain motion information ("The body posture estimation process starts with a plurality of sensor data 204 as input to the artificial neural network 202 … the sensor data 204 can include inertial measurement data from at least one IMU sensor and/or image data from at least one camera," par. 22). Nan et al. teach wherein the deep learning network further comprises a third deep learning subnetwork ("one shared backbone network plus multiple subnetworks," par. 16). Demaster-Smith et al. and Nan et al. do not explicitly teach all of the waist key node. However, Ohashi et al. teach the waist key node ("the human body model used here includes, as illustrated in FIG. 5, a head node P1, a neck node P2, a chest node P3, a waist node," par. 60). Demaster-Smith et al., Bosworth, Nan et al., and Ohashi et al. are combined as per claim 4. Claim 9 Regarding claim 9, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 4 as noted above. Demaster-Smith et al. teach wherein when the motion information of the key node comprises the position information of the key node, ("one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) and the motion information of the head comprises the position information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18), the motion information of the key node and the motion information of the head ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information," par. 27), the position information of the key node and the position information of the head; and the posture information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18). Demaster-Smith et al. do not explicitly teach all of the determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; and determining posture information of another node of the human body based on the posture information of the key node. However, Bosworth teaches the determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; ("The inverse kinematic optimizer 515 may fit the input keypoints based on the muscular-skeletal model 514 to determine if any key point positions or key point relationships are not complying with the muscular-skeletal model and to make adjustment to accordingly to determine the optimal body pose of the user. The muscular-skeletal model 514 may include a number of constraints limiting the possible body pose of the user and these constraints may be applied by the inverse kinematic optimizer 515. As a result, the refined full body pose 516 may provide more accurate body pose estimation results than the initial full body pose," col. 15, line 63) and determining posture information of another node of the human body ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6) based on the posture information of the key node ("the system may use a muscular-skeletal model of human body to (1) infer the positions of the user's body keypoints based on other keypoints," col. 13, line 25). Demaster-Smith et al., Bosworth, Nan et al., and Ohashi et al. are combined as per claim 4. Claim 10 Regarding claim 10, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 5 as noted above. Demaster-Smith et al. teach wherein when the motion information of the key node comprises the position information of the key node, ("one or more calibrated IMUs can be used to provide measurements describing the force, acceleration, and/or angular position experienced by the IMUs, which are associated with a known position relative to the user's frame of reference. This information can be used to infer changes in the positions of the determined anatomical points," par. 27) and the motion information of the head comprises the position information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18), the motion information of the key node and the motion information of the head ("Movement of the anatomical points can be inferred from the user's head movements, and body posture information," par. 27), the position information of the key node and the position information of the head; and the posture information of the head ("the body posture estimation program 112 uses head pose data received from a head tracking system to infer head posture," par. 18). Demaster-Smith et al. do not explicitly teach all of the determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; and determining posture information of another node of the human body based on the posture information of the key node. However, Bosworth teaches the determining, by using an inverse kinematics method, posture information of the human body based on the motion information of the key node, determining, by using the inverse kinematics method, posture information of the key node and posture information of the head; ("The inverse kinematic optimizer 515 may fit the input keypoints based on the muscular-skeletal model 514 to determine if any key point positions or key point relationships are not complying with the muscular-skeletal model and to make adjustment to accordingly to determine the optimal body pose of the user. The muscular-skeletal model 514 may include a number of constraints limiting the possible body pose of the user and these constraints may be applied by the inverse kinematic optimizer 515. As a result, the refined full body pose 516 may provide more accurate body pose estimation results than the initial full body pose," col. 15, line 63) and determining posture information of another node of the human body ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6) based on the posture information of the key node ("the system may use a muscular-skeletal model of human body to (1) infer the positions of the user's body keypoints based on other keypoints," col. 13, line 25). Demaster-Smith et al., Bosworth, Nan et al., and Ohashi et al. are combined as per claim 4. Claim 14 Regarding claim 14, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 1 as noted above. Demaster-Smith et al., Bosworth, and Nan et al. do not explicitly teach all of wherein the motion information of the foot key node is motion information of an ankle of the human body, and/or the motion information of the hand key node is motion information of a wrist of the human body. However, Ohashi et al. teaches wherein the motion information of the foot key node is motion information of an ankle of the human body, and/or the motion information of the hand key node is motion information of a wrist of the human body ("The inverse kinematics model computation section 32, as regards body parts correlated to each other in the information regarding position and posture, such as a hand and a wrist, obtains the posture information regarding either one of them (e.g., wrist) from the information regarding the position and orientation of the other (e.g., hand)," par. 54). Claim 21 Regarding claim 21, Demaster-Smith et al., Bosworth, and Nan et al. teach the electronic device according to claim 19 as noted above. Demaster-Smith et al. teach wherein the second deep learning subnetwork further outputs motion information ("The body posture estimation process starts with a plurality of sensor data 204 as input to the artificial neural network 202 … the sensor data 204 can include inertial measurement data from at least one IMU sensor and/or image data from at least one camera," par. 22). Demaster-Smith et al. do not explicitly teach all of the waist key node. However, Ohashi et al. teach the waist key node ("the human body model used here includes, as illustrated in FIG. 5, a head node P1, a neck node P2, a chest node P3, a waist node," par. 60). Demaster-Smith et al., Bosworth, Nan et al., and Ohashi et al. are combined as per claim 4. Claim 22 Regarding claim 22, Demaster-Smith et al., Bosworth, and Nan et al. teach the electronic device according to claim 19 as noted above. Demaster-Smith et al. teach inputting the human body image to the third deep learning subnetwork to obtain motion information ("The body posture estimation process starts with a plurality of sensor data 204 as input to the artificial neural network 202 … the sensor data 204 can include inertial measurement data from at least one IMU sensor and/or image data from at least one camera," par. 22). Nan et al. teach wherein the deep learning network further comprises a third deep learning subnetwork, and the first position determining module is configured to: ("one shared backbone network plus multiple subnetworks," par. 16). Demaster-Smith et al. and Nan et al. do not explicitly teach all of the waist key node. However, Ohashi et al. teach the waist key node ("the human body model used here includes, as illustrated in FIG. 5, a head node P1, a neck node P2, a chest node P3, a waist node," par. 60). Demaster-Smith et al., Bosworth, Nan et al., and Ohashi et al. are combined as per claim 4. 3rd Claim Rejections - 35 USC § 103 Claims 13 and 26 are rejected under 35 U.S.C. 103 as obvious over US Patent Publication 2024 0215866 A1, (Demaster-Smith et al.), US Patent No. 11,507,203 B1, and US Patent Publication 2024 0193464 A1, (Nan et al.) in view of US Patent Publication 2023 0039549 A1, (Wang). Claim 13 Regarding claim 13, Demaster-Smith et al., Bosworth, and Nan et al. teach the method according to claim 12 as noted above. Bosworth teaches whether the key node is visible ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6). Demaster-Smith et al. and Bosworth do not explicitly teach all of determining a confidence of the key node based on the human body sample image; and determining, based on the confidence of the key node. However, Wang teaches determining a confidence of the key node based on the human body sample image; and determining, based on the confidence of the key node ("when the face images meet preset requirements, the electronic device may acquire an identification code matched with the face images. The preset requirements include key points of a face can be obtained and a confidence of an identification result exceeds a set confidence threshold," par. 137). Therefore, taking the teachings of Demaster-Smith et al., Bosworth, Nan et al., and Wang as a whole, it would have been obvious to a person having ordinary skill in the art before the time of the effective filing date of the claimed invention of the instant application to modify body posture estimation system as taught by Demaster-Smith et al., headset capturing and keypoint estimation techniques as taught by Bosworth, and the shared network backbone as taught by Nan et al. to use the confidence calculation as taught by Wang. The suggestion/motivation for doing so would have been that, “a signal generating module, configured to generate early warning information when it is determined that there is no object matched with the identification code in a designated database” as noted by the Wang disclosure in paragraph [66], which also motivates combination because the combination would predictably have a higher productivity as there is a reasonable expectation that the system would provide a more robust and accurate body posture analysis by automatically generating early warnings when tracking confidence is low or when the detected keypoint does not match sample images; and/or because doing so merely combines prior art elements according to known methods to yield predictable results. Claim 26 Regarding claim 26, Demaster-Smith et al., Bosworth, and Nan et al. teach the electronic device according to claim 25 as noted above. Bosworth teaches whether the key node is visible ("the system may train a ML model to predict keypoints of non-visible body parts based on the keypoints of the visible (or trackable) body parts," col. 12, line 6). Demaster-Smith et al. and Bosworth do not explicitly teach all of wherein the processor is further configured to perform: determining a confidence of the key node based on the human body sample image; and determining, based on the confidence of the key node. However, Wang teaches wherein the processor is further configured to perform determining a confidence of the key node based on the human body sample image; and determining, based on the confidence of the key node ("when the face images meet preset requirements, the electronic device may acquire an identification code matched with the face images. The preset requirements include key points of a face can be obtained and a confidence of an identification result exceeds a set confidence threshold," par. 137). Demaster-Smith et al., Bosworth, Nan et al., and Wang are combined as per claim 13. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Karsten F. Lantz whose telephone number is (571)272-4564. The examiner can normally be reached Monday-Friday 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached on 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Karsten F. Lantz/Examiner, Art Unit 2664 Date: 8/6/2026 /JENNIFER MEHMOOD/Supervisory Patent Examiner, Art Unit 2664
Read full office action

Prosecution Timeline

Apr 19, 2024
Application Filed
Mar 30, 2026
Non-Final Rejection mailed — §103
Jun 30, 2026
Response Filed
Aug 13, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731003
IMAGE DETECTION METHOD BASED ON NEURAL NETWORK MODEL, ELECTRONIC DEVICE, AND STORAGE MEDIUM
3y 11m to grant Granted Sep 08, 2026
Patent 12731227
SYSTEMS, METHODS, STORAGE MEDIUMS FOR IMAGE PROCESSING
2y 9m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 2 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 7m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 8 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month