Prosecution Insights
Last updated: August 17, 2026
Application No. 18/907,042

ELECTRONIC DEVICE AND METHODS FOR OUT OF FIELD OF VIEW HAND TRACKING

Non-Final OA §103
Filed
Oct 04, 2024
Priority
Aug 04, 2023 — IN 202341052639 +1 more
Examiner
MANGIALASCHI, TRACY
Art Unit
Tech Center
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
447 granted / 594 resolved
+15.3% vs TC avg
Strong +27% interview lift
Without
With
+27.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
20 currently pending
Career history
612
Total Applications
across all art units

Statute-Specific Performance

§101
8.6%
-31.4% vs TC avg
§103
55.6%
+15.6% vs TC avg
§102
14.4%
-25.6% vs TC avg
§112
15.8%
-24.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 594 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of the Claims Claims 1-18, as originally filed, are currently pending and have been considered below. Claim Objections Claim 18 is objected to because of the following informalities: Claim 18 recites the limitation “recognizing the hat least one hand gesture” in line 4 of the claim. This limitation appears to contain a typographical error and should recite, i.e., “recognizing the at least one hand gesture.” Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ali Akbarian et al., U.S. Publication No. 2024/0264658, hereinafter, “Akbarian”, and further in view of Yuan, Ye, et al. "GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras." arXiv preprint arXiv:2112.01524v2 (2022), hereinafter, “Yuan”. As per claim 1, Akbarian discloses a method for performing gesture recognition in an electronic device (Akbarian, ¶0025, Given only sparse observations of an articulated entity it is desired to compute body motion of the articulated entity … the body motion is usable for downstream tasks including but not limited to: controlling 3D avatars in mixed-reality applications … and for a variety of applications such as 3D body gesture recognition, computer gaming, mixed-reality, virtual reality and others), comprising: tracking, by the electronic device, a visible trajectory of a hand of a user from a plurality of frames captured by the electronic device (Akbarian, ¶0024, In various examples herein the articulated entity is a person and the capture device is a head mounted display HMD worn by the person. One or more egocentric camera in the HMD capture images of only part of the person due to restricted field of view and occlusions. The person's hands move into and out of the field of view of the egocentric camera; Akbarian, ¶0027, The body motion predictor receives inputs ... In an example the inputs 118 comprise an HMD signal, egocentric images … the inputs comprise at least a pose of a reference joint of an articulated entity for which body motion is to be computed, and an indication of whether a second joint of the articulated entity is observed or unobserved in a current time step; Akbarian, ¶0034, In FIG. 2, where the hand tracking is performed by a HMD worn by the user, it can be seen that at each position of the user from 200 to 208, the hands are in a different position. For example, at position 208, both hands of the user will be in the FoV of the HMD, whilst at position 200, the user's right hand will be in the FoV of the HMD and the user's left hand will be out of the field of view of the HMD; Akbarian, ¶0041, input data 118 to the motion model 102 comprises a pose of a reference joint of the entity and an indication of presence or absence of one or more other joints of the entity in a current observation. In the case that the one or more other joints are observed in the current observation the input data 118 includes the pose of the one or more other joints. In some examples the input data 118 may also comprise per joint changes in joint position between time steps, per joint changes in joint rotation between time steps; Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise); identifying, by the electronic device, the first frame, where the hand of the user has gone out of a Field of View (FOV) of the electronic device, among the plurality of frames (Akbarian, ¶0043, Since some of the observations in the input data 118 may be missing a temporally adaptable mask token mechanism 410 is used. The temporally adaptable mask token mechanism operates in the embedding space of the motion model 102 to generate a mask 412. The mask 412 is either an embedding of a pose of a joint observed at the current time step or a predicted embedding of an unobserved joint, predicted by taking into account the embedding of the reference joint and an embedding of the unobserved joint from a previous time step; Akbarian, ¶0050, FIG. 5 is a schematic diagram of an example of a generative motion model using input data … the person is wearing an HMD with an egocentric camera and the HMD processes the egocentric images to achieve hand tracking to track pose of the hands of the person; Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise. Finally, for all 6-DoF signals, provide the velocity of changes between two consecutive frames; Akbarian, ¶0066, The body motion predictor receives 902 an indication that a second joint of the articulated entity is unobserved or observed. The indication may be a flag such as a binary value. The indication may be received from another entity such as an image recognition system which recognizes particular joints such as hands, feet or other joints in an image; Akbarian, ¶0067, The body motion predictor prompts 910 a motion model which is a trained neural network. The prompt comprises a mask token which is a temporally adaptable mask token; Akbarian, ¶0068, The mask token represents the second joint and is temporally adaptable. In response to receiving an indication that the second joint is unobserved (the negative branch from decision diamond 904), using 906 information about the reference joint pose and a pose of the second joint from a previous time step; Akbarian, ¶0070, In the same time step, operations 902 to 910 of the method may be repeated for a third joint. In response to a determination on whether the third joint is observed, a second mask token is temporally adaptable to perform either operations 906 or 908 for the third joint); identifying, by the electronic device, the second frame, where the hand of the user has come back into the FOV of the electronic device, among the plurality of frames (Akbarian, ¶0069, In response to receiving an indication that the second joint is observed (the positive branch from decision diamond 904), using information 908 about the reference joint pose and a pose of the second joint from the current time step; Akbarian, ¶0070, In the same time step, operations 902 to 910 of the method may be repeated for a third joint. In response to a determination on whether the third joint is observed, a second mask token is temporally adaptable to perform either operations 906 or 908 for the third joint); obtaining, by the electronic device, using an Artificial Intelligence (AI) model, a trajectory of the hand of the user using the one or more frames captured before the first frame among the plurality of frames OR the one or more frames captured after the second frame among the plurality of frames (Akbarian, ¶0071, Prompting the motion model 910 results in the motion model outputting a prediction which is motion parameters 912 comprising a trajectory of the articulated entity and poses of joints of the articulated entity; Akbarian, ¶0072, The process of FIG. 9 repeats for another time step and is able to continue so as to track body motion of an articulated entity over time; Akbarian, ¶0073, FIG. 10 is a flow diagram of a method of training a generative motion model; Akbarian, ¶0074, To train the motion generation model, training data 1000 is used comprising labelled training examples; that is each training example comprises known joint poses of an articulated entity which is moving, a known motion trajectory of the articulated entity and simulated inputs to the motion model; Akbarian, ¶0113-0119, for each of a plurality of time steps: receiving a reference joint pose of an articulated entity; receiving an indication that a second joint of the articulated entity is unobserved or observed; prompting a trained generative motion model using the reference joint pose and a mask token to predict body motion comprising a trajectory of the articulated entity and a pose of a plurality of joints of the articulated entity; wherein the mask token represents the second joint and is temporally adaptable by: in response to receiving an indication that the second joint is unobserved, using information about the reference joint pose and a pose of the second joint from a previous time step; and in response to receiving an indication that the second joint is observed, using information about the reference joint pose and a pose of the second joint from the current time step); and recognizing, by the electronic device, at least one hand gesture performed during the visible trajectory of the hand of the user, and the obtained trajectory of the hand of the user (Akbarian, ¶0122, using the predicted trajectory and the predicted pose of the articulated entity to do any of: animate an avatar representing the articulated entity, recognize gestures made by the articulated entity and/or control motion of the articulated entity). Akbarian does not explicitly disclose the following limitations as further recited however Yuan discloses obtaining, by the electronic device, using an Artificial Intelligence (AI) model, a trajectory of the hand of the user using the one or more frames captured before the first frame among the plurality of frames and the one or more frames captured after the second frame among the plurality of frames (Yuan, page 2, 1. Introduction, To tackle potentially severe occlusions, we propose a deep generative motion infiller that autoregressively infills the local body motions of occluded people based on visible motions … we propose a global trajectory predictor that can generate global human trajectories based on local body motions ... using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence; Yuan, page 4, Figure 3. Left: We autoregressively infill the motion using a sliding window, where the first hc frames are already infilled to serve as context and the last hl frames are look-ahead to guide the ending motion. Frames between the context and look-ahead are infilled. Right: The CVAE-based motion infiller adopts a Transformer-based seq2seq architecture; Yuan, pages 3-4, 3. Method, The input to our framework is a video I = (I1, ..., IT) with T frames, which is captured by a dynamic camera, i.e., the camera poses can change every frame … As outlined in Fig. 2, our framework consists of four stages. In Stage I, we first use multi-object tracking (MOT) and re-identification algorithms ... which is input to a human mesh recovery method ... to extract the motion of each person (including translation) in the camera coordinates. The motion may be incomplete due to various occlusions (e.g., obstruction, missed detection, going outside FoV) ... In Stage II (Sec. 3.1), we propose a generative motion infiller to tackle the occlusions ... In Stage III (Sec. 3.2), we propose a global trajectory predictor that uses the infilled body motion to generate the global trajectory; Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion free ... and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions; Yuan, page 8, Figure 5. Qualitative comparison of GLAMR on Dynamic Human 3.6M. GLAMR can generate natural hand motions for invisible frames). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Yuan and Akbarian because they are in the same field of endeavor. One skilled in the art would have been motivated to include both the first occlusion free frames and later frames with visible poses as taught by Yuan in the system of Akbarian in order to provide an alternate means to determine trajectory while avoiding generating discontinuous motions (Yuan, page 4, 3.1. Generative Motion Infiller). As per claim 2, Akbarian and Yuan disclose the method as claimed in claim 1, further comprising: selecting, by the electronic device, a frame with the presence of the hand of the user from the visible trajectory, among the plurality of frames (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise. Finally, for all 6-DoF signals, provide the velocity of changes between two consecutive frames); and generating, by the electronic device, references of one or more hand landmarks from the frame with the presence of the hand of the user (Akbarian, ¶0025, poses (3D position and orientation) of all specified joints; Akbarian, ¶0035, In FIG. 3, poses 300 are input into a motion generation model 102, the motion generation model comprising temporally adaptable mask tokens, and a body motion estimate is generated. The three poses 300 in FIG. 3 are a head pose 302, a left hand pose 304 and a right hand pose 306. The poses are indicated as three axes originating from a point, where the point is a 3D position of a joint and the axes represent an orientation of a joint). As per claim 3, Akbarian and Yuan disclose the method as claimed in claim 2, further comprising: estimating, by the electronic device, a location of the one or more hand landmarks of the one or more frames captured before the first frame, based on the generated references of the one or more hand landmarks (Akbarian, ¶0035, In FIG. 3, poses 300 are input into a motion generation model 102, the motion generation model comprising temporally adaptable mask tokens, and a body motion estimate is generated. The three poses 300 in FIG. 3 are a head pose 302, a left hand pose 304 and a right hand pose 306. The poses are indicated as three axes originating from a point, where the point is a 3D position of a joint and the axes represent an orientation of a joint); calculating, by the electronic device, one or more kinetic parameters of each hand landmark, using the estimated location of the one or more hand landmarks of consecutive frames, wherein the consecutive frames comprise the one or more frames captured before the first frame (Akbarian, ¶0041, input data 118 to the motion model 102 comprises a pose of a reference joint of the entity and an indication of presence or absence of one or more other joints of the entity in a current observation. In the case that the one or more other joints are observed in the current observation the input data 118 includes the pose of the one or more other joints. In some examples the input data 118 may also comprise per joint changes in joint position between time steps, per joint changes in joint rotation between time steps; Akbarian, ¶0054-0056); and obtaining, by the electronic device, a first trajectory of the hand of the user corresponding to the one or more frames captured before the first frame, using the calculated one or more kinetic parameters of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF; Akbarian, ¶0055, The output of the motion model 102 comprises the pose (including the root orientation) ... represented with axis-angle rotations for the J joints in the body, and the global position in the world ... a root is a reference joint which is used as the root of a kinematic tree of a human skeleton; Akbarian, ¶0056, The sequence of θ0:T and γ0:T is the body motion as well as its global trajectory for the period [0, T]). As per claim 4, Akbarian and Yuan discloses the method as claimed in claim 2, further comprising: reversing, by the electronic device, order of the plurality of frames captured by the electronic device (Yuan, page 8, Figure 5. Qualitative comparison of GLAMR on Dynamic Human 3.6M. GLAMR can generate natural hand motions for invisible frames; Yuan, page 2, 1. Introduction, To tackle potentially severe occlusions, we propose a deep generative motion infiller that autoregressively infills the local body motions of occluded people based on visible motions … we propose a global trajectory predictor that can generate global human trajectories based on local body motions ... using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence; Yuan, page 4, Figure 3. Left: We autoregressively infill the motion using a sliding window, where the first hc frames are already infilled to serve as context and the last hl frames are look-ahead to guide the ending motion. Frames between the context and look-ahead are infilled. Right: The CVAE-based motion infiller adopts a Transformer-based seq2seq architecture; Yuan, pages 3-4, 3. Method, The input to our framework is a video I = (I1, ..., IT) with T frames, which is captured by a dynamic camera, i.e., the camera poses can change every frame … As outlined in Fig. 2, our framework consists of four stages. In Stage I, we first use multi-object tracking (MOT) and re-identification algorithms ... which is input to a human mesh recovery method ... to extract the motion of each person (including translation) in the camera coordinates. The motion may be incomplete due to various occlusions (e.g., obstruction, missed detection, going outside FoV) ... In Stage II (Sec. 3.1), we propose a generative motion infiller to tackle the occlusions ... In Stage III (Sec. 3.2), we propose a global trajectory predictor that uses the infilled body motion to generate the global trajectory; Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion free ... and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions); estimating, by the electronic device, a location of the one or more hand landmarks of the one or more frames captured after the second frame, based on the generated references of the one or more hand landmarks (Akbarian, ¶0035, In FIG. 3, poses 300 are input into a motion generation model 102, the motion generation model comprising temporally adaptable mask tokens, and a body motion estimate is generated. The three poses 300 in FIG. 3 are a head pose 302, a left hand pose 304 and a right hand pose 306. The poses are indicated as three axes originating from a point, where the point is a 3D position of a joint and the axes represent an orientation of a joint); calculating, by the electronic device, one or more kinetic parameters of each hand landmark, using the estimated location of the one or more hand landmarks of consecutive frames, wherein the consecutive frames comprise the one or more frames captured after the second frame (Akbarian, ¶0041, input data 118 to the motion model 102 comprises a pose of a reference joint of the entity and an indication of presence or absence of one or more other joints of the entity in a current observation. In the case that the one or more other joints are observed in the current observation the input data 118 includes the pose of the one or more other joints. In some examples the input data 118 may also comprise per joint changes in joint position between time steps, per joint changes in joint rotation between time steps; Akbarian, ¶0054-0056); and obtaining, by the electronic device, a second trajectory of the hand of the user using the one or more frames captured after the second frame, using the calculated one or more kinetic parameters of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF; Akbarian, ¶0055, The output of the motion model 102 comprises the pose (including the root orientation) ... represented with axis-angle rotations for the J joints in the body, and the global position in the world ... a root is a reference joint which is used as the root of a kinematic tree of a human skeleton; Akbarian, ¶0056, The sequence of θ0:T and γ0:T is the body motion as well as its global trajectory for the period [0, T]). As per claim 5, Akbarian and Yuan disclose the method as claimed in claim 4, wherein the one or more kinetic parameters of the hand of the user comprise at least one of a velocity and an acceleration of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF). As per claim 6, Akbarian and Yuan disclose the method as claimed in claim 4, wherein the method comprises: verifying, by the electronic device, if the hand of the user is in the FOV of the electronic device, after calculating one or more kinetic parameters of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF); and calculating, by the electronic device, a velocity and a position of the one or more hand landmarks in a next frame, after the one or more frames captured after the second frame, using the calculated one or more kinetic parameters of each of the one or more hand landmarks, if the hand of the user is not in the FOV of the electronic device (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF). As per claim 7, Akbarian and Yuan disclose the method as claimed in claim 6, further comprising: verifying, by the electronic device, at least one parameter of the hand of the user, after calculating the velocity and the position of the one or more hand landmarks in the next frame, wherein the at least one parameter comprises at least one of whether a velocity goes to zero, a hand position is beyond a threshold, and the one or more hand landmarks no longer conform to predetermined bio-mechanical constraints of a human hand (Akbarian, ¶0027, The body motion predictor receives inputs ... In an example the inputs 118 comprise an HMD signal, egocentric images ... the inputs comprise at least a pose of a reference joint of an articulated entity for which body motion is to be computed, and an indication of whether a second joint of the articulated entity is observed or unobserved in a current time step. There may be more inputs, such as poses of one or more joints in other coordinate spaces, changes in position of joints between time steps, changes in rotation of joints between time steps); repeating, by the electronic device, a verification of the hand of the user in the FOV of the electronic device, if the at least one parameter of the hand of the user is not satisfied (Akbarian, ¶0044, The temporally adaptable properties of the mask tokens in the present examples, therefore, enable the mask tokens to be updated as shown in FIG. 4, to capture the temporal aspect of the data); and repeating, by the electronic device, estimation of the location of the one or more hand landmarks from a previous frame, previous to the next frame, if the hand of the user is stationary and the at least one parameter of the hand of the user is satisfied (Akbarian, ¶0043, Since some of the observations in the input data 118 may be missing a temporally adaptable mask token mechanism 410 is used. The temporally adaptable mask token mechanism operates in the embedding space of the motion model 102 to generate a mask 412. The mask 412 is … a predicted embedding of an unobserved joint). As per claim 8, Akbarian and Yuan disclose the method as claimed in claim 3, further comprising: checking, by the electronic device, closeness of the one or more hand landmarks for each frame, of the plurality of frames, from the first trajectory and the second trajectory of the hand of the user (Yuan, page 2, 1. Introduction, using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence such as 2D keypoints … We propose a method to generate global human trajectories from local body motions and use the generated trajectories as anchors to constrain global motion and camera optimization; Yuan, page 8, Figure 5. Qualitative comparison of GLAMR on Dynamic Human 3.6M. GLAMR can generate natural hand motions for invisible frames); obtaining, by the electronic device, a spatio-temporal convergence at a frame, of the plurality of frames, where a distance between two extrapolated hand landmarks is below a certain threshold, wherein a trajectory until the frame of the spatio-temporal convergence is considered as the first trajectory, wherein a trajectory after the frame of the spatio-temporal convergence is considered as the second trajectory (Akbarian, ¶0045, The motion model 102 comprises an attention mechanism 402. The attention mechanism 402 is configured to encode information about the reference joint pose and the pose of one or more other joints of the articulated entity over a plurality of the time steps, and to encode information about spatial correlations between the poses of the joints; Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion-free or infilled and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions. Excluding the context and look-ahead frames, only the middle ho = h - hc - hl frames of motion are infilled. We iteratively infill the motion using the sliding window and advance the window by ho frames every step); estimating, by the electronic device, a hand pose by encoding the two extrapolated hand landmarks of the spatio-temporal convergence (Akbarian, ¶0045, The motion model 102 comprises an attention mechanism 402. The attention mechanism 402 is configured to encode information about the reference joint pose and the pose of one or more other joints of the articulated entity over a plurality of the time steps, and to encode information about spatial correlations between the poses of the joints); and recognizing, by the electronic device, the at least one hand gesture, based on a sequence of hand pose information (Akbarian, ¶0122, using the predicted trajectory and the predicted pose of the articulated entity to do any of: animate an avatar representing the articulated entity, recognize gestures made by the articulated entity and/or control motion of the articulated entity). As per claim 9, Akbarian discloses an electronic device, comprising: a memory storing at least one instruction; and, at least one processor configured to execute the at least one instruction stored in the memory; wherein the at least one processor is configured to execute the at least one instruction (Akbarian, ¶0026, FIG. 1 is a schematic diagram of a body motion predictor 100 which is computer implemented and comprises a processor 104 and a memory 106. The body motion predictor 100 comprises a motion model 102) to: track a visible trajectory of a hand of the user from a plurality of frames captured by the electronic device, wherein the plurality of frames comprise a first frame, a second frame, one or more frames captured before the first frame, and one or more frames captured after the second frame (Akbarian, ¶0024, In various examples herein the articulated entity is a person and the capture device is a head mounted display HMD worn by the person. One or more egocentric camera in the HMD capture images of only part of the person due to restricted field of view and occlusions. The person's hands move into and out of the field of view of the egocentric camera; Akbarian, ¶0027, The body motion predictor receives inputs ... In an example the inputs 118 comprise an HMD signal, egocentric images … the inputs comprise at least a pose of a reference joint of an articulated entity for which body motion is to be computed, and an indication of whether a second joint of the articulated entity is observed or unobserved in a current time step; Akbarian, ¶0034, In FIG. 2, where the hand tracking is performed by a HMD worn by the user, it can be seen that at each position of the user from 200 to 208, the hands are in a different position. For example, at position 208, both hands of the user will be in the FoV of the HMD, whilst at position 200, the user's right hand will be in the FoV of the HMD and the user's left hand will be out of the field of view of the HMD; Akbarian, ¶0041, input data 118 to the motion model 102 comprises a pose of a reference joint of the entity and an indication of presence or absence of one or more other joints of the entity in a current observation. In the case that the one or more other joints are observed in the current observation the input data 118 includes the pose of the one or more other joints. In some examples the input data 118 may also comprise per joint changes in joint position between time steps, per joint changes in joint rotation between time steps; Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise); identify the first frame, where the hand of the user has gone out of a Field of View (FOV) of the electronic device, among the plurality of frames (Akbarian, ¶0043, Since some of the observations in the input data 118 may be missing a temporally adaptable mask token mechanism 410 is used. The temporally adaptable mask token mechanism operates in the embedding space of the motion model 102 to generate a mask 412. The mask 412 is either an embedding of a pose of a joint observed at the current time step or a predicted embedding of an unobserved joint, predicted by taking into account the embedding of the reference joint and an embedding of the unobserved joint from a previous time step; Akbarian, ¶0050, FIG. 5 is a schematic diagram of an example of a generative motion model using input data … the person is wearing an HMD with an egocentric camera and the HMD processes the egocentric images to achieve hand tracking to track pose of the hands of the person; Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise. Finally, for all 6-DoF signals, provide the velocity of changes between two consecutive frames; Akbarian, ¶0066, The body motion predictor receives 902 an indication that a second joint of the articulated entity is unobserved or observed. The indication may be a flag such as a binary value. The indication may be received from another entity such as an image recognition system which recognizes particular joints such as hands, feet or other joints in an image; Akbarian, ¶0067, The body motion predictor prompts 910 a motion model which is a trained neural network. The prompt comprises a mask token which is a temporally adaptable mask token; Akbarian, ¶0068, The mask token represents the second joint and is temporally adaptable. In response to receiving an indication that the second joint is unobserved (the negative branch from decision diamond 904), using 906 information about the reference joint pose and a pose of the second joint from a previous time step; Akbarian, ¶0070, In the same time step, operations 902 to 910 of the method may be repeated for a third joint. In response to a determination on whether the third joint is observed, a second mask token is temporally adaptable to perform either operations 906 or 908 for the third joint); identify the second frame, where the hand of the user has come back into the FOV of the electronic device, among the plurality of frames (Akbarian, ¶0069, In response to receiving an indication that the second joint is observed (the positive branch from decision diamond 904), using information 908 about the reference joint pose and a pose of the second joint from the current time step; Akbarian, ¶0070, In the same time step, operations 902 to 910 of the method may be repeated for a third joint. In response to a determination on whether the third joint is observed, a second mask token is temporally adaptable to perform either operations 906 or 908 for the third joint); obtain, using an Artificial Intelligence (AI) model, a trajectory of the hand of the user using the one or more frames captured before the first frame among the plurality of frames, OR the one or more frames captured after the second frame among the plurality of frames (Akbarian, ¶0071, Prompting the motion model 910 results in the motion model outputting a prediction which is motion parameters 912 comprising a trajectory of the articulated entity and poses of joints of the articulated entity; Akbarian, ¶0072, The process of FIG. 9 repeats for another time step and is able to continue so as to track body motion of an articulated entity over time; Akbarian, ¶0073, FIG. 10 is a flow diagram of a method of training a generative motion model; Akbarian, ¶0074, To train the motion generation model, training data 1000 is used comprising labelled training examples; that is each training example comprises known joint poses of an articulated entity which is moving, a known motion trajectory of the articulated entity and simulated inputs to the motion model; Akbarian, ¶0113-0119, for each of a plurality of time steps: receiving a reference joint pose of an articulated entity; receiving an indication that a second joint of the articulated entity is unobserved or observed; prompting a trained generative motion model using the reference joint pose and a mask token to predict body motion comprising a trajectory of the articulated entity and a pose of a plurality of joints of the articulated entity; wherein the mask token represents the second joint and is temporally adaptable by: in response to receiving an indication that the second joint is unobserved, using information about the reference joint pose and a pose of the second joint from a previous time step; and in response to receiving an indication that the second joint is observed, using information about the reference joint pose and a pose of the second joint from the current time step); and recognize at least one hand gesture performed during the visible trajectory of the hand of the user, and the obtained trajectory of the hand of the user (Akbarian, ¶0122, using the predicted trajectory and the predicted pose of the articulated entity to do any of: animate an avatar representing the articulated entity, recognize gestures made by the articulated entity and/or control motion of the articulated entity). Akbarian does not explicitly disclose the following limitations as further recited however Yuan discloses obtain, using an Artificial Intelligence (AI) model, a trajectory of the hand of the user using the one or more frames captured before the first frame among the plurality of frames, and the one or more frames captured after the second frame among the plurality of frames (Yuan, page 2, 1. Introduction, To tackle potentially severe occlusions, we propose a deep generative motion infiller that autoregressively infills the local body motions of occluded people based on visible motions … we propose a global trajectory predictor that can generate global human trajectories based on local body motions ... using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence; Yuan, page 4, Figure 3. Left: We autoregressively infill the motion using a sliding window, where the first hc frames are already infilled to serve as context and the last hl frames are look-ahead to guide the ending motion. Frames between the context and look-ahead are infilled. Right: The CVAE-based motion infiller adopts a Transformer-based seq2seq architecture; Yuan, pages 3-4, 3. Method, The input to our framework is a video I = (I1, ..., IT) with T frames, which is captured by a dynamic camera, i.e., the camera poses can change every frame … As outlined in Fig. 2, our framework consists of four stages. In Stage I, we first use multi-object tracking (MOT) and re-identification algorithms ... which is input to a human mesh recovery method ... to extract the motion of each person (including translation) in the camera coordinates. The motion may be incomplete due to various occlusions (e.g., obstruction, missed detection, going outside FoV) ... In Stage II (Sec. 3.1), we propose a generative motion infiller to tackle the occlusions ... In Stage III (Sec. 3.2), we propose a global trajectory predictor that uses the infilled body motion to generate the global trajectory; Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion free ... and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions; Yuan, page 8, Figure 5. Qualitative comparison of GLAMR on Dynamic Human 3.6M. GLAMR can generate natural hand motions for invisible frames). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Yuan and Akbarian because they are in the same field of endeavor. One skilled in the art would have been motivated to include both the first occlusion free frames and later frames with visible poses as taught by Yuan in the system of Akbarian in order to provide an alternate means to determine trajectory while avoiding generating discontinuous motions (Yuan, page 4, 3.1. Generative Motion Infiller). As per claim 10, Akbarian and Yuan discloses the electronic device as claimed in claim 9, wherein the at least one processor is configured to execute the at least one instruction to: select a frame, among the plurality of frames, with the presence of the hand of the user from the visible trajectory (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise. Finally, for all 6-DoF signals, provide the velocity of changes between two consecutive frames); and generate references of one or more hand landmarks from the frame with the presence of the hand of the user (Akbarian, ¶0025, poses (3D position and orientation) of all specified joints; Akbarian, ¶0035, In FIG. 3, poses 300 are input into a motion generation model 102, the motion generation model comprising temporally adaptable mask tokens, and a body motion estimate is generated. The three poses 300 in FIG. 3 are a head pose 302, a left hand pose 304 and a right hand pose 306. The poses are indicated as three axes originating from a point, where the point is a 3D position of a joint and the axes represent an orientation of a joint). As per claim 11, Akbarian and Yuan disclose the electronic device as claimed in claim 10, wherein the at least one processor is configured to execute the at least one instruction to: estimate a location of the one or more hand landmarks of the one or more frames captured before the first frame, based on the generated references of the one or more hand landmarks (Akbarian, ¶0035, In FIG. 3, poses 300 are input into a motion generation model 102, the motion generation model comprising temporally adaptable mask tokens, and a body motion estimate is generated. The three poses 300 in FIG. 3 are a head pose 302, a left hand pose 304 and a right hand pose 306. The poses are indicated as three axes originating from a point, where the point is a 3D position of a joint and the axes represent an orientation of a joint); calculate one or more kinetic parameters of each hand landmark, using the estimated location of the one or more hand landmarks of consecutive frames, wherein the consecutive frames comprises the one or more frames captured before the first frame (Akbarian, ¶0041, input data 118 to the motion model 102 comprises a pose of a reference joint of the entity and an indication of presence or absence of one or more other joints of the entity in a current observation. In the case that the one or more other joints are observed in the current observation the input data 118 includes the pose of the one or more other joints. In some examples the input data 118 may also comprise per joint changes in joint position between time steps, per joint changes in joint rotation between time steps; Akbarian, ¶0054-0056); and obtain a first trajectory of the hand of the user using the one or more frames captured before the first frame, using the calculated one or more kinetic parameters of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF; Akbarian, ¶0055, The output of the motion model 102 comprises the pose (including the root orientation) ... represented with axis-angle rotations for the J joints in the body, and the global position in the world ... a root is a reference joint which is used as the root of a kinematic tree of a human skeleton; Akbarian, ¶0056, The sequence of θ0:T and γ0:T is the body motion as well as its global trajectory for the period [0, T]). As per claim 12, Akbarian and Yuan disclose the electronic device as claimed in claim 10, wherein the at least one processor is configured to execute the at least one instruction to: reverse the order of the plurality of frames captured by the electronic device (Yuan, page 8, Figure 5. Qualitative comparison of GLAMR on Dynamic Human 3.6M. GLAMR can generate natural hand motions for invisible frames; Yuan, page 2, 1. Introduction, To tackle potentially severe occlusions, we propose a deep generative motion infiller that autoregressively infills the local body motions of occluded people based on visible motions … we propose a global trajectory predictor that can generate global human trajectories based on local body motions ... using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence; Yuan, page 4, Figure 3. Left: We autoregressively infill the motion using a sliding window, where the first hc frames are already infilled to serve as context and the last hl frames are look-ahead to guide the ending motion. Frames between the context and look-ahead are infilled. Right: The CVAE-based motion infiller adopts a Transformer-based seq2seq architecture; Yuan, pages 3-4, 3. Method; Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion free ... and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions); estimate a location of the one or more hand landmarks of the one or more frames captured after the second frame, based on the generated references of the one or more hand landmarks (Akbarian, ¶0035, In FIG. 3, poses 300 are input into a motion generation model 102, the motion generation model comprising temporally adaptable mask tokens, and a body motion estimate is generated. The three poses 300 in FIG. 3 are a head pose 302, a left hand pose 304 and a right hand pose 306. The poses are indicated as three axes originating from a point, where the point is a 3D position of a joint and the axes represent an orientation of a joint); calculate one or more kinetic parameters of each hand landmark, using the estimated location of the one or more hand landmarks of consecutive frames, wherein the consecutive frames comprises the one or more frames captured after the second frame (Akbarian, ¶0041, input data 118 to the motion model 102 comprises a pose of a reference joint of the entity and an indication of presence or absence of one or more other joints of the entity in a current observation. In the case that the one or more other joints are observed in the current observation the input data 118 includes the pose of the one or more other joints. In some examples the input data 118 may also comprise per joint changes in joint position between time steps, per joint changes in joint rotation between time steps; Akbarian, ¶0054-0056); and obtain a second trajectory of the hand of the user corresponding to the one or more frames captured after the second frame, using the calculated one or more kinetic parameters of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF; Akbarian, ¶0055, The output of the motion model 102 comprises the pose (including the root orientation) ... represented with axis-angle rotations for the J joints in the body, and the global position in the world ... a root is a reference joint which is used as the root of a kinematic tree of a human skeleton; Akbarian, ¶0056, The sequence of θ0:T and γ0:T is the body motion as well as its global trajectory for the period [0, T]). As per claim 13, Akbarian and Yuan discloses the electronic device as claimed in claim 12, wherein the one or more kinetic parameters of the hand of the user comprise at least one of a velocity, and an acceleration of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF). As per claim 14, Akbarian and Yuan disclose the electronic device as claimed in claim 12, wherein the at least one processor is configured to execute the at least one instruction to: verify if the hand of the user is in the FOV of the electronic device, after calculating one or more kinetic parameters of each of the one or more hand landmarks (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF); and calculate a velocity and a position of the one or more hand landmarks in a next frame, of the plurality of frames and after the one or more frames captured after the second frame, using the calculated one or more kinetic parameters of each of the one or more hand landmarks, if the hand of the user is not in the FOV of the electronic device (Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise … for all 6-DoF signals, provide the velocity of changes between two consecutive frames. Specifically for translations consider vel (Pt, Pt−1)=Pt−Pt−1=and for rotations consider the geodesic changes in the rotation vel (Rt, Rt−1)=(Rt−1)−1 Rt which together constitute the velocity 6-DoF). As per claim 15, Akbarian and Yuan disclose the electronic device as claimed in claim 14, wherein the at least one processor is configured to execute the at least one instruction to: verify at least one parameter of the hand of the user, after calculating the velocity and the position of the one or more hand landmarks in the next frame, wherein the at least one parameter comprises at least one of whether a velocity goes to zero, a hand position is beyond a threshold, and the one or more hand landmarks no longer conform to predetermined bio-mechanical constraints of a human hand (Akbarian, ¶0027, The body motion predictor receives inputs ... In an example the inputs 118 comprise an HMD signal, egocentric images ... the inputs comprise at least a pose of a reference joint of an articulated entity for which body motion is to be computed, and an indication of whether a second joint of the articulated entity is observed or unobserved in a current time step. There may be more inputs, such as poses of one or more joints in other coordinate spaces, changes in position of joints between time steps, changes in rotation of joints between time steps); repeat a verification of the hand of the user in the FOV of the electronic device, if the at least one parameter of the hand of the user is not satisfied (Akbarian, ¶0044, The temporally adaptable properties of the mask tokens in the present examples, therefore, enable the mask tokens to be updated as shown in FIG. 4, to capture the temporal aspect of the data); and repeat estimation of the location of the one or more hand landmarks from a previous frame, previous to the next frame, if the hand of the user is stationary and the at least one parameter of the hand of the user is satisfied (Akbarian, ¶0043, Since some of the observations in the input data 118 may be missing a temporally adaptable mask token mechanism 410 is used. The temporally adaptable mask token mechanism operates in the embedding space of the motion model 102 to generate a mask 412. The mask 412 is … a predicted embedding of an unobserved joint). As per claim 16, Akbarian and Yuan disclose the electronic device as claimed in claim 11, wherein the at least one processor is configured to execute the at least one instruction to: check closeness of the one or more hand landmarks for each frame, of the plurality of frames, from the first trajectory and the second trajectory of the hand of the user (Yuan, page 2, 1. Introduction, using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence such as 2D keypoints … We propose a method to generate global human trajectories from local body motions and use the generated trajectories as anchors to constrain global motion and camera optimization; Yuan, page 8, Figure 5. Qualitative comparison of GLAMR on Dynamic Human 3.6M. GLAMR can generate natural hand motions for invisible frames); obtain a spatio-temporal convergence at a frame, of the plurality of frames, where a distance between two extrapolated hand landmarks is below a certain threshold, wherein a trajectory until the frame of the spatio-temporal convergence is considered as the first trajectory, wherein a trajectory after the frame of the spatio-temporal convergence is considered as the second trajectory (Akbarian, ¶0045, The motion model 102 comprises an attention mechanism 402. The attention mechanism 402 is configured to encode information about the reference joint pose and the pose of one or more other joints of the articulated entity over a plurality of the time steps, and to encode information about spatial correlations between the poses of the joints; Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion-free or infilled and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions. Excluding the context and look-ahead frames, only the middle ho = h - hc - hl frames of motion are infilled. We iteratively infill the motion using the sliding window and advance the window by ho frames every step); estimate a hand pose by encoding the two extrapolated hand landmarks of the spatio-temporal convergence; and recognize the at least one hand gesture, based on a sequence of hand pose information (Akbarian, ¶0045, The motion model 102 comprises an attention mechanism 402. The attention mechanism 402 is configured to encode information about the reference joint pose and the pose of one or more other joints of the articulated entity over a plurality of the time steps, and to encode information about spatial correlations between the poses of the joints); and recognizing, by the electronic device, the at least one hand gesture, based on a sequence of hand pose information (Akbarian, ¶0122, using the predicted trajectory and the predicted pose of the articulated entity to do any of: animate an avatar representing the articulated entity, recognize gestures made by the articulated entity and/or control motion of the articulated entity). As per claim 17, Akbarian discloses a multi-camera device (Akbarian, ¶0024, In various examples herein the articulated entity is a person and the capture device is a head mounted display HMD worn by the person. One or more egocentric camera in the HMD capture images of only part of the person due to restricted field of view and occlusions. The person's hands move into and out of the field of view of the egocentric camera), comprising: a memory; and, at least one processor configured to execute the at least one instruction stored in the memory, wherein the at least one processor is configured to execute the at least one instruction (Akbarian, ¶0026, FIG. 1 is a schematic diagram of a body motion predictor 100 which is computer implemented and comprises a processor 104 and a memory 106. The body motion predictor 100 comprises a motion model 102) to: obtain an outward trajectory for one or more frames where the hand of the user goes out of a Field of view (FOV) of the multi-camera device (Akbarian, ¶0043, Since some of the observations in the input data 118 may be missing a temporally adaptable mask token mechanism 410 is used. The temporally adaptable mask token mechanism operates in the embedding space of the motion model 102 to generate a mask 412. The mask 412 is either an embedding of a pose of a joint observed at the current time step or a predicted embedding of an unobserved joint, predicted by taking into account the embedding of the reference joint and an embedding of the unobserved joint from a previous time step; Akbarian, ¶0050, FIG. 5 is a schematic diagram of an example of a generative motion model using input data … the person is wearing an HMD with an egocentric camera and the HMD processes the egocentric images to achieve hand tracking to track pose of the hands of the person; Akbarian, ¶0054, In the Hand Tracking (HT) scenario, hands may go in and out of FoV of the HMD, so provide the motion model 102 with hand visibility status for both left and right hand, vlt and vrt, as binary values, 1 being visible and 0 otherwise. Finally, for all 6-DoF signals, provide the velocity of changes between two consecutive frames; Akbarian, ¶0066, The body motion predictor receives 902 an indication that a second joint of the articulated entity is unobserved or observed. The indication may be a flag such as a binary value. The indication may be received from another entity such as an image recognition system which recognizes particular joints such as hands, feet or other joints in an image; Akbarian, ¶0067, The body motion predictor prompts 910 a motion model which is a trained neural network. The prompt comprises a mask token which is a temporally adaptable mask token; Akbarian, ¶0068, The mask token represents the second joint and is temporally adaptable. In response to receiving an indication that the second joint is unobserved (the negative branch from decision diamond 904), using 906 information about the reference joint pose and a pose of the second joint from a previous time step; Akbarian, ¶0070, In the same time step, operations 902 to 910 of the method may be repeated for a third joint. In response to a determination on whether the third joint is observed, a second mask token is temporally adaptable to perform either operations 906 or 908 for the third joint); obtain an inward trajectory from one or more frames where the hand of the user comes back into the FOV of the multi-camera device (Akbarian, ¶0069, In response to receiving an indication that the second joint is observed (the positive branch from decision diamond 904), using information 908 about the reference joint pose and a pose of the second joint from the current time step; Akbarian, ¶0070, In the same time step, operations 902 to 910 of the method may be repeated for a third joint. In response to a determination on whether the third joint is observed, a second mask token is temporally adaptable to perform either operations 906 or 908 for the third joint), estimate a point of the outward trajectory and the inward trajectory in temporal and spatial domains (Akbarian, ¶0045, The motion model 102 comprises an attention mechanism 402. The attention mechanism 402 is configured to encode information about the reference joint pose and the pose of one or more other joints of the articulated entity over a plurality of the time steps, and to encode information about spatial correlations between the poses of the joints); and recognize at least one hand gesture, based on a sequence of hand pose information obtained from the estimated convergence point, wherein a hand pose is estimated by encoding one or more extrapolated hand landmarks of the convergence point (Akbarian, ¶0122, using the predicted trajectory and the predicted pose of the articulated entity to do any of: animate an avatar representing the articulated entity, recognize gestures made by the articulated entity and/or control motion of the articulated entity). Akbarian does not explicitly disclose the following limitations as further recited however Yuan discloses wherein the inward trajectory is predicted by reversing the order of the one or more frames where the hand of the user comes back into the FOV (Yuan, page 8, Figure 5. Qualitative comparison of GLAMR on Dynamic Human 3.6M. GLAMR can generate natural hand motions for invisible frames; Yuan, page 2, 1. Introduction, To tackle potentially severe occlusions, we propose a deep generative motion infiller that autoregressively infills the local body motions of occluded people based on visible motions … we propose a global trajectory predictor that can generate global human trajectories based on local body motions ... using the predicted trajectories as anchors to constrain the solution space, we further propose a global optimization framework that jointly optimizes the global motions and camera poses to match the video evidence; Yuan, page 4, Figure 3. Left: We autoregressively infill the motion using a sliding window, where the first hc frames are already infilled to serve as context and the last hl frames are look-ahead to guide the ending motion. Frames between the context and look-ahead are infilled. Right: The CVAE-based motion infiller adopts a Transformer-based seq2seq architecture; Yuan, pages 3-4, 3. Method, The input to our framework is a video I = (I1, ..., IT) with T frames, which is captured by a dynamic camera, i.e., the camera poses can change every frame … As outlined in Fig. 2, our framework consists of four stages. In Stage I, we first use multi-object tracking (MOT) and re-identification algorithms ... which is input to a human mesh recovery method ... to extract the motion of each person (including translation) in the camera coordinates. The motion may be incomplete due to various occlusions (e.g., obstruction, missed detection, going outside FoV) ... In Stage II (Sec. 3.1), we propose a generative motion infiller to tackle the occlusions ... In Stage III (Sec. 3.2), we propose a global trajectory predictor that uses the infilled body motion to generate the global trajectory; Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion free ... and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions); and estimate a convergence point of the outward trajectory and the inward trajectory in temporal and spatial domains (Yuan, page 4, 3.1. Generative Motion Infiller, Autoregressive Motion Infilling. To ensure that the motion infiller M can handle much longer test motions than the training motions, we propose an autoregressive motion infilling process at test time as illustrated in Fig. 3 (Left). The key idea is to use a sliding window of h frames, where we assume the first hc frames of motion are already occlusion-free or infilled and serve as context, and we also use the last hl frames as look-ahead. The look-ahead is essential to the motion infiller since it may contain visible poses that can guide the ending motion and avoid generating discontinuous motions. Excluding the context and look-ahead frames, only the middle ho = h - hc - hl frames of motion are infilled. We iteratively infill the motion using the sliding window and advance the window by ho frames every step). It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine the teachings of Yuan and Akbarian because they are in the same field of endeavor. One skilled in the art would have been motivated to include both the first occlusion free frames and later frames with visible poses as taught by Yuan in the system of Akbarian in order to provide an alternate means to determine trajectory while avoiding generating discontinuous motions (Yuan, page 4, 3.1. Generative Motion Infiller). As per claim 18, Akbarian and Yuan disclose the multi-camera device according to claim 17, wherein the at least one processor is configured to execute the at least one instruction to control an action to at least one of a video game and a user interface window depending on recognizing the hat least one hand gesture (Akbarian, ¶0025, Given only sparse observations of an articulated entity it is desired to compute body motion of the articulated entity … the body motion is usable for downstream tasks including but not limited to: controlling 3D avatars in mixed-reality applications … and for a variety of applications such as 3D body gesture recognition, computer gaming, mixed-reality, virtual reality and others). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to TRACY MANGIALASCHI whose telephone number is (571)270-5189. The examiner can normally be reached M-F, 9:30AM TO 6:00PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TRACY MANGIALASCHI/Primary Examiner, Art Unit 2668
Read full office action

Prosecution Timeline

Oct 04, 2024
Application Filed
Aug 06, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700132
INFORMATION PROCESSING SYSTEM, INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND RECORDING MEDIUM
3y 6m to grant Granted Aug 04, 2026
Patent 12696893
METHOD AND DEVICE FOR CLASSIFYING PLANTS, AND COMPUTER PROGRAM PRODUCT
3y 4m to grant Granted Aug 04, 2026
Patent 12694710
ARTIFICIAL INTELLIGENCE FOR PASSIVE LIVENESS DETECTION
4y 0m to grant Granted Jul 28, 2026
Patent 12694665
COMPUTER-IMPLEMENTED METHOD FOR CLASSIFYING WASTE
2y 4m to grant Granted Jul 28, 2026
Patent 12684237
METHOD AND APPARATUS WITH IMAGE PROCESSING
3y 8m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
99%
With Interview (+27.3%)
3y 0m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 594 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month