Prosecution Insights
Last updated: October 02, 2026
Application No. 19/034,528

HAND POSE RECOGNITION METHOD AND APPARATUS, DEVICE, STORAGE MEDIUM, AND PROGRAM PRODUCT

Non-Final OA §103
Filed
Jan 22, 2025
Priority
Feb 24, 2023 — CN 202310215949.2 +1 more
Examiner
SATCHER, DION JOHN
Art Unit
Tech Center
Assignee
Tencent Technology (Shenzhen) Company Limited
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
44 granted / 52 resolved
+24.6% vs TC avg
Strong +18% interview lift
Without
With
+17.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
23 currently pending
Career history
81
Total Applications
across all art units

Statute-Specific Performance

§101
14.0%
-26.0% vs TC avg
§103
65.9%
+25.9% vs TC avg
§102
10.2%
-29.8% vs TC avg
§112
9.1%
-30.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 52 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims This communication is in response to the Application Filed on 01/22/2025. Claims 1-20 are pending in this application. Drawings The drawings filed on 01/22/2025are accepted by the Examiner. Information Disclosure Statement The information disclosure statement (IDS) submitted on 01/28/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. Claims 1, 8, 9, 16, 17 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 12573080 B2, hereafter, "Kim") in view of Mirza et al. (US 20210124951 A1, hereafter, "Mirza") further in view of Dal Mutto et al. (US 20190364206 A1, hereafter, "Dal Mutto"). Regarding claim 1, Kim discloses a hand pose recognition method (See Kim, [Col. 6, ln. 27–29], In one embodiment, a 3D pose of the hand H may be recognized using 'hand skeleton detection and tracking technology) performed by a computer device, the method comprising: acquiring a current frame of a multi-lens video of a target object, and the current frame comprising a plurality of views (See Kim, [Col. 8, ln. 60–64], In one embodiment, the electronic device 10 may use the plurality of cameras to accurately estimate a 3D pose of the user's hand H and estimate depths (relative positions on a z-axis) of joints of the hand H from 2D camera images captured at different viewpoints); performing hand detection on a first view of the current frame to obtain a first lens detection result, the first lens detection result including hand detection boxes in the first view of the current frame (See Kim, [Col. 9, ln. 34–36], In operation S220, the electronic device may obtain, from the first image, a first ROI including an image corresponding 35 to the object (e.g., user's hand)); performing hand estimation on a second view of the current frame to obtain a second lens estimation result, the second lens estimation result including hand estimation boxes in the second view of the current frame (See Kim, [Col. 9, ln. 58–60], In operation S240, the electronic device may obtain a second ROI from the second image based on the first skeleton data); [removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes]; performing hand joint point recognition on the remaining hand detection boxes to obtain two-dimensional joint points corresponding to the current frame (See Kim, [Col. 7, ln. 16–18], In operation 130, the electronic device 10 may obtain first skeleton data HCl including at least one keypoint of the hand H from the first ROI ROI1. [Col. 7, ln. 21–25], For example, the first skeleton data HCl may be a data set including coordinate information of at least one keypoint (e.g., joint) of the hand H. Coordinate information of a keypoint of the hand H may include Two-Dimensional (2D). [Col. 8, ln. 21–25], In one embodiment, each of the keypoints included in the second skeleton data HC2 obtained using the third deep learning model may correspond to 2D position coordinates with respect to a preset origin on a plane including the second image IM2); converting the two-dimensional joint points into three-dimensional joint points in a three-dimensional hand coordinate system (See Kim, [Col. 8, ln. 42–49], 3D position coordinates may be obtained for each keypoint in the first skeleton data HC1 by using 2D position coordinates of a keypoint projected onto the first image IM1 and 2D position coordinates of a corresponding keypoint in the second skeleton data HC2. For example, a triangulation method may be used in an operation of obtaining 3D position coordinates by using two 2D position coordinates); and [converting the three-dimensional joint points in the three-dimensional hand coordinate system into three-dimensional joint points of the current frame in a world coordinate system according to pose estimation parameters corresponding to the current frame]. However, Kim fails to teach removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes. converting the three-dimensional joint points in the three-dimensional hand coordinate system into three-dimensional joint points of the current frame in a world coordinate system according to pose estimation parameters corresponding to the current frame. Mirza, working in the same field of endeavor, teaches: removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes (See Mirza, ¶ [0318], For example, region 2222 may be a bounding box determined for the contour 2220 using a non-maximum suppression object-detection algorithm. For instance, the sensor client 105a may determine a plurality of bounding boxes associated with the contour 2220. For each bounding box, the client 105a may calculate a score. The score, for example, may represent an extent to which that bounding box is similar to the other bounding boxes. The sensor client 105a may identify a subset of the bounding boxes with a score that is greater than a threshold value (e.g., 80% or more), and determine region 2222 based on this identified subset). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s reference to removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes based on the method of Mirza’s reference. The suggestion/motivation would have been to reliably distinguish between close regions and accurately determine the best candidate region (See Mirza, ¶ [0323]). However, Kim and Mirza fail to teach converting the three-dimensional joint points in the three-dimensional hand coordinate system into three-dimensional joint points of the current frame in a world coordinate system according to pose estimation parameters corresponding to the current frame. Dal Mutto, working in the same field of endeavor, teaches: converting the three-dimensional joint points in the three-dimensional hand coordinate system into three-dimensional joint points of the current frame in a world coordinate system according to pose estimation parameters corresponding to the current frame (See Dal Mutto, ¶ [0093], For example, the first relative pose may be defined as a three dimensional (3-D) rigid transformation that would map the location and orientation (“pose”) of the second camera (CAM B 100B) onto the pose of the first camera (CAM A 100A). Alternatively, and equivalently, the first relative pose may include two transformations: a transformation from the pose of the first camera (CAM A 100A) to a world coordinate system and a transformation from the pose of the second camera (CAM B 100B) to the same world coordinate system). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s and Mirza’s reference converting the three-dimensional joint points in the three-dimensional hand coordinate system into three-dimensional joint points of the current frame in a world coordinate system according to pose estimation parameters corresponding to the current frame based on the method of Dal Mutto’s reference. The suggestion/motivation would have been to accurately estimate the 3D global pose (See Dal Mutto, ¶ [0005–0008]). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Dal Mutto with Kim and Mirza to obtain the invention as specified in claim 1. Regarding claim 8, Kim in view of Mirza and Dal Mutto teaches the method according to claim 1, wherein the multi-lens video is a video obtained by means of a multi-lens camera performing synchronous image collection on hands of the target object (See Kim, [Col. 8, ln. 60–64], In one embodiment, the electronic device 10 may use the plurality of cameras to accurately estimate a 3D pose of the user's hand H and estimate depths (relative positions on a z-axis) of joints of the hand H from 2D camera images captured at different viewpoints). Regarding claim 9, claim 9 is rejected the same as claim 1 and the arguments similar to that presented above for claim 1 are equally applicable to the claim 9, and all of the other limitations similar to claim 1 are not repeated herein, but incorporated by reference. Furthermore, Kim teaches a computer device, comprising a memory and a processor, the memory having computer-readable instructions stored therein, and the processor, when executing the computer-readable instructions, causing the computer device to implement a hand pose recognition method including (See Kim, [FIG. 10] 1020 Processor, 1030 Storage). Regarding claim 16, claim 16 is rejected the same as claim 8 and the arguments similar to that presented above for claim 8 are equally applicable to the claim 16, and all of the other limitations similar to claim 8 are not repeated herein, but incorporated by reference. Regarding claim 17, claim 17 is rejected the same as claim 1 and the arguments similar to that presented above for claim 1 are equally applicable to the claim 17, and all of the other limitations similar to claim 1 are not repeated herein, but incorporated by reference. Furthermore, Kim teaches a non-transitory computer-readable storage medium, having computer-readable instructions stored therein, wherein the computer-readable instructions, when executed by a processor of a computer device, cause the computer device to perform a hand pose recognition method including (See Kim, [FIG. 10] 1020 Processor, 1030 Storage). Regarding claim 20, claim 20 is rejected the same as claim 8 and the arguments similar to that presented above for claim 8 are equally applicable to the claim 20, and all of the other limitations similar to claim 8 are not repeated herein, but incorporated by reference. Claims 2, 10 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 12573080 B2, hereafter, "Kim") in view of Mirza et al. (US 20210124951 A1, hereafter, "Mirza"), Dal Mutto et al. (US 20190364206 A1, hereafter, "Dal Mutto"), and further in view of Ranjan et al. (US 20180211099 A1, hereafter, "Ranjan"). Regarding claim 2, Kim in view of Mirza and Dal Mutto teaches the method according to claim 1, further comprising: [performing, when the first lens detection result indicates that there are no hand detection boxes in the first view, hand estimation on the first view of the current frame to obtain a first lens estimation result, the first lens estimation result being configured for positioning hand estimation boxes in the first view of the current frame; removing, from the hand estimation boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand estimation boxes]; and performing hand joint point recognition on the remaining hand estimation boxes to obtain the two-dimensional joint points corresponding to the current frame (See Kim, [Col. 7, ln. 16–18], In operation 130, the electronic device 10 may obtain first skeleton data HCl including at least one keypoint of the hand H from the first ROI ROI1. [Col. 7, ln. 21–25], For example, the first skeleton data HCl may be a data set including coordinate information of at least one keypoint (e.g., joint) of the hand H. Coordinate information of a keypoint of the hand H may include Two-Dimensional (2D). [Col. 8, ln. 21–25], In one embodiment, each of the keypoints included in the second skeleton data HC2 obtained using the third deep learning model may correspond to 2D position coordinates with respect to a preset origin on a plane including the second image IM2). However, Kim fail to teach performing, when the first lens detection result indicates that there are no hand detection boxes in the first view, hand estimation on the first view of the current frame to obtain a first lens estimation result, the first lens estimation result being configured for positioning hand estimation boxes in the first view of the current frame; removing, from the hand estimation boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand estimation boxes. Mirza, working in the same field of endeavor, teaches: removing, from the hand estimation boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand estimation boxes (See Mirza, ¶ [0318], For example, region 2222 may be a bounding box determined for the contour 2220 using a non-maximum suppression object-detection algorithm. For instance, the sensor client 105a may determine a plurality of bounding boxes associated with the contour 2220. For each bounding box, the client 105a may calculate a score. The score, for example, may represent an extent to which that bounding box is similar to the other bounding boxes. The sensor client 105a may identify a subset of the bounding boxes with a score that is greater than a threshold value (e.g., 80% or more), and determine region 2222 based on this identified subset). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s reference to removing, from the hand estimation boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand estimation boxes based on the method of Mirza’s reference. The suggestion/motivation would have been to reliably distinguish between close regions and accurately determine the best candidate region (See Mirza, ¶ [0323]). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results. However, Kim, Mirza and Dal Mutto fail to teach performing, when the first lens detection result indicates that there are no hand detection boxes in the first view, hand estimation on the first view of the current frame to obtain a first lens estimation result, the first lens estimation result being configured for positioning hand estimation boxes in the first view of the current frame. Ranjan, working in the same field of endeavor, teaches: performing, when the first lens detection result indicates that there are no hand detection boxes in the first view, hand estimation on the first view of the current frame to obtain a first lens estimation result, the first lens estimation result being configured for positioning hand estimation boxes in the first view of the current frame (See Ranjan, ¶ [0043], The network may fail to detect that face due to low score. In these situations, a candidate box that precisely captures the face may be beneficial. Hence, a new candidate bounding box from the predicted landmark points can be provided, for example using a FaceRectCalculator). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s, Mirza’s and Dal Mutto’s reference to performing, when the first lens detection result indicates that there are no hand detection boxes in the first view, hand estimation on the first view of the current frame to obtain a first lens estimation result, the first lens estimation result being configured for positioning hand estimation boxes in the first view of the current frame based on the method of Ranjan’s reference. The suggestion/motivation would have been to save time calculating another region (See Ranjan, ¶ [0004–0006]). Therefore, it would have been obvious to combine Ranjan with Kim, Mirza and Dal Mutto to obtain the invention as specified in claim 2. Regarding claim 12, claim 12 is rejected the same as claim 2 and the arguments similar to that presented above for claim 2 are equally applicable to the claim 12, and all of the other limitations similar to claim 2 are not repeated herein, but incorporated by reference. Regarding claim 18, claim 18 is rejected the same as claim 2 and the arguments similar to that presented above for claim 2 are equally applicable to the claim 18, and all of the other limitations similar to claim 2 are not repeated herein, but incorporated by reference. Claims 3, 7, 11, 15 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 12573080 B2, hereafter, "Kim") in view of Mirza et al. (US 20210124951 A1, hereafter, "Mirza"), Dal Mutto et al. (US 20190364206 A1, hereafter, "Dal Mutto"), and further in view of Gajria et al. (See NPL attached, "SRDD-Net: A complete pipeline for depth-maps based dynamic hand gesture recognition", hereafter, "Gajria"). Regarding claim 3, Kim in view of Mirza and Dal Mutto teaches the method according to claim 1, wherein the performing hand estimation on a second view of the current frame to obtain a second lens estimation result (See Kim, [Col. 9, ln. 58–60], In operation S240, the electronic device may obtain a second ROI from the second image based on the first skeleton data) comprises: [acquiring three-dimensional joint points of each of previous two frames of the current frame in the world coordinate system; estimating the three-dimensional joint points of the current frame in the world coordinate system according to the three-dimensional joint points of each of the previous two frames in the world coordinate system], and reprojecting estimated three-dimensional joint points of the current frame in the world coordinate system to obtain estimated two-dimensional joint points of the current frame (See Kim, [Col. 9, ln. 58–64], In operation S240, the electronic device may obtain a second ROI from the second image based on the first skeleton data. In one embodiment, an operation of obtaining the second ROI from the second image based on the first skeleton data may include projecting 3D position coordinates of a keypoint included in the first skeleton data to 2D position coordinates on the second image. [Col. 16, ln. 45–49], For example, a weight for combining 3D coordinate values of a first joint with 3D coordinate values of a second joint may be determined based on 3D joint coordinate values in a previous frame on a time axis); and determining the hand estimation boxes in the second view of the current frame according to the estimated two-dimensional joint points of the current frame (See Kim, [Col. 9, ln. 64–67], identifying a block including 2D position coordinates in the second 65 image, and determining the identified block as the second ROI). However, Kim, Mirza and Dal Mutto fail to teach acquiring three-dimensional joint points of each of previous two frames of the current frame in the world coordinate system; estimating the three-dimensional joint points of the current frame in the world coordinate system according to the three-dimensional joint points of each of the previous two frames in the world coordinate system. Gajria, working in the same field of endeavor, teaches: acquiring three-dimensional joint points of each of previous two frames of the current frame in the world coordinate system (See Gajria, [Pg. 1541, Col. 1, ln. 4–5], Our first step is to convert a stream of image frames to a stream of hand poses. [Pg. 1541, Col. 1, ln. 9–11], Our implementation makes use of Stacked Regression Network proposed by [3], which uses depth maps to estimate 21 2D hand points. [Pg. 1542, Col. 1, ln. 13–16], Thus, an equation for conversion of image coordinates to world coordinates, with the help of depth information, is established so as to feed our output of SRN into the next module); estimating the three-dimensional joint points of the current frame in the world coordinate system according to the three-dimensional joint points of each of the previous two frames in the world coordinate system (See Gajria, [Pg. 1541, Col. 1, ln. 4–5], Our first step is to convert a stream of image frames to a stream of hand poses. [Pg. 1541, Col. 1, ln. 9–11], Our implementation makes use of Stacked Regression Network proposed by [3], which uses depth maps to estimate 21 2D hand points. [Pg. 1542, Col. 1, ln. 13–16], Thus, an equation for conversion of image coordinates to world coordinates, with the help of depth information, is established so as to feed our output of SRN into the next module). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s, Mirza’s and Dal Mutto’s reference to acquiring three-dimensional joint points of each of previous two frames of the current frame in the world coordinate system; estimating the three-dimensional joint points of the current frame in the world coordinate system according to the three-dimensional joint points of each of the previous two frames in the world coordinate system, and reprojecting estimated three-dimensional joint points of the current frame in the world coordinate system to obtain estimated two-dimensional joint points of the current frame based on the method of Gajria’s reference. The suggestion/motivation would have been to increase the accuracy of gesture recognition and estimation (See Gajria, [TABLE 1]). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Gajria with Kim, Mirza and Dal Mutto to obtain the invention as specified in claim 3. Regarding claim 7, Kim in view of Mirza and Dal Mutto teaches the method according to claim 1, further comprising: [acquiring the three-dimensional joint points of each of the previous two frames of the current frame in the world coordinate system; calculating pose estimation parameters of each of the previous two frames according to the three-dimensional joint points of each of the previous two frames in the world coordinate system; and performing interpolation on the pose estimation parameters of each of the previous two frames to obtain the pose estimation parameters corresponding to the current frame]. However, Kim, Mirza and Dal Mutto fail to teach acquiring the three-dimensional joint points of each of the previous two frames of the current frame in the world coordinate system; calculating pose estimation parameters of each of the previous two frames according to the three-dimensional joint points of each of the previous two frames in the world coordinate system; and performing interpolation on the pose estimation parameters of each of the previous two frames to obtain the pose estimation parameters corresponding to the current frame. Gajria, working in the same field of endeavor, teaches: acquiring the three-dimensional joint points of each of the previous two frames of the current frame in the world coordinate system (See Gajria, [Pg. 1541, Col. 1, ln. 4–5], Our first step is to convert a stream of image frames to a stream of hand poses. [Pg. 1541, Col. 1, ln. 9–11], Our implementation makes use of Stacked Regression Network proposed by [3], which uses depth maps to estimate 21 2D hand points. [Pg. 1542, Col. 1, ln. 13–16], Thus, an equation for conversion of image coordinates to world coordinates, with the help of depth information, is established so as to feed our output of SRN into the next module); calculating pose estimation parameters of each of the previous two frames according to the three-dimensional joint points of each of the previous two frames in the world coordinate system (See Gajria, [Pg. 1541, Col. 1, ln. 4–5], Our first step is to convert a stream of image frames to a stream of hand poses. [Pg. 1541, Col. 1, ln. 9–11], Our implementation makes use of Stacked Regression Network proposed by [3], which uses depth maps to estimate 21 2D hand points. [Pg. 1542, Col. 1, ln. 13–16], Thus, an equation for conversion of image coordinates to world coordinates, with the help of depth information, is established so as to feed our output of SRN into the next module); and performing interpolation on the pose estimation parameters of each of the previous two frames to obtain the pose estimation parameters corresponding to the current frame (See Gajria, [Pg. 1541, Col. 2, ln. 10 and 11], The world coordinates, when multiplied with the intrinsics matrix, gives us a projection of the world coordinates). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s, Mirza’s and Dal Mutto’s reference to acquiring the three-dimensional joint points of each of the previous two frames of the current frame in the world coordinate system; calculating pose estimation parameters of each of the previous two frames according to the three-dimensional joint points of each of the previous two frames in the world coordinate system; and performing interpolation on the pose estimation parameters of each of the previous two frames to obtain the pose estimation parameters corresponding to the current frame based on the method of Dajria’s reference. The suggestion/motivation would have been to increase the accuracy of gesture recognition and estimation (See Gajria, [TABLE 1]). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Gajria with Kim, Mirza and Dal Mutto to obtain the invention as specified in claim 7. Regarding claim 11, claim 11 is rejected the same as claim 3 and the arguments similar to that presented above for claim 3 are equally applicable to the claim 11, and all of the other limitations similar to claim 3 are not repeated herein, but incorporated by reference. Regarding claim 15, claim 15 is rejected the same as claim 7 and the arguments similar to that presented above for claim 7 are equally applicable to the claim 15, and all of the other limitations similar to claim 7 are not repeated herein, but incorporated by reference. Regarding claim 19, claim 19 is rejected the same as claim 7 and the arguments similar to that presented above for claim 7 are equally applicable to the claim 19, and all of the other limitations similar to claim 7 are not repeated herein, but incorporated by reference. Claim(s) 4 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 12573080 B2, hereafter, "Kim") in view of Mirza et al. (US 20210124951 A1, hereafter, "Mirza"), Dal Mutto et al. (US 20190364206 A1, hereafter, "Dal Mutto"), and further in view of Wang et al. (US 20220405502 A1, hereafter, "Wang") and Kuo et al. (US 9298974 B1, hereafter, "Kuo"). Regarding claim 4, Kim in view of Mirza and Dal Mutto teaches the method according to claim 1, [wherein the removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes comprises: performing hand matching on the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, and selecting, according to a matching result, a pair of boxes matched as the same left hand and a pair of boxes matched as the same right hand; reserving, if there are a plurality of different left hands, a pair of boxes of a left hand corresponding to a box closest to an image center; and reserving, if there are a plurality of different right hands, a pair of boxes of a right hand corresponding to a box closest to the image center]. However, Kim fail to teach wherein the removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes comprises: performing hand matching on the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, and selecting, according to a matching result, a pair of boxes matched as the same left hand and a pair of boxes matched as the same right hand; reserving, if there are a plurality of different left hands, a pair of boxes of a left hand corresponding to a box closest to an image center; and reserving, if there are a plurality of different right hands, a pair of boxes of a right hand corresponding to a box closest to the image center. Mirza, working in the same field of endeavor, teaches: wherein the removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes (See Mirza, ¶ [0323], In these embodiments, the regions 2202c and 2204c are determined using a unique method referred to in this disclosure as “non-minimum suppression.” Non-minimum suppression may involve, for example, determining bounding boxes associated with the contour 2202b, 2204b (e.g., using any appropriate object detection algorithm as appreciated by a person of skilled in the relevant art)). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s reference to wherein the removing, from the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, redundant boxes corresponding to redundant hands to obtain remaining hand detection boxes based on the method of Mirza’s reference. The suggestion/motivation would have been to reliably distinguish between close regions and accurately determine the best candidate region (See Mirza, ¶ [0323]). However, Kim, Mirza and Dal Mutto fail to teach performing hand matching on the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, and selecting, according to a matching result, a pair of boxes matched as the same left hand and a pair of boxes matched as the same right hand; reserving, if there are a plurality of different left hands, a pair of boxes of a left hand corresponding to a box closest to an image center; and reserving, if there are a plurality of different right hands, a pair of boxes of a right hand corresponding to a box closest to the image center. Wang, working in the same field of endeavor, teaches: performing hand matching on the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, and selecting, according to a matching result, a pair of boxes matched as the same left hand and a pair of boxes matched as the same right hand (See Wang, ¶ [0143], The target circumscribed box may include two circumscribed boxes formed by each body and two corresponding hand bounding boxes, such as the left hand bounding box and the right hand bounding box, or may include multiple circumscribed boxes formed by each body and multiple corresponding hand bounding boxes, such as multiple left hand bounding boxes and/or multiple right hand bounding boxes. In this manner, a candidate degree of association consistent with a certain condition is screened based on the target degree of association corresponding to the circumscribed box formed by each body and each hand, and the hand matching each body is further determined). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s, Mirza’s and Dal Mutto’s reference to performing hand matching on the hand detection boxes in the first view and the hand estimation boxes in the second view of the current frame, and selecting, according to a matching result, a pair of boxes matched as the same left hand and a pair of boxes matched as the same right hand based on the method of Wang’s reference. The suggestion/motivation would have been to reliably and accurately detect hand keypoints (See Wang, ¶ [0002–0006]). However, Kim, Mirza, Dal Mutto and Wang fail to teach reserving, if there are a plurality of different left hands, a pair of boxes of a left hand corresponding to a box closest to an image center; and reserving, if there are a plurality of different right hands, a pair of boxes of a right hand corresponding to a box closest to the image center. Kuo, working in the same field of endeavor, teaches: reserving, if there are a plurality of different left hands, a pair of boxes of a left hand corresponding to a box closest to an image center (See Kuo, [Col. 9, ln. 58–63], In another example, a first average distance to center between the first right bounding box and a right image center and the first left bounding box and a left image center is calculated to determine the third selection criterion corresponding to the respective face nearest the center of the first image and the second image); and reserving, if there are a plurality of different right hands, a pair of boxes of a right hand corresponding to a box closest to the image center (See Kuo, [Col. 9, ln. 58–63], In another example, a first average distance to center between the first right bounding box and a right image center and the first left bounding box and a left image center is calculated to determine the third selection criterion corresponding to the respective face nearest the center of the first image and the second image). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s, Mirza’s, Dal Mutto’s and Wang’s reference to reserving, if there are a plurality of different left hands, a pair of boxes of a left hand corresponding to a box closest to an image center; and reserving, if there are a plurality of different right hands, a pair of boxes of a right hand corresponding to a box closest to the image center based on the method of Kuo’s reference. The suggestion/motivation would have been to reliably and accurately track object by minimizing processing and power use (See Kuo, ¶ [Col. 1, ln. 6–27]). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Kuo with Kim, Mirza, Dal Mutto and Wang to obtain the invention as specified in claim 4. Regarding claim 12, claim 12 is rejected the same as claim 4 and the arguments similar to that presented above for claim 4 are equally applicable to the claim 12, and all of the other limitations similar to claim 4 are not repeated herein, but incorporated by reference. Claims 5 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 12573080 B2, hereafter, "Kim") in view of Mirza et al. (US 20210124951 A1, hereafter, "Mirza"), Dal Mutto et al. (US 20190364206 A1, hereafter, "Dal Mutto"), and further in view of Li et al. (US 20210333884 A1, hereafter, "Li") and Kuo et al. (US 9298974 B1, hereafter, "Kuo"). Regarding claim 5, Kim in view of Mirza and Dal Mutto teaches the method according to claim 1, wherein before performing hand joint point recognition (See Kim, [Col. 7, ln. 16–18], In operation 130, the electronic device 10 may obtain first skeleton data HCl including at least one keypoint of the hand H from the first ROI ROI1), the method further comprises: [acquiring a plurality of cached detection boxes, the cached detection boxes being hand detection boxes in images obtained by performing hand detection in historical frames corresponding to the current frame; performing, for each of the hand detection boxes in the current frame, hand matching on the hand detection box and the plurality of cached detection boxes, and determining, from the plurality of cached detection boxes, hand detection boxes belonging to the same hand as the hand detection box in the current frame according to a matching result; and performing voting according to the hand detection boxes belonging to the same hand as the hand detection box in the current frame and the hand detection box in the current frame to obtain a voting result about the hand detection box in the current frame]. However, Kim, Mirza and Dal Mutto fail to teach acquiring a plurality of cached detection boxes, the cached detection boxes being hand detection boxes in images obtained by performing hand detection in historical frames corresponding to the current frame; performing, for each of the hand detection boxes in the current frame, hand matching on the hand detection box and the plurality of cached detection boxes, and determining, from the plurality of cached detection boxes, hand detection boxes belonging to the same hand as the hand detection box in the current frame according to a matching result; and performing voting according to the hand detection boxes belonging to the same hand as the hand detection box in the current frame and the hand detection box in the current frame to obtain a voting result about the hand detection box in the current frame. Li, working in the same field of endeavor, teaches: acquiring a plurality of cached detection boxes, the cached detection boxes being hand detection boxes in images obtained by performing hand detection in historical frames corresponding to the current frame (See Li, ¶ [0073], For example, in some embodiments the gesture data 504 may include hand location data (e.g., absolute location within the frame, location relative to the steering wheel, and/or hand bounding box coordinates), hand movement data, and/or hand location history data (indicating the location of the hand in one more previous frames)). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s, Mirza’s and Dal Mutto’s reference to acquiring a plurality of cached detection boxes, the cached detection boxes being hand detection boxes in images obtained by performing hand detection in historical frames corresponding to the current frame based on the method of Li’s reference. The suggestion/motivation would have been to accurately detect gestures for fine grained control (See Li, ¶ [0005–0007]). However, Kim, Mirza, Dal Mutto and Li fail to teach performing, for each of the hand detection boxes in the current frame, hand matching on the hand detection box and the plurality of cached detection boxes, and determining, from the plurality of cached detection boxes, hand detection boxes belonging to the same hand as the hand detection box in the current frame according to a matching result; and performing voting according to the hand detection boxes belonging to the same hand as the hand detection box in the current frame and the hand detection box in the current frame to obtain a voting result about the hand detection box in the current frame. Wang, working in the same field of endeavor, teaches: performing, for each of the hand detection boxes in the current frame, hand matching on the hand detection box and the plurality of cached detection boxes, and determining, from the plurality of cached detection boxes, hand detection boxes belonging to the same hand as the hand detection box in the current frame according to a matching result (See Wang, ¶ [0143], The target circumscribed box may include two circumscribed boxes formed by each body and two corresponding hand bounding boxes, such as the left hand bounding box and the right hand bounding box, or may include multiple circumscribed boxes formed by each body and multiple corresponding hand bounding boxes, such as multiple left hand bounding boxes and/or multiple right hand bounding boxes. In this manner, a candidate degree of association consistent with a certain condition is screened based on the target degree of association corresponding to the circumscribed box formed by each body and each hand, and the hand matching each body is further determined); and performing voting according to the hand detection boxes belonging to the same hand as the hand detection box in the current frame and the hand detection box in the current frame to obtain a voting result about the hand detection box in the current frame (See Wang, ¶ [0143], The target circumscribed box may include two circumscribed boxes formed by each body and two corresponding hand bounding boxes, such as the left hand bounding box and the right hand bounding box, or may include multiple circumscribed boxes formed by each body and multiple corresponding hand bounding boxes, such as multiple left hand bounding boxes and/or multiple right hand bounding boxes. In this manner, a candidate degree of association consistent with a certain condition is screened based on the target degree of association corresponding to the circumscribed box formed by each body and each hand, and the hand matching each body is further determined. Note: Examiner is interpreting the degree of association as the voting). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s, Mirza’s, Dal Mutto’s and Li’s reference to performing, for each of the hand detection boxes in the current frame, hand matching on the hand detection box and the plurality of cached detection boxes, and determining, from the plurality of cached detection boxes, hand detection boxes belonging to the same hand as the hand detection box in the current frame according to a matching result; and performing voting according to the hand detection boxes belonging to the same hand as the hand detection box in the current frame and the hand detection box in the current frame to obtain a voting result about the hand detection box in the current frame based on the method of Wang’s reference. The suggestion/motivation would have been to reliably and accurately detect hand keypoints (See Wang, ¶ [0002–0006]). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Wang with Kim, Mirza, Dal Mutto and Li to obtain the invention as specified in claim 5. Regarding claim 13, claim 13 is rejected the same as claim 5 and the arguments similar to that presented above for claim 5 are equally applicable to the claim 13, and all of the other limitations similar to claim 5 are not repeated herein, but incorporated by reference. Claims 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 12573080 B2, hereafter, "Kim") in view of Mirza et al. (US 20210124951 A1, hereafter, "Mirza"), Dal Mutto et al. (US 20190364206 A1, hereafter, "Dal Mutto"), and further in view of Xiang et al. (US 20230252670 A1, hereafter, "Xiang"). Regarding claim 6, Kim in view of Mirza and Dal Mutto teaches the method according to claim 1, wherein the two-dimensional joint points represent two-dimensional coordinates of joint points in an image plane coordinate system; the three-dimensional hand joint points represent three-dimensional coordinates of the joint points in the three-dimensional hand coordinate system (See Kim, [Col. 8, ln. 42–49], 3D position coordinates may be obtained for each keypoint in the first skeleton data HC1 by using 2D position coordinates of a keypoint projected onto the first image IM1 and 2D position coordinates of a corresponding keypoint in the second skeleton data HC2. For example, a triangulation method may be used in an operation of obtaining 3D position coordinates by using two 2D position coordinates); and the converting the two-dimensional joint points into three-dimensional joint points in a three-dimensional hand coordinate system (See Kim, [Col. 8, ln. 42–49], 3D position coordinates may be obtained for each keypoint in the first skeleton data HC1 by using 2D position coordinates of a keypoint projected onto the first image IM1 and 2D position coordinates of a corresponding keypoint in the second skeleton data HC2. For example, a triangulation method may be used in an operation of obtaining 3D position coordinates by using two 2D position coordinates) comprises: [taking two-dimensional coordinates of a target joint point in the image plane coordinate system as an origin of a two-dimensional finger coordinate system; determining an adjacent joint point of the target joint point; determining two-dimensional coordinates of the adjacent joint point in the two-dimensional finger coordinate system according to an included angle of digital joints between the adjacent joint point and the target joint point in the two-dimensional finger coordinate system and a length of the digital joints]; and converting the two-dimensional coordinates of the adjacent joint point in the two-dimensional finger coordinate system into the three-dimensional coordinates in the three-dimensional hand coordinate system according to a conversion relationship between the two-dimensional finger coordinate system and the three-dimensional hand coordinate system (See Kim, [Col. 8, ln. 42–49], 3D position coordinates may be obtained for each keypoint in the first skeleton data HC1 by using 2D position coordinates of a keypoint projected onto the first image IM1 and 2D position coordinates of a corresponding keypoint in the second skeleton data HC2. For example, a triangulation method may be used in an operation of obtaining 3D position coordinates by using two 2D position coordinates). However, Kim, Mirza and Dal Mutto fail to teach taking two-dimensional coordinates of a target joint point in the image plane coordinate system as an origin of a two-dimensional finger coordinate system; determining an adjacent joint point of the target joint point; determining two-dimensional coordinates of the adjacent joint point in the two-dimensional finger coordinate system according to an included angle of digital joints between the adjacent joint point and the target joint point in the two-dimensional finger coordinate system and a length of the digital joints. Xiang, working in the same field of endeavor, teaches: taking two-dimensional coordinates of a target joint point in the image plane coordinate system as an origin of a two-dimensional finger coordinate system (See Xiang, [FIG. 6 and FIG. 2], (Note: the wrist joint in the origin of the coordinate system) PNG media_image1.png 456 451 media_image1.png Greyscale PNG media_image2.png 385 314 media_image2.png Greyscale ); determining an adjacent joint point of the target joint point (See Xiang, ¶ [0054], a direction vector of a vector formed by two adjacent hand key points in the hand coordinate system may be determined based on the hand structured connection information); determining two-dimensional coordinates of the adjacent joint point in the two-dimensional finger coordinate system according to an included angle of digital joints between the adjacent joint point and the target joint point in the two-dimensional finger coordinate system and a length of the digital joints (See Xiang, ¶ [0103], The first direction vector of the vector formed by the hand key points in the hand coordinate system is calculated based on the joint bending angles, and the first direction vector is converted into the second direction vector in the world coordinate system based on the Euler angles. The two-dimensional coordinates of multiple hand key points are acquired based on the heat maps to calculate the vector length of the vector). Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Kim’s Mirza’s and Dal Mutto’s reference to taking two-dimensional coordinates of a target joint point in the image plane coordinate system as an origin of a two-dimensional finger coordinate system; determining an adjacent joint point of the target joint point; determining two-dimensional coordinates of the adjacent joint point in the two-dimensional finger coordinate system according to an included angle of digital joints between the adjacent joint point and the target joint point in the two-dimensional finger coordinate system and a length of the digital joints based on the method of Xiang’s reference. The suggestion/motivation would have been to reduce the amount of data required and improve the detection of hand keypoints, (See Xiang, ¶ [0063]). Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results. Therefore, it would have been obvious to combine Xiang with Kim, Mirza and Dal Mutto to obtain the invention as specified in claim 6. Regarding claim 14, claim 14 is rejected the same as claim 6 and the arguments similar to that presented above for claim 6 are equally applicable to the claim 14, and all of the other limitations similar to claim 6 are not repeated herein, but incorporated by reference. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Keh et al. (US 20170084044 A1) teaches an electronic device according to various embodiments of the present disclosure includes: a first image sensor; a second image sensor; and a processor operatively coupled to the first image sensor and the image second sensor, configured to determine at least one Region of Interest (ROI) based on a first information acquired using the first image sensor, acquire second information corresponding to at least a part of the at least one ROI using the second image sensor, identify a motion related to the at least one ROI based on the second information, and perform a function corresponding to the motion. He et al. (US 20200193148 A1) teaches the specification discloses a computer-implemented method for user action determination, comprising: recognizing an item displacement action performed by a user; determining a first time and a first location of the item displacement action; recognizing a target item in a non-stationary state; determining a second time when the target item is in the non-stationary state and a second location where the target item is in the non-stationary state; and in response to determining that the first time matches the second time and the first location matches the second location, determining that the item displacement action of the user is performed with respect to the target item. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DION J SATCHER whose telephone number is (703)756-5849. The examiner can normally be reached Monday - Thursday 5:30 am - 2:30 pm, Friday 5:30 am - 9:30 am PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DION J SATCHER/Patent Examiner, Art Unit 2676 /SHEFALI D GORADIA/Primary Patent Examiner, Art Unit 2676
Read full office action

Prosecution Timeline

Jan 22, 2025
Application Filed
Sep 09, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737911
DEVICE AND METHOD FOR IMAGE PROCESSING INCLUDING VASCULAR IMAGE PROCESSING
3y 11m to grant Granted Sep 15, 2026
Patent 12718525
Classification and sawing of wood shingles using machine vision
3y 8m to grant Granted Aug 25, 2026
Patent 12718591
METHOD AND APPARATUS WITH TRAFFIC LIGHT RECOGNITION MODEL
2y 5m to grant Granted Aug 25, 2026
Patent 12682477
FOOT SHAPE MEASUREMENT APPARATUS AND COMPUTER PROGRAM
3y 2m to grant Granted Jul 14, 2026
Patent 12675857
METHOD FOR EXTENDING DYNAMIC RANGE OF IMAGE AND ELECTRONIC DEVICE
2y 10m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
99%
With Interview (+17.8%)
2y 10m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 52 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month