DETAILED ACTION
Claims 1, 4-8, 11-15, 18-21 are pending in this application. Claims 2, 3, 9, 10, 16, and 17 are cancelled. Claims 1, 4, 8, 11, 15 and 18 have been amended. Claim 21 has been added.
Response to Amendment
The amendment filed July 23rd 2026 in response to the Non-Final Office Action mailed on April 23rd 2026 has been entered.
Claims 1, 4-8, 11-15, 18-21 are currently pending in this application
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Response to Arguments
Applicant's arguments filed July 23rd 2026 are considered moot since there is a new grounds for rejection.
Applicant’s arguments for the remaining limitations have been fully considered but they are not persuasive, however to account for the amendments to the claim, the rejection is changed from 35 U.S.C. § 102(a)(2) to 35 U.S.C. § 103.
Applicant has suggested that Li does not teach the initial limitations of claim 1, prior to the addition to the amendments, however since applicant has not addressed these limitations, the argument is not persuasive, and Examiner maintains the rejection, however has changed the rejection from a 35 U.S.C. § 102(a)(2) to a 35 U.S.C. § 103 to account for the amendments to the claim.
Applicant has amended claims 1, 8 and 15 to add the following additional limitations:
(i) the hand region prediction result is determined based on a hand detection result sequence;
(ii) the hand detection result sequence is determined based on one or more previous images of the target image;
(iii) the hand detection result is a hand three- dimensional joint point corresponding to the previous image of the target image;
(iv) the hand region prediction result is determined by using the hand 3D joint points corresponding to the previous image of the target image to predict the hand 3D joint point corresponding to the target image and projecting the predicted hand 3D joint point onto the target image.
Applicant argues that the initial references of Li and Kim do not teach the additional limitations added to claim 1. Limitation (i) (ii) and (iii) were limitations previously presented in claim 2 and 3 of the application, and Examiner was not persuaded by the arguments therefore the rejection using the Li reference to address these limitations still hold.
Li describes the hand region prediction result is determined based on a hand detection result sequence, paragraph 30 states that the location of the hands [the hand region prediction result] is determined by processing a sequence of frames [tracking system of successive frames tracking the hands [hand detection results sequence].
PNG
media_image1.png
118
1205
media_image1.png
Greyscale
Li describes the hand detection result sequence is determined based on one or more previous images of the target image, paragraph 30 states that processing a sequence of successive frames tracking the hands [hand detection results sequence], involves looking at previous frame of the hands.
Li describes the hand detection result is a hand three-dimensional joint point corresponding to the previous image of the target image; Paragraph 57 states that the hand pose estimation system [hand detection result] uses a plurality of joint locations on previous frames of the hand to identify a pose.
PNG
media_image2.png
85
1085
media_image2.png
Greyscale
Examiner agrees that Li does not explicitly teach the added limitation (iv): the hand region prediction result is determined by using the hand 3D joint points corresponding to the previous image of the target image to predict the hand 3D joint point corresponding to the target image and projecting the predicted hand 3D joint point onto the target image. Applicant argues that Li paragraph 30 does not teach the specific steps of obtaining the prediction predicting 3d hand key points of the target image according to 3d hand key points of previous images, and obtaining the predicted hand region via projection. Examiner agrees that Li does not explicitly teach those limitations, however, there is a new grounds for rejection (Shakar) used in combination with the Li reference that teaches the limitation.
Examiner holds rejection for all dependent claims, as a 35 USC § 103 rejection for claim 1 remains.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4-5, 8, 11-12, 15, 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over US20210074016 (hereinafter referred to as “Li”) in view of US20140232631 (hereinafter referred to as “Shahar”).
Regarding claim 1, Li teaches acquiring a target image wherein the target image is at least one of a plurality of images acquired by multi-view cameras of the extended reality device [See paragraph 0015 which discusses that the image is cropped to target the hands, receives a pair of images, and uses a stereo camera, which is a camera with multiple lenses so that the camera can obtain multiple views of
an image. Additionally see 0003 where this method can be used for augmented reality applications]
PNG
media_image3.png
89
829
media_image3.png
Greyscale
PNG
media_image4.png
36
809
media_image4.png
Greyscale
obtaining a hand region prediction result [See paragraph 0030 where the preprocessing system can crop the region of the image containing a hand. See also paragraph 0047 which discusses a 3D pose prediction module which creates a prediction for the hand based on the results from the machine learning module];
PNG
media_image5.png
61
844
media_image5.png
Greyscale
PNG
media_image1.png
118
1205
media_image1.png
Greyscale
and determining a corresponding hand detection result based on the target image and the hand region prediction result [see paragraph 0017 which identifies the 3D hand pose based on the images of the hand and the result from the hand prediction module once it identifies the joint locations].
PNG
media_image6.png
200
848
media_image6.png
Greyscale
Wherein the hand region prediction result is determined based on a hand detection result sequence; and the hand detection result sequence is determined based on one or more previous images of the target image [See paragraph 0030 where the sequence of image frames are used to estimate the location of the hand in successive frames.]
PNG
media_image1.png
118
1205
media_image1.png
Greyscale
Although Li also uses the hand detection process to determine the 3-D joint points of the hand [see paragraph 0057 where the joint locations of the hand are identified],
PNG
media_image2.png
85
1085
media_image2.png
Greyscale
Li does not explicitly teach obtaining the region of the hand from the image, or that the detection result uses the 3-D joint points from a previous image of the target or projecting the joint points on the target image.
Shakar does teach that the hand region prediction result is determined by using the hand 3D joint points corresponding to the previous image of the target image to predict the hand 3D joint point corresponding to the target image [See Shahar 0052 which teaches that the hand prediction can be made using the pose of the hand in a previous frame to predict the skeleton (joints) for the for the current frame (target frame). See also Fig. 9 where the hand region is segmented].
PNG
media_image7.png
212
980
media_image7.png
Greyscale
and projecting the predicted hand 3D joint point onto the target image [See paragraph 0052 of Shahar above where a 3D hand skeleton model is created for the current frame, depicted in Fig 11 where the joints are visible projected onto the hand image].
Therefore it would have been obvious to one with ordinary skill in the art to combine the detection of images of a hand, and obtain the joints of Li with using the joint position of the previous frame to predict the subsequent frame. The motivation to combine would be to improve the efficiency and accuracy of interaction between user's hands and electronic devices that use real-time tracking of gestures.
PNG
media_image8.png
123
994
media_image8.png
Greyscale
Regarding claim 4, Li teaches the hand detection result sequence comprises a hand detection result determined from an image before N frames; the images before N frames is the plurality of images acquired by the multi-view cameras. [See paragraph 0030 where the hand pose estimation system analyzes the captured images from the cameras continuously to estimate the location of the hand. Since detection is continuous, the results are determined at regular intervals, therefore they would be determined both before and after an N number of frames].
Regarding claim 5, Li teaches that the target image is one of the plurality of images acquired by the multi-view cameras of the extended reality device [See paragraph 0015 above where the target image is acquired by the stereo camera].
Claims 8 and 15 are similarly analyzed to Claim 1.
Claims 11 and 18 are similarly analyzed to Claim 4.
Claims 12 and 19 are similarly analyzed to Claim 5.
Claims 6, 7, 13, 14, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Li in view of Shakar in further view of US 20230186512 (hereafter referred to as "Kim").
Regarding claim 6, Li and Shakar teach the method according to claim 5, and contain multiple cameras from which the target image can be captured (stereo cameras).
Li and Shakar do not explicitly disclose that the camera that the previous target image is sourced from is different.
Kim teaches that the camera from which the target image is sourced is different from a camera from which a previous image of the target image is sourced [See paragraph 00129 of Kim which obtains a first and second image of the same hand]
PNG
media_image9.png
407
1174
media_image9.png
Greyscale
Li, Shakar and Kim are from similar fields of endeavor as they both aim to obtain the position of a hand in an image. It would have been obvious to one with ordinary skill in the art before the effective filing date to combine the pose recognition method of Li with obtaining successive images from different camera in "so that a total computation time required for an operation of obtaining 3D skeleton data of the hand and the amount of power consumed by the electronic device may be reduced" [See paragraph 0130 of Kim].
Regarding claim 7, Li and Shakar teach the method according to claim 5, and contain multiple cameras from which the target image can be captured (stereo cameras).
Li and Shakar do not explicitly disclose that there is a preset camera order for sourcing the target image. Kim teaches that the camera from which the target image is sourced is determined based on a preset camera order. [See paragraph 0117 of Kim which captures four images from four different cameras arranged in an a, b, C, d formation. These images contain the target image of the hand]
PNG
media_image10.png
271
981
media_image10.png
Greyscale
Claims 13 and 20 are similarly analyzed to Claim 6.
Claims 14 and 21 are similarly analyzed to Claim 7.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANUSHA KASHYAPA whose telephone number is (571)272-8766. The examiner can normally be reached Monday-Friday 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at (571) 272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANUSHA KASHYAPA/Examiner, Art Unit 2669 /CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669