Prosecution Insights
Last updated: October 02, 2026
Application No. 19/012,561

SYSTEMS AND METHODS FOR THREE-DIMENSIONAL (3D) POSE ESTIMATION

Non-Final OA §102§103
Filed
Jan 07, 2025
Priority
Apr 26, 2024 — provisional 63/639,391
Examiner
WOLFSON, ETHAN NOAH
Art Unit
Tech Center
Assignee
Samsung Electronics Co., Ltd.
OA Round
1 (Non-Final)
86%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
6 granted / 7 resolved
+25.7% vs TC avg
Strong +50% interview lift
Without
With
+50.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
26 currently pending
Career history
30
Total Applications
across all art units

Statute-Specific Performance

§101
4.7%
-35.3% vs TC avg
§103
75.6%
+35.6% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
8.7%
-31.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 7 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 01/07/2025 is being considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function. Claims 1-2, 5-6, and 19-20 recite limitations that use words like “means” (or “step”) or similar terms with functional language and do invoke 35 U.S.C. 112(f): Claim 1; recites the limitation, “by a computing device…..” [Line 3]. Claim 1; recites the limitation, “by the computing device…..” [Line 9]. Claim 2; recites the limitation, “by a computing device…..” [Line 1-2]. Claim 5; recites the limitation, “by the computing device…..” [Line 2]. Claim 5; recites the limitation, “by a computing device…..” [Line 4]. Claim 6; recites the limitation, “by the computing device…..” [Line 1-2]. Claim 19; recites the limitation, “means for processing…..” [Line 2]. Claim 19; recites the limitation, “by the means for processing…..” [Line 3-4]. Claim 19; recites the limitation, “by the means for processing…..” [Line 5]. Claim 19; recites the limitation, “output of the means for processing…..” [Line 11]. Claim 20; recites the limitation, “by the means for processing…..” [Line 1-2]. Claim 20; recites the limitation, “by the means for processing…..” [Line 2-3]. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 1-2, 5-6, and 19-20; (i) “computing device” (Fig. 1, processor #180. Paragraph [0086]-The system 1 may transmit the 3D pose- estimation data 70 to a computing device (e.g., the second processor and/or the second device 180, including the gesture-recognition circuit 182 and/or to the display circuit 184) for further processing and/or to change a setting on the device 100 (operation 5004). The computing device is illustrated in Fig. 1, as processor #180 thus has sufficient structure or material wherein is a processor.). (ii) “means for processing” (Fig. 4, #420. Paragraph [0066]-The electronic device 401 may communicate with the electronic device 404 via the server408. The electronic device 401 may include the processor 420 (e.g., a processing circuit or a means for processing), the memory 430, an input device 450, a sound output device 455, the display device 460, an audio module 470, a sensor module 476, an interface 477, a haptic module 479, a camera module 480, a power management module 488, a battery 489, a communication module 490, a subscriber identification module (SIM) card 496, or an antenna module 497. The means for processing is illustrated in Fig. 4 as processor #420, and thus has sufficient structure or material wherein is a processor.). If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1, 7, 9-10, 16, and 18-19 are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by LI et al. (US 20230161419 A1), hereinafter referenced as LI. Regarding claim 1, LI explicitly teaches a method for estimating a 3-dimensional (3D) pose (Fig. 4. Paragraph [0014]-LI discloses the following paragraphs describe an end-to-end approach to estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras.), the method comprising: receiving by a computing device (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.): a first input (Fig. 5. Paragraph [0053]-LI discloses in operation 504, the hand pose estimation system 122 identifies, using a neural network, a first set of joint location coordinates in the cropped portion of the given image (wherein the first set of joint location coordinates is a first input).) generated based on first features (Fig. 2. Paragraph [0030]-LI the pre-processing system 202 receives a pair of images from a stereo camera. The pre-processing system 202 identifies a region within each of the images containing a hand. The pre-processing system 202 further crops the region within each image containing the hand. Each image may contain multiple hands. The pre-processing system 202 crops each region within the image containing a hand (wherein the hands are first features).) associated with first image data from a first sensor associated with the computing device (Fig. 1, illustrates a first sensor associated with the computing device (wherein the client device has the first sensor in the stereo camera and is associated with server system #114 which houses the computing device). Paragraph [0018]-LI discloses client device 102 is a device of a given user who would like to capture a pair of images using a stereo camera coupled to the client device 102. The user captures a pair of images using a stereo camera of a three-dimensional hand pose. Further in paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the first view captures first image data).) and based on second image data from a second sensor associated with the computing device (Fig. 1, illustrates a second sensor associated with the computing device (wherein the client device has the second sensor in the stereo camera and is associated with server system #114 which houses the computing device). Paragraph [0018]-LI discloses client device 102 is a device of a given user who would like to capture a pair of images using a stereo camera coupled to the client device 102. The user captures a pair of images using a stereo camera of a three-dimensional hand pose. Further in paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the second view captures second image data).); and a second input generated based on second features (Fig. 5. Paragraph [0056]-LI discloses the first set of joint location coordinates is converted to a third set of joint location coordinates. The third set of joint location coordinates may be computed relative to an uncropped version of the image (e.g., the received stereo images). The third set of joint location coordinates relative to the uncropped version of the image may subsequently be converted to the second set of joint location coordinates (wherein the third set of joint location coordinates is the second input and the first set of joint coordinates is the second features).) associated with the first image data and based on the second image data (Fig. 4. Paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the first view captures first image data and the second view captures second image data).); based on the first input and the second input (Fig. 4. Paragraph [0051]-LI discloses in operation 406, the hand pose estimation system 122 identifies a 3D hand pose based on the plurality of sets of joint location coordinates (wherein the sets of joint location coordinates comprises the first set of joint location coordinates and the third set of joint location coordinates).), generating, by the computing device (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.), 3D pose-estimation data associated with an object (Fig. 4, #406 called identifying a three-dimensional hand pose based on the plurality of sets of joint location coordinates. Paragraph [0051]-LI discloses in operation 406, the hand pose estimation system 122 identifies a 3D hand pose based on the plurality of sets of joint location coordinates. The hand pose estimation system 122 may compare the identified joint locations to a predefined dataset of joint locations. For example, the hand pose estimation system 122 may determine that the hand pose is a “thumbs-up” pose based on the joint locations within the hand (wherein the identified hand pose is 3D pose-estimation data and the hand is an object).) represented in the first image data and represented in the second image data (Fig. 6, illustrates a stereo image captured by a stereo camera with a hand represented in a first image and the second image (wherein the left view is the first image and the right view is the second image). Paragraph [0057]-LI discloses FIG. 6 is an illustration of a stereo image pair of a three-dimensional hand pose according to some example embodiments. In some examples, the stereo image pair is captured by a head-mounted image capture device with cameras mounted on a left and right temple. The hand pose estimation system 122 analyzes the captured images and identifies joint locations in the hand.); and transmitting the 3D pose-estimation data (Fig. 1. Paragraph [0019]-LI discloses in some examples the identified three-dimensional hand pose causes the client device 102 to perform a predefined action. For example, upon identifying a 3D hand pose, the client device 102 may be prompted to capture an image, modify a captured image (e.g., annotate the captured image), annotate an image in real-time or near real-time, transmit a message via one or more applications (e.g., third-party application(s) 108, client application 110), or navigate menus within one or more applications (e.g., third-party application(s) 108, client application 110). The identified 3D hand pose may optionally be used to interact with augmented reality (AR) experiences using the client device 102 (wherein the identified three-dimensional hand pose is 3D pose-estimation data).). Regarding claim 7, LI explicitly teaches the method of claim 1, wherein: LI further explicitly teaches the first sensor is positioned at a first location on an enclosure (Fig. 1. Paragraph [0049]-LI discloses the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. Further in paragraph [0049-LI discloses operation 402 is performed on a wearable client device 102 (e.g., a head-mounted image capture device with a cameras mounted on a left and right temple) (wherein the camera mounted on the left temple of a wearable client device is the first sensor positioned at a first location on an enclosure).) associated with the computing device (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114 and associating itself with the client device #102. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.); the second sensor is positioned at a second location on the enclosure (Fig. 1. Paragraph [0049]-LI discloses the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. Further in paragraph [0049-LI discloses operation 402 is performed on a wearable client device 102 (e.g., a head-mounted image capture device with a cameras mounted on a left and right temple) (wherein the camera mounted on the right temple of a wearable client device is the second sensor positioned at a second location on the enclosure).) associated with the computing device (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114 and associating itself with the client device #102. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.); and the second location is a different location from the first location (Fig. 1. Paragraph [0049]-LI discloses the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. Further in paragraph [0049-LI discloses operation 402 is performed on a wearable client device 102 (e.g., a head-mounted image capture device with a cameras mounted on a left and right temple) (wherein the left and right temple are different locations).). Regarding claim 9, LI explicitly teaches the method of claim 1, wherein: LI further explicitly teaches the object comprises a hand (Fig. 6. Paragraph [0017]-LI discloses the client device 102 may be used to identify a three-dimensional hand pose using the captured pair of stereo images.); and the 3D pose-estimation data comprises 3D hand-joint data (Fig. 4. Paragraph [0014]-LI discloses a 3D hand pose may be defined based on the joint locations within the hand. The joint locations referred to herein may describe a point or articulation, (e.g., a connection) between two or more bones that allow for motion.). Regarding claim 10, LI explicitly teaches a system comprising (Fig. 4. Paragraph [0014]-LI discloses the following paragraphs describe an end-to-end approach to estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras.): one or more processors (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.); and a memory (Fig. 7, #704 called memory.) storing instructions which (Fig. 1 and 7. Paragraph [0060]-LI discloses the instructions 708 may also reside, completely or partially, within the main memory 712, within the static memory 714, within machine-readable medium 718 within the storage unit 716, within at least one of the processors 702 (e.g., within the processor's cache memory), or any suitable combination thereof, during execution thereof by the machine 700.), when executed by the one or more processors (Fig. 7. Paragraph [0059]-LI discloses the term “processor” is intended to include multi-core processors that may comprise two or more independent processors (sometimes referred to as “cores”) that may execute instructions contemporaneously.), cause performance of: receiving by the one or more processors (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.): a first input (Fig. 5. Paragraph [0053]-LI discloses in operation 504, the hand pose estimation system 122 identifies, using a neural network, a first set of joint location coordinates in the cropped portion of the given image (wherein the first set of joint location coordinates is a first input).) generated based on first features (Fig. 2. Paragraph [0030]-LI the pre-processing system 202 receives a pair of images from a stereo camera. The pre-processing system 202 identifies a region within each of the images containing a hand. The pre-processing system 202 further crops the region within each image containing the hand. Each image may contain multiple hands. The pre-processing system 202 crops each region within the image containing a hand (wherein the hands are first features).) associated with first image data from a first sensor (Fig. 1. Paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the first view captures first image data).) and based on second image data from a second sensor (Fig. 1. Paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the second view captures second image data).); and a second input generated based on second features (Fig. 5. Paragraph [0056]-LI discloses the first set of joint location coordinates is converted to a third set of joint location coordinates. The third set of joint location coordinates may be computed relative to an uncropped version of the image (e.g., the received stereo images). The third set of joint location coordinates relative to the uncropped version of the image may subsequently be converted to the second set of joint location coordinates (wherein the third set of joint location coordinates is the second input and the first set of joint coordinates is the second features).) associated with the first image data and based on the second image data (Fig. 4. Paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the first view captures first image data and the second view captures second image data).); generating, based on an output (Fig. 4. Paragraph [0051]-LI discloses in operation 406, the hand pose estimation system 122 identifies a 3D hand pose based on the plurality of sets of joint location coordinates (wherein the sets of joint location coordinates is an output).) of the one or more processors (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses Server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.), 3D pose-estimation data associated with an object (Fig. 4, #406 called identifying a three-dimensional hand pose based on the plurality of sets of joint location coordinates. Paragraph [0051]-LI discloses in operation 406, the hand pose estimation system 122 identifies a 3D hand pose based on the plurality of sets of joint location coordinates. The hand pose estimation system 122 may compare the identified joint locations to a predefined dataset of joint locations. For example, the hand pose estimation system 122 may determine that the hand pose is a “thumbs-up” pose based on the joint locations within the hand (wherein the identified hand pose is 3D pose-estimation data and the hand is an object).) represented in the first image data and represented in the second image data (Fig. 6, illustrates a stereo image captured by a stereo camera with a hand represented in a first image and the second image (wherein the left view is the first image and the right view is the second image). Paragraph [0057]-LI discloses FIG. 6 is an illustration of a stereo image pair of a three-dimensional hand pose according to some example embodiments. In some examples, the stereo image pair is captured by a head-mounted image capture device with cameras mounted on a left and right temple. The hand pose estimation system 122 analyzes the captured images and identifies joint locations in the hand.); and transmitting the 3D pose-estimation data (Fig. 1. Paragraph [0019]-LI discloses in some examples the identified three-dimensional hand pose causes the client device 102 to perform a predefined action. For example, upon identifying a 3D hand pose, the client device 102 may be prompted to capture an image, modify a captured image (e.g., annotate the captured image), annotate an image in real-time or near real-time, transmit a message via one or more applications (e.g., third-party application(s) 108, client application 110), or navigate menus within one or more applications (e.g., third-party application(s) 108, client application 110). The identified 3D hand pose may optionally be used to interact with augmented reality (AR) experiences using the client device 102 (wherein the identified three-dimensional hand pose is 3D pose-estimation data).). Regarding claim 16, LI explicitly teaches the system of claim 10, wherein: LI further explicitly teaches the first sensor is positioned at a first location on the system (Fig. 1. Paragraph [0049]-LI discloses the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. Further in paragraph [0049-LI discloses operation 402 is performed on a wearable client device 102 (e.g., a head-mounted image capture device with a cameras mounted on a left and right temple) (wherein the camera mounted on the left temple of a wearable client device is the first sensor positioned at a first location on a system).); the second sensor is positioned at a second location on the system (Fig. 1. Paragraph [0049]-LI discloses the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. Further in paragraph [0049-LI discloses operation 402 is performed on a wearable client device 102 (e.g., a head-mounted image capture device with a cameras mounted on a left and right temple) (wherein the camera mounted on the right temple of a wearable client device is the second sensor positioned at a second location on the system).); and the second location is a different location from the first location Fig. 1. Paragraph [0049]-LI discloses the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. Further in paragraph [0049-LI discloses operation 402 is performed on a wearable client device 102 (e.g., a head-mounted image capture device with a cameras mounted on a left and right temple) (wherein the left and right temple are different locations).). Regarding claim 18, LI explicitly teaches the system of claim 10, wherein: LI further explicitly teaches the object comprises a hand (Fig. 6. Paragraph [0017]-LI discloses the client device 102 may be used to identify a three-dimensional hand pose using the captured pair of stereo images.); and the 3D pose-estimation data comprises 3D hand-joint data (Fig. 4. Paragraph [0014]-LI discloses a 3D hand pose may be defined based on the joint locations within the hand. The joint locations referred to herein may describe a point or articulation, (e.g., a connection) between two or more bones that allow for motion.). Regarding claim 19, LI explicitly teaches a system comprising (Fig. 4. Paragraph [0014]-LI discloses the following paragraphs describe an end-to-end approach to estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras.): means for processing (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.); and a memory (Fig. 7, #704 called memory.) storing instructions which (Fig. 1 and 7. Paragraph [0060]-LI discloses the instructions 708 may also reside, completely or partially, within the main memory 712, within the static memory 714, within machine-readable medium 718 within the storage unit 716, within at least one of the processors 702 (e.g., within the processor's cache memory), or any suitable combination thereof, during execution thereof by the machine 700.), when executed by the means for processing (Fig. 7. Paragraph [0059]-LI discloses the term “processor” is intended to include multi-core processors that may comprise two or more independent processors (sometimes referred to as “cores”) that may execute instructions contemporaneously.), cause performance of: receiving, by the means for processing (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.): a first input (Fig. 5. Paragraph [0053]-LI discloses in operation 504, the hand pose estimation system 122 identifies, using a neural network, a first set of joint location coordinates in the cropped portion of the given image (wherein the first set of joint location coordinates is a first input).) generated based on first features (Fig. 2. Paragraph [0030]-LI the pre-processing system 202 receives a pair of images from a stereo camera. The pre-processing system 202 identifies a region within each of the images containing a hand. The pre-processing system 202 further crops the region within each image containing the hand. Each image may contain multiple hands. The pre-processing system 202 crops each region within the image containing a hand (wherein the hands are first features).) associated with first image data from a first sensor (Fig. 1. Paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the first view captures first image data).) and based on second image data from a second sensor (Fig. 1. Paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the second view captures second image data).); and a second input generated based on second features (Fig. 5. Paragraph [0056]-LI discloses the first set of joint location coordinates is converted to a third set of joint location coordinates. The third set of joint location coordinates may be computed relative to an uncropped version of the image (e.g., the received stereo images). The third set of joint location coordinates relative to the uncropped version of the image may subsequently be converted to the second set of joint location coordinates (wherein the third set of joint location coordinates is the second input and the first set of joint coordinates is the second features).) associated with the first image data and based on the second image data (Fig. 4. Paragraph [0049]-LI discloses at operation 402, the hand pose estimation system 122 receives a pair of images from a camera. In some examples, the camera is a stereo camera. The pair of images comprise a first view and a second view of a hand. In some examples the first view of the hand replicates a left eye view of the hand and the second view of the hand replicates a right eye view of the hand (wherein the first view captures first image data and the second view captures second image data).); generating, based on an output (Fig. 4. Paragraph [0051]-LI discloses in operation 406, the hand pose estimation system 122 identifies a 3D hand pose based on the plurality of sets of joint location coordinates (wherein the sets of joint location coordinates is an output).) of the means for processing (Fig. 1, illustrates Hand Pose Estimation System #122 as part of Server System 114. Paragraph [0018]-LI discloses Server system 114 receives the image and identifies the three-dimensional hand pose. Further in paragraph [0029]-LI discloses any system of the hand pose estimation system 122 may include software, hardware, or both, that configure an arrangement of one or more processors.), 3D pose-estimation data associated with an object (Fig. 4, #406 called identifying a three-dimensional hand pose based on the plurality of sets of joint location coordinates. Paragraph [0051]-LI discloses in operation 406, the hand pose estimation system 122 identifies a 3D hand pose based on the plurality of sets of joint location coordinates. The hand pose estimation system 122 may compare the identified joint locations to a predefined dataset of joint locations. For example, the hand pose estimation system 122 may determine that the hand pose is a “thumbs-up” pose based on the joint locations within the hand (wherein the identified hand pose is 3D pose-estimation data and the hand is an object).) in the first image data and in the second image data (Fig. 6, illustrates a stereo image captured by a stereo camera with a hand represented in a first image and the second image (wherein the left view is the first image and the right view is the second image). Paragraph [0057]-LI discloses FIG. 6 is an illustration of a stereo image pair of a three-dimensional hand pose according to some example embodiments. In some examples, the stereo image pair is captured by a head-mounted image capture device with cameras mounted on a left and right temple. The hand pose estimation system 122 analyzes the captured images and identifies joint locations in the hand.); and transmitting the 3D pose-estimation data (Fig. 1. Paragraph [0019]-LI discloses in some examples the identified three-dimensional hand pose causes the client device 102 to perform a predefined action. For example, upon identifying a 3D hand pose, the client device 102 may be prompted to capture an image, modify a captured image (e.g., annotate the captured image), annotate an image in real-time or near real-time, transmit a message via one or more applications (e.g., third-party application(s) 108, client application 110), or navigate menus within one or more applications (e.g., third-party application(s) 108, client application 110). The identified 3D hand pose may optionally be used to interact with augmented reality (AR) experiences using the client device 102 (wherein the identified three-dimensional hand pose is 3D pose-estimation data).). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2, 8, 11, 17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over LI et al. (US 20230161419 A1), hereinafter referenced as LI, in view of MAO et al. (US 20180024641 A1), hereinafter referenced as MAO. Regarding claim 2, LI explicitly teaches the method of claim 1, LI fails to explicitly teach further comprising receiving, by the computing device, a third input comprising first sensor parameters associated with the first sensor. However, MAO explicitly teaches further comprising receiving, by the computing device (Fig. 1. Paragraph [0028]-MAO discloses non-transitory computer-readable storage medium 105 may couple to processor 102 and may store instructions that, when executed by processor 102, perform the method(s) or step(s) described below. The instructions may be specialized and may include various machine learning models, inverse kinetics (IK) models, and/or other models and algorithms described in the present disclosure. In order to perform the steps or methods described below, processor 102 and/or the instructions (e.g., the machine learning models, inverse kinetics methods, other models or algorithms, etc.) may be specially trained.), a third input (Fig. 7B, illustrates the IK model. Paragraph [0056]-MAO discloses with known image parameters such as focal lengths, camera positions, and wide angles of the cameras, the 3D location can be calculated (wherein the known image parameters are a third input).) comprising first sensor parameters associated with the first sensor (Fig. 1 and 7B. Paragraph [0025]-MAO discloses the triangulation method may be based at least in part on pairs of 2D joint positions from the two stereo images, focal lengths of cameras respectively capturing the two stereo images, and position information of the cameras (e.g., relative positions between the cameras, relative positions of stereo images to the cameras) (wherein focal lengths of cameras respectively capturing the two stereo images, and position information of the cameras (e.g., relative positions between the cameras, relative positions of stereo images to the cameras) are parameters associated with the first sensor).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a method for estimating a 3-dimensional (3D) pose, the method comprising: receiving by a computing device: a first input generated based on first features associated with first image data from a first sensor associated with the computing device and based on second image data from a second sensor associated with the computing device; and a second input generated based on second features associated with the first image data and based on the second image data; based on the first input and the second input, generating, by the computing device, 3D pose-estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of MAO of further comprising receiving, by the computing device, a third input comprising first sensor parameters associated with the first sensor. Wherein having LI’s method for estimating hand poses from stereo cameras having further comprising receiving, by the computing device, a third input comprising first sensor parameters associated with the first sensor. The motivation behind the modification would have been to obtain a method for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and MAO relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while MAO overcoming the shortcomings in existing technologies and advance the gesture recognition technology, it is desirable to develop fast, robust, and reliable 3D hand tracking systems and methods. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and MAO et al. (US 20180024641 A1), Paragraph [0018]. Regarding claim 8, LI explicitly teaches the method of claim 1, wherein: LI fails to explicitly teach the first image data comprises first gray image data generated by the first sensor; and the second image data comprises second gray image data generated by the second sensor. However, MAO explicitly teaches the first image data comprises first gray image data (Fig. 2. Paragraph [0030]-MAO discloses step 201 is associated with image 201-d, which comprises two stereo images of a hand. The stereo images may be or may be converted to grey scale images, black-and-white images, etc.) generated by the first sensor (Fig. 1, #1012 called cameras. Paragraph [0019]-MAO discloses the one or more images may comprise two stereo images of the portion of the object, and the two stereo images may be captured by two cameras (e.g., infrared cameras).); and the second image data comprises second gray image data (Fig. 2. Paragraph [0030]-MAO discloses step 201 is associated with image 201-d, which comprises two stereo images of a hand. The stereo images may be or may be converted to grey scale images, black-and-white images, etc.) generated by the second sensor (Fig. 1, #1012 called cameras. Paragraph [0019]-MAO discloses the one or more images may comprise two stereo images of the portion of the object, and the two stereo images may be captured by two cameras (e.g., infrared cameras).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a method for estimating a 3-dimensional (3D) pose, the method comprising: receiving by a computing device: a first input generated based on first features associated with first image data from a first sensor associated with the computing device and based on second image data from a second sensor associated with the computing device; and a second input generated based on second features associated with the first image data and based on the second image data; based on the first input and the second input, generating, by the computing device, 3D pose-estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of MAO the first image data comprises first gray image data generated by the first sensor; and the second image data comprises second gray image data generated by the second sensor. Wherein having LI’s method for estimating hand poses from stereo cameras having the first image data comprises first gray image data generated by the first sensor; and the second image data comprises second gray image data generated by the second sensor. The motivation behind the modification would have been to obtain a method for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and MAO relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while MAO overcoming the shortcomings in existing technologies and advance the gesture recognition technology, it is desirable to develop fast, robust, and reliable 3D hand tracking systems and methods. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and MAO et al. (US 20180024641 A1), Paragraph [0018]. Regarding claim 11, LI explicitly teaches the system of claim 10, LI fails to explicitly teach wherein the instructions, when executed by the one or more processors, cause performance of receiving, by the one or more processors, a third input comprising first sensor parameters associated with the first sensor. However, MAO explicitly teaches wherein the instructions, when executed by the one or more processors (Fig. 1. Paragraph [0028]-MAO discloses non-transitory computer-readable storage medium 105 may couple to processor 102 and may store instructions that, when executed by processor 102, perform the method(s) or step(s) described below. The instructions may be specialized and may include various machine learning models, inverse kinetics (IK) models, and/or other models and algorithms described in the present disclosure. In order to perform the steps or methods described below, processor 102 and/or the instructions (e.g., the machine learning models, inverse kinetics methods, other models or algorithms, etc.) may be specially trained.), cause performance of receiving, by the one or more processors (Fig. 1. Paragraph [0028]-MAO discloses non-transitory computer-readable storage medium 105 may couple to processor 102 and may store instructions that, when executed by processor 102, perform the method(s) or step(s) described below. The instructions may be specialized and may include various machine learning models, inverse kinetics (IK) models, and/or other models and algorithms described in the present disclosure. In order to perform the steps or methods described below, processor 102 and/or the instructions (e.g., the machine learning models, inverse kinetics methods, other models or algorithms, etc.) may be specially trained.), a third input (Fig. 7B, illustrates the IK model. Paragraph [0056]-MAO discloses with known image parameters such as focal lengths, camera positions, and wide angles of the cameras, the 3D location can be calculated (wherein the known image parameters are a third input).) comprising first sensor parameters associated with the first sensor (Fig. 1 and 7B. Paragraph [0025]-MAO discloses the triangulation method may be based at least in part on pairs of 2D joint positions from the two stereo images, focal lengths of cameras respectively capturing the two stereo images, and position information of the cameras (e.g., relative positions between the cameras, relative positions of stereo images to the cameras) (wherein focal lengths of cameras respectively capturing the two stereo images, and position information of the cameras (e.g., relative positions between the cameras, relative positions of stereo images to the cameras) are parameters associated with the first sensor).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance of: receiving by the one or more processors: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the one or more processors, 3D pose- estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of MAO of wherein the instructions, when executed by the one or more processors, cause performance of receiving, by the one or more processors, a third input comprising first sensor parameters associated with the first sensor. Wherein having LI’s system for estimating hand poses from stereo cameras having wherein the instructions, when executed by the one or more processors, cause performance of receiving, by the one or more processors, a third input comprising first sensor parameters associated with the first sensor. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and MAO relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while MAO overcoming the shortcomings in existing technologies and advance the gesture recognition technology, it is desirable to develop fast, robust, and reliable 3D hand tracking systems and methods. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and MAO et al. (US 20180024641 A1), Paragraph [0018]. Regarding claim 17, LI explicitly teaches the system of claim 10, wherein: LI fails to explicitly teach the first image data comprises first gray image data generated by the first sensor; and the second image data comprises second gray image data generated by the second sensor. However, MAO explicitly teaches the first image data comprises first gray image data (Fig. 2. Paragraph [0030]-MAO discloses step 201 is associated with image 201-d, which comprises two stereo images of a hand. The stereo images may be or may be converted to grey scale images, black-and-white images, etc.) generated by the first sensor (Fig. 1, #1012 called cameras. Paragraph [0019]-MAO discloses the one or more images may comprise two stereo images of the portion of the object, and the two stereo images may be captured by two cameras (e.g., infrared cameras).); and the second image data comprises second gray image data (Fig. 2. Paragraph [0030]-MAO discloses step 201 is associated with image 201-d, which comprises two stereo images of a hand. The stereo images may be or may be converted to grey scale images, black-and-white images, etc.) generated by the second sensor (Fig. 1, #1012 called cameras. Paragraph [0019]-MAO discloses the one or more images may comprise two stereo images of the portion of the object, and the two stereo images may be captured by two cameras (e.g., infrared cameras).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance of: receiving by the one or more processors: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the one or more processors, 3D pose- estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of MAO of the first image data comprises first gray image data generated by the first sensor; and the second image data comprises second gray image data generated by the second sensor. Wherein having LI’s system for estimating hand poses from stereo cameras having the first image data comprises first gray image data generated by the first sensor; and the second image data comprises second gray image data generated by the second sensor. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and MAO relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while MAO overcoming the shortcomings in existing technologies and advance the gesture recognition technology, it is desirable to develop fast, robust, and reliable 3D hand tracking systems and methods. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and MAO et al. (US 20180024641 A1), Paragraph [0018]. Regarding claim 20, LI explicitly teaches the system of claim 19, LI fails to explicitly teach wherein the instructions, when executed by the means for processing, cause performance of receiving, by the means for processing, a third input comprising first sensor parameters associated with the first sensor. However, MAO explicitly teaches wherein the instructions, when executed by the means for processing (Fig. 1. Paragraph [0028]-MAO discloses non-transitory computer-readable storage medium 105 may couple to processor 102 and may store instructions that, when executed by processor 102, perform the method(s) or step(s) described below. The instructions may be specialized and may include various machine learning models, inverse kinetics (IK) models, and/or other models and algorithms described in the present disclosure. In order to perform the steps or methods described below, processor 102 and/or the instructions (e.g., the machine learning models, inverse kinetics methods, other models or algorithms, etc.) may be specially trained.), cause performance of receiving, by the means for processing (Fig. 1. Paragraph [0028]-MAO discloses non-transitory computer-readable storage medium 105 may couple to processor 102 and may store instructions that, when executed by processor 102, perform the method(s) or step(s) described below. The instructions may be specialized and may include various machine learning models, inverse kinetics (IK) models, and/or other models and algorithms described in the present disclosure. In order to perform the steps or methods described below, processor 102 and/or the instructions (e.g., the machine learning models, inverse kinetics methods, other models or algorithms, etc.) may be specially trained.), a third input (Fig. 7B, illustrates the IK model. Paragraph [0056]-MAO discloses with known image parameters such as focal lengths, camera positions, and wide angles of the cameras, the 3D location can be calculated (wherein the known image parameters are a third input).) comprising first sensor parameters associated with the first sensor (Fig. 1 and 7B. Paragraph [0025]-MAO discloses the triangulation method may be based at least in part on pairs of 2D joint positions from the two stereo images, focal lengths of cameras respectively capturing the two stereo images, and position information of the cameras (e.g., relative positions between the cameras, relative positions of stereo images to the cameras) (wherein focal lengths of cameras respectively capturing the two stereo images, and position information of the cameras (e.g., relative positions between the cameras, relative positions of stereo images to the cameras) are parameters associated with the first sensor).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a system comprising: means for processing; and a memory storing instructions which, when executed by the means for processing, cause performance of: receiving, by the means for processing: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the means for processing, 3D pose- estimation data associated with an object in the first image data and in the second image data; and transmitting the 3D pose-estimation data with the teachings of MAO of wherein the instructions, when executed by the means for processing, cause performance of receiving, by the means for processing, a third input comprising first sensor parameters associated with the first sensor. Wherein having LI’s system for estimating hand poses from stereo cameras having wherein the instructions, when executed by the means for processing, cause performance of receiving, by the means for processing, a third input comprising first sensor parameters associated with the first sensor. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and MAO relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while MAO overcoming the shortcomings in existing technologies and advance the gesture recognition technology, it is desirable to develop fast, robust, and reliable 3D hand tracking systems and methods. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and MAO et al. (US 20180024641 A1), Paragraph [0018]. Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over LI et al. (US 20230161419 A1), hereinafter referenced as LI, in view of HAMPALI et al. (US 20220301304 A1), hereinafter referenced as HAMPALI, and further in view of CHINANANDA et al. (US 20200372246 A1), hereinafter referenced as CHINANANDA. Regarding claim 3, LI explicitly teaches the method of claim 1, LI fails to explicitly teach wherein the 3D pose-estimation data is generated based on output data comprising output features with corresponding attentions, the output data being generated based on: performing a concatenation operation on the first input and the second input; and. However, HAMPALI explicitly teaches wherein the 3D pose-estimation data is generated (Fig. 6, column on the right called output pose. Paragraph [0103]-HAMPALI discloses the joint specific features enable estimation of different joint-related pose parameters, such as joint angle and joint vector. The output poses for the input images are illustrated in the fourth column of FIG. 6.) based on output data comprising output features with corresponding attentions (Fig. 6. Paragraph [0103]-HAMPALI discloses the first column includes input images, and the second column includes the keypoint heatmaps generated by the keypoint determination engine 250 for the input images. For each joint query, the respective colored circles in the third column (labeled “Joint Attention”) indicate the locations of the keypoints attended by the query. The radius of the circle in the Joint Attention column is proportional to the attention weight (wherein the radius proportional to the attention weight is the corresponding attention and the joint attention is output data).), the output data being generated based on (Fig. 6. Paragraph [0103]-HAMPALI discloses the first column includes input images, and the second column includes the keypoint heatmaps generated by the keypoint determination engine 250 for the input images. For each joint query, the respective colored circles in the third column (labeled “Joint Attention”) indicate the locations of the keypoints attended by the query. The radius of the circle in the Joint Attention column is proportional to the attention weight (wherein the radius proportional to the attention weight is the corresponding attention and the joint attention is output data).): performing a concatenation operation on the first input and the second input (Fig. 3. Paragraph [0076]-HAMPALI discloses the feature determination engine 254 can combine the features (e.g., by concatenating the features from the different feature maps) to form a feature representation, such as a feature vector (wherein the first input and second input are different feature maps).); and Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a method for estimating a 3-dimensional (3D) pose, the method comprising: receiving by a computing device: a first input generated based on first features associated with first image data from a first sensor associated with the computing device and based on second image data from a second sensor associated with the computing device; and a second input generated based on second features associated with the first image data and based on the second image data; based on the first input and the second input, generating, by the computing device, 3D pose-estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of HAMPALI of wherein the 3D pose-estimation data is generated based on output data comprising output features with corresponding attentions, the output data being generated based on: performing a concatenation operation on the first input and the second input; and. Wherein having LI’s method for estimating hand poses from stereo cameras having wherein the 3D pose-estimation data is generated based on output data comprising output features with corresponding attentions, the output data being generated based on: performing a concatenation operation on the first input and the second input; and. The motivation behind the modification would have been to obtain a method for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and HAMPALI relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while HAMPALI determining accurate poses of objects can allow a system to generate accurately positioned and oriented representations (e.g., models) of the objects. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and HAMPALI et al. (US 20220301304 A1), Paragraph [0005]. LI in view of HAMPALI fail to explicitly teach performing a second operation on the first input and an operand that is based on a result of the concatenation operation, the second operation being an add operation or a multiplication operation. However, CHIDANANDA explicitly teaches performing a second operation (Fig. 6I. Paragraph [0138]-CHIDANANDA discloses the output of the activation layer 610I and that from the block 619 is joined (e.g., multiplied as in dot product) at 612I and summed at 614I (wherein the second operation is performing multiplying with a dot product and summing). Please see annotated Fig. 6I below.) on the first input (Fig. 6I. Paragraph [0138]-CHIDANANDA discloses the output of the 1×1 convolution layer 608I is sent to another activation layer 610I. The activation function applied at the activation layer transforms the input to the activation layer 610I into the activation of the output for the input (wherein the output of the activation function is the first input). Please see annotated Fig. 6I below.) and an operand that is based on a result of the concatenation operation (Fig. 6I. Paragraph [0136]-CHIDANANDA discloses the concatenated output from the spatial path and the context path is also forwarded to a block 6181 having a convolution layer followed by a batch normalization layer that is further followed by a rectified linear unit (wherein the output after the ReLU layer is an operand based on a result of the concatenation). Please see annotated Fig. 6I below.), the second operation being an add operation or a multiplication operation (Fig. 6I. Paragraph [0138]-CHIDANANDA discloses the output of the activation layer 610I and that from the block 619 is joined (e.g., multiplied as in dot product) at 612I and summed at 614I (wherein the second operation is performing multiplying with a dot product and summing). Please see annotated Fig. 6I below.). PNG media_image1.png 599 721 media_image1.png Greyscale Annotated diagram of CHIDANANDA’s Fig. 6I illustrating performing addition and multiplication on a first input and an operand. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI in view of HAMPALI of a method for estimating a 3-dimensional (3D) pose, the method comprising: receiving by a computing device: a first input generated based on first features associated with first image data from a first sensor associated with the computing device and based on second image data from a second sensor associated with the computing device; and a second input generated based on second features associated with the first image data and based on the second image data; based on the first input and the second input, generating, by the computing device, 3D pose-estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of CHINANANDA of performing a second operation on the first input and an operand that is based on a result of the concatenation operation, the second operation being an add operation or a multiplication operation. Wherein having LI’s method for estimating hand poses from stereo cameras having performing a second operation on the first input and an operand that is based on a result of the concatenation operation, the second operation being an add operation or a multiplication operation. The motivation behind the modification would have been to obtain a method for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and CHINANANDA relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, CHINANANDA the network architecture disclosed herein may be carefully designed to reduce compute/memory overhead and energy consumption. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and CHINANANDA et al. (US 20200372246 A1), Paragraph [0219]. Regarding claim 12, LI explicitly teaches the system of claim 10, LI fails to explicitly teach wherein the output of the one or more processors comprises output features with corresponding attentions generated based on: performing a concatenation operation on the first input and the second input; and. However, HAMPALI explicitly teaches wherein the output (Fig. 6, column on the right called output pose. Paragraph [0103]-HAMPALI discloses the joint specific features enable estimation of different joint-related pose parameters, such as joint angle and joint vector. The output poses for the input images are illustrated in the fourth column of FIG. 6.) of the one or more processors (Fig. 2A. Paragraph [0053]-HAMPALI discloses one or more of the CPU 232, the GPU 234, the DSP 236, and/or the ISP 238 can implement the keypoint determination engine 250, the feature determination engine 254, and/or the pose estimation engine 256.) comprises output features with corresponding attentions generated based on (Fig. 6. Paragraph [0103]-HAMPALI discloses the first column includes input images, and the second column includes the keypoint heatmaps generated by the keypoint determination engine 250 for the input images. For each joint query, the respective colored circles in the third column (labeled “Joint Attention”) indicate the locations of the keypoints attended by the query. The radius of the circle in the Joint Attention column is proportional to the attention weight (wherein the radius proportional to the attention weight is the corresponding attention and the joint attention is output data).): performing a concatenation operation on the first input and the second input (Fig. 3. Paragraph [0076]-HAMPALI discloses the feature determination engine 254 can combine the features (e.g., by concatenating the features from the different feature maps) to form a feature representation, such as a feature vector (wherein the first input and second input are different feature maps).); and Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance of: receiving by the one or more processors: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the one or more processors, 3D pose- estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of HAMPALI of wherein the output of the one or more processors comprises output features with corresponding attentions generated based on: performing a concatenation operation on the first input and the second input; and. Wherein having LI’s system for estimating hand poses from stereo cameras having wherein the output of the one or more processors comprises output features with corresponding attentions generated based on: performing a concatenation operation on the first input and the second input; and. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and HAMPALI relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while HAMPALI determining accurate poses of objects can allow a system to generate accurately positioned and oriented representations (e.g., models) of the objects. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and HAMPALI et al. (US 20220301304 A1), Paragraph [0005]. LI in view of HAMPAILI fail to explicitly teach performing a second operation on the first input and an operand that is based on a result of the concatenation operation, the second operation being an add operation or a multiplication operation. However, CHINANANDA explicitly teaches performing a second operation (Fig. 6I. Paragraph [0138]-CHIDANANDA discloses the output of the activation layer 610I and that from the block 619 is joined (e.g., multiplied as in dot product) at 612I and summed at 614I (wherein the second operation is performing multiplying with a dot product and summing). Please see annotated Fig. 6I below.) on the first input (Fig. 6I. Paragraph [0138]-CHIDANANDA discloses the output of the 1×1 convolution layer 608I is sent to another activation layer 610I. The activation function applied at the activation layer transforms the input to the activation layer 610I into the activation of the output for the input (wherein the output of the activation function is the first input). Please see annotated Fig. 6I below.) and an operand that is based on a result of the concatenation operation (Fig. 6I. Paragraph [0136]-CHIDANANDA discloses the concatenated output from the spatial path and the context path is also forwarded to a block 6181 having a convolution layer followed by a batch normalization layer that is further followed by a rectified linear unit (wherein the output after the ReLU layer is an operand based on a result of the concatenation). Please see annotated Fig. 6I below.), the second operation being an add operation or a multiplication operation (Fig. 6I. Paragraph [0138]-CHIDANANDA discloses the output of the activation layer 610I and that from the block 619 is joined (e.g., multiplied as in dot product) at 612I and summed at 614I (wherein the second operation is performing multiplying with a dot product and summing). Please see annotated Fig. 6I below.). PNG media_image1.png 599 721 media_image1.png Greyscale Annotated diagram of CHIDANANDA’s Fig. 6I illustrating performing addition and multiplication on a first input and an operand. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI in view of HAMPALI of a system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance of: receiving by the one or more processors: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the one or more processors, 3D pose- estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of CHINANANDA of performing a second operation on the first input and an operand that is based on a result of the concatenation operation, the second operation being an add operation or a multiplication operation. Wherein having LI’s system for estimating hand poses from stereo cameras having performing a second operation on the first input and an operand that is based on a result of the concatenation operation, the second operation being an add operation or a multiplication operation. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and CHINANANDA relate to hand pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, CHINANANDA the network architecture disclosed herein may be carefully designed to reduce compute/memory overhead and energy consumption. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and CHINANANDA et al. (US 20200372246 A1), Paragraph [0219]. Claims 4-6 and 13-15 are rejected under 35 U.S.C. 103 as being unpatentable over LI et al. (US 20230161419 A1), hereinafter referenced as LI, in view of TYAGI et al. (US 20250225677 A1), hereinafter referenced as TYAGI. Regarding claim 4, LI explicitly teaches the method of claim 1, wherein: LI fails to explicitly teach the first features comprise first fused-feature data associated with the first image data and the second image data; and the second features comprise second fused-feature data associated with the first image data and the second image data. However, TYAGI explicitly teaches the first features comprise first fused-feature data (Fig. 3. Paragraph [0053]-TYAGI discloses the color feature map 306A and the depth feature map 308A may be fed as inputs to the feature-fusion neural network 102A. Based on the application, a feature fusion map 310A may be generated. Each color feature in the color feature map 306A may be fused with a depth feature of the corresponding depth feature map 308A for the generation of the feature fusion map 310A (wherein the feature fusion map is first fused-feature data).) associated with the first image data (Fig. 3. Paragraph [0049]-TYAGI discloses the circuitry 202 may be configured to acquire color image data. The color image data may include a set of color images 302A . . . 302N. Each color image of the set of color images 302A . . . 302N may include an object (wherein the first image data is color image data).) and the second image data (Fig. 3. Paragraph [0050]-TYAGI discloses the circuitry 202 may be further configured to acquire depth image data. The depth image data may include a set of depth images (or depth maps) 304A . . . 304N. Each depth image (or depth map) of the set of depth images 304A . . . 304N may include the object (wherein the second image data is depth image data).); and the second features comprise second fused-feature data (Fig. 3. Paragraph [0053]-TYAGI discloses the circuitry 202 may apply the feature-fusion neural network 102A on the color feature map 306N and a depth feature map 308N to generate a feature fusion map 310N. Thus, a set of feature fusion maps 310A . . . 310N may be generated.) associated with the first image data (Fig. 3. Paragraph [0049]-TYAGI discloses the circuitry 202 may be configured to acquire color image data. The color image data may include a set of color images 302A . . . 302N. Each color image of the set of color images 302A . . . 302N may include an object (wherein the first image data is color image data).) and the second image data (Fig. 3. Paragraph [0050]-TYAGI discloses the circuitry 202 may be further configured to acquire depth image data. The depth image data may include a set of depth images (or depth maps) 304A . . . 304N. Each depth image (or depth map) of the set of depth images 304A . . . 304N may include the object (wherein the second image data is depth image data).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a method for estimating a 3-dimensional (3D) pose, the method comprising: receiving by a computing device: a first input generated based on first features associated with first image data from a first sensor associated with the computing device and based on second image data from a second sensor associated with the computing device; and a second input generated based on second features associated with the first image data and based on the second image data; based on the first input and the second input, generating, by the computing device, 3D pose-estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of TYAGI of the first features comprise first fused-feature data associated with the first image data and the second image data; and the second features comprise second fused-feature data associated with the first image data and the second image data. Wherein having LI’s method for estimating hand poses from stereo cameras having the first features comprise first fused-feature data associated with the first image data and the second image data; and the second features comprise second fused-feature data associated with the first image data and the second image data. The motivation behind the modification would have been to obtain a method for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and TYAGI relate to pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while TYAGI the output vector representation may enable efficient localization of the set of joints irrespective of the resolution of voxels of the 3D volume or the volumetric model of the object 116. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and TYAGI et al. (US 20250225677 A1), Paragraph [0038]. Regarding claim 5, LI in view of TYAGI explicitly teach the method of claim 4, further comprising: LI fails to explicitly teach receiving, by the computing device, first image features from the first image data as a first feature-fusion input; receiving, by the computing device, second image features from the second image data as a second feature-fusion input; and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input. However, TYAGI explicitly teaches receiving, by the computing device (Fig. 1. Paragraph [0024]-TYAGI discloses each of the feature-fusion neural network 102A and the pose estimation neural network 102B may be implemented using hardware that may include a processor, a microprocessor (e.g., to perform or control performance of one or more operations), an FPGA, or an ASIC.), first image features from the first image data (Fig. 3. Paragraph [0051]-TYAGI discloses based on the application of the neural network on an input that corresponds to the color image data (i.e., the set of color images 302A . . . 302N), a set of color feature maps 306A . . . 306N may be generated as output (wherein the color feature maps are first image features).) as a first feature-fusion input (Fig. 3. [0053]-TYAGI discloses The circuitry 202 may be further configured to apply the feature-fusion neural network 102A on each color feature map of the set of color feature maps 306A . . . 306N and each depth feature map of the set of depth feature maps 308A . . . 308N. Based on such an application, color features in each of the color feature maps of the set of color feature maps 306A . . . 306N may be fused with depth features in each of the corresponding depth feature maps of the set of depth feature maps 308A . . . 308N (wherein the color feature maps input into the feature-fusion neural network are first feature-fusion inputs).); receiving, by the computing device (Fig. 1. Paragraph [0024]-TYAGI discloses each of the feature-fusion neural network 102A and the pose estimation neural network 102B may be implemented using hardware that may include a processor, a microprocessor (e.g., to perform or control performance of one or more operations), an FPGA, or an ASIC.), second image features from the second image data (Fig. 3. Paragraph [0052]-TYAGI discloses based on the application of the neural network on an input that corresponds to the depth image data (i.e., the set of depth images 304A . . . 304N), a set of depth feature maps 308A . . . 308N may be generated as output (wherein the depth feature maps are first images features).) as a second feature-fusion input (Fig. 3. [0053]-TYAGI discloses The circuitry 202 may be further configured to apply the feature-fusion neural network 102A on each color feature map of the set of color feature maps 306A . . . 306N and each depth feature map of the set of depth feature maps 308A . . . 308N. Based on such an application, color features in each of the color feature maps of the set of color feature maps 306A . . . 306N may be fused with depth features in each of the corresponding depth feature maps of the set of depth feature maps 308A . . . 308N (wherein the depth feature maps input into the feature-fusion neural network are second feature-fusion inputs).); and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input (Fig. 3. Paragraph [0053]-TYAGI discloses the color feature map 306A and the depth feature map 308A may be fed as inputs to the feature-fusion neural network 102A. Based on the application, a feature fusion map 310A may be generated. Each color feature in the color feature map 306A may be fused with a depth feature of the corresponding depth feature map 308A for the generation of the feature fusion map 310A (wherein the feature fusion map is first fused-feature data).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a method for estimating a 3-dimensional (3D) pose, the method comprising: receiving by a computing device: a first input generated based on first features associated with first image data from a first sensor associated with the computing device and based on second image data from a second sensor associated with the computing device; and a second input generated based on second features associated with the first image data and based on the second image data; based on the first input and the second input, generating, by the computing device, 3D pose-estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of TYAGI of receiving, by the computing device, first image features from the first image data as a first feature-fusion input; receiving, by the computing device, second image features from the second image data as a second feature-fusion input; and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input. Wherein having LI’s method for estimating hand poses from stereo cameras having receiving, by the computing device, first image features from the first image data as a first feature-fusion input; receiving, by the computing device, second image features from the second image data as a second feature-fusion input; and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input. The motivation behind the modification would have been to obtain a method for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and TYAGI relate to pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while TYAGI the output vector representation may enable efficient localization of the set of joints irrespective of the resolution of voxels of the 3D volume or the volumetric model of the object 116. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and TYAGI et al. (US 20250225677 A1), Paragraph [0038]. Regarding claim 6, LI in view of TYAGI explicitly teach the method of claim 5, LI fails to explicitly teach further comprising performing, by the computing device, a first concatenation operation and a first convolution operation on the first image features and on the second image features. However, TYAGI explicitly teaches further comprising performing, by the computing device (Fig. 1. Paragraph [0024]-TYAGI discloses each of the feature-fusion neural network 102A and the pose estimation neural network 102B may be implemented using hardware that may include a processor, a microprocessor (e.g., to perform or control performance of one or more operations), an FPGA, or an ASIC.), a first concatenation operation (Fig. 3. Paragraph [0054]-TYAGI discloses the input layer (or first layer) of the feature-fusion neural network 102A may receive the color feature map 306A and the depth feature map 308A as inputs. The first layer may include two sub-layers. At the first sub-layer, the depth feature map 308A may be concatenated with the color feature map 306A for generation of a first result. On the other hand, at the second sub-layer, the color feature map 306A may be concatenated with the depth feature map 308A for generation of a second result.) and a first convolution operation (Fig. 5. Paragraph [0054]-TYAGI discloses the second layer may include two sub-layers. At the first sub-layer, a first convolution operation may be applied on the first result to generate a third result. On the other hand, at the second sub-layer, a second convolution operation may be applied on the second result to generate an intermediate result.) on the first image features and on the second image features (Fig. 3. Paragraph [0054]-TYAGI discloses the input layer (or first layer) of the feature-fusion neural network 102A may receive the color feature map 306A and the depth feature map 308A as inputs.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a method for estimating a 3-dimensional (3D) pose, the method comprising: receiving by a computing device: a first input generated based on first features associated with first image data from a first sensor associated with the computing device and based on second image data from a second sensor associated with the computing device; and a second input generated based on second features associated with the first image data and based on the second image data; based on the first input and the second input, generating, by the computing device, 3D pose-estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of TYAGI of further comprising performing, by the computing device, a first concatenation operation and a first convolution operation on the first image features and on the second image features. Wherein having LI’s method for estimating hand poses from stereo cameras having further comprising performing, by the computing device, a first concatenation operation and a first convolution operation on the first image features and on the second image features. The motivation behind the modification would have been to obtain a method for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and TYAGI relate to pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while TYAGI the output vector representation may enable efficient localization of the set of joints irrespective of the resolution of voxels of the 3D volume or the volumetric model of the object 116. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and TYAGI et al. (US 20250225677 A1), Paragraph [0038]. Regarding claim 13, LI explicitly teaches the system of claim 10, wherein: LI fails to explicitly teach the first features comprise first fused-feature data associated with the first image data and the second image data; and the second features comprise second fused-feature data associated with the first image data and the second image data. However, TYAGI explicitly teaches the first features comprise first fused-feature data (Fig. 3. Paragraph [0053]-TYAGI discloses the color feature map 306A and the depth feature map 308A may be fed as inputs to the feature-fusion neural network 102A. Based on the application, a feature fusion map 310A may be generated. Each color feature in the color feature map 306A may be fused with a depth feature of the corresponding depth feature map 308A for the generation of the feature fusion map 310A (wherein the feature fusion map is first fused-feature data).) associated with the first image data (Fig. 3. Paragraph [0049]-TYAGI discloses the circuitry 202 may be configured to acquire color image data. The color image data may include a set of color images 302A . . . 302N. Each color image of the set of color images 302A . . . 302N may include an object (wherein the first image data is color image data).) and the second image data (Fig. 3. Paragraph [0050]-TYAGI discloses the circuitry 202 may be further configured to acquire depth image data. The depth image data may include a set of depth images (or depth maps) 304A . . . 304N. Each depth image (or depth map) of the set of depth images 304A . . . 304N may include the object (wherein the second image data is depth image data).); and the second features comprise second fused-feature data (Fig. 3. Paragraph [0053]-TYAGI discloses the circuitry 202 may apply the feature-fusion neural network 102A on the color feature map 306N and a depth feature map 308N to generate a feature fusion map 310N. Thus, a set of feature fusion maps 310A . . . 310N may be generated.) associated with the first image data (Fig. 3. Paragraph [0049]-TYAGI discloses the circuitry 202 may be configured to acquire color image data. The color image data may include a set of color images 302A . . . 302N. Each color image of the set of color images 302A . . . 302N may include an object (wherein the first image data is color image data).) and the second image data (Fig. 3. Paragraph [0050]-TYAGI discloses the circuitry 202 may be further configured to acquire depth image data. The depth image data may include a set of depth images (or depth maps) 304A . . . 304N. Each depth image (or depth map) of the set of depth images 304A . . . 304N may include the object (wherein the second image data is depth image data).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance of: receiving by the one or more processors: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the one or more processors, 3D pose- estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of TYAGI of the first features comprise first fused-feature data associated with the first image data and the second image data; and the second features comprise second fused-feature data associated with the first image data and the second image data. Wherein having LI’s system for estimating hand poses from stereo cameras having the first features comprise first fused-feature data associated with the first image data and the second image data; and the second features comprise second fused-feature data associated with the first image data and the second image data. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and TYAGI relate to pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while TYAGI the output vector representation may enable efficient localization of the set of joints irrespective of the resolution of voxels of the 3D volume or the volumetric model of the object 116. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and TYAGI et al. (US 20250225677 A1), Paragraph [0038]. Regarding claim 14, LI in view of TYAGI explicitly teach the system of claim 13, LI further explicitly teaches wherein the instructions, when executed by the one or more processors, cause performance of (Fig. 7. Paragraph [0059]-LI discloses the term “processor” is intended to include multi-core processors that may comprise two or more independent processors (sometimes referred to as “cores”) that may execute instructions contemporaneously.): LI fails to explicitly teach receiving, by the one or more processors, first image features from the first image data as a first feature-fusion input; receiving, by the one or more processors, second image features from the second image data as a second feature-fusion input; and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input. However, TYAGI explicitly teaches receiving, by the one or more processors (Fig. 1. Paragraph [0024]-TYAGI discloses each of the feature-fusion neural network 102A and the pose estimation neural network 102B may be implemented using hardware that may include a processor, a microprocessor (e.g., to perform or control performance of one or more operations), an FPGA, or an ASIC.), first image features from the first image data (Fig. 3. Paragraph [0051]-TYAGI discloses based on the application of the neural network on an input that corresponds to the color image data (i.e., the set of color images 302A . . . 302N), a set of color feature maps 306A . . . 306N may be generated as output (wherein the color feature maps are first image features).) as a first feature-fusion input (Fig. 3. [0053]-TYAGI discloses The circuitry 202 may be further configured to apply the feature-fusion neural network 102A on each color feature map of the set of color feature maps 306A . . . 306N and each depth feature map of the set of depth feature maps 308A . . . 308N. Based on such an application, color features in each of the color feature maps of the set of color feature maps 306A . . . 306N may be fused with depth features in each of the corresponding depth feature maps of the set of depth feature maps 308A . . . 308N (wherein the color feature maps input into the feature-fusion neural network are first feature-fusion inputs).); receiving, by the one or more processors (Fig. 1. Paragraph [0024]-TYAGI discloses each of the feature-fusion neural network 102A and the pose estimation neural network 102B may be implemented using hardware that may include a processor, a microprocessor (e.g., to perform or control performance of one or more operations), an FPGA, or an ASIC.), second image features from the second image data (Fig. 3. Paragraph [0052]-TYAGI discloses based on the application of the neural network on an input that corresponds to the depth image data (i.e., the set of depth images 304A . . . 304N), a set of depth feature maps 308A . . . 308N may be generated as output (wherein the depth feature maps are first images features).) as a second feature-fusion input (Fig. 3. [0053]-TYAGI discloses The circuitry 202 may be further configured to apply the feature-fusion neural network 102A on each color feature map of the set of color feature maps 306A . . . 306N and each depth feature map of the set of depth feature maps 308A . . . 308N. Based on such an application, color features in each of the color feature maps of the set of color feature maps 306A . . . 306N may be fused with depth features in each of the corresponding depth feature maps of the set of depth feature maps 308A . . . 308N (wherein the depth feature maps input into the feature-fusion neural network are second feature-fusion inputs).); and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input (Fig. 3. Paragraph [0053]-TYAGI discloses the color feature map 306A and the depth feature map 308A may be fed as inputs to the feature-fusion neural network 102A. Based on the application, a feature fusion map 310A may be generated. Each color feature in the color feature map 306A may be fused with a depth feature of the corresponding depth feature map 308A for the generation of the feature fusion map 310A (wherein the feature fusion map is first fused-feature data).). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance of: receiving by the one or more processors: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the one or more processors, 3D pose- estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of TYAGI of receiving, by the one or more processors, first image features from the first image data as a first feature-fusion input; receiving, by the one or more processors, second image features from the second image data as a second feature-fusion input; and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input. Wherein having LI’s system for estimating hand poses from stereo cameras having receiving, by the one or more processors, first image features from the first image data as a first feature-fusion input; receiving, by the one or more processors, second image features from the second image data as a second feature-fusion input; and generating the first fused-feature data based on the first feature-fusion input and the second feature-fusion input. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and TYAGI relate to pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while TYAGI the output vector representation may enable efficient localization of the set of joints irrespective of the resolution of voxels of the 3D volume or the volumetric model of the object 116. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and TYAGI et al. (US 20250225677 A1), Paragraph [0038]. Regarding claim 15, LI in view of TYAGI explicitly teach the system of claim 14, LI fails to explicitly teach wherein the instructions, when executed by the one or more processors, cause performance of a first concatenation operation and a first convolution operation on the first image features and on the second image features. However, TYAGI explicitly teaches wherein the instructions, when executed by the one or more processors (Fig. 1. Paragraph [0024]-TYAGI discloses each of the feature-fusion neural network 102A and the pose estimation neural network 102B may be implemented using hardware that may include a processor, a microprocessor (e.g., to perform or control performance of one or more operations), an FPGA, or an ASIC.), cause performance of a first concatenation operation (Fig. 3. Paragraph [0054]-TYAGI discloses the input layer (or first layer) of the feature-fusion neural network 102A may receive the color feature map 306A and the depth feature map 308A as inputs. The first layer may include two sub-layers. At the first sub-layer, the depth feature map 308A may be concatenated with the color feature map 306A for generation of a first result. On the other hand, at the second sub-layer, the color feature map 306A may be concatenated with the depth feature map 308A for generation of a second result.) and a first convolution operation (Fig. 5. Paragraph [0054]-TYAGI discloses the second layer may include two sub-layers. At the first sub-layer, a first convolution operation may be applied on the first result to generate a third result. On the other hand, at the second sub-layer, a second convolution operation may be applied on the second result to generate an intermediate result.) on the first image features and on the second image features (Fig. 3. Paragraph [0054]-TYAGI discloses the input layer (or first layer) of the feature-fusion neural network 102A may receive the color feature map 306A and the depth feature map 308A as inputs.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of LI of a system comprising: one or more processors; and a memory storing instructions which, when executed by the one or more processors, cause performance of: receiving by the one or more processors: a first input generated based on first features associated with first image data from a first sensor and based on second image data from a second sensor; and a second input generated based on second features associated with the first image data and based on the second image data; generating, based on an output of the one or more processors, 3D pose- estimation data associated with an object represented in the first image data and represented in the second image data; and transmitting the 3D pose-estimation data with the teachings of TYAGI of wherein the instructions, when executed by the one or more processors, cause performance of a first concatenation operation and a first convolution operation on the first image features and on the second image features. Wherein having LI’s system for estimating hand poses from stereo cameras having wherein the instructions, when executed by the one or more processors, cause performance of a first concatenation operation and a first convolution operation on the first image features and on the second image features. The motivation behind the modification would have been to obtain a system for estimating hand poses from stereo cameras that enhances the accuracy in estimating the hand pose. Since both LI and TYAGI relate to pose estimation, wherein LI is to more accurately estimate full three-dimensional (herein referred to as “3D”) hand poses from stereo cameras for use in human computer interaction, video games, sign language recognition, and augmented reality, while TYAGI the output vector representation may enable efficient localization of the set of joints irrespective of the resolution of voxels of the 3D volume or the volumetric model of the object 116. Please see LI et al. (US 20230161419 A1), Paragraph [0003, 0014, and 0040], and TYAGI et al. (US 20250225677 A1), Paragraph [0038]. Conclusion Listed below are the prior arts made of record and not relied upon but are considered pertinent to applicant’s disclosure. GUO et al. (US 20240161494 A1) – Disclosed herein is a gesture recognition device that includes an input interface configured to receive a sequence of images, each image showing a body part with which a gesture is performed from a viewpoint of a camera. The gesture recognition device also generates a sequence of motion-compensated images from the sequence comprising generating a motion-compensated image for an image of the sequence by compensating the movement of the camera viewpoint from a reference camera viewpoint to the viewpoint from which the image shows the body part based on the image and a motion-compensated image of the sequence generated for a preceding image of the sequence which precedes the image in the sequence and estimate the gesture from the sequence of motion-compensated images…Abstract, Fig. 7. HAN et al. (US 20190392632 A1) – Provided is a method of reconstructing a three-dimensional (3D) model of an object. The method includes sequentially performing, by a camera module, first and second object scanning processes of scanning the same object, reconstructing, by a processor module, a 3D object model, based on a first object image obtained through the first object scanning process, performing pose learning on an object to generate learning data, based on data obtained through a process of reconstructing the 3D object model based on the first object image, and reconstructing, by the processor module, a final 3D object model, based on a second object image obtained through the second object scanning process and the learning data…Abstract, Fig. 2. MANDERS et al. (US 20110299774 A1) – A method and system for detecting and tracking hands in an image. The method for detecting and tracking hands in an image comprises the steps of calculating a first probability map comprising of probabilities that respective pixels in the image corresponds to skin based on colour information associated with the respective pixels; calculating a second probability map comprising of probabilities that the respective pixels in the image corresponds to a part of a hand based on depth information associated with the respective pixels; calculating a joint probability map by combining the first probability map and the second probability map; and detecting and tracking hands in the image using an algorithm with a weight output as a detection threshold applied on the joint probability map…Abstract, Fig. 1. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ETHAN N WOLFSON whose telephone number is (571)272-1898. The examiner can normally be reached Monday - Friday 8:00 am - 5:00 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ETHAN N WOLFSON/Examiner, Art Unit 2673 /CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673
Read full office action

Prosecution Timeline

Jan 07, 2025
Application Filed
Sep 02, 2026
Non-Final Rejection mailed — §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
86%
Grant Probability
99%
With Interview (+50.0%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 7 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month