Prosecution Insights
Last updated: August 18, 2026
Application No. 18/641,711

TWO-HANDED GESTURE INTERPRETATION

Final Rejection §101§102§103§112
Filed
Apr 22, 2024
Priority
May 15, 2023 — provisional 63/466,454 +1 more
Examiner
MERCADO, GABRIEL S
Art Unit
2171
Tech Center
2100 — Computer Architecture & Software
Assignee
Apple Inc.
OA Round
2 (Final)
42%
Grant Probability
Moderate
3-4
OA Rounds
1y 1m
Est. Remaining
69%
With Interview

Examiner Intelligence

Grants 42% of resolved cases
42%
Career Allowance Rate
87 granted / 206 resolved
-12.8% vs TC avg
Strong +27% interview lift
Without
With
+26.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
32 currently pending
Career history
249
Total Applications
across all art units

Statute-Specific Performance

§101
15.0%
-25.0% vs TC avg
§103
47.2%
+7.2% vs TC avg
§102
10.3%
-29.7% vs TC avg
§112
24.3%
-15.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 206 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION This office action is responsive to communication(s) filed on 5/15/2026. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims Status Claims 1-20 are pending and are currently being examined. Claims 1, 18 and 20 are independent, and are newly amended. Claim Interpretation Claims 2-4, 6 and 19 are recite the terms "user centric gesture" and "app centric gesture”. These are not recognized terms of art, the Instant Specification doesn’t provide a special definition, and their scope is broad. The Instant Specification provides examples, see ¶ 5 (as published), wherein the examples are described as what the user centric and app centric gestures “may include” or “may refer to”. The interpretation of the terms are not limited to these examples. In a broad sense, “user centric” gestures include gestures that are identified by using user position references, such as user’s joint positions, and “app centric” as being gestures that don’t require using user position references. User centric gestures can also be interpreted as including gestures that control the system or application-agnostic navigation, while app centric gestures are tied to specific app functions and views. Still, user centric gestures, e.g., in American Sign Language, can also include gestures in the first-person perspective to mimic how a person interacts with an object (e.g., gripping an imaginary steering wheel to drive), and app-centric gestures use a third-person view to map an object's spatial features or shape in the air (e.g., tracing a boxy outline of a car). Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea without significantly more. The representative independent claim 1 recite(s) a method comprising: receiving data corresponding to user activity involving two hands of a user in a three-dimensional (3D) coordinate system, determining gestures performed by the two hands based on the data corresponding to the user activity, identifying actions performed by the two hands based on the determined gestures, each of the two hands performing one of the identified actions, determining whether the identified actions satisfy a criterion for a gesture type based on the data corresponding to the user activity, and in accordance with determining that the identified actions satisfy the criterion for the gesture type, interpreting the identified actions based as input on a reference element corresponding to the gesture type, wherein different gesture types correspond to different reference elements. Each step in the described method can be performed by a human mind—with or without pen and paper—because they represent fundamental cognitive processes of perception, categorization, and logical deduction. Receiving data: A human can observe and record physical movements or coordinates using their eyes and hands. Determining gestures: The mind naturally groups observed movements into recognizable patterns based on memory and visual processing. Identifying actions: A person can logically assign a specific meaning or label to those recognized patterns through simple association. Determining criteria satisfaction: A human can use a checklist or mental rules, to compare observed actions against predefined standards. Interpreting based on reference: The mind can translate a specific action, e.g., as “input” for conversation, into a final conclusion by cross-referencing it (inputting it for a comparison) with a known guide or context. (e.g., American Sign Language). As such the claim recites an abstract idea grouped under “Mental Processes”. This judicial exception is not integrated into a practical application because the claim further includes the limitation: “at a device having a processor and one or more sensors”, and that the input is “to the device” but this is mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, see MPEP 2106.05(f). The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims as a whole include only the abstract idea with the mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, which cannot provide a practical application or satisfy the “significantly more” requirement. As such, representative claim 1 is ineligible under 101. Claim(s) 18 and 20 are directed to a device and computer-readable storage medium for accomplishing the steps of the method in claim 1, and are rejected using similar rationale(s). These claims do nothing more than adding limitations “A device comprising: a non-transitory computer-readable storage medium; and one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations comprising” and “A non-transitory computer-readable storage medium, storing program instructions executable on a device including one or more processors to perform operations comprising”, and that the input is “to the device” (for claim 20) which are reflective mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, see MPEP 2106.05(f). Claims 2-4 and 19 further describe the “user activity” and/or purpose thereof are. This description doesn’t make the related steps any less abstract. Note that whether the user activity is user centric or app centric doesn’t make the claims any less abstract or provide a practical application because these types of user activity predate computers. E.g., in sign language, user-centric gestures may be interpreted as gestures in the first-person perspective to mimic how a person interacts with an object (e.g., gripping an imaginary steering wheel to drive), and app-centric gestures use a third-person view to map an object's spatial features or shape in the air (e.g., tracing a boxy outline of a car). Claim 5 further recites abstract concept of by adding the step of ”associating a pivot point on a body of the user as the reference element”. Associating a pivot point on a user's body as a reference element can be performed mentally with or without pen and paper because it relies on innate proprioception and basic spatial awareness to visualize a joint or body part as a fixed, central anchor point for movement. Claims 6-9, 11, 14 and 15 further recite or further limit the abstract “identifying” and/or ”determining” steps. However, these limitations don’t make the related steps any less abstract. Claim 10 involves further limiting the reference element. This limitation doesn’t make the receiving data step any less abstract. Claims 12-13 involve limitations describing “the 3D environment” mentioned in claim 11. These limitations don’t make the related steps any less abstract. Claims 14, 15 and 17 involve limitation of a step of determining “user intended action” (herein, it interpreted as “identifying actions performed”, see 112B rejection section below). These limitations don’t make the identifying/determining step any less abstract. Claim 16 adds the limitation of “using at least a portion of the identified actions performed by the two hands to directly enable a functional component displayed within an extended reality (XR) environment.” Here, enabling a functional component within an Extended Reality (XR) environment is interpreted as making a digital, interactive element—such as a button, 3D model, slider, or dashboard—fully operational and responsive to user input (gestures, voice, or controllers) while immersed in virtual, augmented, or mixed reality. However, the limitations of “using at least a portion of the identified actions performed by the two hands to directly enable a functional component displayed” and “within an extended reality (XR) +environment” are ineffective to provide a practical application of the abstract idea because they are directed to insignificant extra-solution activity and applying the extra-solution activity to a specific technological environment, respectively, see MPEP 2106.05(h). There is nothing in the claim that adds significantly more than the abstract idea. The limitation of “using at least a portion of the identified actions performed by the two hands to directly enable a functional component displayed” is extra-solution activity and ineffective to show significantly more because it is directed to well-understood, routine, conventional activity. For example, the following references teach this concept: AU; Kin Chung et al. US 20120154313 A1 [0062] Another example of a pop-up interface is a virtual keyboard 800, as shown in FIG. 8. The virtual keyboard 800 is a two-hand interface and can be activated with a two-hand gesture, such as a two-hand five finger tap. El Dokor; Tarek US 8928590 B1 “Users can also lift a single hand or both hands above the keyboard. Thus, the inventive system provides for a delineation of gesture recognition between an active zone that is enabled once the user's hand (or hands) is visible for the cameras and inactive zone when the user is typing on the keyboard or using the mouse”, col 2:26-32. Wong; Yoon Kean US 20110320982 A1 [0002]…Traditionally, a user is required to use two hands and multiple input mechanisms to activate a palmtop computer application. Karlsson; David et al. US 20120131488 A1 [0108]…The controls can accept two-hand inputs, one to activate one of the movable icons and one to carry out a touch gesture to select a choice that carries out a desired interaction associated with the activated alternate interaction icon (block 218). The limitation of “within an extended reality (XR) environment” simply applies the extra-solution activity to a specific technological environment and is therefore also ineffective to provide significantly more than the abstract idea. For the reasons above, claims 2-20 are also ineligible under 101, similar to representative claim 1, for reasons explained above. Claim Rejections - 35 USC § 112(a) or 112(1st) The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 2-4, 6, 14 and 19 are/is rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contain(s) subject matter which was/were not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 14 recites “user activity is associated with two simultaneous two-handed gestures”. However, the specification doesn’t sufficient describe how more than one two-handed gestures can occur simultaneously, when the claimed user activity involves only “two hands” (see claim 1 upon which claim 14 depends). Claim Rejections - 35 USC § 112(b) or 112(2nd) The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2-4, 6, 8, 14, 15, 17 and 19 is/are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claim 6 recites that “wherein determining to associate each of the actions performed by each of the two hands with the user intended action corresponding to the gesture type”. Here, it is unclear whether or not the phrase “determining to associate each of the actions performed by each of the two hands with the user intended action corresponding to the gesture type” is referring to the step of “identifying actions performed by the two hands based on the determined gestures, each of the two hands performing one of the identified actions” in claim 1. The applicant has provided ¶¶ 12 and 160 as alleged support for the clarity of this limitation, see Remarks Pg(s) 17. However, these paragraphs do not clarify a differentiation between the determining and identifying steps because it applies the same app-centric technique and reference orientation to interpret motions and define both functions. Both steps rely on shared spatial logic, such as using user or application orientation to detect gestures, merging the technical boundaries between them. For purposes of compact prosecution only, the examiner interprets the claim’s limitation(s) as referring to the abovementioned “identifying actions…” limitation in claim 1. Correction required. Furthermore, Claim 6 recites “the user intended action” in the abovementioned phrase. There is insufficient antecedent basis for this limitation in the claim. The same issue is present in Claims 14, 15 and 17. Claim 8 recites “the device worn on the head” in the phrase “positioning information associated with a head or the device worn on the head”. There is insufficient antecedent basis for this limitation in the claim. Claim 14 recites “user activity is associated with two simultaneous two-handed gestures”. Here, this is unclear because it can be interpreted as a plurality of two-handed gestures that are performed simultaneously. However, it is unclear how more than one two-handed gestures can occur simultaneously, when the user activity involves only “two hands” (see claim 1 upon which claim 14 depends). For purposes of compact prosecution only, the examiner interprets the limitation(s) as being directed to a single two-handed gesture that can be used control two functions at the same time, e.g., zoom and pan function, see ¶¶ 64 and 162. Correction required. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 1-3, 5-13, and 15-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Iyer; Vivek Viswanathan et al. (hereinafter Iyer – US 20190384408 A1). Independent Claim 1: Iyer teaches: A method comprising: (e.g., fig. 4) at a device having a processor and one or more sensors: (Abstract and systems in figs. 2 and 3) receiving, via the one or more sensors, sensor data corresponding to user activity involving two hands of a user (receiving hand tracking sensor data from camera(s) or optical sensor(s) [via the one or more sensors], ¶¶ 65 and 91, for one or two handed gesture sequences detection, by a processor(s) of the system, based on start of gesture, start of gesture motion, and gesture end sequence recognition data [receiving data], Abstract and fig. 4,404,405,406 and ¶ 89) in a three-dimensional (3D) coordinate system; (tracking a physical location (e.g., Euclidian or Cartesian coordinates x, y, and z) or translation, ¶¶ 52 and 82) determining gestures (start, motion, and end gestures) performed by the two hands based on the sensor data corresponding to the user activity; (one or two handed gesture sequences are detected at 407, by a processor(s) of the system, based on recognizing [determining] start, motion, and end of gestures in a sequence of gesture recognition data, Abstract and fig. 4:404,405,406,407 and ¶¶ 89-90) identifying actions (a sequence of gestures) performed by the two hands based on the determined gestures, (one or two handed gesture sequences are detected at 407, by a processor(s) of the system, based on recognizing [determining] start, motion, and end of gestures in a sequence of gesture recognition data, Abstract and fig. 4:404,405,406,407 and ¶¶ 89-90) each of the two hands performing one of the identified actions; (because the gestures in the sequence are two-handed, ¶ 89, each of the two hands perform at least one of the identified actions. An example of two-handed gestures is illustrated in ¶ 152 and figs. 24A-24B) determining whether the identified actions satisfy a criterion for a gesture type based on the sensor data corresponding to the user activity; (a positive gesture sequence recognition meets an accuracy threshold [satisfy a criterion], ¶ 166, also, the sequences are detected by the specific series of gestures determined to match certain attributes and/or reference images, ¶¶ 86-87, e.g., direction, position, and/or location, ¶ 119, and gesture duration, ¶ 135. Herein, a “gesture type” is broadly interpreted as including one or more candidate/identified gesture sequences) and in accordance with determining that the identified actions satisfy the criterion for the gesture type, interpreting the identified actions as input to the device (a positive gesture sequence recognition meets an accuracy threshold [satisfy a criterion], ¶ 166, also, the sequences are detected by the series of gestures determined to match certain attributes and/or reference images, ¶¶ 86-87, e.g., direction, position, and/or location, ¶ 119. The gestures are used as input to the device. Specifically, a device, such as HMD 102 of system 300, ¶¶ 5 and 65 and figs. 2-3, interprets actions as input by capturing video data of hand movements, identifying specific gesture sequences through software, and mapping those motions to control user interface actions or application functions, ¶¶ 74, 77, and 91 ) based on a reference element (reference parameters [element] assigned to each joint, e.g., coordinates and/or or other parameters specifying a conformation of the body part, e.g., hand open, etc., ¶¶ 81 and 82) corresponding to the gesture type, wherein different gesture types correspond to different reference elements. (By tracking the positional changes, trajectories, and interactions of skeletal joints and segments over time, the system can identify, classify, and interpret specific gestures, actions, or behavioral patterns [corresponding to the gesture type], see ¶¶ 81-82. Because the combination of reference elements, such as points in a trajectory are used to identify specific sequences, different gesture types correspond to different reference elements) Claim 2: The rejection of claim 1 is incorporated. Iyer further teaches: wherein the user activity is determined to provide a user centric gesture as the gesture type. (a gesture can be identified based on fitting a virtual skeleton to the user and analyzing the positional changes of its joints and segments over time, ¶ 82. This can be described as “user centric” due to its focus on modeling the user's physical characteristics.) Claim 3: The rejection of claim 1 is incorporated. Iyer further teaches: wherein the user activity is determined to provide an app centric gesture as the gesture type. (a gesture can be identified without relying on fitting a virtual skeleton to the user and analyzing the positional changes of its joints and segments over time, by sending raw point-cloud data directly to a feature extraction routine within a gesture sequence recognition, ¶¶ 82-83, and this can be seen as “app centric” because it prioritizes direct data processing for the application.) Claim 5: The rejection of claim 1 is incorporated. Iyer further teaches: further comprising: associating a pivot point on a body of the user as the reference element. (each body’s joints [pivot point on a body] is assigned a number parameters, such Cartesian coordinate points, ¶ 81-82) Claim 6: The rejection of claim 1 is incorporated. Iyer further teaches: wherein determining to associate each of the actions performed by each of the two hands with the user intended action corresponding to the gesture type is based on an app centric technique (because the gestures in the sequence are two-handed, ¶ 89, each of the two hands perform at least one of the identified actions, an example of two-handed gestures is found in ¶ 152 and figs. 24A-24B. A gesture can be identified without relying on fitting a virtual skeleton to the user and analyzing the positional changes of its joints and segments over time, by sending raw point-cloud data directly to a feature extraction routine within a gesture sequence recognition, ¶¶ 82-83, and this can be seen as “app centric technique” because it prioritizes direct data processing for the application.) and a reference orientation of a space defined based on positioning information of the user. (using positional tracking devices to map an environment where an HMD [Head-Mounted Device] is located, its orientation, and/or pose, ¶¶ 16 and 47, where Heads-Up Displays (HUDs), and eyeglasses-collectively referred to as “HMDs”, ¶¶ 16 and 46, and the same reference information is used for tracking the user’s state, e.g., orientation and movement, ¶¶ 51 and 52) Claim 7: The rejection of claim 1 is incorporated. Iyer further teaches: wherein determining whether the identified actions satisfy the criterion for the gesture type is based on determining spatial positioning between the user and a user interface. (gestures can control an opened menu, e.g., closing or repositioning a menu [a user interface] by placing hand in a certain way “behind” the opened menu [based on determining spatial positioning between the user and a user interface], ¶¶ 124-125.) Claim 8: The rejection of claim 1 is incorporated. Iyer further teaches: wherein determining whether the identified actions satisfy the criterion for the gesture type is based on determining: positioning information associated with a head or the device worn on the head, (using positional tracking devices to map an environment where an HMD [Head-Mounted Device] is located, its orientation, and/or pose, ¶¶ 16 and 47, where Heads-Up Displays (HUDs), and eyeglasses-collectively referred to as “HMDs”, ¶¶ 16 and 46, and the mapping information is used for tracking the user’s state, e.g., orientation and movement, ¶¶ 51 and 52. As mentioned above, gesture sequences are identified by tracking user movements in the environment, e.g., see ¶ 82, wherein specific sequences mapped to specific functions, e.g., minimizing all workspaces, ¶ 151 and figs. 23A-23B) positioning information associated with a torso, a gaze, or iv) a combination thereof. Claim 9: The rejection of claim 1 is incorporated. Iyer further teaches: wherein determining whether the identified actions satisfy the criterion for the gesture type is based on determining a motion type associate with motion data for each of the two hands. (for a successful sequence identification, the motion characteristics [type] for both hands have to match according to an accuracy threshold, ¶ 166, e.g., concerning motion velocity, ¶ 84) Claim 10: The rejection of claim 1 is incorporated. Iyer further teaches: wherein the reference element comprises a reference point in the 3D coordinate system. (reference parameters [element] assigned to each joint, e.g., coordinates and/or or other parameters specifying a conformation of the body part, e.g., hand open, etc., ¶¶ 81 and 82, e.g., Euclidian or Cartesian coordinates x, y, and z, ¶ 52) Claim 11: The rejection of claim 1 is incorporated. Iyer further teaches: wherein determining whether the identified actions satisfy the criterion for the gesture type is based on determining a context of the user within a 3D environment. ("two-handed gesture sequences for opening and closing files, applications, or workspaces (or for any other opening and closing action, depending on application or context)", ¶ 152 and figs. 24A-25B) Claim 12: The rejection of claim 11 is incorporated. Iyer further teaches: wherein the 3D environment comprises a physical environment. (use distinctive visual characteristics of the physical environment to identify specific images or shapes which are then usable to calculate HMD 102's position and orientation, ¶ 64) Claim 13: The rejection of claim 11 is incorporated. Iyer further teaches: wherein the 3D environment comprises an extended reality (XR) environment. (environment where a virtual, augmented, or mixed reality, ¶ 15. It was well within the capabilities of a person having ordinary skill in the art to have realized that Mixed Reality (MR) is a component of Extended Reality (XR).) Claim 15: The rejection of claim 1 is incorporated. Iyer further teaches: wherein determining the user intended action comprises determining that the user activity is associated with simultaneous one-handed gestures. (As mentioned above for claim 1, the identified gestures are two handed gestures, e.g., see ¶¶ 65 and 89. It was well within the capabilities of a person having ordinary skill in the art to have realized that two-handed gestures are interpreted as being simultaneous one-handed gestures because they consist of either two identical, independent motions or a dominant hand moving relative to a static, supportive, or passive "base" hand. The description of Iyer supports that two-handed gestures are meant to be performed simultaneously by describing a method that compensates for "natural human asynchronicity," the tendency for hands to start at different times, to recognize the intended coordinated two-handed movement, ¶ 104.) Claim 16: The rejection of claim 1 is incorporated. Iyer further teaches: further comprising: using at least a portion of the identified actions performed by the two hands to directly enable a functional component displayed within an extended reality (XR) environment. (the use may place both hands 2107 and 2108 over the object 2106 to capture the area, with palms facing out and all ten fingers extended. After a waiting in position for a predetermined amount of time, the xR object 2106 is highlighted with surround effect 2109 and object menu 2110 appears [enable a functional component], as shown in frame 2103, ¶ 148 and fig. 21) Claim 17: The rejection of claim 1 is incorporated. Iyer further teaches: wherein the user intended action comprises a pan interaction, a zoom interaction, or a rotation of one or more elements of a user interface. (xR application may provide a workspace that enables operations for single VO [virtual objects] or group of VOs in workspace, e.g., rotate [rotation of one or more elements of a user interface], ¶¶ 126 and 128) Independent Claims 18 and 20: Claim(s) 18 and 20 are directed to a device and computer-readable storage medium for accomplishing the steps of the method in claim 1, and are rejected using similar rationale(s). Claim 19: The rejection of claim 18 is incorporated. Claim 19 is directed to a device for accomplishing the steps of the method in claim 2, and is rejected using similar rationale(s). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Iyer (US 20190384408 A1) as applied to claim 1 above, and further in view of Gorski; Ryan Joseph et al. (hereinafter Gorski – US 20240312247 A1). Claim 4: The rejection of claim 1 is incorporated. Iyer further teaches the gestures can be: user centric (a gesture can be identified based on fitting a virtual skeleton to the user and analyzing the positional changes of its joints and segments over time, ¶ 82. This can be described as “user centric” due to its focus on modeling the user's physical characteristics.), or app centric (a gesture can be identified without relying on fitting a virtual skeleton to the user and analyzing the positional changes of its joints and segments over time, by sending raw point-cloud data directly to a feature extraction routine within a gesture sequence recognition, ¶¶ 82-83, and this can be seen as “app centric” because it prioritizes direct data processing for the application) Iyer does not appear to expressly teach, but Gorski teaches: wherein the user activity is determined to provide a hybrid gesture as the gesture type, wherein the hybrid gesture comprises a portion of a user centric gesture and a portion of an app centric gesture (Point cloud data [app centric] combined with a 3D skeleton model [user centric] enables precise tracking of an occupant's articulating joints, gestures, and six-degree-of-freedom head movement to accurately predict gaze direction and indication direction, ¶ 94). Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Iyer to include wherein the user activity is determined to provide a hybrid gesture as the gesture type, wherein the hybrid gesture comprises a portion of a user centric gesture and a portion of an app centric gesture, as taught by Gorski. One would have been motivated to make such a combination in order to improve the functionalities and accuracy of the method, e.g., improved accuracy in tracking of articulating joints, enhanced gesture detection using depth information, and improved prediction of indication direction or gaze direction, Gorski ¶ 94. Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Iyer (US 20190384408 A1) as applied to claim 1 above, and further in view of Rao; Liang et al. (hereinafter Rao – US 20170192668 A1). Claim 14: The rejection of claim 1 is incorporated. Iyer does not appear to expressly teach, but Rao teaches: wherein determining the user intended action comprises determining that the user activity is associated with two simultaneous two-handed gestures. (a single gesture bears two functions…refresh and exit, ¶ 82) Accordingly, it would have been obvious to a person having ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the method of Iyer to include wherein determining the user intended action comprises determining that the user activity is associated with two simultaneous two-handed gestures, as taught by Rao. One would have been motivated to make such a combination in order to which improve user experience and interface consistency offered by the method, Rao ¶ 82. Response to Arguments 101: Applicant's 101 arguments have been fully considered but they are not persuasive. First, for Step 2A (Prong One) analysis for claims 1, 18, 20, which were amended to emphasize sensor related language and to include that actions are interpreted “as input to the device”, the applicant alleges that the claims do not recite a mental process, because the “amendments emphasize that the claims are directed to device-implemented, sensor-based interpretation of two-handed activity as device input, not to observation or mental classification of human movement”, because the Instant Specification “repeatedly characterizes the invention as sensor-based, device-executed interpretation of two-handed activity in a 3D/XR environment”, which “describes the claimed technological environment and technical problem”, and that evidence for such technological environment also includes the “claimed determination of gestures and actions”, “criterion for gesture type”, and ‘interpreting the identified actions “as input to the device.”’. Remarks Pg(s) 7-10. The applicant also cites other portions of the Instant Specification, including ¶ 3, which states that “and-based input systems for identifying user input and displaying content based on the hand-based user input may be improved with respect to providing means for users to create, edit, view, or otherwise use content in an extended reality (XR) environment, especially when detecting user intentions for two handed gestures”, Remarks Pg(s) 9, and that the gesture type may be identified as user-centric or app-centric, Remarks Pg(s) 11. The examiner respectfully disagrees because: Claims 1 and 20 recite “as input to the device”, but claim 18 doesn’t. Neither claim 1, claim 18, nor claim 20 requires an “XR environment”. In their argument(s), the applicant has quoted many sections of the Instant Specification. However, although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). None of the quoted paragraphs provide evidence that the claim doesn’t recite an abstract idea. For example, ¶ 3, does not point to a specific problem in the extended reality environment, it only generally states that improvements “may be” had “especially when detecting user intentions for two handed gestures”. On its face, this is not a problem that is specific to extended reality. E.g., the problem of accuracy in identifying two-hand gestures in American Sign Language may be improved simply by studying and practicing. Whether the user activity is user centric or app centric doesn’t make the claims any less abstract or provide a practical application because these types of user activity predate computers. E.g., in sign language, user-centric gestures may be interpreted as gestures in the first-person perspective to mimic how a person interacts with an object (e.g., gripping an imaginary steering wheel to drive), and app-centric gestures use a third-person view to map an object's spatial features or shape in the air (e.g., tracing a boxy outline of a car). Even though the claims include computer/sensor components, and two-handed activity in a 3D environment, each of the recited steps/functions are performable by a human mind, with and/or without the use of aids, and therefore recite an abstract idea grouped under “Mental Processes”, as explained in the 101 rejection section above. That is, the claims fail the test under Step 2A (Prong One). Second, for Step 2A (Prong Two), the applicant alleges that any alleged judicial exception is integrated into a practical application because “the amended claims do not merely automate the general idea of recognizing gestures, [but] recite a specific implementation for improving how a device interprets two-handed activity in a 3D coordinate system as input to the device by determining whether identified two-hand actions satisfy a criterion for a gesture type and, when they do, interpreting the identified actions based on a reference element corresponding to that gesture type”, again quotes multiple paragraphs of the Instant Specification and arguments provided for Step 2A (Prong One), Remarks Pg(s) 11-14. The examiner respectfully disagrees because: For the one or more of the reason(s) provided above for Step 2A (Prong One). One or more of the arguments/responses provided for Step 2A (Prong One) actually belong in the analysis at Prong Two. Although the Instant Specification (as published) purports an improvement in accuracy of gesture detection and/or protecting sensitive data in an extended reality environment, e.g., by using measuring arc length, ¶¶ 60-63 and 71, using gaze input in a specific way, ¶ 109, and/or measuring a distance between a user an the user interface, ¶ 152, none of these technological improvements are adequately reflected in the claims. The examiner suggests the applicant look into the specific details in the Instant Specification concerning how the accuracy/privacy improvements are achieved and reflect the same in the claims to help overcome the 101 rejection by practical application. Third, the applicant argues that the claims are patent-eligible under the BASCOM "significantly more" (Step 2B) test because they present a non-conventional, non-generic ordered combination of sensor-based operations for interpreting 3D, two-handed gestures. The applicant contends the Office Action failed to prove that this specific, integrated combination of sensor data, gesture-type criteria, and reference elements was well-understood, routine, or conventional. Remarks Pg(s) 15. The examiner respectfully disagrees because: See reasons in response to arguments 1 and 2, and the 101 rejection section above. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Instead, the claims as a whole include only the abstract idea with the mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea, which cannot provide a practical application or satisfy the “significantly more” requirement. The claims at hand are not like Bascom. Bascom was directed to an uncommon method of filtering Internet content, and the present claims do not involve an uncommon method of filtering Internet content. Lastly, in the 101 analysis, an examiner only need to explain why additional claim elements are well-known, routine, and conventional during Step 2B, if at Step 2A (Prong Two), the examiner explained that the elements were extra-solution elements because they are well-known, routine, and conventional. The examiner this not rely on this extra-solution analysis in Step 2A (Prong Two) for the independent claims, and is not required to provide this evidence for these claims. For claim 16, the limitation of “using at least a portion of the identified actions performed by the two hands to directly enable a functional component displayed” is extra-solution activity and ineffective to show significantly more because it is directed to well-understood, routine, conventional activity. However, the applicant has not provided evidence to rebut the prior art evidence already provided by examiner. Fourth, the applicant relies on the argument(s) above to allege patentability of the remaining claims. The examiner respectfully disagrees for similar reason(s). 112(a): The 112(a) rejections of Claims 2-4, 6 and 19 are overcome after further consideration. The applicant’s 112(a) argument concerning claim 14 is fully considered but is unpersuasive. The applicant alleges that the specification supports claim 14 for two simultaneous two-handed gestures by citing text (¶¶ 64 and162) showing pan and zoom can happen at once with two hands, using an algorithm to tell gestures apart, Remarks Pg(s) 18. The examiner respectfully disagrees. The description is insufficient because it seems to treat pan and zoom as a single multi-task gesture rather than defining how two/simultaneous two-handed gestures (that is, involving at least four-hands) are tracked and distinctly generated. Therefore, the Instant Specification doesn’t sufficiently describe the claimed simultaneous two-handed gestures, as it lacks structural details on the algorithm distinguishing these gestures. 112(b): The 112(b) rejections for Claims 2-4, 6 and 19 are overcome after further consideration. Here, the terms "user centric gesture" and "app centric gesture” are not recognized terms of art, and their scope is broad, as explained in the claim interpretation section above. The remaining 112(b) arguments are fully considered but are unpersuasive. First, the applicant alleges that claim 6 is definite because the specification, at ¶ 160, describes a step of determining to associate actions with intended actions and further alleges that “The fact that the Office Action can articulate a compact-prosecution interpretation of claim 6 confirms that the claim can be reasonably understood”, Remarks Pg(s) 20 The examiner respectfully disagrees because: The applicant has quoted ¶ 160 to allege the clarity of claim 6. However, this paragraph does not differentiate between the determining and identifying steps because it applies the same app-centric technique and reference orientation to interpret motions and define both functions. Both steps rely on shared spatial logic, such as using user or application orientation to detect gestures, merging the technical boundaries between them. Furthermore, Claim 6 is unclear for reciting “the user intended action” without sufficient antecedent basis for this limitation in the claim. Pursuant MPEP 2173.05(e), a ‘lack of clarity could arise where a claim refers to "said lever" or "the lever," where the claim contains no earlier recitation or limitation of a lever and where it would be unclear as to what element the limitation was making reference’. Here, there is no previous mention of a “user intended action” in the claim(s). Furthermore, pursuant MPEP 2111.01.II, it is improper to import claim limitations from the specification. Stating that an examiner's ability to force a temporary working interpretation for prior art analysis under MPEP § 2173.06 does not cure a claim's failure to provide the public and a person of ordinary skill in the art clear notice of its absolute boundaries under 35 U.S.C. § 112(b). This argument has no legal basis. The forced interpretations for each 112(b), indefinite limitations, are attempts to promote compact prosecution, not an indication of clarity. Second, the applicant alleges, for Claim 8, that “the device worn on the head” is clear because of recitation claim 1 recites “a device having a processor and one or more sensors” and the Instant Specification recites that a device can be head-worn. Remarks Pg(s) 20. The examiner respectfully disagrees because: As explained before, there is insufficient antecedent basis for this limitation in the claim itself. Specifically, the claim doesn’t previously mention a device worn on the head, and it is improper to import limitations from the Instant Specification. Third, the applicant alleges that the Instant Specification explains, for claim 14, the recited “user activity is associated with two simultaneous two-handed gestures”, and that “claim 14 would be understood as directed to a two-hand user activity that is associated with simultaneous two- handed gesture functions, not to a requirement for additional hands”. Remarks Pg(s) 21 The examiner doesn’t import limitations from the specification. Although the specification mentions two simultaneous two-handed gestures, it doesn’t help clarify how the two simultaneous two-handed gestures can be performed using only two hands. Furthermore, because the claim states the qualifies the two-handed gestures as being “two” and that are “simultaneous[ly]” performed, a person having ordinary skill in the art would understand this as a requirement for at least four hands, because 2 x 2 = 4. For purposes of compact prosecution only, the examiner interprets the limitation(s) as being directed to a single two-handed gesture that can be used control two functions at the same time, e.g., zoom and pan function, see ¶¶ 64 and 162. Fourth, the applicant alleges claims 14, 15 and 17 are clear in view of Instant Specification at ¶¶ 150 and 163. Remarks Pg(s) 21-22. The examiner respectfully disagrees because: Because, as mentioned above concerning claim 6, pursuant MPEP 2173.05(e), a ‘lack of clarity could arise where a claim refers to "said lever" or "the lever," where the claim contains no earlier recitation or limitation of a lever and where it would be unclear as to what element the limitation was making reference’. Here, there is no previous mention of a “user intended action” in the claim(s). Furthermore, pursuant MPEP 2111.01.II, it is improper to import claim limitations from the specification. Fifth, the applicant relies on the argument(s) above to allege patentability of the remaining claims. The examiner respectfully disagrees for similar reason(s). 102/103: First, Iyer’s core disclose is gesture sequence recognition of predefined gestures, whereas claim 1 determine whether already-identified actions satisfy a criterion for input to the device, Remarks Pg(s) 22-23. The examiner respectfully disagrees because: The claimed elements are taught by the cited portions of Iyer, specifically by ¶¶ 81-82, 86-87, 119 and 166, because these portions reflects a system that tracks skeletal joints and segments over time using specific reference elements (such as coordinates, reference images, and parameters like hand shape) to verify that a recognized positive gesture sequence meets a required accuracy threshold for matched directional and positional attributes. Checking the required gesture accuracy threshold must happen after identifying the gesture sequence because the system needs to first recognize and isolate the complete set of matched directional and positional attributes across the tracked skeletal frames before it can compute and verify those values against the specific parameters. As explained in 102 rejection above, the gestures in Iyer are used as input to the device. Specifically, a device, such as HMD 102 of system 300, ¶¶ 5 and 65 and figs. 2-3, interprets actions as input by capturing video data of hand movements, identifying specific gesture sequences through software, and mapping those motions to control user interface actions or application functions, ¶¶ 74, 77, and 91. Second, Iyer doesn’t teach the claimed gesture type, e.g., user-centric, app-centric, and hybrid/blended gesture types. Remarks Pg(s) 23. The examiner respectfully disagrees because: Claim 1 doesn’t require the gesture type to be user-centric, centric, or hybrid/blended. Concerning claims 2-3, the terms user centric and app-centric are terms that the Instant Specification provides examples, see ¶ 5 (as published), but the interpretation of the terms are not limited to these examples. In a broad sense, “user centric” gestures include gestures that are identified by using user position references, such as user’s joint positions, and “app centric” as being gestures that don’t require using user position references. User centric gestures can also be interpreted as including gestures that control the system or application-agnostic navigation, while app centric gestures are tied to specific app functions and views. The 102 rejections over Iyer clearly maps the gestures as being app-centric or user-centric. Concerning claim 4, the hybrid gestures limitation is addressed by the mapping in the 103 rejection Iyer-Gorski combination. Third, Iyer doesn’t teach different reference elements corresponding to the different gesture types. Remarks Pg(s) 23. The examiner respectfully disagrees because: For the reason(s) mentioned for the second argument above. As mentioned in the 102 rejection above, reference elements include points in a trajectory used to identify specific sequences, Iyer ¶¶ 81-81. Fourth, Iyer fails to teach the claimed, ordered combination of using sensor data to determine a gesture-type criterion and interpreting two-hand actions based on a specific reference element. Unlike the claimed invention, Iyer only initiates actions after a gesture sequence is detected, rather than interpreting actions based on sensor-data-driven criteria and reference elements. Remarks Pg(s) 24. The examiner respectfully disagrees because: As mentioned in the 102 rejection above, reference elements include points in a trajectory used to identify specific sequences, Iyer ¶¶ 81-81. As reflected in the mapping above, Iyer teaches interpreting actions using sensor data because it checks if gesture sequences meet a set accuracy number, and it matches the gestures to reference images using specific rules like direction, position, location, and time duration. Specifically, Iyer teaches a positive gesture sequence recognition meets an accuracy threshold [satisfy a criterion], ¶ 166, also, the sequences are detected by the specific series of gestures determined to match certain attributes and/or reference images, ¶¶ 86-87, e.g., direction, position, and/or location, ¶ 119, and gesture duration, ¶ 135. Herein, a “gesture type” is broadly interpreted as including one or more candidate/identified gesture sequences. Fifth, the applicant relies on the arguments above to allege patentability of the remaining claims. The examiner respectfully disagrees because for the reason(s) above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Below is a list of these references, including why they are pertinent: Reville; Brendan et al. US 20110289455 A1, is pertinent to claim 1 for disclosing a method that enables user-interface navigation by recognizing a coordinated, two-handed press gesture, where both hands move away from the body towards a capture device, Reville Claim 10, including utilizing real-time skeletal mapping, adjusting to the user's movements via depth and visual images, ¶ 86 and figs. 3 and 6. Kin; Kenrick Cheng-kuo US 10261595 B1, is pertinent to claim 1 for disclosing a console that detects gestures performed using both of a user’s hands, Abstract and col 14:46-49. Ravasz; Jonathan et al. US 20200387229 A1, is pertinent to claim 1 for disclosing an artificial reality system for rendering, presenting, and controlling element in an artificial reality environment, Abstract, wherein a gesture detector identifies two-handed inputs in addition to one-handed inputs, ¶¶ 43 and 94. Schwesinger; Mark et al. US 20150193107 A1, is pertinent to claim 1 for disclosing a device that detects one or more hands, tracks one or more hands and recognizes gestures, cols 13:2-4. The following reference are pertinent to claim 16 for disclosing that the limitation of “using at least a portion of the identified actions performed by the two hands to directly enable a functional component displayed” is well-understood, routine, conventional activity: AU; Kin Chung et al. US 20120154313 A1 [0062] Another example of a pop-up interface is a virtual keyboard 800, as shown in FIG. 8. The virtual keyboard 800 is a two-hand interface and can be activated with a two-hand gesture, such as a two-hand five finger tap. El Dokor; Tarek US 8928590 B1 “Users can also lift a single hand or both hands above the keyboard. Thus, the inventive system provides for a delineation of gesture recognition between an active zone that is enabled once the user's hand (or hands) is visible for the cameras and inactive zone when the user is typing on the keyboard or using the mouse”, col 2:26-32. Wong; Yoon Kean US 20110320982 A1 [0002]…Traditionally, a user is required to use two hands and multiple input mechanisms to activate a palmtop computer application. Karlsson; David et al. US 20120131488 A1 [0108]…The controls can accept two-hand inputs, one to activate one of the movable icons and one to carry out a touch gesture to select a choice that carries out a desired interaction associated with the activated alternate interaction icon (block 218). Any inquiry concerning this communication or earlier communications from the examiner should be directed to GABRIEL S MERCADO whose telephone number is (408)918-7537. The examiner can normally be reached Mon-Fri 8am-5pm (Eastern Time). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kieu Vu can be reached at (571) 272-4057. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Gabriel Mercado/Primary Examiner, Art Unit 2171
Read full office action

Prosecution Timeline

Apr 22, 2024
Application Filed
Feb 18, 2026
Non-Final Rejection mailed — §101, §102, §103
May 14, 2026
Applicant Interview (Telephonic)
May 14, 2026
Examiner Interview Summary
May 15, 2026
Response Filed
Jul 30, 2026
Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705260
USER-DEFINED GRAPHICAL HIERARCHIES
2y 11m to grant Granted Aug 11, 2026
Patent 12656927
POSITION INPUT TERMINAL WITH Z-POSITION DURATION-BASED AND Z-POSITION RANGE-BASED MODE SWITCHING
2y 9m to grant Granted Jun 16, 2026
Patent 12543983
SYSTEMS AND METHODS FOR EMOTION PREDICTION
3y 1m to grant Granted Feb 10, 2026
Patent 12535942
BLOWOUT PREVENTER SYSTEM WITH DATA PLAYBACK
5y 6m to grant Granted Jan 27, 2026
Patent 12511024
Multi-Application Interaction Method
2y 9m to grant Granted Dec 30, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
42%
Grant Probability
69%
With Interview (+26.6%)
3y 5m (~1y 1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 206 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month