Prosecution Insights
Last updated: August 17, 2026
Application No. 18/827,175

SELECTING IMAGE SENSOR FOR OBJECT CLASSIFICATION

Non-Final OA §102
Filed
Sep 06, 2024
Examiner
ADEDIRAN, ABDUL -SAMAD A
Art Unit
2621
Tech Center
2600 — Communications
Assignee
Snap Inc.
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
2m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
496 granted / 632 resolved
+16.5% vs TC avg
Moderate +14% lift
Without
With
+13.6%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 1m
Avg Prosecution
29 currently pending
Career history
651
Total Applications
across all art units

Statute-Specific Performance

§101
2.2%
-37.8% vs TC avg
§103
46.6%
+6.6% vs TC avg
§102
16.7%
-23.3% vs TC avg
§112
26.8%
-13.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 632 resolved cases

Office Action

§102
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Oath/Declaration Oath/Declaration as filed on August 11, 2025 is noted by the Examiner. Claim Objections Claims 1, 9, and 17 are objected to because of the following informalities: In particular, the limitations “a respective predicted skeleton” in fourth line, seventh line, and sixth line respectively of the claims render the claims indefinite, because the meaning of the coined terms “a respective predicted skeleton” recited in the fourth line, the seventh line, and sixth line of the claims respectively are not apparent in light of the specification. See MPEP § 2173.05(a). Examiner recommends applicant amend the claim, without adding new matter, to positively recite in definite terms more clearly what “a respective predicted skeleton” actually is. Accordingly, any claim(s) dependent on claims 1, 9, or 17 are objected to based on same above reasoning. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-3, 8-11, and 16-19 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Schwarz et al., U.S. Patent Application Publication 2023/0244316 A1 (hereinafter Schwarz). Regarding claim 1, Schwarz teaches a computer-implemented method comprising: obtaining a first image captured by a first image sensor of a device, the device including the first image sensor and one or more second image sensors (305 FIGS. 1-3, paragraph[0044] of Schwarz teaches first neural network 315 may evaluate input data, such as pre-processed sensor data from data pre-processing machines 310, raw sensor data from sensor suite 305, UI data 320, data from secondary device inputs 325, and heuristically evaluated data 330; first neural network 315 may evaluate input data for a sequence of data frames (e.g., a single data frame or a plurality of data frames), and output an indication of a likelihood of gesture interaction 335, such as an indication of a likelihood of the user performing one or more subsequent gesture interactions with a user interface during a predetermined window of one or more data frames; in some examples, a single data frame may provide a clear indication that the user is not intending to make a gesture interaction with their hands in a subsequent data frame (e.g., holding a baby, taking a casserole out of the oven), while in other scenarios cooperatively considering a plurality of sequential frames may allow first neural network 315 to more accurately assess the context of a user's hand movements; in other words, first neural network 315 may infer whether a use is likely to interact with the UI via gesture input, unlikely, not at all likely, already interacting, etc; and for example, a likelihood may be output as a real number between 0 and 1, where 0 represents that the user is not at all likely to perform a gesture in the predetermined window, while 1 represents already interacting with the UI or has already initiated performing a gesture and See also at least paragraphs[0030], [0036]-[0041], [0043], [0045]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches a computing system, which includes at least a processor and memory storing instructions executable by the processor, for implementing a first neural network that evaluates input data, such as sensor data selectively from a sensor suite, for a sequence of data frames and outputs an indication of a likelihood of a gesture interaction)); generating, based on obtaining the first image, a respective predicted skeleton corresponding to a respective view from each of the first image sensor and the one or more second image sensors; selecting, based on generating the respective predicted skeletons, an image sensor from among the first image sensor and the one or more second image sensors, the selected image sensor being used for classifying the first image; and determining, based on the selected image sensor, a classification for the first image (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)). Regarding claim 2, Schwarz teaches the computer-implemented method of claim 1, wherein the first image corresponds to an object, and wherein the classification corresponds to an identification of the object or a position of the object (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)). Regarding claim 3, Schwarz teaches the computer-implemented method of claim 2, wherein the object is a hand, and wherein the classification corresponds to a hand gesture (FIGS. 1-4B, paragraph[0050] of Schwarz teaches in some implementations, a virtual skeleton or other data structure for tracking feature positions (e.g., joints) may be fit to the pixels of depth and/or color video that correspond to the user; FIG. 4A shows an example virtual skeleton 400; the virtual skeleton includes a plurality of skeletal segments 405 pivotally coupled at a plurality of joints 410; in some embodiments, a body-part designation may be assigned to each skeletal segment and/or each joint; in FIG. 4A, the body-part designation of each skeletal segment 405 is represented by an appended letter: A for the head, B for the clavicle, C for the upper arm, D for the forearm, E for the hand, F for the torso, G for the pelvis, H for the thigh, J for the lower leg, and K for the foot; likewise, a body-part designation of each joint 410 is represented by an appended letter: A for the neck, B for the shoulder, C for the elbow, D for the wrist, E for the lower back, F for the hip, G for the knee, and H for the ankle; naturally, the arrangement of skeletal segments and joints shown in FIG. 4A is in no way limiting; and a virtual skeleton consistent with this disclosure may include virtually any type and number of skeletal segments, joints, and/or other features, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0049], [0051]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, and wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton)). Regarding claim 8, Schwarz teaches the computer-implemented method of claim 1, further comprising: obtaining a second image captured by the selected image sensor, wherein determining the classification for the first image is based on the second image (FIGS. 1-4B, paragraph[0077] of Schwarz teaches the use of multiple sequential frames may allow for anticipation or early recognition of some gesture interactions; first neural network 315 may generate predictions based on each frame individually, and/or based on changes in input data across multiple frames; additionally, the sequential frames may be used to smooth predictions, for example, selecting a most frequent prediction over a window of frames and/or tossing out predictions that do not align with those frames before and after; and in some examples, frames with higher confidence scores may be weighed more heavily than frames with lower confidence scores in generating a likelihood of interaction for predetermined window of frames 805, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0045], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton, and wherein multiple sequential frames are utilized for early recognition of the gesture interaction combined with the classification of each pixel)). Regarding claim 9, Schwarz teaches a system comprising: at least one processor; at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: obtaining a first image captured by a first image sensor of a device, the device including the first image sensor and one or more second image sensors (305 FIGS. 1-3, paragraph[0044] of Schwarz teaches first neural network 315 may evaluate input data, such as pre-processed sensor data from data pre-processing machines 310, raw sensor data from sensor suite 305, UI data 320, data from secondary device inputs 325, and heuristically evaluated data 330; first neural network 315 may evaluate input data for a sequence of data frames (e.g., a single data frame or a plurality of data frames), and output an indication of a likelihood of gesture interaction 335, such as an indication of a likelihood of the user performing one or more subsequent gesture interactions with a user interface during a predetermined window of one or more data frames; in some examples, a single data frame may provide a clear indication that the user is not intending to make a gesture interaction with their hands in a subsequent data frame (e.g., holding a baby, taking a casserole out of the oven), while in other scenarios cooperatively considering a plurality of sequential frames may allow first neural network 315 to more accurately assess the context of a user's hand movements; in other words, first neural network 315 may infer whether a use is likely to interact with the UI via gesture input, unlikely, not at all likely, already interacting, etc; and for example, a likelihood may be output as a real number between 0 and 1, where 0 represents that the user is not at all likely to perform a gesture in the predetermined window, while 1 represents already interacting with the UI or has already initiated performing a gesture and See also at least paragraphs[0030], [0036]-[0041], [0043], [0045]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches a computing system, which includes at least a processor and memory storing instructions executable by the processor, for implementing a first neural network that evaluates input data, such as sensor data selectively from a sensor suite, for a sequence of data frames and outputs an indication of a likelihood of a gesture interaction)); generating, based on obtaining the first image, a respective predicted skeleton corresponding to a respective view from each of the first image sensor and the one or more second image sensors; selecting, based on generating the respective predicted skeletons, an image sensor from among the first image sensor and the one or more second image sensors, the selected image sensor being used for classifying the first image; and determining, based on the selected image sensor, a classification for the first image (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)). Regarding claim 10, Schwarz teaches the system of claim 9, wherein the first image corresponds to an object, and wherein the classification corresponds to an identification of the object or a position of the object (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)). Regarding claim 11, Schwarz teaches the system of claim 10, wherein the object is a hand, and wherein the classification corresponds to a hand gesture (FIGS. 1-4B, paragraph[0050] of Schwarz teaches in some implementations, a virtual skeleton or other data structure for tracking feature positions (e.g., joints) may be fit to the pixels of depth and/or color video that correspond to the user; FIG. 4A shows an example virtual skeleton 400; the virtual skeleton includes a plurality of skeletal segments 405 pivotally coupled at a plurality of joints 410; in some embodiments, a body-part designation may be assigned to each skeletal segment and/or each joint; in FIG. 4A, the body-part designation of each skeletal segment 405 is represented by an appended letter: A for the head, B for the clavicle, C for the upper arm, D for the forearm, E for the hand, F for the torso, G for the pelvis, H for the thigh, J for the lower leg, and K for the foot; likewise, a body-part designation of each joint 410 is represented by an appended letter: A for the neck, B for the shoulder, C for the elbow, D for the wrist, E for the lower back, F for the hip, G for the knee, and H for the ankle; naturally, the arrangement of skeletal segments and joints shown in FIG. 4A is in no way limiting; and a virtual skeleton consistent with this disclosure may include virtually any type and number of skeletal segments, joints, and/or other features, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0049], [0051]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, and wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton)). Regarding claim 16, Schwarz teaches the system of claim 9, the operations further comprising: obtaining a second image captured by the selected image sensor, wherein determining the classification for the first image is based on the second image (FIGS. 1-4B, paragraph[0077] of Schwarz teaches the use of multiple sequential frames may allow for anticipation or early recognition of some gesture interactions; first neural network 315 may generate predictions based on each frame individually, and/or based on changes in input data across multiple frames; additionally, the sequential frames may be used to smooth predictions, for example, selecting a most frequent prediction over a window of frames and/or tossing out predictions that do not align with those frames before and after; and in some examples, frames with higher confidence scores may be weighed more heavily than frames with lower confidence scores in generating a likelihood of interaction for predetermined window of frames 805, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0045], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton, and wherein multiple sequential frames are utilized for early recognition of the gesture interaction combined with the classification of each pixel)). Regarding claim 17, Schwarz teaches a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: obtaining a first image captured by a first image sensor of a device, the device including the first image sensor and one or more second image sensors (305 FIGS. 1-3, paragraph[0044] of Schwarz teaches first neural network 315 may evaluate input data, such as pre-processed sensor data from data pre-processing machines 310, raw sensor data from sensor suite 305, UI data 320, data from secondary device inputs 325, and heuristically evaluated data 330; first neural network 315 may evaluate input data for a sequence of data frames (e.g., a single data frame or a plurality of data frames), and output an indication of a likelihood of gesture interaction 335, such as an indication of a likelihood of the user performing one or more subsequent gesture interactions with a user interface during a predetermined window of one or more data frames; in some examples, a single data frame may provide a clear indication that the user is not intending to make a gesture interaction with their hands in a subsequent data frame (e.g., holding a baby, taking a casserole out of the oven), while in other scenarios cooperatively considering a plurality of sequential frames may allow first neural network 315 to more accurately assess the context of a user's hand movements; in other words, first neural network 315 may infer whether a use is likely to interact with the UI via gesture input, unlikely, not at all likely, already interacting, etc; and for example, a likelihood may be output as a real number between 0 and 1, where 0 represents that the user is not at all likely to perform a gesture in the predetermined window, while 1 represents already interacting with the UI or has already initiated performing a gesture and See also at least paragraphs[0030], [0036]-[0041], [0043], [0045]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches a computing system, which includes at least a processor and memory storing instructions executable by the processor, for implementing a first neural network that evaluates input data, such as sensor data selectively from a sensor suite, for a sequence of data frames and outputs an indication of a likelihood of a gesture interaction)); generating, based on obtaining the first image, a respective predicted skeleton corresponding to a respective view from each of the first image sensor and the one or more second image sensors; selecting, based on generating the respective predicted skeletons, an image sensor from among the first image sensor and the one or more second image sensors, the selected image sensor being used for classifying the first image; and determining, based on the selected image sensor, a classification for the first image (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)). Regarding claim 18, Schwarz teaches the non-transitory computer-readable storage medium of claim 17, wherein the first image corresponds to an object, and wherein the classification corresponds to an identification of the object or a position of the object (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)). Regarding claim 19, Schwarz teaches the non-transitory computer-readable storage medium of claim 18, wherein the object is a hand, and wherein the classification corresponds to a hand gesture (FIGS. 1-4B, paragraph[0050] of Schwarz teaches in some implementations, a virtual skeleton or other data structure for tracking feature positions (e.g., joints) may be fit to the pixels of depth and/or color video that correspond to the user; FIG. 4A shows an example virtual skeleton 400; the virtual skeleton includes a plurality of skeletal segments 405 pivotally coupled at a plurality of joints 410; in some embodiments, a body-part designation may be assigned to each skeletal segment and/or each joint; in FIG. 4A, the body-part designation of each skeletal segment 405 is represented by an appended letter: A for the head, B for the clavicle, C for the upper arm, D for the forearm, E for the hand, F for the torso, G for the pelvis, H for the thigh, J for the lower leg, and K for the foot; likewise, a body-part designation of each joint 410 is represented by an appended letter: A for the neck, B for the shoulder, C for the elbow, D for the wrist, E for the lower back, F for the hip, G for the knee, and H for the ankle; naturally, the arrangement of skeletal segments and joints shown in FIG. 4A is in no way limiting; and a virtual skeleton consistent with this disclosure may include virtually any type and number of skeletal segments, joints, and/or other features, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0049], [0051]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, and wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton)). Potentially Allowable Subject Matter Claims 4-7, 12-15, and 20 would be allowable if rewritten to overcome applicable objection(s) indicated above, and if rewritten in independent form including all of the limitations of the base claim and any intervening, because for each of the 4-7, 12-15, and 20 the prior art references of record do not teach the combination of all element limitations as presently claimed. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDUL-SAMAD A ADEDIRAN whose telephone number is (571)272-3128. The examiner can normally be reached on Monday through Thursday, 8:00 am to 5:00 pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amr Awad can be reached on 571-272-7764. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ABDUL-SAMAD A ADEDIRAN/Primary Examiner, Art Unit 2621
Read full office action

Prosecution Timeline

Sep 06, 2024
Application Filed
Jun 09, 2026
Non-Final Rejection mailed — §102 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12701898
DISPLAY DEVICE
1y 8m to grant Granted Aug 04, 2026
Patent 12687935
METHOD AND APPARATUS FOR A THREE DIMENSIONAL INTERFACE
2y 5m to grant Granted Jul 21, 2026
Patent 12687922
METHODS FOR DISPLAYING AND REARRANGING OBJECTS IN AN ENVIRONMENT
2y 2m to grant Granted Jul 21, 2026
Patent 12682664
CONSTRUCTING COMPACT THREE-DIMENSIONAL BUILDING MODELS
1y 7m to grant Granted Jul 14, 2026
Patent 12670829
DISPLAY DEVICE
1y 5m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
92%
With Interview (+13.6%)
2y 1m (~2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 632 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month