DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Oath/Declaration
Oath/Declaration as filed on August 11, 2025 is noted by the Examiner.
Claim Objections
Claims 1, 9, and 17 are objected to because of the following informalities:
In particular, the limitations “a respective predicted skeleton” in fourth line, seventh line, and sixth line respectively of the claims render the claims indefinite, because the meaning of the coined terms “a respective predicted skeleton” recited in the fourth line, the seventh line, and sixth line of the claims respectively are not apparent in light of the specification. See MPEP § 2173.05(a). Examiner recommends applicant amend the claim, without adding new matter, to positively recite in definite terms more clearly what “a respective predicted skeleton” actually is. Accordingly, any claim(s) dependent on claims 1, 9, or 17 are objected to based on same above reasoning.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 8-11, and 16-19 are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Schwarz et al., U.S. Patent Application Publication 2023/0244316 A1 (hereinafter Schwarz).
Regarding claim 1, Schwarz teaches a computer-implemented method comprising: obtaining a first image captured by a first image sensor of a device, the device including the first image sensor and one or more second image sensors (305 FIGS. 1-3, paragraph[0044] of Schwarz teaches first neural network 315 may evaluate input data, such as pre-processed sensor data from data pre-processing machines 310, raw sensor data from sensor suite 305, UI data 320, data from secondary device inputs 325, and heuristically evaluated data 330; first neural network 315 may evaluate input data for a sequence of data frames (e.g., a single data frame or a plurality of data frames), and output an indication of a likelihood of gesture interaction 335, such as an indication of a likelihood of the user performing one or more subsequent gesture interactions with a user interface during a predetermined window of one or more data frames; in some examples, a single data frame may provide a clear indication that the user is not intending to make a gesture interaction with their hands in a subsequent data frame (e.g., holding a baby, taking a casserole out of the oven), while in other scenarios cooperatively considering a plurality of sequential frames may allow first neural network 315 to more accurately assess the context of a user's hand movements; in other words, first neural network 315 may infer whether a use is likely to interact with the UI via gesture input, unlikely, not at all likely, already interacting, etc; and for example, a likelihood may be output as a real number between 0 and 1, where 0 represents that the user is not at all likely to perform a gesture in the predetermined window, while 1 represents already interacting with the UI or has already initiated performing a gesture and See also at least paragraphs[0030], [0036]-[0041], [0043], [0045]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches a computing system, which includes at least a processor and memory storing instructions executable by the processor, for implementing a first neural network that evaluates input data, such as sensor data selectively from a sensor suite, for a sequence of data frames and outputs an indication of a likelihood of a gesture interaction));
generating, based on obtaining the first image, a respective predicted skeleton corresponding to a respective view from each of the first image sensor and the one or more second image sensors; selecting, based on generating the respective predicted skeletons, an image sensor from among the first image sensor and the one or more second image sensors, the selected image sensor being used for classifying the first image; and determining, based on the selected image sensor, a classification for the first image (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)).
Regarding claim 2, Schwarz teaches the computer-implemented method of claim 1, wherein the first image corresponds to an object, and wherein the classification corresponds to an identification of the object or a position of the object (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)).
Regarding claim 3, Schwarz teaches the computer-implemented method of claim 2, wherein the object is a hand, and wherein the classification corresponds to a hand gesture (FIGS. 1-4B, paragraph[0050] of Schwarz teaches in some implementations, a virtual skeleton or other data structure for tracking feature positions (e.g., joints) may be fit to the pixels of depth and/or color video that correspond to the user; FIG. 4A shows an example virtual skeleton 400; the virtual skeleton includes a plurality of skeletal segments 405 pivotally coupled at a plurality of joints 410; in some embodiments, a body-part designation may be assigned to each skeletal segment and/or each joint; in FIG. 4A, the body-part designation of each skeletal segment 405 is represented by an appended letter: A for the head, B for the clavicle, C for the upper arm, D for the forearm, E for the hand, F for the torso, G for the pelvis, H for the thigh, J for the lower leg, and K for the foot; likewise, a body-part designation of each joint 410 is represented by an appended letter: A for the neck, B for the shoulder, C for the elbow, D for the wrist, E for the lower back, F for the hip, G for the knee, and H for the ankle; naturally, the arrangement of skeletal segments and joints shown in FIG. 4A is in no way limiting; and a virtual skeleton consistent with this disclosure may include virtually any type and number of skeletal segments, joints, and/or other features, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0049], [0051]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, and wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton)).
Regarding claim 8, Schwarz teaches the computer-implemented method of claim 1, further comprising: obtaining a second image captured by the selected image sensor, wherein determining the classification for the first image is based on the second image (FIGS. 1-4B, paragraph[0077] of Schwarz teaches the use of multiple sequential frames may allow for anticipation or early recognition of some gesture interactions; first neural network 315 may generate predictions based on each frame individually, and/or based on changes in input data across multiple frames; additionally, the sequential frames may be used to smooth predictions, for example, selecting a most frequent prediction over a window of frames and/or tossing out predictions that do not align with those frames before and after; and in some examples, frames with higher confidence scores may be weighed more heavily than frames with lower confidence scores in generating a likelihood of interaction for predetermined window of frames 805, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0045], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton, and wherein multiple sequential frames are utilized for early recognition of the gesture interaction combined with the classification of each pixel)).
Regarding claim 9, Schwarz teaches a system comprising: at least one processor; at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: obtaining a first image captured by a first image sensor of a device, the device including the first image sensor and one or more second image sensors (305 FIGS. 1-3, paragraph[0044] of Schwarz teaches first neural network 315 may evaluate input data, such as pre-processed sensor data from data pre-processing machines 310, raw sensor data from sensor suite 305, UI data 320, data from secondary device inputs 325, and heuristically evaluated data 330; first neural network 315 may evaluate input data for a sequence of data frames (e.g., a single data frame or a plurality of data frames), and output an indication of a likelihood of gesture interaction 335, such as an indication of a likelihood of the user performing one or more subsequent gesture interactions with a user interface during a predetermined window of one or more data frames; in some examples, a single data frame may provide a clear indication that the user is not intending to make a gesture interaction with their hands in a subsequent data frame (e.g., holding a baby, taking a casserole out of the oven), while in other scenarios cooperatively considering a plurality of sequential frames may allow first neural network 315 to more accurately assess the context of a user's hand movements; in other words, first neural network 315 may infer whether a use is likely to interact with the UI via gesture input, unlikely, not at all likely, already interacting, etc; and for example, a likelihood may be output as a real number between 0 and 1, where 0 represents that the user is not at all likely to perform a gesture in the predetermined window, while 1 represents already interacting with the UI or has already initiated performing a gesture and See also at least paragraphs[0030], [0036]-[0041], [0043], [0045]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches a computing system, which includes at least a processor and memory storing instructions executable by the processor, for implementing a first neural network that evaluates input data, such as sensor data selectively from a sensor suite, for a sequence of data frames and outputs an indication of a likelihood of a gesture interaction));
generating, based on obtaining the first image, a respective predicted skeleton corresponding to a respective view from each of the first image sensor and the one or more second image sensors; selecting, based on generating the respective predicted skeletons, an image sensor from among the first image sensor and the one or more second image sensors, the selected image sensor being used for classifying the first image; and determining, based on the selected image sensor, a classification for the first image (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)).
Regarding claim 10, Schwarz teaches the system of claim 9, wherein the first image corresponds to an object, and wherein the classification corresponds to an identification of the object or a position of the object (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)).
Regarding claim 11, Schwarz teaches the system of claim 10, wherein the object is a hand, and wherein the classification corresponds to a hand gesture (FIGS. 1-4B, paragraph[0050] of Schwarz teaches in some implementations, a virtual skeleton or other data structure for tracking feature positions (e.g., joints) may be fit to the pixels of depth and/or color video that correspond to the user; FIG. 4A shows an example virtual skeleton 400; the virtual skeleton includes a plurality of skeletal segments 405 pivotally coupled at a plurality of joints 410; in some embodiments, a body-part designation may be assigned to each skeletal segment and/or each joint; in FIG. 4A, the body-part designation of each skeletal segment 405 is represented by an appended letter: A for the head, B for the clavicle, C for the upper arm, D for the forearm, E for the hand, F for the torso, G for the pelvis, H for the thigh, J for the lower leg, and K for the foot; likewise, a body-part designation of each joint 410 is represented by an appended letter: A for the neck, B for the shoulder, C for the elbow, D for the wrist, E for the lower back, F for the hip, G for the knee, and H for the ankle; naturally, the arrangement of skeletal segments and joints shown in FIG. 4A is in no way limiting; and a virtual skeleton consistent with this disclosure may include virtually any type and number of skeletal segments, joints, and/or other features, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0049], [0051]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, and wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton)).
Regarding claim 16, Schwarz teaches the system of claim 9, the operations further comprising: obtaining a second image captured by the selected image sensor, wherein determining the classification for the first image is based on the second image (FIGS. 1-4B, paragraph[0077] of Schwarz teaches the use of multiple sequential frames may allow for anticipation or early recognition of some gesture interactions; first neural network 315 may generate predictions based on each frame individually, and/or based on changes in input data across multiple frames; additionally, the sequential frames may be used to smooth predictions, for example, selecting a most frequent prediction over a window of frames and/or tossing out predictions that do not align with those frames before and after; and in some examples, frames with higher confidence scores may be weighed more heavily than frames with lower confidence scores in generating a likelihood of interaction for predetermined window of frames 805, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0045], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton, and wherein multiple sequential frames are utilized for early recognition of the gesture interaction combined with the classification of each pixel)).
Regarding claim 17, Schwarz teaches a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: obtaining a first image captured by a first image sensor of a device, the device including the first image sensor and one or more second image sensors (305 FIGS. 1-3, paragraph[0044] of Schwarz teaches first neural network 315 may evaluate input data, such as pre-processed sensor data from data pre-processing machines 310, raw sensor data from sensor suite 305, UI data 320, data from secondary device inputs 325, and heuristically evaluated data 330; first neural network 315 may evaluate input data for a sequence of data frames (e.g., a single data frame or a plurality of data frames), and output an indication of a likelihood of gesture interaction 335, such as an indication of a likelihood of the user performing one or more subsequent gesture interactions with a user interface during a predetermined window of one or more data frames; in some examples, a single data frame may provide a clear indication that the user is not intending to make a gesture interaction with their hands in a subsequent data frame (e.g., holding a baby, taking a casserole out of the oven), while in other scenarios cooperatively considering a plurality of sequential frames may allow first neural network 315 to more accurately assess the context of a user's hand movements; in other words, first neural network 315 may infer whether a use is likely to interact with the UI via gesture input, unlikely, not at all likely, already interacting, etc; and for example, a likelihood may be output as a real number between 0 and 1, where 0 represents that the user is not at all likely to perform a gesture in the predetermined window, while 1 represents already interacting with the UI or has already initiated performing a gesture and See also at least paragraphs[0030], [0036]-[0041], [0043], [0045]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches a computing system, which includes at least a processor and memory storing instructions executable by the processor, for implementing a first neural network that evaluates input data, such as sensor data selectively from a sensor suite, for a sequence of data frames and outputs an indication of a likelihood of a gesture interaction));
generating, based on obtaining the first image, a respective predicted skeleton corresponding to a respective view from each of the first image sensor and the one or more second image sensors; selecting, based on generating the respective predicted skeletons, an image sensor from among the first image sensor and the one or more second image sensors, the selected image sensor being used for classifying the first image; and determining, based on the selected image sensor, a classification for the first image (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)).
Regarding claim 18, Schwarz teaches the non-transitory computer-readable storage medium of claim 17, wherein the first image corresponds to an object, and wherein the classification corresponds to an identification of the object or a position of the object (FIGS. 1-4B, paragraph[0045] of Schwarz teaches likelihood of gesture interaction 335 may then be fed as an input to second neural network 340, which may be trained to recognize features indicative of whether the user is currently performing one or more of the plurality of subsequent gesture interactions; in some examples, second neural network 340 may be one of a plurality of neural networks, each trained to recognize a different gesture interaction or set of gesture interaction; each of these neural networks may be provided with the likelihood of gesture interaction 335; the gesture recognition parameters 345 of second neural network 340 are then adjusted based on likelihood of gesture interaction 335; the nodes of second neural network may be associated with adjustable parameters that when changed, alter the likelihoods of certain outputs of second neural network 340; gesture recognition parameters 345 may include node coefficients, connection weights, gradients, etc; and as such, different output data may be produced based on the values of the adjustable parameters even though the same input data is being evaluated by second neural network 340, and See also at least ABSTRACT, paragraphs[0036]-[0041], [0043]-[0044], [0046]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map)).
Regarding claim 19, Schwarz teaches the non-transitory computer-readable storage medium of claim 18, wherein the object is a hand, and wherein the classification corresponds to a hand gesture (FIGS. 1-4B, paragraph[0050] of Schwarz teaches in some implementations, a virtual skeleton or other data structure for tracking feature positions (e.g., joints) may be fit to the pixels of depth and/or color video that correspond to the user; FIG. 4A shows an example virtual skeleton 400; the virtual skeleton includes a plurality of skeletal segments 405 pivotally coupled at a plurality of joints 410; in some embodiments, a body-part designation may be assigned to each skeletal segment and/or each joint; in FIG. 4A, the body-part designation of each skeletal segment 405 is represented by an appended letter: A for the head, B for the clavicle, C for the upper arm, D for the forearm, E for the hand, F for the torso, G for the pelvis, H for the thigh, J for the lower leg, and K for the foot; likewise, a body-part designation of each joint 410 is represented by an appended letter: A for the neck, B for the shoulder, C for the elbow, D for the wrist, E for the lower back, F for the hip, G for the knee, and H for the ankle; naturally, the arrangement of skeletal segments and joints shown in FIG. 4A is in no way limiting; and a virtual skeleton consistent with this disclosure may include virtually any type and number of skeletal segments, joints, and/or other features, and See also at least ABSTRACT, paragraphs[0035]-[0041], [0043]-[0049], [0051]-[0055], and [0088]-[0105] of Schwarz (i.e., Schwarz teaches the computing system, which includes the processor and memory, for implementing the first neural network that evaluates input data, such as sensor data selectively from the sensor suite, for the sequence of data frames and outputs the indication of the likelihood of the gesture interaction that is then fed to a second neural network that is trained to recognize different gesture interaction or a set of gesture interaction, wherein the neural networks are even capable of being utilized as components of a gesture-recognition machine to analyze pixels of a depth map (i.e., an image) that correspond to a user, in order to determine what part of the user’s body, for example a hand, each pixel corresponds to and classify each pixel wherein a virtual skeleton for tracking features may be fit to the pixels of the depth map, and wherein gestures including hand gestures of the user are capable of being determined based on analyzing positional change in various skeletal joints of the virtual skeleton)).
Potentially Allowable Subject Matter
Claims 4-7, 12-15, and 20 would be allowable if rewritten to overcome applicable objection(s) indicated above, and if rewritten in independent form including all of the limitations of the base claim and any intervening, because for each of the 4-7, 12-15, and 20 the prior art references of record do not teach the combination of all element limitations as presently claimed.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ABDUL-SAMAD A ADEDIRAN whose telephone number is (571)272-3128. The examiner can normally be reached on Monday through Thursday, 8:00 am to 5:00 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amr Awad can be reached on 571-272-7764. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ABDUL-SAMAD A ADEDIRAN/Primary Examiner, Art Unit 2621