DETAILED ACTION
This action is responsive to Applicant’s reply filed July 17, 2026. This action is made final.
Status of the Claims
Claims 1-2, 4-5, 10-11, 13-14 and 19-20 are amended.
Claim status is currently pending and under examination for claims 1-5, 7-14 and 16-20 of which independent claims are 1, 10 and 19.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Applicant’s arguments regarding the art rejections are moot in view of the new grounds of rejection necessitated by applicant’s amendment.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Nere (US 20160335534 A1)
Deisher (US 20180121796 A1)
Mann (US 8438122 B1)
Kassner (US 20210041701 A1)
T. Du, X. Ren and H. Li, "Gesture recognition method based on deep learning," 2018 33rd Youth Academic Annual Conference of Chinese Association of Automation (YAC), Nanjing, China, 2018, pp. 782-787, doi: 10.1109/YAC.2018.8406477.
Claims 1-5, 7, 9-14, 16, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Nere / Deisher / Mann / Kassner / Du.
With respect to claim 1, Nere teaches:
A method for generating a neural network model for a wearable device, the method comprising ([Abstract] “Systems and methods for a sensor hub system that accurately and efficiently performs sensory analysis across a broad range of users and sensors and is capable of recognizing a broad set of sensor-based events of interest using flexible and modifiable neural networks are disclosed”
[0003] “For other devices, such as IoT devices and wearable devices, where space and battery life are even more limited, the sensor hub processor may be the only processing hardware on the device.”):
generating a trained model using gesture training data acquired from the wearable device ([0023] “the neural-network algorithms used in this embodiment can be trained to accurately detect sensory events of interest across all types of sensors, including, but not limited to, accelerometers, gyroscopes, … the NSHS can be configured to detect a motion-based gesture using the gyroscope”
[Abstract] “The output of the one or more sensors is converted into a spike signal, and the neural network takes the spike signal as input and determines whether a sensory event of interest has occurred.”
[0026-0027] “FIG. 2 illustrates a typical network architecture of an LSM that can be used in a NSHS according to the present invention. LSMs are typically composed of a number of modeled neurons … a subset of the neurons in the LSM receives a time-varying input from an external source. … Typically, in LSMs, linear output units, or readout units (205) as shown in FIG. 2 can be trained to classify the unique spatiotemporal patterns generated by the LSM … the architecture of the LSM may also be trained or adapted to aid in accurate recognition and classification of patterns of interest”
Sensor data is used to detect gestures and is converted into spike signals. A LSM model receives spike signals used in recognition training, therefore sensor data is ‘gesture training data’.),
wherein the wearable device is an eyewear device having an accelerometer and a gyroscope ([0023] “the neural-network algorithms used in this embodiment can be trained to accurately detect sensory events of interest across all types of sensors, including, but not limited to, accelerometers, gyroscopes”),
the gesture training data is collected from the accelerometer and the gyroscope for a gesture … ([0023] “the neural-network algorithms used in this embodiment can be trained to accurately detect sensory events of interest across all types of sensors, including, but not limited to, accelerometers, gyroscopes, … the NSHS can be configured to detect a motion-based gesture using the gyroscope, a spoken command word using the microphone, or human activity/exercise using the accelerometer”
[0043] “the same neural network algorithm can be used to analyze accelerometer data for detecting motion-based gestures as well as audio data for detecting spoken “hot words”. The most straightforward approach would be that these two tasks are handled by two separate instantiations of the same neural network algorithm, and each instantiation then has its own independent set of neural network parameters, such as trained weights, thresholds, etc. However, it is not a requirement that these different tasks, with their different sensory data, be performed on two separate instantiations”),
extracting a calibrated set of [data] from the trained model (Nere discloses incoming data from sensors may be pre-processed (‘calibrated’) before being encoded to spike signals, “It should also be noted that prior to encoding sensor data into spikes, other pre-processing may be utilized to modify the incoming data. For example, accelerometer data may pass through a low-pass filter to remove high frequency noise before encoding. Other types of data pre-processing may include, but is not limited to, taking the Fast-Fourier Transform (FFT) of the incoming data, calculating standard deviation, mean, minimum and maximum values over a window of data, other filtering schemes, scaling, or calculating the integrals or derivatives of the incoming data. Any of these filtering, scaling, or manipulation techniques may occur before the data is encoded as spikes to be processed by the LSM neural network” [0029].);
generating a trained … code model based on the trained model ([0044] “one clear advantage of the NSHS system is code re-usability. In the example provided, an audio analysis and an accelerometer analysis application both utilize the same neural network, which, in the NSHS, can be the same underlying source code. While the data structures for the two instantiations are separate, the source code is the same. … the clear advantage of the NSHS is that a single algorithm, and single source code, can be used for a broad variety of applications and tasks”);
compiling the trained … code model into a trained machine … code model ([0045] “alternatively, rather than using a single source code with an individual data structure for each neural network instantiation, the neural networks in the NSHS can be “compiled” directly. That is, rather than having multiple data structures, such as an array, which is populated to describe the connectivity of the neural network, the connections, weights, and other parameters, are simply part of the source code. That is, the structure of Neural Network A is directly included in its source code, while the same is true for Neural Network B. The key here is the “compiled” version means larger code blocks, but less RAM utilization, while the previous approach means smaller code blocks but greater RAM utilization”);
quantizing the calibrated set of [data] ([0029] “it is assumed that the input to the NSHS is a single-axis 8-bit digital accelerometer (300) operating at a sampling rate of 20 Hz. Every 50 ms, a signed 8-bit value (301) is produced by the accelerometer. The sign of the incoming value is determined (302) to decide if the data will create spikes in the positive encoder bank or the negative encoder bank. Positive-sign values are 7-bit values between 0 and 127 (303) as the sign of the signal is no longer needed. For negative-sign values, the sign bit is discarded and the magnitude is determined (304) and is also represented with a 7-bit value (305). … It should also be noted that prior to encoding sensor data into spikes, other pre-processing may be utilized to modify the incoming data”. See Figure 3 illustrating how incoming 8-bit signals are converted to a fixed 7-bit signal.);
and loading the trained … machine code model and the quantized calibrated set of [data] onto the wearable device ([0049] “another memory optimization, which Sensor Hub Processor B can use when memory resources are sparse, is efficient “packing” of neural network and the readout/output weights of the corresponding neural network. The readout/output units may be a linear output unit, a perceptron, a multi-layered perceptron, or one of many other classification algorithms used to categorize the current state of the LSM. Typically the readout/output units have weights associated with each element of the LSM”
[0003] “For other devices, such as IoT devices and wearable devices, where space and battery life are even more limited, the sensor hub processor may be the only processing hardware on the device.”).
Nere teaches calibrating and quantizing signal data, however, Nere does not teach calibrating and quantizing a set of weights, which is taught by Deisher:
extracting a calibrated set of weights from the trained model ([0101-0102] “logic blocks have hardware logic elements arranged to alternatively use either 8-bit weights or 16-bit weights providing a NN developer flexibility to use the 16-bit weights for better quality results or alternatively to use the 8-bit weights for more efficiency, e.g. less computational load, faster processing, and so forth. Other or more weight bit sizes could be used as well. The process 600 may include “applying a scale factor to weights provided in at least the first bit length, and to omit the scale factor when weights are provided in at least the second bit length” 606. Thus, to compensate for dynamic range loss, the hardware logic elements are arranged to apply a scale factor to the smaller weights, here the 8-bit weights, so that a much greater range of values for the weights is available (0 to 65,535) while calculating neural network value, which in this case may be a weighted input sum of an affine transform to be provided to an activation function to thereby determine a final output for a node by one example”
[0113] “the process 700 may include “scale entire weight matrix” 716. Here, the entire weight matrix is scaled for both weight sizes to increase the dynamic range of the entire matrix”);
quantizing the calibrated set of weights ([0115] “the process 700 may include “quantize weight matrix” 720, where the weights are quantized by converting the weights from floating point to integer. Then, the process 700 may include “quantize bias vector” 722, and this operation may include determining scaled bias values for the vector and converting the scaled bias values from floating point values to integers as well”. See also pseudocode for quantizing weights on P. 14. See also Figure 7A, Step 716 and Figure 7B.);
Deisher teaches scaling (‘calibrating’) and quantizing neural network weights is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to combine the neural network algorithm of Nere with the technique disclosed by Deisher to prevent bias and reduce model size. By adjusting the weights of a neural network to be on the same scale, the contribution of smaller weights can be balanced and impactful, thereby reducing bias towards larger weights during training. By reducing training bias, a trained neural network can make more accurate predictions and increase model performance. Quantizing weights can reduce the size of a neural network model since integers replace the floating-point weights that take up more storage space. By reducing the size of a neural network model, the model can be trained faster thereby requiring less computational resources.
Furthermore, the combined neural network algorithm of Nere / Deisher does not teach generating and compiling a trained inferencing code model, which is taught by Mann:
generating a trained inferencing code model based on the trained model (Mann discloses “in other implementations, the client computing system can use code (provided by the client computing system or otherwise) that is configured to make a request to the predictive modeling server system 206 to generate a predictive output using the trained model 218. By way of example, the code can be a command line program (e.g., using cURL) or a program written in a compiled language (e.g., C, C++, Java) or an interpreted language (e.g., Python). In some implementations, the trained model can be made accessible to the client computing system by an API through a hosted execution platform, e.g., AppEngine” (Col. 8, line 60 to Col. 9, line 3).);
compiling the trained inferencing code model into a trained machine inferencing code model (Mann discloses “in other implementations, the client computing system can use code (provided by the client computing system or otherwise) that is configured to make a request to the predictive modeling server system 206 to generate a predictive output using the trained model 218. By way of example, the code can be a command line program (e.g., using cURL) or a program written in a compiled language (e.g., C, C++, Java) or an interpreted language (e.g., Python). In some implementations, the trained model can be made accessible to the client computing system by an API through a hosted execution platform, e.g., AppEngine” (Col. 8, line 60 to Col. 9, line 3).);
Mann teaches configuring compiled code to use a trained machine learning model is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined neural network algorithm of Nere / Deisher with the method disclosed by Mann for faster code execution. By using a compiled language, low-level instructions can more directly interact and be understood by a processor, thereby decreasing code execution time and increasing performance and efficiency.
Furthermore, the combination of Nere/ Deisher/ Mann does not teach collecting and labeling training data from an eyewear device for gestures corresponding to a right and left side of an eyewear device, which is taught by Kassner:
wherein the wearable device is an eyewear device having an accelerometer ([Abstract] “A head-wearable spectacles device for determining one or more gaze-related parameters of a user is disclosed. In one example, the device includes a spectacles body having a middle plane and configured for being wearable on a user's head”
[0303] “The head wearable device 620 may also include components that allow determining the device orientation in 3D space, accelerometers, GPS functionality and the like”),
the gesture training data is collected from the accelerometer … for a gesture on a right side of the eyewear device along with an indication of the right side and for the gesture on a left side of the eyewear device along with an indication of the left side of the eyewear device ([0087] “a user wearing a head-wearable device as shown in FIG. 1A to FIG. 1C can be asked to look at a particular marker point or object in space, the coordinates of which within the video images recorded by a scene camera connected to the device can be precisely determined. The image or images recorded by one or more optical sensors (cameras) facing the eye(s) of the person then represent the input data which encodes the information about the person's gaze direction, while said coordinates represent the ground truth. Having the person look at markers in many different directions and distances thus produces ground truth for all possible gaze directions. Collecting large amounts of labelled data, also called training data, thus forms the basis for training a learning algorithm”
Kassner discloses “The head-wearable device includes a first camera and a second camera. … When the first user is expected to respond to the first stimulus or expected to have responded to the first stimulus, the first camera of the head-wearable device is used to generate a first left image of at least a portion of the left eye of the first user, and a second camera of the head-wearable device is used to generate a first right image of at least a portion of the right eye of the first user. A data connection is established between the head-wearable device and the database. A first dataset including the first left image, the first right image and a first representation of a gaze-related parameter is generated” [0010].
Kassner discloses a left image, right image, and gaze-related parameters make up a labelled dataset (training data), “a dataset which includes a left image of at least a portion of the left eye, a right image of at least a portion of the right eye and a corresponding actual or ground truth value of one or more a gaze-related parameters such as the gaze point or gaze direction is also referred to as labelled dataset” [0125].
Kassner discloses a gaze-related parameter is a pair of 3D gaze directions, “The term “gaze-related parameter” as used within this specification intends to describe … a pair of 3D gaze directions (left and right eye)” [0121].
Kassner discloses “in response to a respective stimulus, e.g. a visual stimulus and/or an acoustical stimulus, the respective user is caused to gaze/gazes at a given respective object defining a respective given gaze direction relative to a co-ordinate system fixed with the respective head-wearable device and/or a respective given gaze point in the co-ordinate system” [0129].
An accelerometer can be used to determine an orientation of a head-wearable device (spectacles) in 3D space. A gaze direction is relative to a coordinate system fixed with the head-wearable device, therefore a gaze direction relies on head/device orientation (and therefore a 3D gaze direction uses an accelerometer for device orientation).
A first camera on the head-wearable device is used to take an image of a user’s left eye. A second camera on the head-wearable device is used to take an image of a user’s right eye. When a user looks towards a stimulus, the cameras capture a left and right image, respectively, to represent the user’s gaze (gesture) and gaze direction for each eye. Gaze-related parameters (a pair of 3D gaze directions) correspond to the captured images.
A pair of 3D gaze directions represents one 3D gaze direction for the left eye and one 3D gaze direction for the right eye (and therefore this pair can be used to distinguish the gazes (gestures) from each side of the head-wearable device (since each eye corresponds to a left or right side of the spectacles)). Therefore, the 3D gaze direction for the left eye in the pair represents training data collected from the accelerometer for a gesture (gaze) on a left side of the eyewear device along with an indication of the left side of the eyewear device, and the 3D gaze direction for the right eye in the pair represents training data collected from the accelerometer for a gesture on a right side of the eyewear device along with an indication of the right side).)
and wherein generating the trained model includes creating labels for each element of the gesture training data (Kassner discloses left and right images represent input data and gaze directions (coordinates) represent the ground truth, and the inputs and ground truths form training data used for training a learning algorithm, see [0087] above.
Gaze-related parameters are ground-truth values (labels), see [0125] above.),
the labels comprising a first label corresponding to the gesture on the left side of the eyewear device and a second label corresponding to the gesture on the right side of the eyewear device (Gaze-related parameters are used as ground-truth values (labels) in training data. A pair of 3D gaze directions is a gaze-related parameter (therefore pairs of 3D gaze directions are labels). A pair of 3D gaze directions represents one 3D gaze direction for the left eye and one 3D gaze direction for the right eye (and therefore this pair can be used to distinguish the gazes (gestures) from each side of the head-wearable device). Therefore, in a pair of 3D gaze directions, the 3D gaze direction for the left eye is a first label corresponding to the gesture on the left side of the eyewear device, and the 3D gaze direction for the right eye is a second label corresponding to the gesture on the right side of the eyewear device.);
Kassner teaches generating training data from spectacles and determining ground truths corresponding to each side of the spectacles is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined neural network algorithm of Nere / Deisher/ Mann with the training data creation technique disclosed by Kassner to train a model with labeled data corresponding to each side of an eyewear device. Training a model with labeled data corresponding to each side of an eyewear device enhances task accuracy by allowing the model to learn and compare data from both sides, thereby providing extra context from the relationships between features.
Furthermore, the combination of Nere/ Deisher/ Mann / Kassner does not teach gesture training data includes acceleration measurements, acceleration time coordinates, rotation measurements, rotation time coordinates, and motion interrupt time coordinates, which is taught by Du:
the gesture training data includes acceleration measurements and acceleration time coordinates, rotation measurements and rotation time coordinates, and motion interrupt time coordinates (The Examiner interprets “motion interrupt time coordinates” according to its broadest reasonable interpretation (BRI) in view of the Applicant’s specification as encompassing start points. This interpretation is consistent with the illustrative descriptions in the Applicant’s specification at [0087], (see excerpt below).
Applicant’s written description at [0087] “Wear training data 351A-N also includes motion interrupt time coordinates 384A-N (e.g., times when motion is detected).”
(P. 785, Sec. IV, ¶1) “16 experimenters (8 males and 8 females) were selected to meet the requirements of the gesture definition, and each gesture was operated 10 times in the manner and intensity of their respective habits. Each of our actions received 160 sets of data, and a total of 1,600 sets of data samples were obtained. We randomly selected 1200 sets of data as training samples”
(P. 784, Sec. C, ¶2) “There are two major changes to the arm when people make gestures. One is the movement of the arm and the other is the rotation of the arm. The acceleration and angular velocity data collected by the sensor can effectively reflect the movement of the arm. However, the attitude of the movement in the world coordinate system X, Y, Z axis rotation angle can be very good identification of arm rotation. Therefore, the data information we collect includes acceleration data, angular velocity data and angle data. we collect 9-dimensional data to identify actions.”
“The first step in gesture recognition is to accurately detect the start and end points of the gesture so that an effective gesture signal segment can be intercepted in real time [15]. … When there is no gesture, those two signals are relatively stable; when the gesture action starts, those two signals change intensely. The difference between the signals shows how violently the signal changes. Therefore, it is used to detect the start and end of a gesture online in real time. The calculation method is shown in Eqs. (1), (2), (3) and (4), where
△
α
k
is the acceleration difference at the k point”
(P. 784, Sec. III, ¶1) “Sensor data is a time series, so this article selects the RNN, LSTM and GRU models which have great success in the timing problem as the gesture recognition model.”
Rotation angle is represented in a X, Y, Z coordinate system. Rotation angles are detected from a sensor and sensor data is a time series, therefore a rotation angle is tracked over time and represents ‘rotation time coordinates’ and ‘rotation measurements’ (since rotation data is collected from a sensor). Acceleration data, angular velocity, and angle data are 9-dimensional and rotation angles are in a X, Y, Z coordinate system, therefore acceleration must be represented as coordinates. Acceleration is collected by a sensor and sensor data is time series data, therefore acceleration data is tracked over time and represents ‘acceleration time coordinates’ and ‘acceleration measurements’ (since acceleration is collected from a sensor). The difference between acceleration signals can be used to detect when a gesture starts (start point) and ends (end point), therefore start points in a time-series are times when motion (gestures) is detected (and therefore start points are ‘motion interrupt time coordinates’).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined neural network algorithm of Nere / Deisher/ Mann / Kassner with the training technique disclosed by Du to train a model to detect gestures by using acceleration, rotation, and gesture start point data. By training a model to detect gestures by using acceleration, rotation, and gesture start point data, a model can be taught how to use acceleration and rotation data to determine when a gesture begins and ends, thereby enabling real-time gesture recognition and improving gesture recognition accuracy.
With respect to claims 2, 11 and 20, the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du teaches:
the method of claim 1, wherein the trained model is a trained scripting language model (Mann discloses “Components of the client computing system 202 and/or the predictive modeling system 206, e.g., the model training module 212, model selection module 210 and trained model 218, can be realized by instructions that upon execution cause one or more computers to carry out the operations described above. Such instructions can comprise, for example, interpreted instructions, such as script instructions, e.g., JavaScript or ECMAScript instructions, or executable code, or other instructions stored in a computer readable medium” (Col. 9, lines 15-28).)
and wherein generating the trained model comprises: generating the trained scripting language model in an interpreted programming language based on the acquired gesture training data (Nere discloses sensor data is used to detect gestures (see [0023]), sensor data is converted into spike signals (see [Abstract]), and a LSM model receives spike signals used in recognition training (see [0026-0027]), therefore sensor data is ‘gesture training data’.
Mann discloses “the client computing system 202 uploads training data to the predictive modeling server system 206 over the network 204 (Step 302) … the client computing system can send input data and a prediction request to the trained model (Step 306). In response, the client computing system receives a predictive output generated by the trained model from the input data” (Col. 5, lines 5-21).).
Mann teaches training a machine learning model implemented with interpreted programming instructions is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du with the method disclosed by Mann to decrease development time. The high-level instructions of interpreted languages decrease development time due to their simple syntax and access to extensive pre-built libraries. By decreasing development time, more time can be spent retraining or optimizing the trained machine learning model.
With respect to claims 3 and 12, the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du teaches:
the method of claim 2, wherein the inferencing comprises: generating C language code based on the trained scripting language model (Mann discloses “in other implementations, the client computing system can use code (provided by the client computing system or otherwise) that is configured to make a request to the predictive modeling server system 206 to generate a predictive output using the trained model 218. By way of example, the code can be a command line program (e.g., using cURL) or a program written in a compiled language (e.g., C, C++, Java) or an interpreted language (e.g., Python). In some implementations, the trained model can be made accessible to the client computing system by an API through a hosted execution platform, e.g., AppEngine” (Col. 8, line 60 to Col. 9, line 3).).
With respect to claims 4 and 13, the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du teaches:
the method of claim 2, wherein the acquired gesture training data includes at least one of: tracked movement of the wearable device over time intervals occurring while the wearable device is being carried and while the wearable device is not being carried; or tracked movement of the wearable device over time intervals occurring while the wearable device is being worn and while the wearable device is not being worn (Nere discloses, “Even across different motion-based gestures, all using accelerometer data, different algorithms and software approaches may be used. For example, the circle gesture described above may last for 1-2 seconds, and can be implemented with an algorithm that samples the accelerometer at a fairly slow rate (e.g. 10 Hz). However, another gesture, “double tapping” the side of the device, is very quick—much less than 1 second in duration. The accelerometer signatures of the “taps” are so brief that they require a much faster sampling rate (e.g. 100 Hz or more)” [0009].
Nere further discloses “the sensors themselves are capable of generating a large amount of data, especially when they are kept “always on” and/or utilizing a high sampling rate.” [0005].).
With respect to claims 5 and 14, the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du teaches:
the method of claim 2, wherein the acquired gesture training data includes tracked movement of the wearable device over time intervals in response to known and unclassified gestures or activities occurring while the wearable device is being worn and while the wearable device is not being worn (Nere discloses “an accelerometer sensor (401) propagates sensory data via an I2C bus (402) to one neural network instantiation (403), which performs gesture recognition on the accelerometer. Outputs (404) may then be communicated to a user, an application processor, a data log, or some other device or component. The outputs (404) may be a classification (e.g. Gesture A just happened, or Gesture B just happened), a confidence level (e.g. Gesture A happened with 70% certainty), or multiple simultaneous classifications and/or confidence levels (e.g. Gesture A happened with 70% confidence and Gesture B happened with 90% confidence)” [0038].
Nere further discloses “Different Neural Networks may use different sampling rates of the same sensor. For Example, Neural Network 1 may sample accelerometer data at 100 Hz to detect taps or double taps for example, while Neural Network 2 may sample data at 10 Hz to detect slow gestures like lifting the phone to the ear” [0041].
Nere further discloses “the sensors themselves are capable of generating a large amount of data, especially when they are kept “always on” and/or utilizing a high sampling rate” [0005].).
With respect to claims 7 and 16, the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du teaches:
the method of claim 1, wherein the calibrated set of weights are extracted in a floating point format (Deisher discloses, “logic blocks have hardware logic elements arranged to alternatively use either 8-bit weights or 16-bit weights providing a NN developer flexibility to use the 16-bit weights for better quality results or alternatively to use the 8-bit weights for more efficiency, e.g. less computational load, faster processing, and so forth. … Thus, to compensate for dynamic range loss, the hardware logic elements are arranged to apply a scale factor to the smaller weights, here the 8-bit weights, so that a much greater range of values for the weights is available (0 to 65,535) while calculating neural network value, which in this case may be a weighted input sum of an affine transform to be provided to an activation function to thereby determine a final output for a node by one example” [0101-0102].)
and wherein the quantizing comprises: converting the calibrated set of weights from the floating point format to a fixed number of bits format (Deisher discloses “the process 700 may include “quantize weight matrix” 720, where the weights are quantized by converting the weights from floating point to integer. Then, the process 700 may include “quantize bias vector” 722, and this operation may include determining scaled bias values for the vector and converting the scaled bias values from floating point values to integers as well” [0115].
See also pseudocode for quantizing weights on P. 14. See also Figure 7A, Step 716 and Figure 7B.).
With respect to claims 9 and 18, the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du teaches:
the method of claim 2, wherein the wearable device comprise flash memory and random access memory (RAM) (Nere discloses “smartwatches contain an ever-growing number of sensors … To enable a more continuous, or “always on”, sensory processing capability, many device manufacturers have opted to include a dedicated coprocessor, or a sensor hub processor, in their designs … Typical sensor hub systems include between 8 KB and 128 KB of RAM, and between 32 KB and 512 KB of flash memory, with peak operating frequencies of 100 MHz or less” [0011-0012].)
and wherein the loading comprises: loading the trained inferencing machine code model and the quantized calibrated set of weights into at least one of the flash memory or RAM of the wearable device (Nere discloses “the neural networks in the NSHS can be “compiled” directly. That is, rather than having multiple data structures, such as an array, which is populated to describe the connectivity of the neural network, the connections, weights, and other parameters, are simply part of the source code. That is, the structure of Neural Network A is directly included in its source code, while the same is true for Neural Network B. The key here is the “compiled” version means larger code blocks, but less RAM utilization, while the previous approach means smaller code blocks but greater RAM utilization. At the same time, the “compiled” version may execute faster than the non-compiled approach because it need not access and interpret data structures to obtain neural network connections, weights, and other parameters” [0045].).
With respect to claim 10, the rejection of claim 1 is incorporated. The difference in scope being
A wearable device comprising a neural network model, the neural network model generated by (Nere discloses “smartwatches contain an ever-growing number of sensors … To enable a more continuous, or “always on”, sensory processing capability, many device manufacturers have opted to include a dedicated coprocessor, or a sensor hub processor, in their designs” [0011-0012].
Nere discloses “another memory optimization, which Sensor Hub Processor B can use when memory resources are sparse, is efficient “packing” of neural network and the readout/output weights of the corresponding neural network” [0049].):
wherein the neural network model comprises the trained inferencing machine code model and the quantized calibrated set of weights (Nere discloses “the neural networks in the NSHS can be “compiled” directly. That is, rather than having multiple data structures, such as an array, which is populated to describe the connectivity of the neural network, the connections, weights, and other parameters, are simply part of the source code. That is, the structure of Neural Network A is directly included in its source code, while the same is true for Neural Network B. The key here is the “compiled” version means larger code blocks, but less RAM utilization, while the previous approach means smaller code blocks but greater RAM utilization” [0045]
Nere further discloses “another memory optimization, which Sensor Hub Processor B can use when memory resources are sparse, is efficient “packing” of neural network and the readout/output weights of the corresponding neural network. The readout/output units may be a linear output unit, a perceptron, a multi-layered perceptron, or one of many other classification algorithms used to categorize the current state of the LSM. Typically the readout/output units have weights associated with each element of the LSM” [0049].).
With respect to claim 19, the rejection of claim 10 is incorporated. The difference in scope being:
A non-transitory computer readable media storing instructions … ([0045] “the neural networks in the NSHS can be “compiled” directly. … The key here is the “compiled” version means larger code blocks, but less RAM utilization”), the instructions, when executed by a processor perform functions comprising functions to ([0038] “the software is executed on a sensor hub microprocessor”).
The following are the references relied upon in the rejections below:
Craddock (US 20170161604 A1)
Claims 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Nere / Deisher / Mann / Kassner / Du / Craddock.
With respect to claims 8 and 17, the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du teaches the method of claim 2, however the combination does not teach statically allocated memory, which Craddock does:
wherein the wearable device comprises statically allocated memory (Craddock discloses “another example aspect of the present disclosure is directed to a computing system to transform dynamically allocated execution of a neural network into statically allocated execution” [0007].
Craddock further discloses “at (310), method (300) can include storing the order of execution and the memory allocation for a future execution of the convolutional neural network. At (312), method (300) can include providing data indicative of the order of execution and/or memory allocation to a remote computing device configured to execute the convolutional neural network. For instance, the remote computing device may be a user device, such as a smartphone, tablet, laptop computer, desktop computer, wearable computing device, etc. Upon receiving the data indicative of the order of execution and/or memory allocation, the remote computing device can execute the convolutional neural network in accordance with the order of execution and memory allocation” [0068].)
and wherein the loading comprises: loading the trained inferencing machine code model and the quantized calibrated set of weights into the statically allocated memory of the wearable device (Craddock discloses “determine a static memory allocation for a computing task represented by a graph of operators. For instance, the graph of operators may be associated with various suitable types of neural networks, such as for instance, convolutional neural networks, long short-term memory neural networks, etc. In particular, the systems and methods of the present disclosure can determine an execution order for a neural network that satisfies various memory constraints associated with execution of the neural network within a constrained memory space” [0019].
Craddock further discloses “once determined, data indicative of the order of execution of the convolutional neural network and/or the memory allocation associated with the execution of the convolutional neural network can be stored for use in a future execution of the convolutional neural network. In one example, such data can then be provided to a remote computing device configured to execute the convolutional neural network. For instance, the remote computing device can be a wearable image capture device configured to execute the convolutional neural network.).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined neural network algorithm of Nere / Deisher/ Mann / Kassner / Du with the method disclosed by Craddock to predict memory usage. By determining and using an optimal execution order for a neural network, it can be predicted if memory usage satisfies constrained memory requirements in future executions. By predicting memory usage, insufficient memory errors can be avoided and consistent performance across devices can be ensured.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Huang et al. (CN 102982315 A) teaches capturing gesture actions using a camera and sensors to collect three-axis angular speed and three-axis acceleration data. Rotation angle and acceleration are characteristics extracted from the collected data and converted into feature vectors used to train a gesture action identification model.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PEDRO J MORALES/Examiner, Art Unit 2124 /MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124