DETAILED ACTION
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1, 2, 13, 14-15, and 20 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by JP2017224161A.
Regarding claim 1, JP2017224161A discloses a biometric payment processing method, performed by an electronic device (pg. 4 This control unit 6 has a movement speed threshold and a movement amount threshold 63, and determines that a pay gesture has been made if the amount of movement (determined movement amount 57) while the movement speed 50 is equal to or greater than the movement speed threshold is equal to or greater than the movement), the method comprising:
obtaining image data, the image data comprising a plurality of images of an organism that are successively acquired (pg. 4 As shown in Figures 1(a) to 2(b), the gesture determination device 1 is generally configured to include: an imaging unit 3 that images the hand 9 and generates a plurality of images 35; pg. 5 The imaging unit 3 periodically images the imaging area 30 and generates image information S1 as information of the captured images, which is output to the control unit 6. This image information S1 includes, for example, position information in a three-dimensional coordinate system. The imaging unit 3 captures, for example, 30 images per second);
detecting a target part in an image in the image data (pg. 5 ¶22 The movement amount calculation unit 4 is configured to calculate the movement amount 40 of the feature points 90 of the hand 9 based on the image information S1 acquired from the imaging unit 3. These feature points 90 are, for example, the centroid coordinates or fingertip coordinates of the hand 9, as shown in Figure 2(a). In this embodiment, the feature points 90 are the centroid coordinates; ¶23 Figure 2(b) shows the movement of feature points 90 for each image (one frame). The movement amount calculation unit 4 calculates the movement amount 40 as the length connecting, for example, the feature point 90 of one frame and the feature point 90 of the next frame), the target part being a part in the organism to which a biometric payment function is bound (pg. 8 ¶59 If the control unit 6 confirms that the feature point 90 has been stationary for at least time T1 or longer based on the amount of movement information S2 and the speed of movement information S3 , it determines that a trigger for starting the determination has been obtained and starts determining whether to make a pay gesture);
determining, in response to that the target part is detected from the plurality of images in the image data, a movement speed corresponding to the target part in the plurality of images (pg. 5 ¶24-25 movement amount calculation unit 4 outputs the calculated movement amount 40 of the feature point 90 and the movement amount information of each component as movement amount information S2 for each frame); and
performing a payment operation based on the target part in response to the movement speed being less than a speed threshold (pg. 8 ¶59 If the control unit 6 confirms that the feature point 90 has been stationary for at least time T1 or longer (e.g. less than a speed threshold) based on the amount of movement information S2 and the speed of movement information S3 , it determines that a trigger for starting the determination has been obtained and starts determining whether to make a pay gesture (Step 2: Yes).
Regarding claim 2, JP2017224161A discloses the method according to claim 1, wherein the determining a movement speed corresponding to the target part in the plurality of images comprises:
performing key point detection processing on the target part to obtain a plurality of key points comprised in the target part (pg. 5 ¶22 The movement amount calculation unit 4 is configured to calculate the movement amount 40 of the feature points 90 of the hand 9 based on the image information S1 acquired from the imaging unit 3. These feature points 90 are, for example, the centroid coordinates or fingertip coordinates of the hand 9, as shown in Figure 2(a). In this embodiment, the feature points 90 are the centroid coordinates; ¶23 Figure 2(b) shows the movement of feature points 90 for each image (one frame). The movement amount calculation unit 4 calculates the movement amount 40 as the length connecting, for example, the feature point 90 of one frame and the feature point 90 of the next frame);
determining a movement speed corresponding to each key point in the plurality of images (pg. 5 ¶24-25 movement amount calculation unit 4 outputs the calculated movement amount 40 of the feature point 90 and the movement amount information of each component as movement amount information S2 for each frame); and
determining, based on movement speeds respectively corresponding to the plurality of key points, the movement speed corresponding to the target part in the plurality of images (pg. 8 ¶59 If the control unit 6 confirms that the feature point 90 has been stationary for at least time T1 or longer based on the amount of movement information S2 and the speed of movement information S3 , it determines that a trigger for starting the determination has been obtained and starts determining whether to make a pay gesture).
Regarding claim 13, JP2017224161A discloses the method according to claim 1, wherein a type of the target part comprises: a palm (abstract gesture determination device (1) includes an imaging unit (3) that captures a hand and generates a plurality of images), a finger, a wrist, or a face.
Regarding claim(s) 14-15 (drawn to an apparatus):
The rejection/proposed combination of JP2017224161A, explained in the rejection of method claim(s) 1-2, anticipates/renders obvious the steps of the apparatus of claim(s) 14-15 because these steps occur in the operation of the proposed combination as discussed above. Thus, the arguments similar to that presented above for claim(s) 1-2 is/are equally applicable to claim(s) 14-15.
Regarding claim(s) 20 (drawn to a CRM):
The rejection/proposed combination of JP2017224161A, explained in the rejection of method claim(s) 1, anticipates/renders obvious the steps of the computer readable medium of claim(s) 20 because these steps occur in the operation of the proposed combination as discussed above. Thus, the arguments similar to that presented above for claim(s) 1 is/are equally applicable to claim(s) 20.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 3 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over JP2017224161A as applied to claim 2 and 16 above, and further in view of CN115170870A.
Regarding claim 3, JP2017224161A discloses the method according to claim 2, but fails to teach where CN115170870A teaches wherein the determining a movement speed corresponding to each key point in the plurality of images comprises: performing the following processing for the each key point (pg. 6 Step 3.1: Preprocessing of human key points. The coordinates of each frame of the human body key points in the video form the human body key point sequence):
selecting a first image and a second image from the plurality of images, and determining a time difference between an acquisition time of the first image and an acquisition time of the second image (pg. 6 Step 3.2: Calculate the features of human key points. The number of frames per second (FPS) of the video is stored in the video file, and the inverse of the FPS is the time difference between two frames of images.);
determining a distance between first coordinates and second coordinates, the first coordinates being coordinates of the key point in the first image, and the second coordinates being coordinates of the key point in the second image (pg. 6 Step 3.2: determining the movement distance of the human key points between the two frames of images); and
determining a result of dividing the distance by the time difference as the movement speed corresponding to the key point in the plurality of images (pg. 6 Step 3.2: The movement distance of the human key points between the two frames of images is divided by the time difference to form the speed of the key points. Therefore, for each key point of the human body, four features can be calculated, which are the abscissa, ordinate, moving distance, and speed of the key point.).
Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein the determining a movement speed corresponding to each key point in the plurality of images comprises: performing the following processing for the each key point, selecting a first image and a second image from the plurality of images, and determining a time difference between an acquisition time of the first image and an acquisition time of the second image, determining a distance between first coordinates and second coordinates, the first coordinates being coordinates of the key point in the first image, and the second coordinates being coordinates of the key point in the second image, and determining a result of dividing the distance by the time difference as the movement speed corresponding to the key point in the plurality of images from CN115170870A into the method as disclosed by JP2017224161A. The motivation for doing this is to improve methods and systems for classifying features based on deep learning.
Regarding claim(s) 16 (drawn to an apparatus):
The rejection/proposed combination of JP2017224161A and CN115170870A, explained in the rejection of method claim(s) 3, anticipates/renders obvious the steps of the apparatus of claim(s) 16 because these steps occur in the operation of the proposed combination as discussed above. Thus, the arguments similar to that presented above for claim(s) 3 is/are equally applicable to claim(s) 16.
Claim(s) 5 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over JP2017224161A as applied to claim 2 above, and further in view of Huelsdunk et al (US 20210192783).
Regarding claim 5, JP2017224161A discloses the method according to claim 2, but fails to teach where Huelsdunk teaches wherein the determining, based on movement speeds respectively corresponding to the plurality of key points (¶309 movement speeds of different body joints may be computed e.g. to “diagnose” an athlete's kicking speed or punching speed, based on the absolute 3D location of her or his wrists and ankles as computed using the methods herein), the movement speed corresponding to the target part in the plurality of images comprises: determining an average movement speed of the plurality of movement speeds in a one-to-one correspondence with the plurality of key points (¶311 A person's speed may be computed by computing the average speed of all or some of her or his joints); and determining the average movement speed as the movement speed corresponding to the target part in the plurality of images (¶310 Joint speed may be computed conventionally, by computing distances between a joint's (typically absolute) 3D locations over time and dividing by the time period separating sequential locations; ¶311 A person's speed may be computed by computing the average speed of all or some of her or his joints).
Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein the determining, based on movement speeds respectively corresponding to the plurality of key points, the movement speed corresponding to the target part in the plurality of images comprises: determining an average movement speed of the plurality of movement speeds in a one-to-one correspondence with the plurality of key points, and determining the average movement speed as the movement speed corresponding to the target part in the plurality of images from Huelsdunk into the method as disclosed by JP2017224161A. The motivation for doing this is to improve estimating an absolute 3D location of at least one object x imaged by a single camera.
Regarding claim(s) 18 (drawn to an apparatus):
The rejection/proposed combination of JP2017224161A and Huelsdunk, explained in the rejection of method claim(s) 5, anticipates/renders obvious the steps of the apparatus of claim(s) 18 because these steps occur in the operation of the proposed combination as discussed above. Thus, the arguments similar to that presented above for claim(s) 5 is/are equally applicable to claim(s) 18.
Claim(s) 6 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over JP2017224161A as applied to claim 2 and 5 above, and further in view of CN112132099A.
Regarding claim 6, JP2017224161A discloses the method according to claim 2, but fails to teach where CN112132099A teaches wherein the performing key point detection processing on the target part to obtain a plurality of key points comprised in the target part comprises: calling a key point detection model to detect the plurality of key points comprised in the target part, the key point detection model being obtained through training based on a sample part of a sample organism and key points annotated for the sample part (pg. 6 S204, the sample palmprint image is input into the palmprint key point detection model to be trained, and the palmprint key point detection model uses the feature extraction layer to perform feature extraction on the sample palmprint image to obtain training palmprint features, and based on the training palmprint features Determine the predicted palmprint key point positions corresponding to each palmprint key point respectively; pg. 7 The location of the standard palmprint key points can be determined by manual annotation, or detected by a trained neural network model; pg. 8 S212: Adjust model parameters in the palmprint key point detection model based on the target loss value to obtain a trained palmprint key point detection model.).
Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein the performing key point detection processing on the target part to obtain a plurality of key points comprised in the target part comprises: calling a key point detection model to detect the plurality of key points comprised in the target part, the key point detection model being obtained through training based on a sample part of a sample organism and key points annotated for the sample part from CN112132099A into the method as disclosed by JP2017224161A. The motivation for doing this is to reliably determine the positions of the tracked hand/palm keypoints from successive images.
Regarding claim(s) 19 (drawn to an apparatus):
The rejection/proposed combination of JP2017224161A and CN112132099A, explained in the rejection of method claim(s) 6, anticipates/renders obvious the steps of the apparatus of claim(s) 19 because these steps occur in the operation of the proposed combination as discussed above. Thus, the arguments similar to that presented above for claim(s) 6 is/are equally applicable to claim(s) 19.
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over JP2017224161A and CN112132099 as applied to claim 6 above, and further in view of WO-2022027768-A1.
Regarding claim 7, the combination of JP2017224161A and CN112132099 discloses the method according to claim 6, but fails to teach where WO-2022027768-A1 teaches cropping the image to obtain a region image of a region in which the target part is located (pg. 2 extract multiple partial images corresponding to the multiple target detection frames, the multiple target detection frames correspond to the same human body area; pg. 4 the above-mentioned human body region may specifically include human body regions with several key points, such as a hand region and a face region in the overall human body region,), and zooming in the region image (pg. 4 the resolution of the video frame is correspondingly enlarged); and the calling a key point detection model comprises: calling the key point detection model to perform key point detection processing on the zoomed-in region image (pg. 4 S102: Perform key point feature detection of the human body region on the multiple video frames respectively, so as to obtain multiple key point features corresponding to the multiple video frames respectively.).
Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of cropping the image to obtain a region image of a region in which the target part is located, and zooming in the region image, and the calling a key point detection model comprises: calling the key point detection model to perform key point detection processing on the zoomed-in region image from WO-2022027768-A1 into the method as disclosed by the combination of JP2017224161A and CN112132099. The motivation for doing this is to improve methods for dynamic gesture recognition.
Claim(s) 11-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over JP2017224161A as applied to claim 1 above, and further in view of Redmon et al (NPL You Only Look Once: Unified, Real-Time Object Detection).
Regarding claim 11, JP2017224161A discloses the method according to claim 1, but fails to teach where Redmon teaches wherein the detecting a target part in an image in the image data comprises: calling an object detection model to detect the target part in the image in the image data (abstract A single neural network predicts bounding boxes and class probabilities directly from full images in one evaluation), the object detection model being obtained through training based on a sample image and a sample part annotated for the sample image (pg. 781-782 2.2. Training: We pretrain our convolutional layers on the ImageNet 1000-class competition dataset [29]. For pretraining we use the first 20 convolutional layers from Figure 3 followed by a average-pooling layer and a fully connected layer. We train this network for approximately a week and achieve a single crop top-5 accuracy of 88% on the ImageNet 2012 validation set, comparable to the GoogLeNet models in Caffe’s Model Zoo [24]; YOLO predicts multiple bounding boxes per grid cell. At training time we only want one bounding box predictor to be responsible for each object. We assign one predictor to be “responsible” for predicting an object based on which prediction has the highest current IOU with the ground truth. This leads to specialization between the bounding box predictors. Each predictor gets better at predicting certain sizes, aspect ratios, or classes of object, improving overall recall.).
Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein the detecting a target part in an image in the image data comprises: calling an object detection model to detect the target part in the image in the image data, the object detection model being obtained through training based on a sample image and a sample part annotated for the sample image from Redmon into the method as disclosed by JP2017224161A. The motivation for doing this is to improve models for object detection.
Regarding claim 12, the combination of JP2017224161A and Redmon disclose the method according to claim 11, wherein the calling an object detection model to detect the target part in the image in the image data comprises:
for the image in the image data, calling the object detection model to perform the following processing: determining a plurality of bounding boxes in the image and a confidence score corresponding to each bounding box, the confidence score being configured for representing a probability that the bounding box comprises the target part (Redmon pg. 780 2. Unified Detection: Each grid cell predicts B bounding boxes and confidence scores for those boxes. These confidence scores reflect how confident the model is that the box contains an object and also how accurate it thinks the box is that it predicts. Formally we define confidence as Pr(Object) ∗ IOUtruth pred. If no object exists in that cell, the confidence scores should be zero. Otherwise we want the confidence score to equal the intersection over union (IOU) between the predicted box and the ground truth);
classifying the each bounding box based on the confidence score according to whether the each bounding box comprises the target part (Redmon pg. 780 2. Unified Detection: Each grid cell also predicts C conditional class probabilities, Pr(Classi|Object). These probabilities are conditioned on the grid cell containing an object); and
performing regression processing on a target bounding box determined to comprise the target part, to obtain a corrected position of the target bounding box (Redmon abstract we frame object detection as a regression problem to spatially separated bounding boxes and associated class probabilities.).
Therefore, it would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to have implemented the teaching of wherein the calling an object detection model to detect the target part in the image in the image data comprises: for the image in the image data, calling the object detection model to perform the following processing: determining a plurality of bounding boxes in the image and a confidence score corresponding to each bounding box, the confidence score being configured for representing a probability that the bounding box comprises the target part, classifying the each bounding box based on the confidence score according to whether the each bounding box comprises the target part, and performing regression processing on a target bounding box determined to comprise the target part, to obtain a corrected position of the target bounding box from Redmon into the method as disclosed by JP2017224161A. The motivation for doing this is to improve models for object detection.
Allowable Subject Matter
Claims 4, 8, 9-10, and 17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding claim 4, and similarly regarding claim 17, the prior art of record, alone or in combination, fails to teach at least “wherein the selecting a first image and a second image from the plurality of images comprises: selecting, from the plurality of images, the first image and the second image each with a quality parameter greater than a quality parameter threshold, the time difference between the acquisition time of the first image and the acquisition time of the second image being greater than a time difference threshold.”
Regarding claim 8, and similarly regarding claim 17, the prior art of record, alone or in combination, fails to teach at least “wherein the key point detection model comprises a plurality of cascaded convolutional layers and a plurality of cascaded fully connected layers; and the calling a key point detection model to detect the plurality of key points comprised in the target part comprises: performing convolution processing on feature information corresponding to the target part through the first convolutional layer in the plurality of cascaded convolutional layers; inputting a convolution result outputted by the first convolutional layer to a subsequent cascaded convolutional layer, and continuing to perform convolution processing through the subsequent cascaded convolutional layer until the last convolutional layer; performing, through the first fully connected layer in the plurality of cascaded fully connected layers, fully connected processing on a convolution result outputted by the last convolutional layer; inputting a fully connected result outputted by the first fully connected layer to a subsequent cascaded fully connected layer, and continuing to perform fully connected processing through the subsequent cascaded fully connected layer until a last fully connected layer; and determining a plurality of points respectively corresponding to a plurality of coordinates outputted by the last fully connected layer in the image as the plurality of key points comprised in the target part.”.
Regarding claim 9, the prior art of record, alone or in combination, fails to teach at least wherein the determining a movement speed corresponding to the target part in the plurality of images comprises: dividing the plurality of images into a plurality of image groups according to a set frame interval; and determining a movement speed corresponding to the target part in each image group; and the performing a payment operation based on the target part in response to that the movement speed is less than a speed threshold comprises: performing the payment operation based on the target part in response to that a plurality of movement speeds respectively corresponding to the target part in the plurality of image groups are less than the speed threshold. Claim 10 depends on claim 9 and would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEVIN KY whose telephone number is (571)272-7648. The examiner can normally be reached Monday-Friday 9-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KEVIN KY/ Primary Examiner, Art Unit 2671