DETAILED ACTION
Notice of AIA Status
The present application is being examined under the AIA the first inventor to file provisions.
Priority
Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file.
Response to Arguments
Applicant’s arguments see remarks, filed 06/23/2026, with respect to the claim 1-18 have been fully considered but are not persuasive.
The applicant argues on page 12, “In contrast to Luo, aspects of presently amended claim 1 include a camera that is being kept fixed, whereas Luo is mainly premised on an intelligent phone (i.e., smartphone) as an image capturing device.”
In response, the office does not find this argument to be persuasive. Based on the breadth of the claim language the prior art by Ng et al. (US 20200211154 A1) explicitly teaches comprising: acquiring an image of the user captured by a camera being kept fixed (Fig. 1 Paragraph [0049]- Ng discloses these multiple standalone embedded fall-detection vision sensors can be installed at multiple fixed locations different from one another, wherein each of the multiple embedded fall-detection vision sensors can include at least one camera for capturing video images and various software and hardware modules for processing the captured video images and generating corresponding fall-detection output including fall alarms/notifications based on the captured video images.);
The applicant argues on page 12, “Secondly, aspects of the presently amended claim 1 are distinguished from Luo in that aspects of presently amended claim 1 disclose extracting, from estimated nodes, a predetermined detectable node which is detectable by the camera. Luo fails to execute such extraction. While the Office Action relies on paragraph [0075] of Luo as disclosing such features, the relied-on paragraph of Luo merely discloses a training phase of the visibility determining model and fails to disclose an inference phase of the visibility determining model. Moreover, the visibility determining model in Luo is a model to output visibility probability of each key point of an input target object, and thus, cannot reasonably be understood as a model that extracts, from a plurality of key points, a predetermined key point as a key point falling within a capturing range of a capturing device.”
In response, the office does not find this argument to be persuasive. Based on the breadth of the claim language the prior art by Luo et al. (US 20210281744 A1) explicitly teaches extracting, from the estimated nodes, a predetermined detectable node which is detectable by the camera (Fig. 4, Paragraph [0075]- Luo discloses the outputs of the model are probabilities each indicating whether a key point is visible. If the probability is greater than a threshold, it is determined that a visibility of the key point indicated by the probability is consistent with the mark of the key point on the image. (wherein the predetermined node is a node above a threshold value)),
the detectable node falling within a capturing range of the camera (Fig. 1, Paragraph [0076]- Luo discloses then the probability is outputted through an activation function, where 1 is outputted for a visible key point, and 0 is outputted for an invisible key point. In this way, a vector of values of visibility attributes is obtained. The vector is a 1*N vector where N represents the number of the key points. A value of an element in the vector is equal to 0 or 1. 0 represents that a corresponding key point is invisible and 1 represents that a corresponding key point is visible (wherein if a node is visible it is determined to be detectable and since it is in the image it is within a capturing range of the camera).);
The applicant argues on page 12, “Luo inherently discloses that its intelligent phone or smartphone operates as the image capturing device, and thus, cannot reasonably extract, from a plurality of key points, a predetermined key point as a key point falling within a capturing range of the image capturing device.”
In response, the office does not find this argument to be persuasive base on the same reasons set forth above and the rejection below.
The applicant argues on page 13, “Thirdly, presently amended claim 1 discloses determining one or more candidate actions from a plurality of target actions by comparing reference reliability of a detectable node predetermined for each of the target actions with reliability of the extracted detectable node, while Luo calculates reliability (visibility probability) not for only an extracted detectable node (key point) but for all the nodes (key points) and determines an action by using the calculated reliability (visibility probability). See e.g., paragraphs [0022] and [0023]. Moreover, at least since Luo cannot reasonably disclose "extracting, from the estimated nodes, a predetermined detectable node which is detectable by the image capturing device" features of amended claim 1 for the reasons noted above, Luo similarly cannot be relied upon to disclose or render obvious "determining one or more candidate actions from a plurality of target actions by comparing reference reliability of the detectable node predetermined for each of the target actions with reliability of the extracted detectable node" features of amended claim 1.”
In response, the office does not find this argument to be persuasive. Based on the breadth of the claim language the prior art by Luo et al. (US 20210281744 A1) explicitly teaches an action recognition method for an action recognition device that recognizes an action of a user (Fig. 8, Paragraph [0031]- Luo discloses a method and a device for recognizing an action of a target object and an electronic apparatus are provided according to the present disclosure.),
by a processor included in the action recognition device (Fig. 8, Paragraph [0143]- Luo discloses as shown in FIG. 8, the electronic apparatus 800 may include a processing device (for example, a central processing unit, a graphics processing unit and the like) 801.),
estimating a plurality of nodes of the user and reliability of each of the nodes from the image (Fig. 1, Paragraph [0022]- Luo discloses the visibility probability determining module is configured to output, by using the visibility determining model, a visibility probability of each of the multiple key points.);
extracting, from the estimated nodes, a predetermined detectable node which is detectable by the camera (Fig. 4, Paragraph [0075]- Luo discloses the outputs of the model are probabilities each indicating whether a key point is visible. If the probability is greater than a threshold, it is determined that a visibility of the key point indicated by the probability is consistent with the mark of the key point on the image. (wherein the predetermined node is a node above a threshold value)),
the detectable node falling within a capturing range of the camera (Fig. 1, Paragraph [0076]- Luo discloses then the probability is outputted through an activation function, where 1 is outputted for a visible key point, and 0 is outputted for an invisible key point. In this way, a vector of values of visibility attributes is obtained. The vector is a 1*N vector where N represents the number of the key points. A value of an element in the vector is equal to 0 or 1. 0 represents that a corresponding key point is invisible and 1 represents that a corresponding key point is visible (wherein if a node is visible it is determined to be detectable and since it is in the image it is within a capturing range of the camera).);
determining one or more candidate actions from a plurality of target actions by comparing reference reliability of the detectable node predetermined for each of the target actions with reliability of the extracted detectable node (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.);
determining the action of the user from the one or more candidate actions (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.);
Although Luo teaches Luo comprising: acquiring an image of the user captured by a camera Luo fails to explicitly teach comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
However, Ng explicitly teaches comprising: acquiring an image of the user captured by a camera being kept fixed (Fig. 1 Paragraph [0049]- Ng discloses these multiple standalone embedded fall-detection vision sensors can be installed at multiple fixed locations different from one another, wherein each of the multiple embedded fall-detection vision sensors can include at least one camera (wherein the camera is the camera being kept fixed at a location.) for capturing video images and various software and hardware modules for processing the captured video images and generating corresponding fall-detection output including fall alarms/notifications based on the captured video images.);
outputting an action label indicating the determined action (Fig. 1, Paragraph [0097]- Ng discloses that action-recognition module 108 is followed by fall-detection module 110, which receives the outputs from both pose-estimation module 106 (i.e., human keypoints 122) and action-recognition module 108 (i.e., the action labels/classifications 124)… action-recognition module 108 can then use cropped image 132 and/or keypoints 122 of a detected person to generate frame-by-frame action labels/classifications 124 for the detected person.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo of a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Ng comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
Wherein having Luo’s system for action recognition wherein comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
The motivation behind the modification would have been to allow for more accurate detection and classification of actions to be obtained, since both Luo and Ng are systems that use images to determine actions. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Ng’s system wherein further improved accuracy and robustness of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Ng et al. (US 20200211154 A1) Paragraph [0004-7].
Luo in view of Ng fails to explicitly teach and being set at an initial setting.
However, Fu explicitly teaches and being set at an initial setting (Fig. 2, Paragraph [0039]- Fu discloses In turn, an indication of limb orientation in the current frame is generated by refining the initial prediction of limb orientation at each initial prediction of limb location in the current frame using indications of limb orientations from the previous frame (wherein the initial prediction created using the previous frame is seen as an initial setting).);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Fu and being set at an initial setting.
Wherein having Luo’s system for action recognition wherein and being set at an initial setting.
The motivation behind the modification would have been to allow for a more consistent system prediction, since both Luo and Fu are systems that determine location of joints in an image. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Fu’s system wherein further improved accuracy and consistency of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Fu et al. (US 20220254157 A1) Paragraph [0083].
The applicant argues on page 13-14, “Although the above noted disclosure in Ng may be relevant to "outputting an action label indicating the determined action" features of claim 1, such disclosure is irrelevant to the above noted distinctions noted with respect to Luo. Accordingly, Ng fails to reasonably fail to cure the above noted deficiencies of Luo for establishing a primafacie case of obviousness under 35 USC 103.”
In response, the office does not find this argument to be persuasive based on the same reasons set forth above and the rejection below.
The applicant argues on page 14, “In addition to the above, Applicant further notes that Kechichian appears to disclose a technology which may arguably be relevant to original claims 2 and 3 of the present application. Hu appears to disclose a technology which may arguably be relevant to original claim 4 of the present application. Luo 2 appears to disclose a technology which may arguably be relevant to claim 9 of the present application. Gu appears to disclose a technology which may arguably be relevant to claims 1 to 13 of the present application. Lastly, Fu appears to disclose a technology which may arguably be relevant to original claims 14, 15 of the present application. However, all these cited documents fail to disclose the above noted distinctions with respect to Luo and thus fail to cure the above noted deficiencies of Luo for establishing a prima facie case of obviousness under 35 USC 103.”
In response, the office does not find this argument to be persuasive based on the same reasons set forth above and the rejection below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 5-8, 10, 14-15, and 17-18 are rejected under 35 U.S.C 103 as being unpatentable over Luo et al. (US 20210281744 A1) hereafter referenced as Luo in view of Ng et al. (US 20200211154 A1) hereafter referenced as Ng and Fu et al. (US 20220254157 A1) hereafter referenced as Fu.
Regarding claim 1, Luo explicitly teaches an action recognition method for an action recognition device that recognizes an action of a user (Fig. 8, Paragraph [0031]- Luo discloses a method and a device for recognizing an action of a target object and an electronic apparatus are provided according to the present disclosure.),
by a processor included in the action recognition device (Fig. 8, Paragraph [0143]- Luo discloses as shown in FIG. 8, the electronic apparatus 800 may include a processing device (for example, a central processing unit, a graphics processing unit and the like) 801.),
estimating a plurality of nodes of the user and reliability of each of the nodes from the image (Fig. 1, Paragraph [0022]- Luo discloses the visibility probability determining module is configured to output, by using the visibility determining model, a visibility probability of each of the multiple key points.);
extracting, from the estimated nodes, a predetermined detectable node which is detectable by the camera (Fig. 4, Paragraph [0075]- Luo discloses the outputs of the model are probabilities each indicating whether a key point is visible. If the probability is greater than a threshold, it is determined that a visibility of the key point indicated by the probability is consistent with the mark of the key point on the image. (wherein the predetermined node is a node above a threshold value)),
the detectable node falling within a capturing range of the camera (Fig. 1, Paragraph [0076]- Luo discloses then the probability is outputted through an activation function, where 1 is outputted for a visible key point, and 0 is outputted for an invisible key point. In this way, a vector of values of visibility attributes is obtained. The vector is a 1*N vector where N represents the number of the key points. A value of an element in the vector is equal to 0 or 1. 0 represents that a corresponding key point is invisible and 1 represents that a corresponding key point is visible (wherein if a node is visible it is determined to be detectable and since it is in the image it is within a capturing range of the camera).);
determining one or more candidate actions from a plurality of target actions by comparing reference reliability of the detectable node predetermined for each of the target actions with reliability of the extracted detectable node (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.);
determining the action of the user from the one or more candidate actions (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.);
Although Luo teaches Luo comprising: acquiring an image of the user captured by a camera Luo fails to explicitly teach comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
However, Ng explicitly teaches comprising: acquiring an image of the user captured by a camera being kept fixed (Fig. 1 Paragraph [0049]- Ng discloses these multiple standalone embedded fall-detection vision sensors can be installed at multiple fixed locations different from one another, wherein each of the multiple embedded fall-detection vision sensors can include at least one camera (wherein the camera is the camera being kept fixed at a location.) for capturing video images and various software and hardware modules for processing the captured video images and generating corresponding fall-detection output including fall alarms/notifications based on the captured video images.);
outputting an action label indicating the determined action (Fig. 1, Paragraph [0097]- Ng discloses that action-recognition module 108 is followed by fall-detection module 110, which receives the outputs from both pose-estimation module 106 (i.e., human keypoints 122) and action-recognition module 108 (i.e., the action labels/classifications 124)… action-recognition module 108 can then use cropped image 132 and/or keypoints 122 of a detected person to generate frame-by-frame action labels/classifications 124 for the detected person.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo of a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Ng comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
Wherein having Luo’s system for action recognition wherein comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
The motivation behind the modification would have been to allow for more accurate detection and classification of actions to be obtained, since both Luo and Ng are systems that use images to determine actions. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Ng’s system wherein further improved accuracy and robustness of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Ng et al. (US 20200211154 A1) Paragraph [0004-7].
Luo in view of Ng fails to explicitly teach and being set at an initial setting.
However, Fu explicitly teaches and being set at an initial setting (Fig. 2, Paragraph [0039]- Fu discloses In turn, an indication of limb orientation in the current frame is generated by refining the initial prediction of limb orientation at each initial prediction of limb location in the current frame using indications of limb orientations from the previous frame (wherein the initial prediction created using the previous frame is seen as an initial setting).);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Fu and being set at an initial setting.
Wherein having Luo’s system for action recognition wherein and being set at an initial setting.
The motivation behind the modification would have been to allow for a more consistent system prediction, since both Luo and Fu are systems that determine location of joints in an image. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Fu’s system wherein further improved accuracy and consistency of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Fu et al. (US 20220254157 A1) Paragraph [0083].
Regarding claim 5, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo further teaches wherein, in the determining of the action, each of the one or more candidate actions is determined to be the action (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.).
Regarding claim 6, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo further teaches wherein, in the determining of the one or more candidate actions, a similarity between a distribution of reliability of a plurality of detectable nodes and a distribution of reference reliability of the detectable nodes is calculated for each of the target actions (Fig. 1, Paragraph [0082]- Luo discloses Then the combined value of visibility attributes is compared with each of the reference values. In the present disclosure, the comparison may be performed by calculating a similarity between two vectors. The similarity between two vectors may be calculated by using the Pearson correlation coefficient, the Euclidean distance, the cosine similarity, the Manhattan distance and the like.),
and the one or more candidate actions are determined based on the similarity calculated for each of the target actions (Fig. 1, Paragraph [0082]- Luo discloses if a similarity is greater than a predetermined second threshold, it is determined that the combined value matches the reference value. The action corresponding to the reference value is outputted as a recognized action of the target object.).
Regarding claim 7, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 6, Luo further teaches wherein the similarity represents a total value of respective differences between the reliability and the reference reliability calculated for each of the detectable nodes (Fig. 1, Paragraph [0082]- Luo discloses Then the combined value of visibility attributes is compared with each of the reference values. In the present disclosure, the comparison may be performed by calculating a similarity between two vectors. The similarity between two vectors may be calculated by using the Pearson correlation coefficient, the Euclidean distance, the cosine similarity, the Manhattan distance and the like.).
Regarding claim 8, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 6, Luo further teaches wherein the reference reliability includes true reliability given to a detectable node having preliminarily estimated reliability exceeding a threshold (Fig. 5, Paragraph [0082]- Luo discloses the reference value may also be a vector, for example, a 1*N vector. Futher in Fig. 4, Paragraph [0076]- Luo discloses the vector is a 1*N vector where N represents the number of the key points. A value of an element in the vector is equal to 0 or 1. 0 represents that a corresponding key point is invisible and 1 represents that a corresponding key point is visible. (Wherein 1 or visible is considered true)),
and false reliability given to a detectable node having preliminarily estimated reliability falling below the threshold (Fig. 5, Paragraph [0082]- Luo discloses the reference value may also be a vector, for example, a 1*N vector. Futher in Fig. 4, Paragraph [0076]- Luo discloses the vector is a 1*N vector where N represents the number of the key points. A value of an element in the vector is equal to 0 or 1. 0 represents that a corresponding key point is invisible and 1 represents that a corresponding key point is visible. (Wherein 0 or invisible is considered false)),
the action recognition method further comprising: giving the true reliability to a detectable node whose reliability estimated from the image exceeds the threshold (Fig. 1, Paragraph [0076]- Luo discloses the visibility probability is compared with the predetermined first threshold. The threshold may be equal to 0.8. That is, in a case that an outputted probability of a key point is greater than 0.8, it is determined that the key point is visible. In a case that an outputted probability of a key point is less than 0.8, it is determined that the key point is invisible. Then the probability is outputted through an activation function, where 1 is outputted for a visible key point, and 0 is outputted for an invisible key point. (Wherein 1 or visible is considered true)),
and giving the false reliability to the detectable node whose reliability estimated from the image falls below the threshold (Fig. 1, Paragraph [0076]- Luo discloses the visibility probability is compared with the predetermined first threshold. The threshold may be equal to 0.8. That is, in a case that an outputted probability of a key point is greater than 0.8, it is determined that the key point is visible. In a case that an outputted probability of a key point is less than 0.8, it is determined that the key point is invisible. Then the probability is outputted through an activation function, where 1 is outputted for a visible key point, and 0 is outputted for an invisible key point. (Wherein 0 or invisible is considered false)),
wherein the similarity is based on a number of values of reliability where truth of the reliability and truth of the reference reliability agree with each other, and false of the reliability and false of the reference reliability agree with each other on each of the detectable nodes (Fig. 5, Paragraph [0080-81]- Luo discloses the combined value of the visibility attributes is compared with the reference value. In step S503, if the combined value matches the reference value, the action corresponding to the reference value is outputted as the recognized action of the target object.).
Regarding claim 10, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo fails to explicitly teach wherein the nodes and the reliability are estimated by inputting the image into a learned model obtained through machine learning of a relation between the image and the node.
However, Ng explicitly teaches wherein the nodes and the reliability are estimated by inputting the image into a learned model obtained through machine learning of a relation between the image and the node (Fig. 1, Paragraph [0071]- Ng discloses here the probability of a detected keypoint represents a confidence score assigned to the detected keypoint by the pose-estimation model.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo of a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Ng wherein the nodes and the reliability are estimated by inputting the image into a learned model obtained through machine learning of a relation between the image and the node.
Wherein having Luo’s system for action recognition wherein the nodes and the reliability are estimated by inputting the image into a learned model obtained through machine learning of a relation between the image and the node.
The motivation behind the modification would have been to allow for more accurate detection and classification of actions to be obtained, since both Luo and Ng are systems that use images to determine actions. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Ng’s system wherein further improved accuracy and robustness of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Ng et al. (US 20200211154 A1) Paragraph [0004-7].
Regarding claim 14, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo in view of Ng fails to explicitly teach wherein the detectable node is preliminarily determined based on a result of analysis of the image of the user captured by the camera at an initial setting.
However, Fu explicitly teaches wherein the detectable node is preliminarily determined based on a result of analysis of the image of the user captured by the camera at an initial setting (Fig. 2, Paragraph [0038]- Fu discloses generating the indications of the joint and limb locations in the current frame 222 includes processing the initial prediction of joint locations in the current frame and the indications of joint locations from the previous frame with a first deep convolutional neural network to generate the indication of joint locations in the current frame. (wherein the previous frame is considered the initial setting)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Fu wherein the detectable node is preliminarily determined based on a result of analysis of the image of the user captured by the camera at an initial setting.
Wherein having Luo’s system for action recognition wherein the detectable node is preliminarily determined based on a result of analysis of the image of the user captured by the camera at an initial setting.
The motivation behind the modification would have been to allow for a more consistent system prediction, since both Luo and Fu are systems that determine location of joints in an image. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Fu’s system wherein further improved accuracy and consistency of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Fu et al. (US 20220254157 A1) Paragraph [0083].
Regarding claim 15, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo in view of Ng fails to explicitly teach wherein the reference reliability is preliminarily calculated based on the reliability of each node estimated from an image of the user having made each of the target actions, the image being captured by camera at an initial setting.
However, Fu explicitly teaches wherein the reference reliability is preliminarily calculated based on the reliability of each node estimated from an image of the user having made each of the target actions, the image being captured by camera at an initial setting (Fig. 4, Paragraph [0058]- Fu discloses the prediction can also be seen as a confidence map with size H×W×2q, where q is the number of limbs defined. To prepare the ground-truth confidence map for limb prediction, e.g., the ground truth predictions 441b and 442b discussed hereinbelow in relation to FIG. 4, an embodiment first defines q limbs between a pair of joints indicating meaningful human limbs (or limbs of any object being detected) such as head, neck, body, trunk and forearm, which will form a skeleton of a human body in the pose association part.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Fu wherein the reference reliability is preliminarily calculated based on the reliability of each node estimated from an image of the user having made each of the target actions, the image being captured by camera at an initial setting.
Wherein having Luo’s system for action recognition wherein the reference reliability is preliminarily calculated based on the reliability of each node estimated from an image of the user having made each of the target actions, the image being captured by camera at an initial setting.
The motivation behind the modification would have been to allow for a more consistent system prediction, since both Luo and Fu are systems that determine location of joints in an image. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Fu’s system wherein further improved accuracy and consistency of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Fu et al. (US 20220254157 A1) Paragraph [0083].
Regarding claim 17, Luo explicitly teaches an action recognition device for recognizing an action of a user (Fig. 8, Paragraph [0031]- Luo discloses a method and a device for recognizing an action of a target object and an electronic apparatus are provided according to the present disclosure.),
comprising: a processor that (Fig. 8, Paragraph [0028]- Luo discloses the electronic apparatus includes a memory and a processor.):
estimates a plurality of nodes of the user and reliability of each of the nodes from the image (Fig. 1, Paragraph [0022]- Luo discloses the visibility probability determining module is configured to output, by using the visibility determining model, a visibility probability of each of the multiple key points.);
extracts, from the estimated nodes, a predetermined detectable node which is detectable by the camera (Fig. 4, Paragraph [0075]- Luo discloses the outputs of the model are probabilities each indicating whether a key point is visible. If the probability is greater than a threshold, it is determined that a visibility of the key point indicated by the probability is consistent with the mark of the key point on the image. (wherein the predetermined node is a node above a threshold value)),
the detectable node falling within a capturing range of the camera (Fig. 1, Paragraph [0076]- Luo discloses then the probability is outputted through an activation function, where 1 is outputted for a visible key point, and 0 is outputted for an invisible key point. In this way, a vector of values of visibility attributes is obtained. The vector is a 1*N vector where N represents the number of the key points. A value of an element in the vector is equal to 0 or 1. 0 represents that a corresponding key point is invisible and 1 represents that a corresponding key point is visible (wherein if a node is visible it is determined to be detectable and since it is in the image it is within a capturing range of the camera).);
determines one or more candidate actions from a plurality of target actions by comparing reference reliability of the detectable node predetermined for each of the target actions with reliability of the extracted detectable node (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.);
Although Luo teaches acquires an image of the user captured by a camera Luo fails to explicitly teach acquires an image of the user captured by a camera being kept fixed; outputs an action label indicating the determined action.
However, Ng explicitly teaches acquires an image of the user captured by a camera being kept fixed (Fig. 1 Paragraph [0049]- Ng discloses these multiple standalone embedded fall-detection vision sensors can be installed at multiple fixed locations different from one another, wherein each of the multiple embedded fall-detection vision sensors can include at least one camera (wherein the camera is the camera being kept fixed at a location.) for capturing video images and various software and hardware modules for processing the captured video images and generating corresponding fall-detection output including fall alarms/notifications based on the captured video images.);
outputs an action label indicating the determined action (Fig. 1, Paragraph [0097]- Ng discloses that action-recognition module 108 is followed by fall-detection module 110, which receives the outputs from both pose-estimation module 106 (i.e., human keypoints 122) and action-recognition module 108 (i.e., the action labels/classifications 124)… action-recognition module 108 can then use cropped image 132 and/or keypoints 122 of a detected person to generate frame-by-frame action labels/classifications 124 for the detected person.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo of an action recognition device for recognizing an action of a user, comprising: an acquisition part that acquires an image of the user captured by an image capturing device with the teachings of Ng acquires an image of the user captured by a camera being kept fixed; outputs an action label indicating the determined action.
Wherein having Luo’s system for action recognition wherein acquires an image of the user captured by a camera being kept fixed; outputs an action label indicating the determined action.
The motivation behind the modification would have been to allow for more accurate detection and classification of actions to be obtained, since both Luo and Ng are systems that use images to determine actions. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Ng’s system wherein further improved accuracy and robustness of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Ng et al. (US 20200211154 A1) Paragraph [0004-7].
Luo in view of Ng fails to explicitly teach and being set at an initial setting.
However, Fu explicitly teaches and being set at an initial setting (Fig. 2, Paragraph [0039]- Fu discloses In turn, an indication of limb orientation in the current frame is generated by refining the initial prediction of limb orientation at each initial prediction of limb location in the current frame using indications of limb orientations from the previous frame (wherein the initial prediction created using the previous frame is seen as an initial setting).);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo an action recognition device for recognizing an action of a user, comprising: an acquisition part that acquires an image of the user captured by an image capturing device with the teachings of Fu and being set at an initial setting.
Wherein having Luo’s system for action recognition wherein and being set at an initial setting.
The motivation behind the modification would have been to allow for a more consistent system prediction, since both Luo and Fu are systems that determine location of joints in an image. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Fu’s system wherein further improved accuracy and consistency of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Fu et al. (US 20220254157 A1) Paragraph [0083].
Regarding claim 18, Luo explicitly teaches a non-transitory computer readable recording medium storing an action recognition program for causing a computer to execute an action recognition method for recognizing an action of a user, by the computer (Fig. 8, Paragraph [0030]- Luo discloses a computer readable storage media is provided. The computer readable storage media is configured to store non-transient computer readable instructions that when being executed by a computer, cause the computer to perform steps in any one of the above methods.),
estimating a plurality of nodes of the user and reliability of each of the nodes from the image (Fig. 1, Paragraph [0022]- Luo discloses the visibility probability determining module is configured to output, by using the visibility determining model, a visibility probability of each of the multiple key points.);
extracting, from the estimated nodes, a predetermined detectable node which is detectable by the camera (Fig. 4, Paragraph [0075]- Luo discloses the outputs of the model are probabilities each indicating whether a key point is visible. If the probability is greater than a threshold, it is determined that a visibility of the key point indicated by the probability is consistent with the mark of the key point on the image. (wherein the predetermined node is a node above a threshold value));
the detectable node falling within a capturing range of the camera (Fig. 1, Paragraph [0076]- Luo discloses then the probability is outputted through an activation function, where 1 is outputted for a visible key point, and 0 is outputted for an invisible key point. In this way, a vector of values of visibility attributes is obtained. The vector is a 1*N vector where N represents the number of the key points. A value of an element in the vector is equal to 0 or 1. 0 represents that a corresponding key point is invisible and 1 represents that a corresponding key point is visible (wherein if a node is visible it is determined to be detectable and since it is in the image it is within a capturing range of the camera).);
determining one or more candidate actions from a plurality of target actions by comparing reference reliability of the detectable node predetermined for each of the target actions with reliability of the extracted detectable node (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.);
determining the action of the user from the one or more candidate actions (Fig. 7, Paragraph [0128]- Luo discloses the first recognizing module is configured to output, if the combined value matches the reference value, the action corresponding to the reference value as the recognized action of the target object.);
Although Luo teaches Luo comprising: acquiring an image of the user captured by a camera Luo fails to explicitly teach comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
However, Ng explicitly teaches comprising: acquiring an image of the user captured by a camera being kept fixed (Fig. 1 Paragraph [0049]- Ng discloses these multiple standalone embedded fall-detection vision sensors can be installed at multiple fixed locations different from one another, wherein each of the multiple embedded fall-detection vision sensors can include at least one camera (wherein the camera is the camera being kept fixed at a location.) for capturing video images and various software and hardware modules for processing the captured video images and generating corresponding fall-detection output including fall alarms/notifications based on the captured video images.);
outputting an action label indicating the determined action (Fig. 1, Paragraph [0097]- Ng discloses that action-recognition module 108 is followed by fall-detection module 110, which receives the outputs from both pose-estimation module 106 (i.e., human keypoints 122) and action-recognition module 108 (i.e., the action labels/classifications 124)… action-recognition module 108 can then use cropped image 132 and/or keypoints 122 of a detected person to generate frame-by-frame action labels/classifications 124 for the detected person.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo of an a non-transitory computer readable recording medium storing an action recognition program for causing a computer to execute an action recognition method for recognizing an action of a user, by the computer, comprising: acquiring an image of the user captured by an image capturing device with the teachings of Ng comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
Wherein having Luo’s system for action recognition wherein comprising: acquiring an image of the user captured by a camera being kept fixed; outputting an action label indicating the determined action.
The motivation behind the modification would have been to allow for more accurate detection and classification of actions to be obtained, since both Luo and Ng are systems that use images to determine actions. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Ng’s system wherein further improved accuracy and robustness of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Ng et al. (US 20200211154 A1) Paragraph [0004-7].
Luo in view of Ng fails to explicitly teach and being set at an initial setting.
However, Fu explicitly teaches and being set at an initial setting (Fig. 2, Paragraph [0039]- Fu discloses In turn, an indication of limb orientation in the current frame is generated by refining the initial prediction of limb orientation at each initial prediction of limb location in the current frame using indications of limb orientations from the previous frame (wherein the initial prediction created using the previous frame is seen as an initial setting).);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng a non-transitory computer readable recording medium storing an action recognition program for causing a computer to execute an action recognition method for recognizing an action of a user, by the computer, comprising: acquiring an image of the user captured by an image capturing device with the teachings of Fu and being set at an initial setting.
Wherein having Luo’s system for action recognition wherein and being set at an initial setting.
The motivation behind the modification would have been to allow for a more consistent system prediction, since both Luo and Fu are systems that determine location of joints in an image. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Fu’s system wherein further improved accuracy and consistency of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Fu et al. (US 20220254157 A1) Paragraph [0083].
Claims 2-3 are rejected under 35 U.S.C 103 as being unpatentable over Luo et al. (US 20210281744 A1) hereafter referenced as Luo in view of Ng et al. (US 20200211154 A1) hereafter referenced as Ng, Fu et al. (US 20220254157 A1) hereafter referenced as Fu, and Kechichian et al. (US 20210349122 A1) hereafter referenced as Kechichian.
Regarding claim 2, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo in view of Ng fails to explicitly teach wherein the action includes an action of the user using an appliance or equipment arranged in a facility.
However, Kechichian explicitly teaches wherein the action includes an action of the user using an appliance or equipment arranged in a facility (Fig. 1, Paragraph [0067]- Kechichian discloses some users make use of walking aids, including walkers, canes and wheelchairs. Using mobility aids changes the direction of the wrist during walking and may result in lower trigger rates for the detection algorithm, leading to delays in correctly estimating the wearing location.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Kechichian wherein the action includes an action of the user using an appliance or equipment arranged in a facility.
Wherein having Luo’s system for action recognition wherein the action includes an action of the user using an appliance or equipment arranged in a facility.
The motivation behind the modification would have been to allow for more information to be obtained, since both Luo and Kechichian are systems that determine location of joints. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Kechichian’s system wherein further improved accuracy and efficiency of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Kechichian et al. (US 20210349122 A1) Paragraph [0025].
Regarding claim 3, Luo in view of Fu, Ng, and Kechichian explicitly teaches the action recognition method according to claim 2,
Luo in view of Ng fails to explicitly teach wherein the equipment includes a rod for assisting a motion of the user, and the appliance includes a stand or a chair for assisting a motion of the user.
However, Kechichian explicitly teaches wherein the equipment includes a rod for assisting a motion of the user (Fig. 1, Paragraph [0067]- Kechichian discloses some users make use of walking aids, including walkers, canes and wheelchairs. Using mobility aids changes the direction of the wrist during walking and may result in lower trigger rates for the detection algorithm, leading to delays in correctly estimating the wearing location. (Wherein a Cane is considered a rod for assisting a motion of the user)),
and the appliance includes a stand or a chair for assisting a motion of the user (Fig. 1, Paragraph [0067]- Keechichian discloses some users make use of walking aids, including walkers, canes and wheelchairs. Using mobility aids changes the direction of the wrist during walking and may result in lower trigger rates for the detection algorithm, leading to delays in correctly estimating the wearing location. (Wherein a Walker is considered a stand for assisting a motion of the user and a Wheelchair a chair for assisting user motion)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Fu, Ng, and Keechichian of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Kechichian wherein the equipment includes a rod for assisting a motion of the user, and the appliance includes a stand or a chair for assisting a motion of the user.
Wherein having Luo’s system for action recognition wherein the equipment includes a rod for assisting a motion of the user, and the appliance includes a stand or a chair for assisting a motion of the user.
The motivation behind the modification would have been to allow for more information to be obtained, since both Luo and Kechichian are systems that determine location of joints. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Kechichian’s system wherein further improved accuracy and efficiency of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Kechichian et al. (US 20210349122 A1) Paragraph [0025].
Claim 4 and 16 are rejected under 35 U.S.C 103 as being unpatentable over Luo et al. (US 20210281744 A1) hereafter referenced as Luo in view of Fu et al. (US 20220254157 A1) hereafter referenced as Fu, Ng et al. (US 20200211154 A1) hereafter referenced as Ng, and Hu et al. (US 20210097101 A1) hereafter referenced as Hu.
Regarding claim 4, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo in view of Ng fails to explicitly teach wherein, in the determining of the action, in connection with each of the one or more candidate actions, a distance between a coordinate of the extracted detectable node and a reference coordinate of the detectable node is calculated for each of the target actions, and the action is determined based on the distance calculated for each of the target actions.
However, Hu explicitly teaches wherein, in the determining of the action, in connection with each of the one or more candidate actions, a distance between a coordinate of the extracted detectable node and a reference coordinate of the detectable node is calculated for each of the target actions (Fig. 6, Pargraph [0049]- Hu discloses FIG. 6 shows a schematic flowchart of a process for calculating the pose similarity between the candidate person and the reference person according to an embodiment of the present disclosure. As shown in FIG. 6, at block 610, based on the reference keypoint data of the reference person and the candidate keypoint data of the candidate person, the pose distance L between the candidate person and the reference person is calculated according to Equation (1)),
and the action is determined based on the distance calculated for each of the target actions (Fig. 5, Paragraph [0053]- Hu discloses it is determined whether the pose similarity is greater than the predetermined threshold. If the pose similarity is greater than or equal to the predetermined threshold, at block 520, it is determined that the candidate person has the pose similar to the reference person.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Hu wherein, in the determining of the action, in connection with each of the one or more candidate actions, a distance between a coordinate of the extracted detectable node and a reference coordinate of the detectable node is calculated for each of the target actions, and the action is determined based on the distance calculated for each of the target actions.
Wherein having Luo’s system for action recognition wherein, in the determining of the action, in connection with each of the one or more candidate actions, a distance between a coordinate of the extracted detectable node and a reference coordinate of the detectable node is calculated for each of the target actions, and the action is determined based on the distance calculated for each of the target actions.
The motivation behind the modification would have been to allow for a more accurate system, since both Luo and Hu are systems that determine a pose using location of joints. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Hu’s system wherein further improved accuracy and speed of system. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Hu et al. (US 20210097101 A1) Paragraph [0003-4].
Regarding claim 16, Luo in view of Fu, Ng, and Hu explicitly teaches the action recognition method according to claim 4, Luo in view of Ng and Hu fails to explicitly teach wherein the reference coordinate is preliminarily calculated based on a coordinate of each node estimated from an image of the user having taken each of the target actions, the image being captured by the camera at an initial setting.
However, Fu explicitly teaches wherein the reference coordinate is preliminarily calculated based on a coordinate of each node estimated from an image of the user having taken each of the target actions, the image being captured by the camera at an initial setting (Fig. 2, Paragraph [0038]- Fu discloses generating the indications of the joint and limb locations in the current frame 222 includes processing the initial prediction of joint locations in the current frame and the indications of joint locations from the previous frame with a first deep convolutional neural network to generate the indication of joint locations in the current frame. (wherein the previous frame is considered the initial setting)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Fu, Ng, and Hu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Fu wherein the reference coordinate is preliminarily calculated based on a coordinate of each node estimated from an image of the user having taken each of the target actions, the image being captured by the camera at an initial setting.
Wherein having Luo’s system for action recognition wherein the reference coordinate is preliminarily calculated based on a coordinate of each node estimated from an image of the user having taken each of the target actions, the image being captured by the camera at an initial setting.
The motivation behind the modification would have been to allow for a more consistent system prediction, since both Luo and Fu are systems that determine location of joints in an image. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Fu’s system wherein further improved accuracy and consistency of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Fu et al. (US 20220254157 A1) Paragraph [0083].
Claim 9 is rejected under 35 U.S.C 103 as being unpatentable over Luo et al. (US 20210281744 A1) hereafter referenced as Luo in view of Ng et al. (US 20200211154 A1) hereafter referenced as Ng, Fu et al. (US 20220254157 A1) hereafter referenced as Fu, and Luo et al. (US 20210271892 A1) hereafter referenced as Luo2.
Regarding claim 9, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 6, Luo in view of Ng fails to explicitly teach wherein, in the determining of the one or more candidate actions, a target action in a top N-th similarity, where “N” is an integer equal to or greater than one, is determined to be each of the one or more candidate actions.
However, Luo2 explicitly teaches wherein, in the determining of the one or more candidate actions, a target action in a top N-th similarity, where “N” is an integer equal to or greater than one, is determined to be each of the one or more candidate actions (Fig. 1, Paragraph [0179]- Luo2 discloses the electronic device obtains a maximum similarity among the plurality of similarities as a first similarity (e.g., a highest first similarity). When the first similarity is greater than a first preset threshold, it indicates that a similarity between a dynamic action included in a first target window and the preset dynamic action is relatively high, and the dynamic action included in the first target window corresponding to the first similarity can be considered as the preset dynamic action.).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Luo2 wherein, in the determining of the one or more candidate actions, a target action in a top N-th similarity, where “N” is an integer equal to or greater than one, is determined to be each of the one or more candidate actions.
Wherein having Luo’s system for action recognition wherein, in the determining of the one or more candidate actions, a target action in a top N-th similarity, where “N” is an integer equal to or greater than one, is determined to be each of the one or more candidate actions.
The motivation behind the modification would have been to allow for more accurate and flexible system, since both Luo and Luo2 are systems for action recognition. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Luo2’s system wherein further improved accuracy and flexibility of prediction results. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Luo2 et al. (US 20210271892 A1) Paragraph [0230].
Claims 11-13 are rejected under 35 U.S.C 103 as being unpatentable over Luo et al. (US 20210281744 A1) hereafter referenced as Luo in view of Ng et al. (US 20200211154 A1) hereafter referenced as Ng, Fu et al. (US 20220254157 A1) hereafter referenced as Fu, and Gu et al. (Multi-Person Pose Estimation using an Orientation and Occlusion Aware Deep Learning Network) hereafter referenced as Gu.
Regarding claim 11, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo in view of Ng fails to explicitly teach wherein, in the extracting of the detectable node, the detectable node is extracted with reference to a first database defining information indicating whether each of the nodes represents the detectable node.
However, Gu explicitly teaches wherein, in the extracting of the detectable node, the detectable node is extracted with reference to a first database defining information indicating whether each of the nodes represents the detectable node (Fig. 12, Section 3.1 Paragraph [0001]- Gu discloses the images of the COCO keypoint dataset were captured in human’s daily life, such as a party, meeting room and sport field. The images include multi-person with different scales, orientations, occlusions and postures. Figure 12 gives a demonstration of this task, each person in the image has a set of annotations including segmentation, class name (in this task it is “person”), bounding box position, the number of labeled keypoints and a list of body joint information (labeled in the format (x,y,v) with 17 joints, where x,y is the location of joint and v is the visibility of joint that v = 0 means not labeled, v = 1 means labeled but not visible and v = 2 means labeled and visible) (wherein the visible label represents if the node is detectable).).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Gu wherein, in the extracting of the detectable node, the detectable node is extracted with reference to a first database defining information indicating whether each of the nodes represents the detectable node.
Wherein having Luo’s system for action recognition wherein, in the extracting of the detectable node, the detectable node is extracted with reference to a first database defining information indicating whether each of the nodes represents the detectable node.
The motivation behind the modification would have been to allow for more accurate predictions to be obtained, since both Luo and Gu are systems that determine joint locations in images. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Gu’s system wherein further improved accuracy of joint estimation. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Gu et al. (Multi-Person Pose Estimation using an Orientation and Occlusion Aware Deep Learning Network) Section 3.6 Paragraph [0005].
Regarding claim 12, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo in view of Ng fails to explicitly teach wherein, in the determining of the one or more candidate actions, the one or more candidate actions are determined with reference to a second database defining the reference reliability of the detectable node for each of the target actions.
However, Gu explicitly teaches wherein, in the determining of the one or more candidate actions, the one or more candidate actions are determined with reference to a second database defining the reference reliability of the detectable node for each of the target actions (Fig. 12, Section 3.1 Paragraph [0001]- Gu discloses the images of the COCO keypoint dataset were captured in human’s daily life, such as a party, meeting room and sport field. The images include multi-person with different scales, orientations, occlusions and postures. Figure 12 gives a demonstration of this task, each person in the image has a set of annotations including segmentation, class name (in this task it is “person”), bounding box position, the number of labeled keypoints and a list of body joint information (labeled in the format (x,y,v) with 17 joints, where x,y is the location of joint and v is the visibility of joint that v = 0 means not labeled, v = 1 means labeled but not visible and v = 2 means labeled and visible)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Gu wherein, in the determining of the one or more candidate actions, the one or more candidate actions are determined with reference to a second database defining the reference reliability of the detectable node for each of the target actions.
Wherein having Luo’s system for action recognition wherein, in the determining of the one or more candidate actions, the one or more candidate actions are determined with reference to a second database defining the reference reliability of the detectable node for each of the target actions.
The motivation behind the modification would have been to allow for more accurate predictions to be obtained, since both Luo and Gu are systems that determine joint locations in images. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Gu’s system wherein further improved accuracy of joint estimation. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Gu et al. (Multi-Person Pose Estimation using an Orientation and Occlusion Aware Deep Learning Network) Section 3.6 Paragraph [0005].
Regarding claim 13, Luo in view of Ng and Fu explicitly teaches the action recognition method according to claim 1, Luo in view of Ng fails to explicitly teach wherein, in the determining of the action, the action is determined with reference to a third database defining a reference coordinate of the detectable node for each of the target actions.
However, Gu explicitly teaches wherein, in the determining of the action, the action is determined with reference to a third database defining a reference coordinate of the detectable node for each of the target actions (Fig. 12, Section 3.1 Paragraph [0001]- Gu discloses the images of the COCO keypoint dataset were captured in human’s daily life, such as a party, meeting room and sport field. The images include multi-person with different scales, orientations, occlusions and postures. Figure 12 gives a demonstration of this task, each person in the image has a set of annotations including segmentation, class name (in this task it is “person”), bounding box position, the number of labeled keypoints and a list of body joint information (labeled in the format (x,y,v) with 17 joints, where x,y is the location of joint and v is the visibility of joint that v = 0 means not labeled, v = 1 means labeled but not visible and v = 2 means labeled and visible). (wherein the visible label represents if the node is detectable)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Luo in view of Ng and Fu of an a an action recognition method for an action recognition device that recognizes an action of a user, by a processor included in the action recognition device, comprising: acquiring an image of the user captured by an image capturing device views with the teachings of Gu wherein, in the determining of the action, the action is determined with reference to a third database defining a reference coordinate of the detectable node for each of the target actions.
Wherein having Luo’s system for action recognition wherein, in the determining of the action, the action is determined with reference to a third database defining a reference coordinate of the detectable node for each of the target actions.
The motivation behind the modification would have been to allow for more accurate predictions to be obtained, since both Luo and Gu are systems that determine joint locations in images. Wherein Luo’s system wherein improved accuracy of complex action recognition, while Gu’s system wherein further improved accuracy of joint estimation. Please see Luo et al. (US 20210281744 A1), Paragraph [0031] and Gu et al. (Multi-Person Pose Estimation using an Orientation and Occlusion Aware Deep Learning Network) Section 3.6 Paragraph [0005].
Conclusion
Listed below are the prior arts made of record and not relied upon but are considered
pertinent to applicant`s disclosure.
Shen et al. (US 11367313 B2)- Embodiments of the present disclosure disclose a method and apparatus for recognizing a body movement. A specific embodiment of the method includes: sampling an input to-be-recognized video to obtain a sampled image frame sequence of the to-be-recognized video; performing key point detection on the sampled image frame sequence by using a trained body key point detection model, to obtain a body key point position heat map of each sampled image frame in the sampled image frame sequence, the body key point position heat map being used to represent a probability feature of a position of a preset body key point; and inputting body key point position heat maps of the sampled image frame sequence into a trained movement classification model to perform classification, to obtain a body movement recognition result corresponding to the to-be-recognized video.......................Please see Fig. 1. Abstract.
Lan et al. (US 20190080176 A1)- In implementations of the subject matter described herein, an action detection scheme using a recurrent neural network (RNN) is proposed. Representation information of an incoming frame of a video and a predefined action label for the frame are obtained to train a learning network including RNN elements and a classification element. The representation information represents an observed entity in the frame. Specifically, parameters for the RNN elements are determined based on the representation information and the predefined action label. With the determined parameters, the RNN elements are caused to extract features for the frame based on the representation information and features for a preceding frame. Parameters for the classification element are determined based on the extracted features and the predefined action label. The classification element with the determined parameters generates a probability of the frame being associated with the predefined action label. The parameters for the RNN elements are updated according to the probability.....................Please see Fig. 1. Abstract.
KRUEGER et al. (US 20220405922 A1)- Disclosed herein is a medical instrument (100, 300). Execution of the machine executable instructions causes a processor (106) to: receive (206) a set of joint location coordinates (128) for a subject (118) reposing on a subject support (120), receive (207) a body orientation (132) in response to inputting the set of joint location coordinates into a predetermined logic module (130), calculate (208) a torso aspect ratio (134) from set of joint location coordinates. If (210) the torso aspect ratio is greater than a predetermined threshold (136) then (212) the body pose of the subject is a decubitus pose. Execution of the machine executable instructions further cause the processor to assign (220) the body pose as being a supine pose if the subject is face up on the subject support or assign (222) the body pose as being a prone pose if the subject is face down on the subject support if the torso aspect ratio is less than or equal to the predetermined threshold. Execution of the machine executable instructions further cause the processor to generate (216) a subject pose label (142) .......................Please see Fig. 1. Abstract.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LUCIUS C.G. ALLEN whose telephone number is (703)756-5987. The examiner can normally be reached Mon - Fri 8-5pm (EST).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571)272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LUCIUS CAMERON GREEN ALLEN/Examiner, Art Unit 2673
/CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673