Prosecution Insights
Last updated: October 01, 2026
Application No. 19/011,120

ACTION RECOGNITION DEVICE, ACTION RECOGNITION METHOD, AND NON-TRANSITORY COMPUTER READABLE RECORDING MEDIUM

Non-Final OA §103
Filed
Jan 06, 2025
Priority
Jul 07, 2022 — JP 2022-109937 +1 more
Examiner
YANG, WEI WEN
Art Unit
Tech Center
Assignee
Panasonic Holdings Corporation
OA Round
1 (Non-Final)
82%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
560 granted / 684 resolved
+21.9% vs TC avg
Moderate +12% lift
Without
With
+11.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
31 currently pending
Career history
705
Total Applications
across all art units

Statute-Specific Performance

§101
7.8%
-32.2% vs TC avg
§103
75.0%
+35.0% vs TC avg
§102
9.3%
-30.7% vs TC avg
§112
7.8%
-32.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 684 resolved cases

Office Action

§103
DETAILED ACTION Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-14 are rejected under 35 U.S.C. 103 as being unpatentable over ZHANG (CN 113378649 A), and in view of WANG (WO 2021185317 A1), and further in view of Matsugu (US 20050157908 A1). Re Claim 1, ZHANG discloses an action recognition device that recognizes an action of a user (see ZHANG: e.g., -- The motion recognition module is configured to identify an action of the person based on a sequence of video frames including an image of the person.--, in abstract), comprising: an acquisition part that acquires an image (see ZHANG: e.g., -- the position identifying module is used for identifying the position of the person based on the image of the person obtained from a plurality of cameras fixed at the relative position; The motion recognition module is configured to identify an action of the person based on a sequence of video frames including an image of the person.--, in abstract, and, -- camera devices fixed at the relative position--, in page 2 of 18 of English version of CN 113378649 A); an estimation part that estimates node coordinates of the user from the image acquired by the acquisition part (see ZHANG: e.g., --using open-pose single posture estimation algorithm to perform two-dimensional pixel coordinate positioning to the image of the person obtained by each imaging device in the plurality of imaging devices; obtaining the two-dimensional pixel coordinate of human body in each image; …… based on the two-dimensional pixel coordinate of the human body in each image and the inner and outer parameters of the plurality of camera devices, reconstructing the human three-dimensional position coordinate, so as to identify the position of the person.--, in page 3 of 18 of English version of CN 113378649 A; and, --The obtained human body image is used for combining the inner and outer parameters of the camera, reconstructing the human body node three-dimensional coordinate, calculating the area of the human body. .. based on human body two-dimensional pixel coordinate and camera inner and outer parameter reconstruction human joint three-dimensional coordinate, calculating to obtain the human body position based on human body node three-dimensional coordinate. … obtaining camera inner and outer parameter by camera calibration, the premise of human body node from two-dimensional to three-dimensional conversion.--, in pages 8-9 of 18 of English version of CN 113378649 A); a calculation part that calculates, on the basis of the node coordinates, time-series feature vectors (see Zhang: e.g., -- obtaining the front face image; in order to identify the identity of the person, using the ResNet-29 network in D11b as the feature extraction model of the system; converting the face image into the 128-dimensional feature vector… the identity identification module according to the embodiment by OpenPose comprises the human faces frame of the image characteristic point capture based on the feature point of the capture to rotate the face to the horizontal position, the image after rotating is divided into five parts input TP-GAN to obtain the front human faces image; using ResNet-29 network in D11b as feature extraction model to extract the feature vector; performing similarity judgement to the extracted vector and the feature vector in the database; identifying the identity information… The identity identification module provided by another embodiment of the invention may include a gait recognition module; the gait recognition is identified by the posture of people walking; it has the advantages of non-contact, long-distance identification and not easy to disguising; it has obvious advantages in the intelligent video monitoring field. Because the pedestrian has a certain difference in muscle strength, tendon and bone length, bone density, gravity center and so on; based on these differences, it can uniquely mark one person; using these characteristics can build human body motion model or directly extract features from the human body profile to realize gait recognition. The gait recognition module mainly comprises gait collection, gait segmentation, characteristic extraction, characteristic ratio peer-to-peer module. the input of the gait recognition module is a section of walking video image sequence; the video sequence captures the continuous change of the pedestrian in the walking process; the gait recognition algorithm excavate the gait characteristic of the pedestrian; and comparing the gait characteristic obtained by the newly obtained gait characteristic with the gait characteristic stored in the data set, so as to finish the identification..--, in in pages 7-8 of 18 of English version of CN 113378649 A); Zhang however does not explicitly disclose that the feature vector each indicating a feature vector connecting a trunk of the user and a head of the user to each other in time series; WANG discloses the feature vector each indicating a feature vector connecting a trunk of the user and a head of the user to each other in time series (see WANG: e.g., -- determining the first coordinate value of the face position of the person on the feature map; according to a preset vector and The first coordinate value determines the second coordinate value; wherein the preset vector is a vector that points from the position of the human face to the position of the human body; and the second coordinate value is used as the reference human body position. In some optional embodiments, the associating the face position and the human body position belonging to the same person according to the reference human body position and the at least one human body position includes: linking with the The human body position with the smallest reference human body position distance is associated with the face position corresponding to the reference human body position the at least one character included in the scene image and the target action type of each character in the at least one character are determined according to the associated face position and the human body position , Including: for each character in at least one character, determining a plurality of feature vectors according to the face position and the human body position associated with the character; Describe the target action type. In some optional embodiments, the determining a plurality of feature vectors according to the face position and the human body position associated with the person includes: determining that they correspond to at least one preset action type and are determined by the person. The face position points to multiple feature vectors of the associated human body position. In some optional embodiments, the determining the target action type of each character in the at least one character based on the plurality of feature vectors includes: normalizing the plurality of feature vectors corresponding to the character respectively The normalized value of each feature vector is obtained; the feature vector corresponding to the maximum normalized value is used as the target feature vector of the person; the action type corresponding to the target feature vector is used as the person’s Target action type.--, in page 3/20 of English version of WO-2021185317-A1; also see: -- there is provided an action recognition method, the method comprising: acquiring a scene image; performing detection of different parts of an object on the scene image, association of different parts in the same object, and motion recognition of the object , Determining at least one object included in the scene image and a target action type of each object in the at least one object… the object includes a person, and different parts of the object include the face and the human body of the person; the scene image is detected for different parts of the object, the association of different parts in the same object, and Object action recognition, determining at least one object included in the scene image and the target action type of each object in the at least one object includes: extracting features of the scene image to obtain a feature map; determining the feature map At least one human face position and at least one human body position; determine at least one person included in the scene image according to the at least one human face position and/or the at least one human body position; Associating with the position of the human body; and determining the target action type of each character in the at least one character in the scene image according to the associated face position and the human body position.--, in page 2/20 of English version of WO-2021185317-A1; also see: -- an external camera is set in the classroom. After the external camera collects the scene image in the classroom, it is sent to the cloud server through a router or gateway. The cloud server detects different parts of the object and detects different parts of the same object on the scene image. Associate and recognize the action of the object, and determine the at least one object included in the scene image and the target action type of each object in the at least one object. Further, the cloud server can feed back the above results to the corresponding teaching task analysis server as required, so as to remind the teacher to adjust the teaching content so as to better carry out the teaching activities…. the terminal device or the cloud server can also determine the cumulative detection result of each object included in the scene image within a set time period that matches the target action type … it is possible to determine how many times each teaching object has raised his hand, paid attention to the blackboard, the length of time to write with his head down, and the number of times he stood up to answer questions during the time period during which the teacher was teaching, for example, the time period of a class. --, in pages 7-8/20 of English version of WO-2021185317-A1.); Zhang and Wang are combinable as they are in the same field of endeavor: action recognition based on feature vectors and the calculations of nodes of a human head and body. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify ZHANG’s device using WANG’s teachings by including the feature vector each indicating a feature vector connecting a trunk of the user and a head of the user to each other in time series to ZHANG’s feature vectors in order to determine the action and action change of the user (see WANG: e.g., in in pages 2-3/20, and pages 7-8 of English version of WO-2021185317-A1); ZHANG as modified by WANG further disclose reference feature vectors of associating human face and human body, and action character and type (see WANG: e.g., -- the associating the position of the face and the position of the human body belonging to the same person includes: for each of the at least one person, determining the position corresponding to the position of the person's face Reference human body position; according to the reference human body position and the at least one human body position, the face position and the human body position belonging to the same person are associated. In some optional embodiments, the determining the reference human body position corresponding to each face position includes: determining the first coordinate value of the face position of the person on the feature map; according to a preset vector and The first coordinate value determines the second coordinate value; wherein the preset vector is a vector that points from the position of the human face to the position of the human body; and the second coordinate value is used as the reference human body position. In some optional embodiments, the associating the face position and the human body position belonging to the same person according to the reference human body position and the at least one human body position includes: linking with the The human body position with the smallest reference human body position distance is associated with the face position corresponding to the reference human body position. In some optional embodiments, the at least one character included in the scene image and the target action type of each character in the at least one character are determined according to the associated face position and the human body position , Including: for each character in at least one character, determining a plurality of feature vectors according to the face position and the human body position associated with the character; Describe the target action type. In some optional embodiments, the determining a plurality of feature vectors according to the face position and the human body position associated with the person includes: determining that they correspond to at least one preset action type and are determined by the person. The face position points to multiple feature vectors of the associated human body position.--, in in page 2/20, and pages 7-8 of English version of WO-2021185317-A1; and, -- the association submodule includes: a first determining unit, configured to determine, for each of at least one character, a reference human body position corresponding to the position of the person's face; the association unit uses According to the reference human body position and the at least one human body position, the face position and the human body position belonging to the same person are associated. In some optional embodiments, the first determining unit includes: determining, on the scene image, the first coordinate value of the person's face position on the feature map; and according to a preset vector and the first coordinate value A coordinate value to determine a second coordinate value respectively; wherein the preset vector is a vector that points from the position of the human face to the position of the human body; and the second coordinate value is used as the reference human body position. In some optional embodiments, the associating unit includes: associating the human body position with the smallest distance from the reference human body position and the face position corresponding to the reference human body position. In some optional embodiments, the second determining sub-module includes: a second determining unit, configured to, for each of the at least one character, determine according to the position of the face and the human body associated with the character Position, determining multiple feature vectors; a third determining unit, configured to determine the target action type of each of the at least one person based on the multiple feature vectors. In some optional embodiments, the second determining unit includes: determining multiple feature vectors respectively corresponding to at least one preset action type and pointing from the face position to the associated human body position.-, in page 3/20 of English version of WO-2021185317-A1); ZHANG as modified by WANG however do not explicitly disclose a storage part that stores reference time-series feature vectors being the time-series feature vectors to be a reference; Matsugu discloses a storage part that stores reference time-series feature vectors being the time-series feature vectors to be a reference (see Matsugu: e.g., -- an image of the head acquired by detecting the upper body, in particular, the head image of a person and actions of the person. The learning unit 5 pre-stores in the recording unit (not shown) model data of heads associated with a plurality of postures of a particular person, as well as model data related to action patterns specific to the person. It is sufficient to discretely sample, in predetermined angle steps, postures of the person for whom model data is generated. It is not necessary to prepare such a large amount of model data that continuously covers most postures.--, in [0051]; and, -- the association of the state category with time-series feature data is performed interactively using the dictionary data generation/recording unit 50 in the learning unit 5, or it is presumed that for the state category, the storage means of the dictionary data generation/recording unit 50 pre-stores time-series image data of the feature quantity related to the action pattern corresponding to a standard state category. Next, a subject person is detected from the input moving image in order to perform subject recognition processing in the same manner as with the above-described embodiment based on an action pattern or an appearance exhibited by the person (step S602)…. it is determined whether the time-series image data can be associated with the state category in the dictionary (step S605). If the time-series image data cannot be associated with any of state category in the dictionary or such an association is difficult, the time-series image data is learned as a new state category (step S606), and the flow proceeds to step S607. In contrast, if the time-series image data can be associated with the state category in the dictionary in step S605, the flow proceeds to step S607. Furthermore, the time-series image data is stored in a predetermined storage device along with one of the place and the time slot associated with the state category (step S607).--, in [0109]-[0111]; also see: -- [0035] Referring to FIG. 1, the action recognition apparatus includes an image input unit 1 for inputting moving image data from an imaging unit or a database (not shown); a moving-object detection unit 2 for detecting and outputting a dynamic (moving) area based on motion vectors and other features of the input moving image data; a subject recognition unit 3 for recognizing a subject, such as a person or an object, located in the detected dynamic area and outputting the recognition result; a state detection unit 4 for detecting an action and a state of a person serving as the subject recognized by the subject recognition unit 3; a learning unit 5 for learning person-specific meaning of the action or state detected by the state detection unit 4 and for storing the meaning in predetermined storage means; and a processing control unit 6 for controlling the startup and stop of processing by the moving-object detection unit 2, the subject recognition unit 3, the state detection unit 4, and the learning unit 5 and changes in processing mode of these units.--, in [0035]; and, -- [0050] Next, the subject recognition unit 3 inputs the feature of the action pattern extracted from the state detection unit 4 and evaluates the degree of similarity between the feature of the action pattern and the action-related model data (step S203). The subject recognition unit 3 then calculates the confidence coefficients about the feature of the image data and the feature of the action pattern (step S204), and applies weighting to similarities of the features according to the respective calculated confidence coefficients and outputs the result (step S205). The subject recognition unit 3 then detects the output similarity to determine whether the maximum detection level in input image is a threshold or more (step S206)… it is determined whether the subject estimated based on the image data has an action feature peculiar thereto (e.g., walking pattern, posture). In this case, the subject recognition unit 3 refers to the dictionary data in the learning unit 5, and if the subject does not have a characteristic action feature, the subject recognition unit 3 outputs information indicating that the identification of the subject is difficult. In response to this, the processing control unit 6 aborts the subsequent processing. [0054] The confidence coefficient about the feature of the image data is evaluated by applying weighting according to the detection level of similarity of the feature required to detect a potential subject category, the observation position (e.g., whether the frontal position or dorsal position) corresponding to the image data, and the orientation of the detected subject, where the largest weighting factor is applied to the observation position or direction optimal for the recognition of the subject and a lower weighting factor is applied to an image from an observation position not suitable for recognition..--, in [0046]-[0055]; and, -- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]); Zhang (as modified by Wang) and Matsugu are combinable as they are in the same field of endeavor: action recognition based on feature vectors and the calculations of nodes of a human head and body. Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify ZHANG (as modified by Wang)’s device using Matsugu’s teachings by including a storage part that stores reference time-series feature vectors being the time-series feature vectors to be a reference to ZHANG (as modified by Wang)’s reference feature vectors in order to detect an action and a state of a person, and stored as reference for further detection and recognition (see Matsugu: e.g., in [0035], [0051], [0046]-[0057], and [0109]-[0111]); Zhang as modified by Wang and Matsugu further disclose a decision part that decides the action of the user and a change in the action of the user by comparing input time-series feature vectors being the time-series feature vectors calculated by the calculation part with the reference time-series feature vectors (see Matsugu: e.g., ---- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]; also see see Zhang: e.g., -- obtaining the front face image; in order to identify the identity of the person, using the ResNet-29 network in D11b as the feature extraction model of the system; converting the face image into the 128-dimensional feature vector… the identity identification module according to the embodiment by OpenPose comprises the human faces frame of the image characteristic point capture based on the feature point of the capture to rotate the face to the horizontal position, the image after rotating is divided into five parts input TP-GAN to obtain the front human faces image; using ResNet-29 network in D11b as feature extraction model to extract the feature vector; performing similarity judgement to the extracted vector and the feature vector in the database; identifying the identity information… The identity identification module provided by another embodiment of the invention may include a gait recognition module; the gait recognition is identified by the posture of people walking; it has the advantages of non-contact, long-distance identification and not easy to disguising; it has obvious advantages in the intelligent video monitoring field. Because the pedestrian has a certain difference in muscle strength, tendon and bone length, bone density, gravity center and so on; based on these differences, it can uniquely mark one person; using these characteristics can build human body motion model or directly extract features from the human body profile to realize gait recognition. The gait recognition module mainly comprises gait collection, gait segmentation, characteristic extraction, characteristic ratio peer-to-peer module. the input of the gait recognition module is a section of walking video image sequence; the video sequence captures the continuous change of the pedestrian in the walking process; the gait recognition algorithm excavate the gait characteristic of the pedestrian; and comparing the gait characteristic obtained by the newly obtained gait characteristic with the gait characteristic stored in the data set, so as to finish the identification..--, in in pages 7-8 of 18 of English version of CN 113378649 A). Re Claim 2, Zhang as modified by Wang and Matsugu further disclose in the comparing by the decision part, a feature vector constituting the input time-series feature vectors is compared with a feature vector constituting the reference time-series feature vectors in terms of one of a length of the vector, an angle of the vector, a chronological change of the length, and a chronological change of the angle (see Matsugu: e.g., -- [0120] If, as shown in FIG. 12B, the observation position is behind the person A and an object to be gazed at by the person is not hidden by the person A, an object that is shown in the image and that exists near (as viewed from behind) the person A is a possible object category. In general, a search range X of this gazed object varies depending on the relative distance between the person A and the object viewed from the observation position of the action recognition apparatus. According to the present embodiment, the distance to the person A is estimated based on the size of the person's head viewed from the observation position of the action recognition apparatus, and a predetermined range of viewing angle with respect to the person is presumed to be the search range X. [0121] In this case, if possible, the action recognition apparatus can make a rough estimation of the face orientation by means of, for example, pre-association processing (calibration) based on an image of the back of the head to presume that a predetermined range of viewing angle with respect to the estimated face orientation of the person A is the search range X--, in [0120]-[0121]; also see Zhang: e.g., -- obtaining the front face image; in order to identify the identity of the person, using the ResNet-29 network in D11b as the feature extraction model of the system; converting the face image into the 128-dimensional feature vector… the identity identification module according to the embodiment by OpenPose comprises the human faces frame of the image characteristic point capture based on the feature point of the capture to rotate the face to the horizontal position, the image after rotating is divided into five parts input TP-GAN to obtain the front human faces image; using ResNet-29 network in D11b as feature extraction model to extract the feature vector; performing similarity judgement to the extracted vector and the feature vector in the database; identifying the identity information… The identity identification module provided by another embodiment of the invention may include a gait recognition module; the gait recognition is identified by the posture of people walking; it has the advantages of non-contact, long-distance identification and not easy to disguising; it has obvious advantages in the intelligent video monitoring field. Because the pedestrian has a certain difference in muscle strength, tendon and bone length, bone density, gravity center and so on; based on these differences, it can uniquely mark one person; using these characteristics can build human body motion model or directly extract features from the human body profile to realize gait recognition. The gait recognition module mainly comprises gait collection, gait segmentation, characteristic extraction, characteristic ratio peer-to-peer module. the input of the gait recognition module is a section of walking video image sequence; the video sequence captures the continuous change of the pedestrian in the walking process; the gait recognition algorithm excavate the gait characteristic of the pedestrian; and comparing the gait characteristic obtained by the newly obtained gait characteristic with the gait characteristic stored in the data set, so as to finish the identification..--, in in pages 7-8 of 18 of English version of CN 113378649 A). Re Claim 3, Zhang as modified by Wang and Matsugu further disclose wherein the time-series feature vectors are normalized in such a manner that a time is a reference value (see WANG: e.g., -- determining the first coordinate value of the face position of the person on the feature map; according to a preset vector and The first coordinate value determines the second coordinate value; wherein the preset vector is a vector that points from the position of the human face to the position of the human body; and the second coordinate value is used as the reference human body position. In some optional embodiments, the associating the face position and the human body position belonging to the same person according to the reference human body position and the at least one human body position includes: linking with the The human body position with the smallest reference human body position distance is associated with the face position corresponding to the reference human body position the at least one character included in the scene image and the target action type of each character in the at least one character are determined according to the associated face position and the human body position , Including: for each character in at least one character, determining a plurality of feature vectors according to the face position and the human body position associated with the character; Describe the target action type. In some optional embodiments, the determining a plurality of feature vectors according to the face position and the human body position associated with the person includes: determining that they correspond to at least one preset action type and are determined by the person. The face position points to multiple feature vectors of the associated human body position. In some optional embodiments, the determining the target action type of each character in the at least one character based on the plurality of feature vectors includes: normalizing the plurality of feature vectors corresponding to the character respectively The normalized value of each feature vector is obtained; the feature vector corresponding to the maximum normalized value is used as the target feature vector of the person; the action type corresponding to the target feature vector is used as the person’s Target action type.--, in page 3/20 of English version of WO-2021185317-A1). Re Claim 4, Zhang as modified by Wang and Matsugu further disclose wherein the node coordinates estimated by the estimation part include reliability indicating an estimation accuracy (see Matsugu: e.g., --a subject person can be identified based on the confidence coefficients of both image data and an action pattern, and the subject person can be identified with high accuracy even if the person exists at far away position or the lighting conditions vary greatly. Furthermore, device control can be personalized by learning the meaning of an action such that the action is associated with a particular person serving as a subject. In addition, objects-for-daily-use and habitual behavior patterns of a particular person can be detected and learned easily.--, in [0125]), and the calculation part calculates a coordinate of the trunk and a coordinate of the head by weighting the node coordinates by using the reliability, and calculates the feature vector on the basis of the calculated coordinate of the trunk and the calculated coordinate of the head (see Zhang: e.g., -- obtaining the front face image; in order to identify the identity of the person, using the ResNet-29 network in D11b as the feature extraction model of the system; converting the face image into the 128-dimensional feature vector… the identity identification module according to the embodiment by OpenPose comprises the human faces frame of the image characteristic point capture based on the feature point of the capture to rotate the face to the horizontal position, the image after rotating is divided into five parts input TP-GAN to obtain the front human faces image; using ResNet-29 network in D11b as feature extraction model to extract the feature vector; performing similarity judgement to the extracted vector and the feature vector in the database; identifying the identity information… The identity identification module provided by another embodiment of the invention may include a gait recognition module; the gait recognition is identified by the posture of people walking; it has the advantages of non-contact, long-distance identification and not easy to disguising; it has obvious advantages in the intelligent video monitoring field. Because the pedestrian has a certain difference in muscle strength, tendon and bone length, bone density, gravity center and so on; based on these differences, it can uniquely mark one person; using these characteristics can build human body motion model or directly extract features from the human body profile to realize gait recognition. The gait recognition module mainly comprises gait collection, gait segmentation, characteristic extraction, characteristic ratio peer-to-peer module. the input of the gait recognition module is a section of walking video image sequence; the video sequence captures the continuous change of the pedestrian in the walking process; the gait recognition algorithm excavate the gait characteristic of the pedestrian; and comparing the gait characteristic obtained by the newly obtained gait characteristic with the gait characteristic stored in the data set, so as to finish the identification..--, in in pages 7-8 of 18 of English version of CN 113378649 A). Re Claim 5, Zhang as modified by Wang and Matsugu further disclose wherein the reference time-series feature vectors include a plurality of first reference time-series feature vectors each associated with an action label indicating a kind of an action (See Matsugu: e.g., -- [0111] Subsequently, it is determined whether the time-series image data can be associated with the state category in the dictionary (step S605). If the time-series image data cannot be associated with any of state category in the dictionary or such an association is difficult, the time-series image data is learned as a new state category (step S606), and the flow proceeds to step S607. In contrast, if the time-series image data can be associated with the state category in the dictionary in step S605, the flow proceeds to step S607. Furthermore, the time-series image data is stored in a predetermined storage device along with one of the place and the time slot associated with the state category (step S607). [0112] The learning processing in step S606 is performed by the learning unit 5 that generates dictionary data in which the time-series feature quantity data is associated with the action category and state category. The learning processing can be performed via an interaction between the user and the system.--, in [0111]-[0112]; and, -- [0113] Finally, the frequency at which the extracted state category is observed is calculated. States or actions with a frequency of, for example, three or more times per week are regarded as habitual behavior when the determination in step S608 is made, and the current processing ends. [0114] According to the processing in FIG. 11, feature-quantity time-series image data of the action pattern of the identified person is detected (step S603), and the state category is learned if it is determined that the time-series image data cannot be associated with a state category in the dictionary or if such an association is difficult (step S606). This enables the state category to be learned and the frequency of the state category to be calculated to determine whether the state category is habitual. [0115] The detection processing of the gazed-object category 25 in FIGS. 10A and 10B will now be described. First, the eye direction or the face orientation of the subject person is detected. If the detected direction or orientation is maintained within a viewing angle of about .+-.10.degree. for more than a predetermined period of time, the person is determined to be in the state "gazing". It is noted that if the eye direction is not constant regardless of a constant face orientation, the person should not be determined to be in the state "gazing".--, in [0113]-[0115], and the decision part decides an action label showing the action of the user by comparing the input time-series feature vectors with the first reference time-series feature vectors, calculates an average time-series feature vector by averaging the first reference time-series feature vectors associated with the decided action label, and decides a change in the action of the user by comparing the input time-series feature vectors with the average time-series feature vector (see Matsugu: e.g., --0047] Referring to FIG. 3, geometrical features or the color-related features are extracted from the image data of the dynamic area (step S201). Then, the degree of similarity between the extracted feature of the image data and the model data pre-set for each subject person is evaluated (step S202). This degree of similarity is evaluated as a correlation coefficient with the model data or the saliency of the particular subject category. The saliency is evaluated based on, for example, the ratio of Cmax to Cnext, where Cmax is the maximum correlation coefficient with respect to the model data and Cnext is the second largest correlation coefficient, or based on the kurtosis of the correlation coefficient distribution classified according to subject category. The kurtosis of distribution "a" is calculated based on the expression shown below: kurtosis(a)=(1/N).multidot.(a-)4/.sigma.(a)4-3 [0048] where .rho. is the variance of distribution "a", "a" is the average of distribution "a", and N is the number of samples. [0049] The degree of similarity of geometrical features is evaluated as follows. That is, template matching is performed at each level of a predetermined hierarchical structure of geometrical features detected in a local area (e.g., see Japanese Patent Publication No. 3078166 by the assignee). More specifically, the hierarchical structure includes, for example, a first level of relatively simple patterns (e.g., line segments with specific direction components; and V-shaped, L-shaped, and T-shaped corner patterns generated by combining line segments), a second level of intermediately complicated patterns expressed as a combination of simple patterns or the relationship among the relative positions of the simple patterns, and a third level of highly complicated patterns generated by a combination of intermediately complicated patterns., and, -- [0050] Next, the subject recognition unit 3 inputs the feature of the action pattern extracted from the state detection unit 4 and evaluates the degree of similarity between the feature of the action pattern and the action-related model data (step S203). The subject recognition unit 3 then calculates the confidence coefficients about the feature of the image data and the feature of the action pattern (step S204), and applies weighting to similarities of the features according to the respective calculated confidence coefficients and outputs the result (step S205). The subject recognition unit 3 then detects the output similarity to determine whether the maximum detection level in input image is a threshold or more (step S206)… it is determined whether the subject estimated based on the image data has an action feature peculiar thereto (e.g., walking pattern, posture). In this case, the subject recognition unit 3 refers to the dictionary data in the learning unit 5, and if the subject does not have a characteristic action feature, the subject recognition unit 3 outputs information indicating that the identification of the subject is difficult. In response to this, the processing control unit 6 aborts the subsequent processing. [0054] The confidence coefficient about the feature of the image data is evaluated by applying weighting according to the detection level of similarity of the feature required to detect a potential subject category, the observation position (e.g., whether the frontal position or dorsal position) corresponding to the image data, and the orientation of the detected subject, where the largest weighting factor is applied to the observation position or direction optimal for the recognition of the subject and a lower weighting factor is applied to an image from an observation position not suitable for recognition..--, in [0046]-[0055]; and, -- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]). Re Claim 6, Zhang as modified by Wang and Matsugu further disclose wherein a feature vector calculated by the calculation part indicates a coordinate of the trunk and a coordinate of the head by using coordinates of the image (see WANG: e.g., -- determining the first coordinate value of the face position of the person on the feature map; according to a preset vector and The first coordinate value determines the second coordinate value; wherein the preset vector is a vector that points from the position of the human face to the position of the human body; and the second coordinate value is used as the reference human body position. In some optional embodiments, the associating the face position and the human body position belonging to the same person according to the reference human body position and the at least one human body position includes: linking with the The human body position with the smallest reference human body position distance is associated with the face position corresponding to the reference human body position the at least one character included in the scene image and the target action type of each character in the at least one character are determined according to the associated face position and the human body position , Including: for each character in at least one character, determining a plurality of feature vectors according to the face position and the human body position associated with the character; Describe the target action type. In some optional embodiments, the determining a plurality of feature vectors according to the face position and the human body position associated with the person includes: determining that they correspond to at least one preset action type and are determined by the person. The face position points to multiple feature vectors of the associated human body position. In some optional embodiments, the determining the target action type of each character in the at least one character based on the plurality of feature vectors includes: normalizing the plurality of feature vectors corresponding to the character respectively The normalized value of each feature vector is obtained; the feature vector corresponding to the maximum normalized value is used as the target feature vector of the person; the action type corresponding to the target feature vector is used as the person’s Target action type.--, in page 3/20 of English version of WO-2021185317-A1, and the decision part executes a parallel translation of the feature vector SO that the coordinate of the trunk on the feature vector calculated by the calculation part meets an origin of a coordinate system of the image, and shows the input time-series feature vectors by using the feature vector obtained by the parallel translation (see ZHANG: e.g., --In the algorithm 1, the video frame sequence (f1, f2, ..., fn) represents each action video V, convenient for using image processing method indirectly processing the video data. Specifically, algorithm 1 the video frame sequence in each image according to the original order, in a given range in the horizontal direction of the translation (translation unit length and direction is random, n= -1, representing the left translation; n= 1, representing to the right). if one action video comprises 50 frames, and generating the range of the random number is (-6, 6), then the data enhancement processing, at most can expand the data 600 times. at the same time, translating the video frame image in the horizontal direction; it also can relieve the edge information loss problem caused by central cutting.--, in pages 10-11 of 18 of English version of CN 113378649 A; and, -- using data enhancement algorithm of the data pre-processing process and the video frame sampling algorithm of the video frame sampling process of the data processing sub-module, using the characteristic extraction sub-module of the residual network structure of the concentration; using two layers of LSTM action classification sub-module, which can effectively increase the diversity of the data, improve the quality of the data, enhancing model to extract the discriminatory feature, finally better identifying the action of the human body. Compared with the existing method, the embodiment of the invention claims an attention mechanism of the deep learning more effectively extract the space information and the time information of the video data, has better identification effect. Ubuntu16.04 in the operating system; deep learning frame pytorch1.6.0; general parallel computing architecture cuda10.2; deep neural network GPU acceleration library cudnn7.6.5; a display 11GB GeForce RTX 2080Ti the display card drives nvidia450.80; Under the environment of hard disk 512GB the action identification method provided by the invention is verified.--, in page 13 of 18 of English version of CN 113378649; also see WANG: e.g., -- determining the first coordinate value of the face position of the person on the feature map; according to a preset vector and The first coordinate value determines the second coordinate value; wherein the preset vector is a vector that points from the position of the human face to the position of the human body; and the second coordinate value is used as the reference human body position. In some optional embodiments, the associating the face position and the human body position belonging to the same person according to the reference human body position and the at least one human body position includes: linking with the The human body position with the smallest reference human body position distance is associated with the face position corresponding to the reference human body position the at least one character included in the scene image and the target action type of each character in the at least one character are determined according to the associated face position and the human body position , Including: for each character in at least one character, determining a plurality of feature vectors according to the face position and the human body position associated with the character; Describe the target action type. In some optional embodiments, the determining a plurality of feature vectors according to the face position and the human body position associated with the person includes: determining that they correspond to at least one preset action type and are determined by the person. The face position points to multiple feature vectors of the associated human body position. In some optional embodiments, the determining the target action type of each character in the at least one character based on the plurality of feature vectors includes: normalizing the plurality of feature vectors corresponding to the character respectively The normalized value of each feature vector is obtained; the feature vector corresponding to the maximum normalized value is used as the target feature vector of the person; the action type corresponding to the target feature vector is used as the person’s Target action type.--, in page 3/20 of English version of WO-2021185317-A1). Re Claim 7, Zhang as modified by Wang and Matsugu further disclose wherein the decision part normalizes feature vectors so that each of a vertical length and a horizontal length of the image has 1, and shows the time-series feature vectors by using the normalized feature vectors (see Zhang: e.g., -- obtaining the front face image; in order to identify the identity of the person, using the ResNet-29 network in D11b as the feature extraction model of the system; converting the face image into the 128-dimensional feature vector… the identity identification module according to the embodiment by OpenPose comprises the human faces frame of the image characteristic point capture based on the feature point of the capture to rotate the face to the horizontal position, the image after rotating is divided into five parts input TP-GAN to obtain the front human faces image; using ResNet-29 network in D11b as feature extraction model to extract the feature vector; performing similarity judgement to the extracted vector and the feature vector in the database; identifying the identity information… The identity identification module provided by another embodiment of the invention may include a gait recognition module; the gait recognition is identified by the posture of people walking; it has the advantages of non-contact, long-distance identification and not easy to disguising; it has obvious advantages in the intelligent video monitoring field. Because the pedestrian has a certain difference in muscle strength, tendon and bone length, bone density, gravity center and so on; based on these differences, it can uniquely mark one person; using these characteristics can build human body motion model or directly extract features from the human body profile to realize gait recognition. The gait recognition module mainly comprises gait collection, gait segmentation, characteristic extraction, characteristic ratio peer-to-peer module. the input of the gait recognition module is a section of walking video image sequence; the video sequence captures the continuous change of the pedestrian in the walking process; the gait recognition algorithm excavate the gait characteristic of the pedestrian; and comparing the gait characteristic obtained by the newly obtained gait characteristic with the gait characteristic stored in the data set, so as to finish the identification..--, in in pages 7-8 of 18 of English version of CN 113378649 A; and, --In the algorithm 1, the video frame sequence (f1, f2, ..., fn) represents each action video V, convenient for using image processing method indirectly processing the video data. Specifically, algorithm 1 the video frame sequence in each image according to the original order, in a given range in the horizontal direction of the translation (translation unit length and direction is random, n= -1, representing the left translation; n= 1, representing to the right). if one action video comprises 50 frames, and generating the range of the random number is (-6, 6), then the data enhancement processing, at most can expand the data 600 times. at the same time, translating the video frame image in the horizontal direction; it also can relieve the edge information loss problem caused by central cutting.--, in pages 10-11 of 18 of English version of CN 113378649 A). Re Claim 8, Zhang as modified by Wang and Matsugu further disclose in a case where a period during which time-series feature vectors are continuously calculated is not shorter than a predetermined period, the calculation part determines the feature vectors within the period as the input time-series feature vectors (see Matsugu: e.g., -- an image of the head acquired by detecting the upper body, in particular, the head image of a person and actions of the person. The learning unit 5 pre-stores in the recording unit (not shown) model data of heads associated with a plurality of postures of a particular person, as well as model data related to action patterns specific to the person. It is sufficient to discretely sample, in predetermined angle steps, postures of the person for whom model data is generated. It is not necessary to prepare such a large amount of model data that continuously covers most postures.--, in [0051]; and, -- the association of the state category with time-series feature data is performed interactively using the dictionary data generation/recording unit 50 in the learning unit 5, or it is presumed that for the state category, the storage means of the dictionary data generation/recording unit 50 pre-stores time-series image data of the feature quantity related to the action pattern corresponding to a standard state category. Next, a subject person is detected from the input moving image in order to perform subject recognition processing in the same manner as with the above-described embodiment based on an action pattern or an appearance exhibited by the person (step S602)…. it is determined whether the time-series image data can be associated with the state category in the dictionary (step S605). If the time-series image data cannot be associated with any of state category in the dictionary or such an association is difficult, the time-series image data is learned as a new state category (step S606), and the flow proceeds to step S607. In contrast, if the time-series image data can be associated with the state category in the dictionary in step S605, the flow proceeds to step S607. Furthermore, the time-series image data is stored in a predetermined storage device along with one of the place and the time slot associated with the state category (step S607).--, in [0109]-[0111]; also see: -- [0035] Referring to FIG. 1, the action recognition apparatus includes an image input unit 1 for inputting moving image data from an imaging unit or a database (not shown); a moving-object detection unit 2 for detecting and outputting a dynamic (moving) area based on motion vectors and other features of the input moving image data; a subject recognition unit 3 for recognizing a subject, such as a person or an object, located in the detected dynamic area and outputting the recognition result; a state detection unit 4 for detecting an action and a state of a person serving as the subject recognized by the subject recognition unit 3; a learning unit 5 for learning person-specific meaning of the action or state detected by the state detection unit 4 and for storing the meaning in predetermined storage means; and a processing control unit 6 for controlling the startup and stop of processing by the moving-object detection unit 2, the subject recognition unit 3, the state detection unit 4, and the learning unit 5 and changes in processing mode of these units.--, in [0035]; and, -- [0050] Next, the subject recognition unit 3 inputs the feature of the action pattern extracted from the state detection unit 4 and evaluates the degree of similarity between the feature of the action pattern and the action-related model data (step S203). The subject recognition unit 3 then calculates the confidence coefficients about the feature of the image data and the feature of the action pattern (step S204), and applies weighting to similarities of the features according to the respective calculated confidence coefficients and outputs the result (step S205). The subject recognition unit 3 then detects the output similarity to determine whether the maximum detection level in input image is a threshold or more (step S206)… it is determined whether the subject estimated based on the image data has an action feature peculiar thereto (e.g., walking pattern, posture). In this case, the subject recognition unit 3 refers to the dictionary data in the learning unit 5, and if the subject does not have a characteristic action feature, the subject recognition unit 3 outputs information indicating that the identification of the subject is difficult. In response to this, the processing control unit 6 aborts the subsequent processing. [0054] The confidence coefficient about the feature of the image data is evaluated by applying weighting according to the detection level of similarity of the feature required to detect a potential subject category, the observation position (e.g., whether the frontal position or dorsal position) corresponding to the image data, and the orientation of the detected subject, where the largest weighting factor is applied to the observation position or direction optimal for the recognition of the subject and a lower weighting factor is applied to an image from an observation position not suitable for recognition..--, in [0046]-[0055]; and, -- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]). Re Claim 9, Zhang as modified by Wang and Matsugu further disclose wherein the reference time-series feature vectors represent input time-series feature vectors calculated by the calculation part in past (see Matsugu: e.g., -- an image of the head acquired by detecting the upper body, in particular, the head image of a person and actions of the person. The learning unit 5 pre-stores in the recording unit (not shown) model data of heads associated with a plurality of postures of a particular person, as well as model data related to action patterns specific to the person. It is sufficient to discretely sample, in predetermined angle steps, postures of the person for whom model data is generated. It is not necessary to prepare such a large amount of model data that continuously covers most postures.--, in [0051]; and, -- the association of the state category with time-series feature data is performed interactively using the dictionary data generation/recording unit 50 in the learning unit 5, or it is presumed that for the state category, the storage means of the dictionary data generation/recording unit 50 pre-stores time-series image data of the feature quantity related to the action pattern corresponding to a standard state category. Next, a subject person is detected from the input moving image in order to perform subject recognition processing in the same manner as with the above-described embodiment based on an action pattern or an appearance exhibited by the person (step S602)…. it is determined whether the time-series image data can be associated with the state category in the dictionary (step S605). If the time-series image data cannot be associated with any of state category in the dictionary or such an association is difficult, the time-series image data is learned as a new state category (step S606), and the flow proceeds to step S607. In contrast, if the time-series image data can be associated with the state category in the dictionary in step S605, the flow proceeds to step S607. Furthermore, the time-series image data is stored in a predetermined storage device along with one of the place and the time slot associated with the state category (step S607).--, in [0109]-[0111]; also see: -- [0035] Referring to FIG. 1, the action recognition apparatus includes an image input unit 1 for inputting moving image data from an imaging unit or a database (not shown); a moving-object detection unit 2 for detecting and outputting a dynamic (moving) area based on motion vectors and other features of the input moving image data; a subject recognition unit 3 for recognizing a subject, such as a person or an object, located in the detected dynamic area and outputting the recognition result; a state detection unit 4 for detecting an action and a state of a person serving as the subject recognized by the subject recognition unit 3; a learning unit 5 for learning person-specific meaning of the action or state detected by the state detection unit 4 and for storing the meaning in predetermined storage means; and a processing control unit 6 for controlling the startup and stop of processing by the moving-object detection unit 2, the subject recognition unit 3, the state detection unit 4, and the learning unit 5 and changes in processing mode of these units.--, in [0035]; and, -- [0050] Next, the subject recognition unit 3 inputs the feature of the action pattern extracted from the state detection unit 4 and evaluates the degree of similarity between the feature of the action pattern and the action-related model data (step S203). The subject recognition unit 3 then calculates the confidence coefficients about the feature of the image data and the feature of the action pattern (step S204), and applies weighting to similarities of the features according to the respective calculated confidence coefficients and outputs the result (step S205). The subject recognition unit 3 then detects the output similarity to determine whether the maximum detection level in input image is a threshold or more (step S206)… it is determined whether the subject estimated based on the image data has an action feature peculiar thereto (e.g., walking pattern, posture). In this case, the subject recognition unit 3 refers to the dictionary data in the learning unit 5, and if the subject does not have a characteristic action feature, the subject recognition unit 3 outputs information indicating that the identification of the subject is difficult. In response to this, the processing control unit 6 aborts the subsequent processing. [0054] The confidence coefficient about the feature of the image data is evaluated by applying weighting according to the detection level of similarity of the feature required to detect a potential subject category, the observation position (e.g., whether the frontal position or dorsal position) corresponding to the image data, and the orientation of the detected subject, where the largest weighting factor is applied to the observation position or direction optimal for the recognition of the subject and a lower weighting factor is applied to an image from an observation position not suitable for recognition..--, in [0046]-[0055]; and, -- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]). Re Claim 10, Zhang as modified by Wang and Matsugu further disclose wherein the reference time-series feature vectors and the input time-series feature vectors belong to the same user (see Matsugu: e.g., -- an image of the head acquired by detecting the upper body, in particular, the head image of a person and actions of the person. The learning unit 5 pre-stores in the recording unit (not shown) model data of heads associated with a plurality of postures of a particular person, as well as model data related to action patterns specific to the person. It is sufficient to discretely sample, in predetermined angle steps, postures of the person for whom model data is generated. It is not necessary to prepare such a large amount of model data that continuously covers most postures.--, in [0051]; and, -- the association of the state category with time-series feature data is performed interactively using the dictionary data generation/recording unit 50 in the learning unit 5, or it is presumed that for the state category, the storage means of the dictionary data generation/recording unit 50 pre-stores time-series image data of the feature quantity related to the action pattern corresponding to a standard state category. Next, a subject person is detected from the input moving image in order to perform subject recognition processing in the same manner as with the above-described embodiment based on an action pattern or an appearance exhibited by the person (step S602)…. it is determined whether the time-series image data can be associated with the state category in the dictionary (step S605). If the time-series image data cannot be associated with any of state category in the dictionary or such an association is difficult, the time-series image data is learned as a new state category (step S606), and the flow proceeds to step S607. In contrast, if the time-series image data can be associated with the state category in the dictionary in step S605, the flow proceeds to step S607. Furthermore, the time-series image data is stored in a predetermined storage device along with one of the place and the time slot associated with the state category (step S607).--, in [0109]-[0111]; also see: -- [0035] Referring to FIG. 1, the action recognition apparatus includes an image input unit 1 for inputting moving image data from an imaging unit or a database (not shown); a moving-object detection unit 2 for detecting and outputting a dynamic (moving) area based on motion vectors and other features of the input moving image data; a subject recognition unit 3 for recognizing a subject, such as a person or an object, located in the detected dynamic area and outputting the recognition result; a state detection unit 4 for detecting an action and a state of a person serving as the subject recognized by the subject recognition unit 3; a learning unit 5 for learning person-specific meaning of the action or state detected by the state detection unit 4 and for storing the meaning in predetermined storage means; and a processing control unit 6 for controlling the startup and stop of processing by the moving-object detection unit 2, the subject recognition unit 3, the state detection unit 4, and the learning unit 5 and changes in processing mode of these units.--, in [0035]; and, -- [0050] Next, the subject recognition unit 3 inputs the feature of the action pattern extracted from the state detection unit 4 and evaluates the degree of similarity between the feature of the action pattern and the action-related model data (step S203). The subject recognition unit 3 then calculates the confidence coefficients about the feature of the image data and the feature of the action pattern (step S204), and applies weighting to similarities of the features according to the respective calculated confidence coefficients and outputs the result (step S205). The subject recognition unit 3 then detects the output similarity to determine whether the maximum detection level in input image is a threshold or more (step S206)… it is determined whether the subject estimated based on the image data has an action feature peculiar thereto (e.g., walking pattern, posture). In this case, the subject recognition unit 3 refers to the dictionary data in the learning unit 5, and if the subject does not have a characteristic action feature, the subject recognition unit 3 outputs information indicating that the identification of the subject is difficult. In response to this, the processing control unit 6 aborts the subsequent processing. [0054] The confidence coefficient about the feature of the image data is evaluated by applying weighting according to the detection level of similarity of the feature required to detect a potential subject category, the observation position (e.g., whether the frontal position or dorsal position) corresponding to the image data, and the orientation of the detected subject, where the largest weighting factor is applied to the observation position or direction optimal for the recognition of the subject and a lower weighting factor is applied to an image from an observation position not suitable for recognition..--, in [0046]-[0055]; and, -- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]). Re Claim 11, Zhang as modified by Wang and Matsugu further disclose wherein the decision part decides an occurrence of a change in the action of the user when a statistical value of correlations between feature vectors of the input time-series feature vectors and feature vectors of the reference time- series feature vectors in time-series association exceeds a threshold (see Matsugu: e.g., -- an image of the head acquired by detecting the upper body, in particular, the head image of a person and actions of the person. The learning unit 5 pre-stores in the recording unit (not shown) model data of heads associated with a plurality of postures of a particular person, as well as model data related to action patterns specific to the person. It is sufficient to discretely sample, in predetermined angle steps, postures of the person for whom model data is generated. It is not necessary to prepare such a large amount of model data that continuously covers most postures.--, in [0051]; and, -- the association of the state category with time-series feature data is performed interactively using the dictionary data generation/recording unit 50 in the learning unit 5, or it is presumed that for the state category, the storage means of the dictionary data generation/recording unit 50 pre-stores time-series image data of the feature quantity related to the action pattern corresponding to a standard state category. Next, a subject person is detected from the input moving image in order to perform subject recognition processing in the same manner as with the above-described embodiment based on an action pattern or an appearance exhibited by the person (step S602)…. it is determined whether the time-series image data can be associated with the state category in the dictionary (step S605). If the time-series image data cannot be associated with any of state category in the dictionary or such an association is difficult, the time-series image data is learned as a new state category (step S606), and the flow proceeds to step S607. In contrast, if the time-series image data can be associated with the state category in the dictionary in step S605, the flow proceeds to step S607. Furthermore, the time-series image data is stored in a predetermined storage device along with one of the place and the time slot associated with the state category (step S607).--, in [0109]-[0111]; also see: -- [0035] Referring to FIG. 1, the action recognition apparatus includes an image input unit 1 for inputting moving image data from an imaging unit or a database (not shown); a moving-object detection unit 2 for detecting and outputting a dynamic (moving) area based on motion vectors and other features of the input moving image data; a subject recognition unit 3 for recognizing a subject, such as a person or an object, located in the detected dynamic area and outputting the recognition result; a state detection unit 4 for detecting an action and a state of a person serving as the subject recognized by the subject recognition unit 3; a learning unit 5 for learning person-specific meaning of the action or state detected by the state detection unit 4 and for storing the meaning in predetermined storage means; and a processing control unit 6 for controlling the startup and stop of processing by the moving-object detection unit 2, the subject recognition unit 3, the state detection unit 4, and the learning unit 5 and changes in processing mode of these units.--, in [0035]; and, -- [0050] Next, the subject recognition unit 3 inputs the feature of the action pattern extracted from the state detection unit 4 and evaluates the degree of similarity between the feature of the action pattern and the action-related model data (step S203). The subject recognition unit 3 then calculates the confidence coefficients about the feature of the image data and the feature of the action pattern (step S204), and applies weighting to similarities of the features according to the respective calculated confidence coefficients and outputs the result (step S205). The subject recognition unit 3 then detects the output similarity to determine whether the maximum detection level in input image is a threshold or more (step S206)… it is determined whether the subject estimated based on the image data has an action feature peculiar thereto (e.g., walking pattern, posture). In this case, the subject recognition unit 3 refers to the dictionary data in the learning unit 5, and if the subject does not have a characteristic action feature, the subject recognition unit 3 outputs information indicating that the identification of the subject is difficult. In response to this, the processing control unit 6 aborts the subsequent processing. [0054] The confidence coefficient about the feature of the image data is evaluated by applying weighting according to the detection level of similarity of the feature required to detect a potential subject category, the observation position (e.g., whether the frontal position or dorsal position) corresponding to the image data, and the orientation of the detected subject, where the largest weighting factor is applied to the observation position or direction optimal for the recognition of the subject and a lower weighting factor is applied to an image from an observation position not suitable for recognition..--, in [0046]-[0055]; and, -- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]). . Re Claim 12, Zhang as modified by Wang and Matsugu further disclose an output part that outputs the action and the change in the action that are decided by the decision part (see Matsugu: e.g., -- an image of the head acquired by detecting the upper body, in particular, the head image of a person and actions of the person. The learning unit 5 pre-stores in the recording unit (not shown) model data of heads associated with a plurality of postures of a particular person, as well as model data related to action patterns specific to the person. It is sufficient to discretely sample, in predetermined angle steps, postures of the person for whom model data is generated. It is not necessary to prepare such a large amount of model data that continuously covers most postures.--, in [0051]; and, -- the association of the state category with time-series feature data is performed interactively using the dictionary data generation/recording unit 50 in the learning unit 5, or it is presumed that for the state category, the storage means of the dictionary data generation/recording unit 50 pre-stores time-series image data of the feature quantity related to the action pattern corresponding to a standard state category. Next, a subject person is detected from the input moving image in order to perform subject recognition processing in the same manner as with the above-described embodiment based on an action pattern or an appearance exhibited by the person (step S602)…. it is determined whether the time-series image data can be associated with the state category in the dictionary (step S605). If the time-series image data cannot be associated with any of state category in the dictionary or such an association is difficult, the time-series image data is learned as a new state category (step S606), and the flow proceeds to step S607. In contrast, if the time-series image data can be associated with the state category in the dictionary in step S605, the flow proceeds to step S607. Furthermore, the time-series image data is stored in a predetermined storage device along with one of the place and the time slot associated with the state category (step S607).--, in [0109]-[0111]; also see: -- [0035] Referring to FIG. 1, the action recognition apparatus includes an image input unit 1 for inputting moving image data from an imaging unit or a database (not shown); a moving-object detection unit 2 for detecting and outputting a dynamic (moving) area based on motion vectors and other features of the input moving image data; a subject recognition unit 3 for recognizing a subject, such as a person or an object, located in the detected dynamic area and outputting the recognition result; a state detection unit 4 for detecting an action and a state of a person serving as the subject recognized by the subject recognition unit 3; a learning unit 5 for learning person-specific meaning of the action or state detected by the state detection unit 4 and for storing the meaning in predetermined storage means; and a processing control unit 6 for controlling the startup and stop of processing by the moving-object detection unit 2, the subject recognition unit 3, the state detection unit 4, and the learning unit 5 and changes in processing mode of these units.--, in [0035]; and, -- [0050] Next, the subject recognition unit 3 inputs the feature of the action pattern extracted from the state detection unit 4 and evaluates the degree of similarity between the feature of the action pattern and the action-related model data (step S203). The subject recognition unit 3 then calculates the confidence coefficients about the feature of the image data and the feature of the action pattern (step S204), and applies weighting to similarities of the features according to the respective calculated confidence coefficients and outputs the result (step S205). The subject recognition unit 3 then detects the output similarity to determine whether the maximum detection level in input image is a threshold or more (step S206)… it is determined whether the subject estimated based on the image data has an action feature peculiar thereto (e.g., walking pattern, posture). In this case, the subject recognition unit 3 refers to the dictionary data in the learning unit 5, and if the subject does not have a characteristic action feature, the subject recognition unit 3 outputs information indicating that the identification of the subject is difficult. In response to this, the processing control unit 6 aborts the subsequent processing. [0054] The confidence coefficient about the feature of the image data is evaluated by applying weighting according to the detection level of similarity of the feature required to detect a potential subject category, the observation position (e.g., whether the frontal position or dorsal position) corresponding to the image data, and the orientation of the detected subject, where the largest weighting factor is applied to the observation position or direction optimal for the recognition of the subject and a lower weighting factor is applied to an image from an observation position not suitable for recognition..--, in [0046]-[0055]; and, -- obtaining data of the appearance as viewed from an averaged head (by gender) image vector or the feature vector after feature extraction and calculating the eigenvector from the covariance matrix related to the feature vector, and then extracting the saliency as a deviation (absolute value of differences) from the mean vector of the coordinate values in the eigenvector space. [0057] On the other hand, the confidence coefficient about the feature of action pattern is evaluated based on, for example, the SN ratio of time-series feature data, the ratio of the threshold to the degree of similarity between the model data and the time-series feature data, or the variance of the maximum degree of similarity calculated in a predetermined period of time. Furthermore, an index (i.e., degree of personality) representing whether the detected action pattern is peculiar to the subject person may be pre-evaluated for the action feature category, so that the confidence coefficient is multiplied by the index.--, in [0056]-[0057]). . Re Claim 13, claim 13 is the corresponding method claim to claim 1, respectively. Claim 13 is rejected for the similar reasons for claim 1. See above discussions with regard to claim 1 respectively. Zhang as modified by Wang and Matsugu further disclose action recognition method for an action recognition device that recognizes an action of a user (see ZHANG: e.g., -- the position identifying module is used for identifying the position of the person based on the image of the person obtained from a plurality of cameras fixed at the relative position; The motion recognition module is configured to identify an action of the person based on a sequence of video frames including an image of the person.--, in abstract). Re Claim 14, claim 14 is the corresponding medium claim to claim 1, respectively. Claim 14 is rejected for the similar reasons for claim 1. See above discussions with regard to claim 1 respectively. Zhang as modified by Wang and Matsugu further disclose anon-transitory computer readable recording medium storing an action recognition program for causing a computer to serve as an action recognition device that recognizes an action of a user, the action recognition program comprising: causing the computer to perform the method (see ZHANG: e.g., -- the position identifying module is used for identifying the position of the person based on the image of the person obtained from a plurality of cameras fixed at the relative position; The motion recognition module is configured to identify an action of the person based on a sequence of video frames including an image of the person.--, in abstract; also see Matsugu: e.g., --0126] The present invention may also be realized by supplying a system or an apparatus with a recording (storage) medium storing software program code, i.e., computer-executable process steps, for realizing the function of the present invention, and then causing the computer (CPU or MPU) of the system or the apparatus to read and execute the supplied program code. [0127] In this case, the program code itself read from the recording medium realizes the function of the present invention, and therefore, the recording medium storing the program code is also covered by the present invention. [0128] As described above, the function of the present invention is achieved with the execution of the program code read by the computer. In addition, the function of the present invention may also be achieved by, for example, the OS running on the computer that performs all or part of the processing according to the commands of the program code. [0129] Furthermore, the function of the present invention may also be achieved such that the program code read from a recording medium is written to a memory provided in an expansion card disposed in the computer or an expansion unit connected to the computer, and then, for example, the CPU provided on the expansion card or the expansion unit performs all or part of the processing based on commands in the program code.--, in [0126]-[0130]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIWEN YANG whose telephone number is (571)270-5670. The examiner can normally be reached on Monday-Friday 8:30am-4:30pm east. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on 571-272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /WEI WEN YANG/Primary Examiner, Art Unit 2662
Read full office action

Prosecution Timeline

Jan 06, 2025
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743775
BIOMARKERS OF COLLAGEN FIBER ARCHITECTURE IN EPITHELIAL OVERIAN CANCER (EOC) PATIENTS
2y 11m to grant Granted Sep 22, 2026
Patent 12737879
Machine Learning for Detection of Diseases from External Anterior Eye Images
3y 9m to grant Granted Sep 15, 2026
Patent 12737880
SYSTEMS AND METHODS OF ANALYZING MICROBIOMES USING ARTIFICIAL INTELLIGENCE
3y 4m to grant Granted Sep 15, 2026
Patent 12738057
CUT-PASTE TRAINING AUGMENTATION FOR MACHINE LEARNING MODELS
2y 7m to grant Granted Sep 15, 2026
Patent 12737851
ENHANCED QUALITY BOREHOLE IMAGE GENERATION AND METHOD
2y 3m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
82%
Grant Probability
93%
With Interview (+11.5%)
2y 5m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 684 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month